From 019b014f732b01911e110889087c16e541ba4034 Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:46:54 +0100 Subject: [PATCH 001/311] read-width experiment (gate 1): load classes W=4/16/64, per-load width mix, scratch RMW variant behind a generator flag; 20 packs; CPU verifier and acceptance mirror; emulator shims Nothing changes for the default class: the pinned packs are byte-identical (tests/packs.rs), the v2 draw stream is untouched. LoadClass {mix, load_slots, scratch}: fixed widths w16, w64, w64x4 (4 loads of 64 B), era mixes 50/35/15 and 25/50/25 drawn per load with one extra below(100) roll, and the scratch variant scr0/2/4/8 (persistent warps, 1 MiB per warp, tagged lazy fill, measurement only). A wide load reads the W-aligned address and folds every word: x = dst ^ w0; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]. Program ids carry the class. proto-opencl/host.c taken from opencl-rdna4 a08c371 (--memprobe, select read-back). Co-Authored-By: Claude Fable 5.1 --- igneum-pow/src/accept.rs | 89 +++- igneum-pow/src/emit.rs | 390 ++++++++++++-- igneum-pow/src/generator.rs | 348 ++++++++++++- igneum-pow/src/main.rs | 43 +- igneum-pow/src/memhard.rs | 28 + igneum-pow/src/verify.rs | 213 +++++++- proto-cuda/emu/cuda_runtime.h | 3 + proto-cuda/packs-readwidth/mixA-0/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/mixA-0/kernel.cu | 164 ++++++ .../packs-readwidth/mixA-0/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/mixA-0/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/mixA-0/memhard.h | 108 ++++ .../packs-readwidth/mixA-0/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/mixA-0/program.h | 60 +++ .../packs-readwidth/mixA-0/program.json | 127 +++++ .../packs-readwidth/mixA-0/program.metal | 109 ++++ .../mixA-0/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/mixA-0/vectors.h | 57 ++ .../packs-readwidth/mixA-0/vectors.json | 36 ++ proto-cuda/packs-readwidth/mixA-1/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/mixA-1/kernel.cu | 164 ++++++ .../packs-readwidth/mixA-1/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/mixA-1/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/mixA-1/memhard.h | 108 ++++ .../packs-readwidth/mixA-1/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/mixA-1/program.h | 60 +++ .../packs-readwidth/mixA-1/program.json | 127 +++++ .../packs-readwidth/mixA-1/program.metal | 109 ++++ .../mixA-1/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/mixA-1/vectors.h | 57 ++ .../packs-readwidth/mixA-1/vectors.json | 36 ++ proto-cuda/packs-readwidth/mixA-2/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/mixA-2/kernel.cu | 164 ++++++ .../packs-readwidth/mixA-2/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/mixA-2/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/mixA-2/memhard.h | 108 ++++ .../packs-readwidth/mixA-2/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/mixA-2/program.h | 60 +++ .../packs-readwidth/mixA-2/program.json | 127 +++++ .../packs-readwidth/mixA-2/program.metal | 109 ++++ .../mixA-2/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/mixA-2/vectors.h | 57 ++ .../packs-readwidth/mixA-2/vectors.json | 36 ++ proto-cuda/packs-readwidth/mixA-3/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/mixA-3/kernel.cu | 164 ++++++ .../packs-readwidth/mixA-3/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/mixA-3/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/mixA-3/memhard.h | 108 ++++ .../packs-readwidth/mixA-3/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/mixA-3/program.h | 60 +++ .../packs-readwidth/mixA-3/program.json | 127 +++++ .../packs-readwidth/mixA-3/program.metal | 109 ++++ .../mixA-3/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/mixA-3/vectors.h | 57 ++ .../packs-readwidth/mixA-3/vectors.json | 36 ++ proto-cuda/packs-readwidth/mixA-4/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/mixA-4/kernel.cu | 164 ++++++ .../packs-readwidth/mixA-4/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/mixA-4/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/mixA-4/memhard.h | 108 ++++ .../packs-readwidth/mixA-4/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/mixA-4/program.h | 60 +++ .../packs-readwidth/mixA-4/program.json | 127 +++++ .../packs-readwidth/mixA-4/program.metal | 109 ++++ .../mixA-4/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/mixA-4/vectors.h | 57 ++ .../packs-readwidth/mixA-4/vectors.json | 36 ++ proto-cuda/packs-readwidth/mixA-5/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/mixA-5/kernel.cu | 164 ++++++ .../packs-readwidth/mixA-5/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/mixA-5/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/mixA-5/memhard.h | 108 ++++ .../packs-readwidth/mixA-5/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/mixA-5/program.h | 60 +++ .../packs-readwidth/mixA-5/program.json | 127 +++++ .../packs-readwidth/mixA-5/program.metal | 109 ++++ .../mixA-5/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/mixA-5/vectors.h | 57 ++ .../packs-readwidth/mixA-5/vectors.json | 36 ++ proto-cuda/packs-readwidth/mixB-0/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/mixB-0/kernel.cu | 164 ++++++ .../packs-readwidth/mixB-0/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/mixB-0/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/mixB-0/memhard.h | 108 ++++ .../packs-readwidth/mixB-0/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/mixB-0/program.h | 60 +++ .../packs-readwidth/mixB-0/program.json | 127 +++++ .../packs-readwidth/mixB-0/program.metal | 109 ++++ .../mixB-0/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/mixB-0/vectors.h | 57 ++ .../packs-readwidth/mixB-0/vectors.json | 36 ++ proto-cuda/packs-readwidth/mixB-1/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/mixB-1/kernel.cu | 164 ++++++ .../packs-readwidth/mixB-1/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/mixB-1/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/mixB-1/memhard.h | 108 ++++ .../packs-readwidth/mixB-1/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/mixB-1/program.h | 60 +++ .../packs-readwidth/mixB-1/program.json | 127 +++++ .../packs-readwidth/mixB-1/program.metal | 109 ++++ .../mixB-1/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/mixB-1/vectors.h | 57 ++ .../packs-readwidth/mixB-1/vectors.json | 36 ++ proto-cuda/packs-readwidth/mixB-2/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/mixB-2/kernel.cu | 164 ++++++ .../packs-readwidth/mixB-2/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/mixB-2/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/mixB-2/memhard.h | 108 ++++ .../packs-readwidth/mixB-2/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/mixB-2/program.h | 60 +++ .../packs-readwidth/mixB-2/program.json | 127 +++++ .../packs-readwidth/mixB-2/program.metal | 109 ++++ .../mixB-2/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/mixB-2/vectors.h | 57 ++ .../packs-readwidth/mixB-2/vectors.json | 36 ++ proto-cuda/packs-readwidth/mixB-3/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/mixB-3/kernel.cu | 164 ++++++ .../packs-readwidth/mixB-3/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/mixB-3/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/mixB-3/memhard.h | 108 ++++ .../packs-readwidth/mixB-3/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/mixB-3/program.h | 60 +++ .../packs-readwidth/mixB-3/program.json | 127 +++++ .../packs-readwidth/mixB-3/program.metal | 109 ++++ .../mixB-3/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/mixB-3/vectors.h | 57 ++ .../packs-readwidth/mixB-3/vectors.json | 36 ++ proto-cuda/packs-readwidth/mixB-4/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/mixB-4/kernel.cu | 164 ++++++ .../packs-readwidth/mixB-4/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/mixB-4/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/mixB-4/memhard.h | 108 ++++ .../packs-readwidth/mixB-4/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/mixB-4/program.h | 60 +++ .../packs-readwidth/mixB-4/program.json | 127 +++++ .../packs-readwidth/mixB-4/program.metal | 109 ++++ .../mixB-4/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/mixB-4/vectors.h | 57 ++ .../packs-readwidth/mixB-4/vectors.json | 36 ++ proto-cuda/packs-readwidth/mixB-5/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/mixB-5/kernel.cu | 164 ++++++ .../packs-readwidth/mixB-5/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/mixB-5/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/mixB-5/memhard.h | 108 ++++ .../packs-readwidth/mixB-5/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/mixB-5/program.h | 60 +++ .../packs-readwidth/mixB-5/program.json | 127 +++++ .../packs-readwidth/mixB-5/program.metal | 109 ++++ .../mixB-5/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/mixB-5/vectors.h | 57 ++ .../packs-readwidth/mixB-5/vectors.json | 36 ++ proto-cuda/packs-readwidth/scr0/kernel.cl | 291 +++++++++++ proto-cuda/packs-readwidth/scr0/kernel.cu | 177 +++++++ .../packs-readwidth/scr0/kernel_bound.cl | 393 ++++++++++++++ .../packs-readwidth/scr0/kernel_bound.cu | 136 +++++ proto-cuda/packs-readwidth/scr0/memhard.h | 108 ++++ proto-cuda/packs-readwidth/scr0/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/scr0/program.h | 67 +++ proto-cuda/packs-readwidth/scr0/program.json | 129 +++++ proto-cuda/packs-readwidth/scr0/program.metal | 126 +++++ .../packs-readwidth/scr0/program_bound.metal | 128 +++++ proto-cuda/packs-readwidth/scr0/vectors.h | 57 ++ proto-cuda/packs-readwidth/scr0/vectors.json | 36 ++ proto-cuda/packs-readwidth/scr2/kernel.cl | 291 +++++++++++ proto-cuda/packs-readwidth/scr2/kernel.cu | 177 +++++++ .../packs-readwidth/scr2/kernel_bound.cl | 393 ++++++++++++++ .../packs-readwidth/scr2/kernel_bound.cu | 136 +++++ proto-cuda/packs-readwidth/scr2/memhard.h | 108 ++++ proto-cuda/packs-readwidth/scr2/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/scr2/program.h | 67 +++ proto-cuda/packs-readwidth/scr2/program.json | 129 +++++ proto-cuda/packs-readwidth/scr2/program.metal | 126 +++++ .../packs-readwidth/scr2/program_bound.metal | 128 +++++ proto-cuda/packs-readwidth/scr2/vectors.h | 57 ++ proto-cuda/packs-readwidth/scr2/vectors.json | 36 ++ proto-cuda/packs-readwidth/scr4/kernel.cl | 291 +++++++++++ proto-cuda/packs-readwidth/scr4/kernel.cu | 177 +++++++ .../packs-readwidth/scr4/kernel_bound.cl | 393 ++++++++++++++ .../packs-readwidth/scr4/kernel_bound.cu | 136 +++++ proto-cuda/packs-readwidth/scr4/memhard.h | 108 ++++ proto-cuda/packs-readwidth/scr4/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/scr4/program.h | 67 +++ proto-cuda/packs-readwidth/scr4/program.json | 129 +++++ proto-cuda/packs-readwidth/scr4/program.metal | 126 +++++ .../packs-readwidth/scr4/program_bound.metal | 128 +++++ proto-cuda/packs-readwidth/scr4/vectors.h | 57 ++ proto-cuda/packs-readwidth/scr4/vectors.json | 36 ++ proto-cuda/packs-readwidth/scr8/kernel.cl | 291 +++++++++++ proto-cuda/packs-readwidth/scr8/kernel.cu | 177 +++++++ .../packs-readwidth/scr8/kernel_bound.cl | 393 ++++++++++++++ .../packs-readwidth/scr8/kernel_bound.cu | 136 +++++ proto-cuda/packs-readwidth/scr8/memhard.h | 108 ++++ proto-cuda/packs-readwidth/scr8/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/scr8/program.h | 67 +++ proto-cuda/packs-readwidth/scr8/program.json | 129 +++++ proto-cuda/packs-readwidth/scr8/program.metal | 126 +++++ .../packs-readwidth/scr8/program_bound.metal | 128 +++++ proto-cuda/packs-readwidth/scr8/vectors.h | 57 ++ proto-cuda/packs-readwidth/scr8/vectors.json | 36 ++ proto-cuda/packs-readwidth/w16/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/w16/kernel.cu | 164 ++++++ .../packs-readwidth/w16/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/w16/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/w16/memhard.h | 108 ++++ proto-cuda/packs-readwidth/w16/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/w16/program.h | 60 +++ proto-cuda/packs-readwidth/w16/program.json | 127 +++++ proto-cuda/packs-readwidth/w16/program.metal | 109 ++++ .../packs-readwidth/w16/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/w16/vectors.h | 57 ++ proto-cuda/packs-readwidth/w16/vectors.json | 36 ++ proto-cuda/packs-readwidth/w4/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/w4/kernel.cu | 164 ++++++ proto-cuda/packs-readwidth/w4/kernel_bound.cl | 372 +++++++++++++ proto-cuda/packs-readwidth/w4/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/w4/memhard.h | 108 ++++ proto-cuda/packs-readwidth/w4/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/w4/program.h | 51 ++ proto-cuda/packs-readwidth/w4/program.json | 121 +++++ proto-cuda/packs-readwidth/w4/program.metal | 109 ++++ .../packs-readwidth/w4/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/w4/vectors.h | 57 ++ proto-cuda/packs-readwidth/w4/vectors.json | 36 ++ proto-cuda/packs-readwidth/w64/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/w64/kernel.cu | 164 ++++++ .../packs-readwidth/w64/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/w64/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/w64/memhard.h | 108 ++++ proto-cuda/packs-readwidth/w64/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/w64/program.h | 60 +++ proto-cuda/packs-readwidth/w64/program.json | 127 +++++ proto-cuda/packs-readwidth/w64/program.metal | 109 ++++ .../packs-readwidth/w64/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/w64/vectors.h | 57 ++ proto-cuda/packs-readwidth/w64/vectors.json | 36 ++ proto-cuda/packs-readwidth/w64x4/kernel.cl | 278 ++++++++++ proto-cuda/packs-readwidth/w64x4/kernel.cu | 164 ++++++ .../packs-readwidth/w64x4/kernel_bound.cl | 372 +++++++++++++ .../packs-readwidth/w64x4/kernel_bound.cu | 123 +++++ proto-cuda/packs-readwidth/w64x4/memhard.h | 108 ++++ .../packs-readwidth/w64x4/memhard.metal | 106 ++++ proto-cuda/packs-readwidth/w64x4/program.h | 60 +++ proto-cuda/packs-readwidth/w64x4/program.json | 127 +++++ .../packs-readwidth/w64x4/program.metal | 109 ++++ .../packs-readwidth/w64x4/program_bound.metal | 111 ++++ proto-cuda/packs-readwidth/w64x4/vectors.h | 57 ++ proto-cuda/packs-readwidth/w64x4/vectors.json | 36 ++ proto-opencl/README.md | 20 + proto-opencl/WAVEFRONT.md | 9 + proto-opencl/emu/emu_opencl.h | 6 + proto-opencl/host.c | 491 +++++++++++++++++- proto-opencl/test-host.sh | 10 + proto-opencl/test_host.c | 71 +++ 253 files changed, 35043 insertions(+), 95 deletions(-) create mode 100644 proto-cuda/packs-readwidth/mixA-0/kernel.cl create mode 100644 proto-cuda/packs-readwidth/mixA-0/kernel.cu create mode 100644 proto-cuda/packs-readwidth/mixA-0/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/mixA-0/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/mixA-0/memhard.h create mode 100644 proto-cuda/packs-readwidth/mixA-0/memhard.metal create mode 100644 proto-cuda/packs-readwidth/mixA-0/program.h create mode 100644 proto-cuda/packs-readwidth/mixA-0/program.json create mode 100644 proto-cuda/packs-readwidth/mixA-0/program.metal create mode 100644 proto-cuda/packs-readwidth/mixA-0/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/mixA-0/vectors.h create mode 100644 proto-cuda/packs-readwidth/mixA-0/vectors.json create mode 100644 proto-cuda/packs-readwidth/mixA-1/kernel.cl create mode 100644 proto-cuda/packs-readwidth/mixA-1/kernel.cu create mode 100644 proto-cuda/packs-readwidth/mixA-1/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/mixA-1/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/mixA-1/memhard.h create mode 100644 proto-cuda/packs-readwidth/mixA-1/memhard.metal create mode 100644 proto-cuda/packs-readwidth/mixA-1/program.h create mode 100644 proto-cuda/packs-readwidth/mixA-1/program.json create mode 100644 proto-cuda/packs-readwidth/mixA-1/program.metal create mode 100644 proto-cuda/packs-readwidth/mixA-1/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/mixA-1/vectors.h create mode 100644 proto-cuda/packs-readwidth/mixA-1/vectors.json create mode 100644 proto-cuda/packs-readwidth/mixA-2/kernel.cl create mode 100644 proto-cuda/packs-readwidth/mixA-2/kernel.cu create mode 100644 proto-cuda/packs-readwidth/mixA-2/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/mixA-2/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/mixA-2/memhard.h create mode 100644 proto-cuda/packs-readwidth/mixA-2/memhard.metal create mode 100644 proto-cuda/packs-readwidth/mixA-2/program.h create mode 100644 proto-cuda/packs-readwidth/mixA-2/program.json create mode 100644 proto-cuda/packs-readwidth/mixA-2/program.metal create mode 100644 proto-cuda/packs-readwidth/mixA-2/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/mixA-2/vectors.h create mode 100644 proto-cuda/packs-readwidth/mixA-2/vectors.json create mode 100644 proto-cuda/packs-readwidth/mixA-3/kernel.cl create mode 100644 proto-cuda/packs-readwidth/mixA-3/kernel.cu create mode 100644 proto-cuda/packs-readwidth/mixA-3/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/mixA-3/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/mixA-3/memhard.h create mode 100644 proto-cuda/packs-readwidth/mixA-3/memhard.metal create mode 100644 proto-cuda/packs-readwidth/mixA-3/program.h create mode 100644 proto-cuda/packs-readwidth/mixA-3/program.json create mode 100644 proto-cuda/packs-readwidth/mixA-3/program.metal create mode 100644 proto-cuda/packs-readwidth/mixA-3/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/mixA-3/vectors.h create mode 100644 proto-cuda/packs-readwidth/mixA-3/vectors.json create mode 100644 proto-cuda/packs-readwidth/mixA-4/kernel.cl create mode 100644 proto-cuda/packs-readwidth/mixA-4/kernel.cu create mode 100644 proto-cuda/packs-readwidth/mixA-4/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/mixA-4/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/mixA-4/memhard.h create mode 100644 proto-cuda/packs-readwidth/mixA-4/memhard.metal create mode 100644 proto-cuda/packs-readwidth/mixA-4/program.h create mode 100644 proto-cuda/packs-readwidth/mixA-4/program.json create mode 100644 proto-cuda/packs-readwidth/mixA-4/program.metal create mode 100644 proto-cuda/packs-readwidth/mixA-4/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/mixA-4/vectors.h create mode 100644 proto-cuda/packs-readwidth/mixA-4/vectors.json create mode 100644 proto-cuda/packs-readwidth/mixA-5/kernel.cl create mode 100644 proto-cuda/packs-readwidth/mixA-5/kernel.cu create mode 100644 proto-cuda/packs-readwidth/mixA-5/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/mixA-5/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/mixA-5/memhard.h create mode 100644 proto-cuda/packs-readwidth/mixA-5/memhard.metal create mode 100644 proto-cuda/packs-readwidth/mixA-5/program.h create mode 100644 proto-cuda/packs-readwidth/mixA-5/program.json create mode 100644 proto-cuda/packs-readwidth/mixA-5/program.metal create mode 100644 proto-cuda/packs-readwidth/mixA-5/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/mixA-5/vectors.h create mode 100644 proto-cuda/packs-readwidth/mixA-5/vectors.json create mode 100644 proto-cuda/packs-readwidth/mixB-0/kernel.cl create mode 100644 proto-cuda/packs-readwidth/mixB-0/kernel.cu create mode 100644 proto-cuda/packs-readwidth/mixB-0/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/mixB-0/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/mixB-0/memhard.h create mode 100644 proto-cuda/packs-readwidth/mixB-0/memhard.metal create mode 100644 proto-cuda/packs-readwidth/mixB-0/program.h create mode 100644 proto-cuda/packs-readwidth/mixB-0/program.json create mode 100644 proto-cuda/packs-readwidth/mixB-0/program.metal create mode 100644 proto-cuda/packs-readwidth/mixB-0/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/mixB-0/vectors.h create mode 100644 proto-cuda/packs-readwidth/mixB-0/vectors.json create mode 100644 proto-cuda/packs-readwidth/mixB-1/kernel.cl create mode 100644 proto-cuda/packs-readwidth/mixB-1/kernel.cu create mode 100644 proto-cuda/packs-readwidth/mixB-1/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/mixB-1/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/mixB-1/memhard.h create mode 100644 proto-cuda/packs-readwidth/mixB-1/memhard.metal create mode 100644 proto-cuda/packs-readwidth/mixB-1/program.h create mode 100644 proto-cuda/packs-readwidth/mixB-1/program.json create mode 100644 proto-cuda/packs-readwidth/mixB-1/program.metal create mode 100644 proto-cuda/packs-readwidth/mixB-1/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/mixB-1/vectors.h create mode 100644 proto-cuda/packs-readwidth/mixB-1/vectors.json create mode 100644 proto-cuda/packs-readwidth/mixB-2/kernel.cl create mode 100644 proto-cuda/packs-readwidth/mixB-2/kernel.cu create mode 100644 proto-cuda/packs-readwidth/mixB-2/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/mixB-2/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/mixB-2/memhard.h create mode 100644 proto-cuda/packs-readwidth/mixB-2/memhard.metal create mode 100644 proto-cuda/packs-readwidth/mixB-2/program.h create mode 100644 proto-cuda/packs-readwidth/mixB-2/program.json create mode 100644 proto-cuda/packs-readwidth/mixB-2/program.metal create mode 100644 proto-cuda/packs-readwidth/mixB-2/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/mixB-2/vectors.h create mode 100644 proto-cuda/packs-readwidth/mixB-2/vectors.json create mode 100644 proto-cuda/packs-readwidth/mixB-3/kernel.cl create mode 100644 proto-cuda/packs-readwidth/mixB-3/kernel.cu create mode 100644 proto-cuda/packs-readwidth/mixB-3/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/mixB-3/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/mixB-3/memhard.h create mode 100644 proto-cuda/packs-readwidth/mixB-3/memhard.metal create mode 100644 proto-cuda/packs-readwidth/mixB-3/program.h create mode 100644 proto-cuda/packs-readwidth/mixB-3/program.json create mode 100644 proto-cuda/packs-readwidth/mixB-3/program.metal create mode 100644 proto-cuda/packs-readwidth/mixB-3/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/mixB-3/vectors.h create mode 100644 proto-cuda/packs-readwidth/mixB-3/vectors.json create mode 100644 proto-cuda/packs-readwidth/mixB-4/kernel.cl create mode 100644 proto-cuda/packs-readwidth/mixB-4/kernel.cu create mode 100644 proto-cuda/packs-readwidth/mixB-4/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/mixB-4/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/mixB-4/memhard.h create mode 100644 proto-cuda/packs-readwidth/mixB-4/memhard.metal create mode 100644 proto-cuda/packs-readwidth/mixB-4/program.h create mode 100644 proto-cuda/packs-readwidth/mixB-4/program.json create mode 100644 proto-cuda/packs-readwidth/mixB-4/program.metal create mode 100644 proto-cuda/packs-readwidth/mixB-4/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/mixB-4/vectors.h create mode 100644 proto-cuda/packs-readwidth/mixB-4/vectors.json create mode 100644 proto-cuda/packs-readwidth/mixB-5/kernel.cl create mode 100644 proto-cuda/packs-readwidth/mixB-5/kernel.cu create mode 100644 proto-cuda/packs-readwidth/mixB-5/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/mixB-5/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/mixB-5/memhard.h create mode 100644 proto-cuda/packs-readwidth/mixB-5/memhard.metal create mode 100644 proto-cuda/packs-readwidth/mixB-5/program.h create mode 100644 proto-cuda/packs-readwidth/mixB-5/program.json create mode 100644 proto-cuda/packs-readwidth/mixB-5/program.metal create mode 100644 proto-cuda/packs-readwidth/mixB-5/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/mixB-5/vectors.h create mode 100644 proto-cuda/packs-readwidth/mixB-5/vectors.json create mode 100644 proto-cuda/packs-readwidth/scr0/kernel.cl create mode 100644 proto-cuda/packs-readwidth/scr0/kernel.cu create mode 100644 proto-cuda/packs-readwidth/scr0/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/scr0/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/scr0/memhard.h create mode 100644 proto-cuda/packs-readwidth/scr0/memhard.metal create mode 100644 proto-cuda/packs-readwidth/scr0/program.h create mode 100644 proto-cuda/packs-readwidth/scr0/program.json create mode 100644 proto-cuda/packs-readwidth/scr0/program.metal create mode 100644 proto-cuda/packs-readwidth/scr0/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/scr0/vectors.h create mode 100644 proto-cuda/packs-readwidth/scr0/vectors.json create mode 100644 proto-cuda/packs-readwidth/scr2/kernel.cl create mode 100644 proto-cuda/packs-readwidth/scr2/kernel.cu create mode 100644 proto-cuda/packs-readwidth/scr2/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/scr2/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/scr2/memhard.h create mode 100644 proto-cuda/packs-readwidth/scr2/memhard.metal create mode 100644 proto-cuda/packs-readwidth/scr2/program.h create mode 100644 proto-cuda/packs-readwidth/scr2/program.json create mode 100644 proto-cuda/packs-readwidth/scr2/program.metal create mode 100644 proto-cuda/packs-readwidth/scr2/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/scr2/vectors.h create mode 100644 proto-cuda/packs-readwidth/scr2/vectors.json create mode 100644 proto-cuda/packs-readwidth/scr4/kernel.cl create mode 100644 proto-cuda/packs-readwidth/scr4/kernel.cu create mode 100644 proto-cuda/packs-readwidth/scr4/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/scr4/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/scr4/memhard.h create mode 100644 proto-cuda/packs-readwidth/scr4/memhard.metal create mode 100644 proto-cuda/packs-readwidth/scr4/program.h create mode 100644 proto-cuda/packs-readwidth/scr4/program.json create mode 100644 proto-cuda/packs-readwidth/scr4/program.metal create mode 100644 proto-cuda/packs-readwidth/scr4/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/scr4/vectors.h create mode 100644 proto-cuda/packs-readwidth/scr4/vectors.json create mode 100644 proto-cuda/packs-readwidth/scr8/kernel.cl create mode 100644 proto-cuda/packs-readwidth/scr8/kernel.cu create mode 100644 proto-cuda/packs-readwidth/scr8/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/scr8/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/scr8/memhard.h create mode 100644 proto-cuda/packs-readwidth/scr8/memhard.metal create mode 100644 proto-cuda/packs-readwidth/scr8/program.h create mode 100644 proto-cuda/packs-readwidth/scr8/program.json create mode 100644 proto-cuda/packs-readwidth/scr8/program.metal create mode 100644 proto-cuda/packs-readwidth/scr8/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/scr8/vectors.h create mode 100644 proto-cuda/packs-readwidth/scr8/vectors.json create mode 100644 proto-cuda/packs-readwidth/w16/kernel.cl create mode 100644 proto-cuda/packs-readwidth/w16/kernel.cu create mode 100644 proto-cuda/packs-readwidth/w16/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/w16/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/w16/memhard.h create mode 100644 proto-cuda/packs-readwidth/w16/memhard.metal create mode 100644 proto-cuda/packs-readwidth/w16/program.h create mode 100644 proto-cuda/packs-readwidth/w16/program.json create mode 100644 proto-cuda/packs-readwidth/w16/program.metal create mode 100644 proto-cuda/packs-readwidth/w16/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/w16/vectors.h create mode 100644 proto-cuda/packs-readwidth/w16/vectors.json create mode 100644 proto-cuda/packs-readwidth/w4/kernel.cl create mode 100644 proto-cuda/packs-readwidth/w4/kernel.cu create mode 100644 proto-cuda/packs-readwidth/w4/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/w4/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/w4/memhard.h create mode 100644 proto-cuda/packs-readwidth/w4/memhard.metal create mode 100644 proto-cuda/packs-readwidth/w4/program.h create mode 100644 proto-cuda/packs-readwidth/w4/program.json create mode 100644 proto-cuda/packs-readwidth/w4/program.metal create mode 100644 proto-cuda/packs-readwidth/w4/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/w4/vectors.h create mode 100644 proto-cuda/packs-readwidth/w4/vectors.json create mode 100644 proto-cuda/packs-readwidth/w64/kernel.cl create mode 100644 proto-cuda/packs-readwidth/w64/kernel.cu create mode 100644 proto-cuda/packs-readwidth/w64/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/w64/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/w64/memhard.h create mode 100644 proto-cuda/packs-readwidth/w64/memhard.metal create mode 100644 proto-cuda/packs-readwidth/w64/program.h create mode 100644 proto-cuda/packs-readwidth/w64/program.json create mode 100644 proto-cuda/packs-readwidth/w64/program.metal create mode 100644 proto-cuda/packs-readwidth/w64/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/w64/vectors.h create mode 100644 proto-cuda/packs-readwidth/w64/vectors.json create mode 100644 proto-cuda/packs-readwidth/w64x4/kernel.cl create mode 100644 proto-cuda/packs-readwidth/w64x4/kernel.cu create mode 100644 proto-cuda/packs-readwidth/w64x4/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/w64x4/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/w64x4/memhard.h create mode 100644 proto-cuda/packs-readwidth/w64x4/memhard.metal create mode 100644 proto-cuda/packs-readwidth/w64x4/program.h create mode 100644 proto-cuda/packs-readwidth/w64x4/program.json create mode 100644 proto-cuda/packs-readwidth/w64x4/program.metal create mode 100644 proto-cuda/packs-readwidth/w64x4/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/w64x4/vectors.h create mode 100644 proto-cuda/packs-readwidth/w64x4/vectors.json create mode 100755 proto-opencl/test-host.sh create mode 100644 proto-opencl/test_host.c diff --git a/igneum-pow/src/accept.rs b/igneum-pow/src/accept.rs index a3a59d7a4..7f8c4438a 100644 --- a/igneum-pow/src/accept.rs +++ b/igneum-pow/src/accept.rs @@ -14,9 +14,9 @@ //! costs about a millisecond on one core. The census (section 7.3) checked on 100,000 programs that the //! closed-form verdict agrees with the memory-hard one on all but 39 threshold-edge cases. -use crate::generator::{Instr, Op, Program, INSTR_COUNT, ITERATIONS, LANES}; +use crate::generator::{Instr, Op, Program, INSTR_COUNT, ITERATIONS, LANES, SCRATCH_SLOT_MASK}; use crate::seed::{fnv1a64, SplitMix64}; -use crate::verify::{dataset_elem, splitmix32}; +use crate::verify::{dataset_elem, fold_words, splitmix32, ScratchModel}; /// Units (32-lane warps) the dynamic test interprets. pub const ACCEPT_UNITS: usize = 64; @@ -33,6 +33,12 @@ pub const BIAS_TOLERANCE: u32 = 136; /// Distinct addresses per lane per evaluation, summed over 2,048 evaluations, must exceed this (mean above 120). pub const MIN_DISTINCT_SUM: u64 = 245_760; +/// The distinct-address bound for a program with `loads` loads per hash: the same 120 of 128 ratio, so +/// [`MIN_DISTINCT_SUM`] for the lottery hash and `loads x 1,920` for the read-width classes with other counts. +pub fn min_distinct_sum(loads: usize) -> u64 { + loads as u64 * ACCEPT_HASHES as u64 * 120 / 128 +} + /// Why a candidate was rejected. The verdict (accept or reject) is what consensus depends on; the reason is the /// first failing test in the order of the module table. #[derive(Clone, Copy, Debug, PartialEq, Eq)] @@ -67,7 +73,7 @@ impl std::fmt::Display for Reject { Reject::Saturated { count } => write!(f, "(c) {count} of 16384 final register values saturated (limit 163)"), Reject::OutputBias { bit, ones } => write!(f, "(c) output bit {bit} set in {ones} of 2048 hashes"), Reject::DistinctAddresses { sum } => { - write!(f, "(c) distinct addresses {sum} over 2048 hashes (mean {:.2}, needs above 120)", *sum as f64 / 2048.0) + write!(f, "(c) distinct addresses {sum} over 2048 hashes (mean {:.2}, needs above 120 of 128 loads)", *sum as f64 / 2048.0) } } } @@ -183,12 +189,28 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut } let mut idx = [0u32; LANES]; let mut nload = 0usize; + let mut scratch = if p.has_scratch() { Some(ScratchModel::new()) } else { None }; for it in 0..ITERATIONS { let sel = r[0]; for (k, ins) in p.instrs.iter().enumerate() { let d = ins.dst as usize; let a = ins.src as usize; match ins.op { + Op::Scratch => { + // Variant 5: the slot stands in for the address (bit 31 set so it never aliases a dataset word). + let m = scratch.as_mut().expect("a scratch op needs a scratch class"); + for lane in 0..LANES { + idx[lane] = r[a][lane] & SCRATCH_SLOT_MASK; + } + if idx.iter().all(|&x| x == idx[0]) { + return Err(Reject::LaneConstantSite { iteration: it as u8, instr: k as u8, unit: unit as u8 }); + } + for lane in 0..LANES { + r[d][lane] = m.rmw(&p.seed, base, lane, idx[lane], r[d][lane]); + lane_addrs[lane * loads + nload] = 0x8000_0000 | idx[lane]; + } + nload += 1; + } Op::Add => { let (imm, imm2, bit) = (ins.imm, ins.imm2, ins.bit as u32); let src = r[a]; @@ -255,14 +277,26 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut } } Op::Load => { + // Read-width experiment: a load of `width` words reads from the aligned address and folds every + // word (verify::fold_words); width 1 is the lottery hash's xor of one word. + let width = ins.width as usize; + let align = !(ins.width as u32 - 1); for lane in 0..LANES { - idx[lane] = r[a][lane] & mask; + idx[lane] = (r[a][lane] & mask) & align; } if idx.iter().all(|&x| x == idx[0]) { return Err(Reject::LaneConstantSite { iteration: it as u8, instr: k as u8, unit: unit as u8 }); } for lane in 0..LANES { - r[d][lane] ^= dataset_elem(idx[lane], d0, d1); + if width == 1 { + r[d][lane] ^= dataset_elem(idx[lane], d0, d1); + } else { + let mut w = [0u32; 16]; + for j in 0..width { + w[j] = dataset_elem(idx[lane] + j as u32, d0, d1); + } + r[d][lane] = fold_words(r[d][lane], &w[..width]); + } lane_addrs[lane * loads + nload] = idx[lane]; } nload += 1; @@ -333,7 +367,7 @@ pub fn check_dynamic(p: &Program) -> Result { } bias_max = bias_max.max(d); } - if acc.distinct_sum <= MIN_DISTINCT_SUM { + if acc.distinct_sum <= min_distinct_sum(loads) { return Err(Reject::DistinctAddresses { sum: acc.distinct_sum }); } Ok(AcceptReport { distinct_sum: acc.distinct_sum, saturated: acc.saturated, bias_max }) @@ -348,9 +382,50 @@ pub fn check(p: &Program) -> Result { #[cfg(test)] mod tests { use super::*; - use crate::generator::{candidate, generate, GeneratorConfig, generate_v1}; + use crate::generator::{candidate, candidate_class, generate, generate_class, GeneratorConfig, generate_v1, LoadClass}; use crate::verify::{DatasetMode, DatasetSource}; + #[test] + fn distinct_bound_scales_with_the_load_count() { + assert_eq!(min_distinct_sum(128), MIN_DISTINCT_SUM); + assert_eq!(min_distinct_sum(32), 61_440); + } + + /// The read-width classes pass the rule at about the version 2 rate, and the instrumented interpreter agrees + /// with `verify.rs` on every class (the fold is shared, the addresses are aligned the same way). + #[test] + fn classes_pass_and_match_verify() { + for name in ["w16", "w64", "w64x4", "50,35,15", "25,50,25", "scr2", "scr8"] { + let c = LoadClass::parse(name).unwrap(); + let p = generate_class("igneum-genesis", c); + assert!(check(&p).is_ok(), "{name}"); + let mut rejected = 0; + for i in 0..60u32 { + let s = format!("igneum-rw-accept/{i}"); + let q = candidate_class(&s, s.as_bytes(), 0, c); + if check(&q).is_err() { + rejected += 1; + } + } + assert!(rejected < 15, "{name}: {rejected} of 60 rejected"); + let ds = DatasetSource::from_key(p.seed, DatasetMode::ClosedForm, ACCEPT_DATASET_LOG2); + let bases = accept_base_nonces(&p.seed); + let loads = p.loads_per_hash(); + let mut acc = Acc { and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 }; + let mut la = vec![0u32; LANES * loads]; + let mut ones = [0u32; 64]; + for (u, &b) in bases.iter().enumerate() { + run_unit(&p, u, b, &mut acc, &mut la).unwrap(); + for h in crate::verify::hash_warp(&p, b, &ds) { + for j in 0..64 { + ones[j] += ((h >> j) & 1) as u32; + } + } + } + assert_eq!(acc.bit_ones, ones, "{name}: bit counts match the reference interpreter"); + } + } + /// The instrumented interpreter agrees with `verify.rs` on the closed-form dataset keyed by the seed words. #[test] fn instrumented_interpreter_matches_verify() { diff --git a/igneum-pow/src/emit.rs b/igneum-pow/src/emit.rs index 717c4d452..4ea7a9f03 100644 --- a/igneum-pow/src/emit.rs +++ b/igneum-pow/src/emit.rs @@ -9,13 +9,207 @@ //! One deliberate difference from the Swift: `program_json` writes the cache line mask inside the `"item"` string //! as a bare `0x003fffff`. The Swift writes it quoted (`jhex`), which is not valid JSON. -use crate::generator::{Op, Program, GENERATOR_VERSION, INSTR_COUNT, ITERATIONS, LOAD_SLOTS}; +use crate::generator::{Instr, Op, Program, GENERATOR_VERSION, INSTR_COUNT, ITERATIONS, LOAD_SLOTS}; use crate::memhard::{ MixParams, CACHE_LINES_PER_SEGMENT, CACHE_LINE_MASK, CACHE_LOG2_WORDS, CACHE_SEGMENTS, CACHE_SEGMENT_LOG2_LINES, CACHE_TAG, CACHE_WORDS, CHACHA_ROUNDS, CHACHA_SIGMA, ITEM_ROUNDS, }; use crate::seed::SplitMix64; -use crate::verify::{DatasetMode, DatasetSource, Epoch}; +use crate::generator::{SCRATCH_BYTES_PER_WARP, SCRATCH_SLOTS, SCRATCH_SLOT_MASK, SCRATCH_WORDS_PER_LANE}; +use crate::verify::{DatasetMode, DatasetSource, Epoch, FOLD_MUL, FOLD_ROT}; + +/// Where the words of a wide load come from (read-width experiment). +#[derive(Clone, Copy, PartialEq, Eq)] +enum WideSource { + /// `dataset`/`ds`: vector loads from the stored dataset. + Stored, + /// the closed form per word (Metal inline shortcut kernel). + InlineClosed, + /// one `mh_item` derivation per load, words taken from it (Metal inline memory-hard kernel). + InlineMemhard, +} + +/// One wide `load` as a single statement block (read-width experiment, 5 October 2026): `width` words from the +/// address aligned down to `width` words, folded into `dst` as `verify::fold_words`. The emitted text is the +/// same shape in the three dialects: the vector loads differ (`uint4` pointer on Metal and CUDA, `vload4` on +/// OpenCL C 1.2). For `width == 1` the caller emits the lottery hash's one-word form instead. +fn wide_load_stmt(dialect: CoreDialect, d: &str, a: &str, width: u8, src: WideSource, closed: Option<(u32, u32)>) -> String { + debug_assert!(width == 4 || width == 16); + let (u, mask, base_ptr) = match dialect { + CoreDialect::Metal => ("uint", "MASK", "dataset"), + CoreDialect::Cuda => ("uint32_t", "mask", "ds"), + CoreDialect::OpenCl => ("uint", "mask", "ds"), + }; + let vectors = width as usize / 4; + let mut s = String::with_capacity(400); + s.push_str(&format!("{{ {u} b_ = ({a} & {mask}) & ~{}u; ", width as u32 - 1)); + match src { + WideSource::Stored => match dialect { + CoreDialect::Metal => s.push_str(&format!("device const uint4* l_ = (device const uint4*)({base_ptr} + b_); ")), + CoreDialect::Cuda => s.push_str(&format!("const uint4* l_ = (const uint4*)({base_ptr} + b_); ")), + CoreDialect::OpenCl => {} + }, + WideSource::InlineClosed => {} + WideSource::InlineMemhard => s.push_str("uint s_[16]; mh_item(cache, b_ >> 4u, s_); "), + } + let word = |j: usize| -> String { + match src { + WideSource::Stored => format!("v{}_.{}", j / 4, ["x", "y", "z", "w"][j % 4]), + WideSource::InlineClosed => { + let (d0, d1) = closed.expect("closed-form words need d0, d1"); + format!("ds_elem(b_ + {j}u, {}, {})", hex(d0), hex(d1)) + } + WideSource::InlineMemhard => format!("s_[(b_ & 15u) + {j}u]"), + } + }; + if src == WideSource::Stored { + for v in 0..vectors { + match dialect { + CoreDialect::OpenCl => s.push_str(&format!("uint4 v{v}_ = vload4({v}u, {base_ptr} + b_); ")), + _ => s.push_str(&format!("uint4 v{v}_ = l_[{v}]; ")), + } + } + } + s.push_str(&format!("{u} x_ = {d} ^ {}; ", word(0))); + for j in 1..width as usize { + s.push_str(&format!("x_ = (rotl_imm(x_, {FOLD_ROT}u) * {}) ^ {}; ", hex(FOLD_MUL), word(j))); + } + s.push_str(&format!("{d} = x_; }}")); + s +} + +/// The load class lines of program.h (empty for the lottery hash, so the pinned packs do not change). +fn class_header_lines(p: &Program) -> String { + if p.class.is_v2() { + return String::new(); + } + let mut s = String::new(); + s.push_str("// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +"); + s.push_str("// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +"); + s.push_str(&format!("#define IGNEUM_LOAD_CLASS {} +", jstr(&p.class.name()))); + s.push_str(&format!("#define IGNEUM_LOAD_SLOTS {} +", p.class.load_slots)); + s.push_str(&format!("#define IGNEUM_LOAD_MIX {{ {}, {}, {} }} +", p.class.mix[0], p.class.mix[1], p.class.mix[2])); + let c = p.width_counts(); + s.push_str(&format!("#define IGNEUM_LOAD_WIDTH_COUNTS {{ {}, {}, {} }} // loads of 4, 16, 64 bytes per program +", c[0], c[1], c[2])); + s.push_str(&format!("#define IGNEUM_BYTES_PER_HASH {} +", p.bytes_per_hash())); + s.push_str(&format!("#define IGNEUM_FOLD_ROT {FOLD_ROT} +")); + s.push_str(&format!("#define IGNEUM_FOLD_MUL {} +", hex(FOLD_MUL))); + s +} + +/// The width of a load instruction's statement, for the emitters (1 for every non-load op). +fn load_width(ins: &Instr) -> u8 { + if ins.op == Op::Load { + ins.width + } else { + 1 + } +} + +/// Variant 5 prelude: `scr_fill(gbase, lane, slot, j)`, the fill word of a scratch slot (`verify::scratch_fill`), +/// with the program's seed words as literals. +fn scratch_prelude(p: &Program, dialect: CoreDialect) -> String { + if !p.has_scratch() { + return String::new(); + } + let (u, fn_) = match dialect { + CoreDialect::Metal => ("uint", "inline"), + CoreDialect::Cuda => ("uint32_t", "__device__ __forceinline__"), + CoreDialect::OpenCl => ("uint", "static inline"), + }; + let mut s = String::new(); + s.push_str(&format!("// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, {} slots of +", SCRATCH_SLOTS)); + s.push_str("// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +"); + s.push_str("// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +"); + s.push_str(&format!( + "{fn_} {u} scr_fill({u} gbase, {u} lane, {u} slot, {u} j) {{ {u} sw = (j == 0u) ? {} : ((j == 1u) ? {} : {}); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); }} +", + hex(p.seed[0]), + hex(p.seed[1]), + hex(p.seed[2]) + )); + s +} + +/// Variant 5: one scratch read-modify-write as a statement block. `arena`, `tag`, `gbase` and `lane` are in scope +/// (the persistent prologue). Reads 16 bytes, folds the three data words into dst, rewrites the slot behind the tag. +fn scratch_stmt(dialect: CoreDialect, d: &str, a: &str) -> String { + let (u, load, store) = match dialect { + CoreDialect::Metal => ("uint", "uint4 v_ = *(device const uint4*)(arena + s_ * 4u);", "*(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_);"), + CoreDialect::Cuda => ("uint32_t", "uint4 v_ = *(const uint4*)(arena + s_ * 4u);", "*(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_);"), + CoreDialect::OpenCl => ("uint", "uint4 v_ = vload4(s_, arena);", "vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena);"), + }; + format!( + "{{ {u} s_ = {a} & {}u; {load} {u} m_ = (v_.x == tag) ? 0xffffffffu : 0u; {u} w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); {u} w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); {u} w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); {u} x_ = {d} ^ w0_; x_ = (rotl_imm(x_, {FOLD_ROT}u) * {k}) ^ w1_; x_ = (rotl_imm(x_, {FOLD_ROT}u) * {k}) ^ w2_; {d} = x_; {store} }}", + SCRATCH_SLOT_MASK, + k = hex(FOLD_MUL) + ) +} + +/// Variant 5: the persistent-warp prologue. The kernel is launched with N warps (the resident count, the host's +/// choice); warp `w` owns arena `w` and runs the units `w, w + N, w + 2N, ...` of the launch. Inside the loop the +/// lottery hash's text is unchanged: `gid` is the unit's first output index plus the lane. The host MUST launch +/// `groups` as a multiple of N (a uniform trip count: the OpenCL local-memory exchange carries a barrier). +fn persistent_prologue(dialect: CoreDialect) -> String { + let (u, tid, nthreads, ptr) = match dialect { + CoreDialect::Metal => ("uint", "tid", "nthreads", "device uint*"), + CoreDialect::Cuda => ("uint32_t", "(blockIdx.x * blockDim.x + threadIdx.x)", "(gridDim.x * blockDim.x)", "uint32_t*"), + CoreDialect::OpenCl => ("uint", "(uint)get_global_id(0)", "(uint)get_global_size(0)", "__global uint*"), + }; + let mut s = String::new(); + s.push_str(&format!(" {u} lane = {tid} & 31u; +")); + s.push_str(&format!(" {u} warp_ = {tid} >> 5; +")); + s.push_str(&format!(" {u} nwarps_ = {nthreads} >> 5; +")); + s.push_str(&format!(" {ptr} arena = scratch + ((size_t)warp_ * 32u + lane) * {}u; +", SCRATCH_WORDS_PER_LANE)); + s.push_str(&format!(" for ({u} g_ = warp_; g_ < groups; g_ += nwarps_) {{ +")); + s.push_str(&format!(" {u} gid = g_ * 32u + lane; +")); + s.push_str(&format!(" {u} gbase = baseNonce + g_ * 32u; +")); + s.push_str(&format!(" {u} tag = salt + g_; +")); + s +} + +/// The scratch lines of program.h (variant 5). +fn scratch_header_lines(p: &Program) -> String { + if !p.has_scratch() { + return String::new(); + } + let mut s = String::new(); + s.push_str("// Variant 5: persistent warps, a 1 MiB scratch per launched warp (the host launches N warps and passes scratch, +"); + s.push_str("// groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). +"); + s.push_str("#define IGNEUM_PERSISTENT_WARPS 1 +"); + s.push_str(&format!("#define IGNEUM_SCRATCH_OPS {} // scratch read-modify-writes per program ({} per hash) +", p.class.scratch_slots(), p.scratch_ops_per_hash())); + s.push_str(&format!("#define IGNEUM_SCRATCH_SLOTS {SCRATCH_SLOTS}u +")); + s.push_str(&format!("#define IGNEUM_SCRATCH_WORDS_PER_LANE {SCRATCH_WORDS_PER_LANE}u +")); + s.push_str(&format!("#define IGNEUM_SCRATCH_BYTES_PER_WARP {SCRATCH_BYTES_PER_WARP}u +")); + s +} pub fn hex(v: u32) -> String { format!("0x{v:08x}u") @@ -259,6 +453,7 @@ fn metal_program_impl(p: &Program, dataset_log2: u32, source: LoadSource, bound: s.push('\n'); buffer0 = "device const uint* cache [[buffer(0)]]"; } + s.push_str(&scratch_prelude(p, CoreDialect::Metal)); if bound { s.push_str("// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.\n"); s.push_str(&format!("kernel void igneum_hash_bound({buffer0},\n")); @@ -270,7 +465,17 @@ fn metal_program_impl(p: &Program, dataset_log2: u32, source: LoadSource, bound: if bound { s.push_str(" constant uint* initw [[buffer(3)]],\n"); } - s.push_str(" uint gid [[thread_position_in_grid]]) {\n"); + if p.has_scratch() { + let b = if bound { 4 } else { 3 }; + s.push_str(&format!(" device uint* scratch [[buffer({b})]],\n")); + s.push_str(&format!(" constant uint& groups [[buffer({})]],\n", b + 1)); + s.push_str(&format!(" constant uint& salt [[buffer({})]],\n", b + 2)); + s.push_str(" uint tid [[thread_position_in_grid]],\n"); + s.push_str(" uint nthreads [[threads_per_grid]]) {\n"); + s.push_str(&persistent_prologue(CoreDialect::Metal)); + } else { + s.push_str(" uint gid [[thread_position_in_grid]]) {\n"); + } s.push_str(" uint nonce = baseNonce + gid;\n"); s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n"); if p.has_wide() { @@ -319,8 +524,17 @@ fn metal_program_impl(p: &Program, dataset_log2: u32, source: LoadSource, bound: Op::Rotr => format!("{d} = rotr_var({d}, {a});"), Op::Mad => format!("{d} = {a} * {b} + {d};"), Op::Shfl => format!("{d} = {d} ^ simd_shuffle_xor({a}, (ushort){});", ins.mask), + Op::Load if load_width(ins) > 1 => { + let (src, closed) = match &source { + LoadSource::Stored => (WideSource::Stored, None), + LoadSource::InlineClosed(d0, d1) => (WideSource::InlineClosed, Some((*d0, *d1))), + LoadSource::InlineMemhard(_) => (WideSource::InlineMemhard, None), + }; + wide_load_stmt(CoreDialect::Metal, &d, &a, ins.width, src, closed) + } Op::Load => format!("{d} = {d} ^ {};", fetch(word_index(&a, false))), Op::WLoad => format!("{d} = {d} ^ {};", fetch(word_index(&a, true))), + Op::Scratch => scratch_stmt(CoreDialect::Metal, &d, &a), }; s.push_str(&format!(" {line} // {k}\n")); } @@ -328,6 +542,9 @@ fn metal_program_impl(p: &Program, dataset_log2: u32, source: LoadSource, bound: s.push_str(" uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);\n"); s.push_str(" uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);\n"); s.push_str(" out[gid] = ((ulong)hi << 32) | (ulong)lo;\n"); + if p.has_scratch() { + s.push_str(" }\n"); + } s.push_str("}\n"); s } @@ -379,8 +596,10 @@ fn cuda_instr_lines(p: &Program) -> String { Op::Rotr => format!("{d} = rotr_var({d}, {a});"), Op::Mad => format!("{d} = {a} * {b} + {d};"), Op::Shfl => format!("{d} = {d} ^ __shfl_xor_sync(0xffffffffu, {a}, {});", ins.mask), + Op::Load if load_width(ins) > 1 => wide_load_stmt(CoreDialect::Cuda, &d, &a, ins.width, WideSource::Stored, None), Op::Load => format!("{d} = {d} ^ ds[{a} & mask];"), Op::WLoad => format!("{d} = {d} ^ ds[(__shfl_sync(0xffffffffu, {a}, 0) & wmask) + lane];"), + Op::Scratch => scratch_stmt(CoreDialect::Cuda, &d, &a), }; s.push_str(&format!(" {line} // {k} {}\n", ins.op.name())); } @@ -448,8 +667,14 @@ pub fn cuda_kernel(p: &Program, memhard: Option<&MixParams>) -> String { s.push_str("// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every\n"); s.push_str("// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a\n"); s.push_str("// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.\n"); - s.push_str("__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {\n"); - s.push_str(" uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;\n"); + s.push_str(&scratch_prelude(p, CoreDialect::Cuda)); + let scratch_args = if p.has_scratch() { ", uint32_t* scratch, uint32_t groups, uint32_t salt" } else { "" }; + s.push_str(&format!("__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask{scratch_args}) {{\n")); + if p.has_scratch() { + s.push_str(&persistent_prologue(CoreDialect::Cuda)); + } else { + s.push_str(" uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;\n"); + } s.push_str(" uint32_t nonce = baseNonce + gid;\n"); s.push_str(" uint32_t r0, r1, r2, r3, r4, r5, r6, r7;\n"); if p.has_wide() { @@ -464,6 +689,9 @@ pub fn cuda_kernel(p: &Program, memhard: Option<&MixParams>) -> String { s.push_str(" uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);\n"); s.push_str(" uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);\n"); s.push_str(" out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;\n"); + if p.has_scratch() { + s.push_str(" }\n"); + } s.push_str("}\n"); s.push('\n'); s.push_str("// Host-side launch wrappers. Declared in program.h, called from host.cu.\n"); @@ -494,16 +722,28 @@ pub fn cuda_kernel(p: &Program, memhard: Option<&MixParams>) -> String { s.push_str("}\n"); s.push('\n'); } - s.push_str( - "cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,\n", - ); - s.push_str(" uint32_t nonces, uint32_t blockWarps) {\n"); - s.push_str(" if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;\n"); - s.push_str(" uint32_t block = 32u * blockWarps;\n"); - s.push_str(" if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;\n"); - s.push_str(" igneum_hash<<>>(ds, out, baseNonce, mask);\n"); - s.push_str(" return cudaGetLastError();\n"); - s.push_str("}\n"); + if p.has_scratch() { + s.push_str("// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it).\n"); + s.push_str("cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,\n"); + s.push_str(" uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) {\n"); + s.push_str(" if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue;\n"); + s.push_str(" uint32_t block = 32u * blockWarps;\n"); + s.push_str(" if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue;\n"); + s.push_str(" igneum_hash<<>>(ds, out, baseNonce, mask, scratch, nonces / 32u, salt);\n"); + s.push_str(" return cudaGetLastError();\n"); + s.push_str("}\n"); + } else { + s.push_str( + "cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,\n", + ); + s.push_str(" uint32_t nonces, uint32_t blockWarps) {\n"); + s.push_str(" if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;\n"); + s.push_str(" uint32_t block = 32u * blockWarps;\n"); + s.push_str(" if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;\n"); + s.push_str(" igneum_hash<<>>(ds, out, baseNonce, mask);\n"); + s.push_str(" return cudaGetLastError();\n"); + s.push_str("}\n"); + } s.push('\n'); s.push_str("cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {\n"); s.push_str(" cudaFuncAttributes attr;\n"); @@ -548,8 +788,14 @@ pub fn cuda_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String { s.push_str("__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }\n"); s.push('\n'); let _ = memhard; // the bound kernel reads the stored dataset in both constructions - s.push_str("__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {\n"); - s.push_str(" uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;\n"); + s.push_str(&scratch_prelude(p, CoreDialect::Cuda)); + let scratch_args = if p.has_scratch() { ", uint32_t* scratch, uint32_t groups, uint32_t salt" } else { "" }; + s.push_str(&format!("__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw{scratch_args}) {{\n")); + if p.has_scratch() { + s.push_str(&persistent_prologue(CoreDialect::Cuda)); + } else { + s.push_str(" uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;\n"); + } s.push_str(" uint32_t nonce = baseNonce + gid;\n"); s.push_str(" uint32_t r0, r1, r2, r3, r4, r5, r6, r7;\n"); if p.has_wide() { @@ -568,18 +814,33 @@ pub fn cuda_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String { s.push_str(" uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);\n"); s.push_str(" uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);\n"); s.push_str(" out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;\n"); + if p.has_scratch() { + s.push_str(" }\n"); + } s.push_str("}\n"); s.push('\n'); - s.push_str( - "cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,\n", - ); - s.push_str(" IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {\n"); - s.push_str(" if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;\n"); - s.push_str(" uint32_t block = 32u * blockWarps;\n"); - s.push_str(" if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;\n"); - s.push_str(" igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw);\n"); - s.push_str(" return cudaGetLastError();\n"); - s.push_str("}\n"); + if p.has_scratch() { + s.push_str("// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it).\n"); + s.push_str("cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,\n"); + s.push_str(" IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) {\n"); + s.push_str(" if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue;\n"); + s.push_str(" uint32_t block = 32u * blockWarps;\n"); + s.push_str(" if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue;\n"); + s.push_str(" igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw, scratch, nonces / 32u, salt);\n"); + s.push_str(" return cudaGetLastError();\n"); + s.push_str("}\n"); + } else { + s.push_str( + "cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,\n", + ); + s.push_str(" IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {\n"); + s.push_str(" if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;\n"); + s.push_str(" uint32_t block = 32u * blockWarps;\n"); + s.push_str(" if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;\n"); + s.push_str(" igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw);\n"); + s.push_str(" return cudaGetLastError();\n"); + s.push_str("}\n"); + } s.push('\n'); s.push_str("cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {\n"); s.push_str(" cudaFuncAttributes attr;\n"); @@ -614,8 +875,10 @@ fn opencl_instr_lines(p: &Program) -> String { Op::Rotr => format!("{d} = rotr_var({d}, {a});"), Op::Mad => format!("{d} = {a} * {b} + {d};"), Op::Shfl => format!("{{ uint t_; IGNEUM_SHFL_XOR(t_, {a}, {}u); {d} = {d} ^ t_; }}", ins.mask), + Op::Load if load_width(ins) > 1 => wide_load_stmt(CoreDialect::OpenCl, &d, &a, ins.width, WideSource::Stored, None), Op::Load => format!("{d} = {d} ^ ds[{a} & mask];"), Op::WLoad => format!("{{ uint t_; IGNEUM_BCAST0(t_, {a}); {d} = {d} ^ ds[(t_ & wmask) + lane]; }}"), + Op::Scratch => scratch_stmt(CoreDialect::OpenCl, &d, &a), }; s.push_str(&format!(" {line} // {k} {}\n", ins.op.name())); } @@ -631,8 +894,13 @@ pub fn opencl_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String { s.push_str( "// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.\n", ); - s.push_str("IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {\n"); - s.push_str(" uint gid = (uint)get_global_id(0);\n"); + let scratch_args = if p.has_scratch() { ", __global uint* scratch, uint groups, uint salt" } else { "" }; + s.push_str(&format!("IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw{scratch_args}) {{\n")); + if p.has_scratch() { + s.push_str(&persistent_prologue(CoreDialect::OpenCl)); + } else { + s.push_str(" uint gid = (uint)get_global_id(0);\n"); + } s.push_str(" uint lid = (uint)get_local_id(0);\n"); s.push_str(" uint nonce = baseNonce + gid;\n"); s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n"); @@ -659,6 +927,9 @@ pub fn opencl_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String { s.push_str(" uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);\n"); s.push_str(" uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);\n"); s.push_str(" out[gid] = ((ulong)hi << 32) | (ulong)lo;\n"); + if p.has_scratch() { + s.push_str(" }\n"); + } s.push_str("}\n"); s } @@ -686,6 +957,9 @@ pub fn opencl_kernel(p: &Program, memhard: Option<&MixParams>) -> String { s.push_str("#ifdef cl_khr_subgroups\n#pragma OPENCL EXTENSION cl_khr_subgroups : enable\n#endif\n"); s.push_str("#ifdef cl_khr_subgroup_shuffle\n#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable\n#endif\n"); s.push_str("#elif IGNEUM_EXCHANGE == 2\n#pragma OPENCL EXTENSION cl_intel_subgroups : enable\n#endif\n"); + if p.has_scratch() { + s.push_str("#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d)))\n"); + } s.push_str("#else\n"); s.push_str("// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.\n"); s.push_str("#include \"emu_opencl.h\"\n#endif\n"); @@ -751,8 +1025,14 @@ pub fn opencl_kernel(p: &Program, memhard: Option<&MixParams>) -> String { s.push_str("// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the\n"); s.push_str("// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and\n"); s.push_str("// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).\n"); - s.push_str("IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {\n"); - s.push_str(" uint gid = (uint)get_global_id(0);\n"); + s.push_str(&scratch_prelude(p, CoreDialect::OpenCl)); + let scratch_args = if p.has_scratch() { ", __global uint* scratch, uint groups, uint salt" } else { "" }; + s.push_str(&format!("IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask{scratch_args}) {{\n")); + if p.has_scratch() { + s.push_str(&persistent_prologue(CoreDialect::OpenCl)); + } else { + s.push_str(" uint gid = (uint)get_global_id(0);\n"); + } s.push_str(" uint lid = (uint)get_local_id(0);\n"); s.push_str(" uint nonce = baseNonce + gid;\n"); s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n"); @@ -774,6 +1054,9 @@ pub fn opencl_kernel(p: &Program, memhard: Option<&MixParams>) -> String { s.push_str(" uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);\n"); s.push_str(" uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);\n"); s.push_str(" out[gid] = ((ulong)hi << 32) | (ulong)lo;\n"); + if p.has_scratch() { + s.push_str(" }\n"); + } s.push_str("}\n"); s.push('\n'); s.push_str("#if IGNEUM_EXCHANGE != 0\n"); @@ -828,6 +1111,8 @@ pub fn program_header(p: &Program, day: &str, ds: &DatasetSource) -> String { s.push_str(&format!("#define IGNEUM_LOADS_PER_HASH {}\n", p.loads_per_hash())); s.push_str(&format!("#define IGNEUM_WIDE_LOADS_PER_HASH {}\n", p.wide_loads_per_hash())); s.push_str(&format!("#define IGNEUM_OP_MIX {}\n", jstr(&p.op_mix()))); + s.push_str(&class_header_lines(p)); + s.push_str(&scratch_header_lines(p)); s.push_str("// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)\n"); s.push_str(&format!("#define IGNEUM_DATASET_MODE {}\n", if memhard.is_some() { 1 } else { 0 })); s.push('\n'); @@ -854,10 +1139,15 @@ pub fn program_header(p: &Program, day: &str, ds: &DatasetSource) -> String { s.push_str("// Defined in kernel.cu. Both launch on the default stream and return cudaGetLastError().\n"); s.push_str("cudaError_t igneum_launch_fill(uint32_t* ds, uint32_t nWords, uint32_t d0, uint32_t d1);\n"); } - s.push_str( - "cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,\n", - ); - s.push_str(" uint32_t nonces, uint32_t blockWarps);\n"); + if p.has_scratch() { + s.push_str("cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,\n"); + s.push_str(" uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt);\n"); + } else { + s.push_str( + "cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,\n", + ); + s.push_str(" uint32_t nonces, uint32_t blockWarps);\n"); + } s.push_str("cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);\n"); s.push_str("#endif\n"); s @@ -1004,6 +1294,19 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String { s.push_str(&format!(" \"iterations\": {ITERATIONS},\n")); s.push_str(&format!(" \"instruction_count\": {INSTR_COUNT},\n")); s.push_str(&format!(" \"loads_per_hash\": {},\n", p.loads_per_hash())); + if !p.class.is_v2() { + let c = p.width_counts(); + s.push_str(&format!(" \"load_class\": {},\n", jstr(&p.class.name()))); + s.push_str(&format!(" \"load_slots\": {},\n", p.class.load_slots)); + s.push_str(&format!(" \"load_mix_percent_4_16_64\": [{}, {}, {}],\n", p.class.mix[0], p.class.mix[1], p.class.mix[2])); + s.push_str(&format!(" \"load_width_counts_4_16_64\": [{}, {}, {}],\n", c[0], c[1], c[2])); + s.push_str(&format!(" \"bytes_per_hash\": {},\n", p.bytes_per_hash())); + if p.has_scratch() { + s.push_str(&format!(" \"scratch_ops_per_hash\": {},\n", p.scratch_ops_per_hash())); + s.push_str(&format!(" \"scratch\": \"variant 5 (measurement only): persistent warps; a 1 MiB scratch per warp of {SCRATCH_SLOTS} 16-byte slots per lane (lane-major); slot = src & 0x{SCRATCH_SLOT_MASK:x}; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)\",\n")); + } + s.push_str(&format!(" \"wide_load\": \"read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, {FOLD_ROT}) * 0x{FOLD_MUL:08x}) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots\",\n")); + } s.push_str(&format!( " \"op_mix\": {{{}}},\n", p.histogram().iter().map(|(n, c)| format!("{}: {c}", jstr(n))).collect::>().join(", ") @@ -1071,6 +1374,23 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String { s.push_str(" \"instructions\": [\n"); let n = p.instrs.len(); for (k, ins) in p.instrs.iter().enumerate() { + if !p.class.is_v2() { + s.push_str(&format!( + " {{\"i\": {k}, \"op\": {}, \"dst\": {}, \"src\": {}, \"src2\": {}, \"imm\": {}, \"imm2\": {}, \"rot\": {}, \"bit\": {}, \"mask\": {}, \"width\": {}}}", + jstr(ins.op.name()), + ins.dst, + ins.src, + ins.src2, + jhex(ins.imm), + jhex(ins.imm2), + ins.rot, + ins.bit, + ins.mask, + ins.width + )); + s.push_str(if k + 1 < n { ",\n" } else { "\n" }); + continue; + } s.push_str(&format!( " {{\"i\": {k}, \"op\": {}, \"dst\": {}, \"src\": {}, \"src2\": {}, \"imm\": {}, \"imm2\": {}, \"rot\": {}, \"bit\": {}, \"mask\": {}}}", jstr(ins.op.name()), diff --git a/igneum-pow/src/generator.rs b/igneum-pow/src/generator.rs index c3d812307..0092818e9 100644 --- a/igneum-pow/src/generator.rs +++ b/igneum-pow/src/generator.rs @@ -18,6 +18,13 @@ //! The retired version 1 generator (op rolled per instruction with a 25 percent load weight, no acceptance) is //! kept as [`generate_v1`] for the census tool and the lever measurements of `proto-metal/MEMHARD.md`. Its //! programs are not the lottery hash and no pack or vector of version 1 is current. +//! +//! Read-width experiment (5 October 2026, gate 1, `docs/plans/read-width.md`; NOT the lottery hash, behind +//! [`LoadClass`]): a program class whose `load` reads `W` bytes (4, 16 or 64: 1, 4 or 16 words, aligned to `W`) +//! and folds every word into `dst` (`verify::fold_words`), with the width fixed per class or drawn per load from +//! an era-fixed mix. The default class [`LoadClass::V2`] is the generator above, draw for draw and byte for byte; +//! every other class takes one extra draw per instruction (the width roll), so its program stream differs from +//! version 2 and its program id carries the class. use crate::accept::{check, Reject}; use crate::seed::{fnv1a64, program_rng, seed_words_from_bytes}; @@ -53,6 +60,9 @@ pub enum Op { Load, /// Warp-coalesced load (lever b of the version 1 generator). Never emitted by version 2. WLoad, + /// Scratch read-modify-write (read-width experiment, variant 5, 5 October 2026): a 16-byte slot of the lane's + /// own 32 KiB of the warp's 1 MiB scratch, read, folded into dst, rewritten. Never emitted by version 2. + Scratch, } impl Op { @@ -71,6 +81,7 @@ impl Op { Op::Shfl => "shfl", Op::Load => "load", Op::WLoad => "wload", + Op::Scratch => "scratch", } } @@ -88,6 +99,7 @@ impl Op { "shfl" => Op::Shfl, "load" => Op::Load, "wload" => Op::WLoad, + "scratch" => Op::Scratch, _ => return None, }) } @@ -95,11 +107,13 @@ impl Op { /// An injecting op: bijective in `dst` and bringing another register (or the dataset) in. The acceptance /// rule's part (b) requires one such write per register. pub fn injects(self) -> bool { - matches!(self, Op::Add | Op::Sub | Op::Xor | Op::Mad | Op::Shfl | Op::Load | Op::WLoad) + matches!(self, Op::Add | Op::Sub | Op::Xor | Op::Mad | Op::Shfl | Op::Load | Op::WLoad | Op::Scratch) } + /// A memory operation: the fresh-source rule, the acceptance tests and the load count treat the scratch + /// read-modify-write as a load (it is one of the program's 128 memory operations). pub fn is_load(self) -> bool { - matches!(self, Op::Load | Op::WLoad) + matches!(self, Op::Load | Op::WLoad | Op::Scratch) } } @@ -124,6 +138,9 @@ pub struct Instr { pub bit: u8, /// Shuffle xor mask: 1, 2, 4, 8 or 16. pub mask: u8, + /// Words read by a `load`: 1 (the lottery hash, 4 bytes), 4 or 16 (the read-width experiment). 1 on every + /// other op. + pub width: u8, } #[derive(Clone, Debug, PartialEq, Eq)] @@ -139,9 +156,144 @@ pub struct Program { pub generator: u32, /// Attempt index: 0 for the bare seed, `k` for the k-th re-derivation after rejections. pub attempt: u32, + /// The load class: [`LoadClass::V2`] for the lottery hash, another for the read-width experiment. + pub class: LoadClass, pub instrs: Vec, } +/// The widths a `load` may read, in words: 4, 16 and 64 bytes. +pub const WIDTH_WORDS: [u8; 3] = [1, 4, 16]; + +/// The load class of a program (read-width experiment, 5 October 2026). `mix` holds the percent weights of the +/// three widths of [`WIDTH_WORDS`] (sum 100); `load_slots` the number of `load` instructions per program. +#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)] +pub struct LoadClass { + pub mix: [u8; 3], + pub load_slots: u8, + /// Variant 5: `Some(k)` gives the program a 1 MiB per-warp scratch (the kernels run persistent warps) and + /// turns `k` of the load slots into scratch read-modify-writes. `None` for every other class. + pub scratch: Option, +} + +/// Scratch geometry (variant 5): 2^11 slots of 16 bytes per lane (32 KiB), 32 lanes per warp (1 MiB), lane-major. +pub const SCRATCH_SLOT_BITS: u32 = 11; +pub const SCRATCH_SLOTS: usize = 1 << SCRATCH_SLOT_BITS; +pub const SCRATCH_SLOT_MASK: u32 = SCRATCH_SLOTS as u32 - 1; +pub const SCRATCH_WORDS_PER_LANE: usize = SCRATCH_SLOTS * 4; +pub const SCRATCH_BYTES_PER_WARP: usize = SCRATCH_WORDS_PER_LANE * 4 * LANES; + +impl LoadClass { + /// Generator version 2 as adopted on 4 October 2026: 16 loads of one word. The lottery hash. + pub const V2: LoadClass = LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None }; + + /// A fixed width (1, 4 or 16 words) with `load_slots` loads per program. + pub fn fixed(width_words: u8, load_slots: u8) -> LoadClass { + let mut mix = [0u8; 3]; + let i = WIDTH_WORDS.iter().position(|&w| w == width_words).expect("width must be 1, 4 or 16 words"); + mix[i] = 100; + LoadClass { mix, load_slots, scratch: None } + } + + /// Per-load width drawn from `mix` (percent for 4, 16, 64 bytes), 16 loads per program. + pub fn mixed(mix: [u8; 3]) -> LoadClass { + assert_eq!(mix.iter().map(|&m| m as u32).sum::(), 100, "the mix must sum to 100"); + LoadClass { mix, load_slots: LOAD_SLOTS as u8, scratch: None } + } + + /// Variant 5: version 2 widths, 16 memory operations of which `k` are scratch read-modify-writes. + pub fn scratch(k: u8) -> LoadClass { + assert!(k as usize <= LOAD_SLOTS); + LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: Some(k) } + } + + /// Parse "p4,p16,p64" or one of the names of [`LoadClass::name`]. + pub fn parse(s: &str) -> Option { + if let Some(k) = s.strip_prefix("scr") { + let k: u8 = k.parse().ok()?; + if k as usize > LOAD_SLOTS { + return None; + } + return Some(LoadClass::scratch(k)); + } + let (mix_s, slots) = match s.split_once("x") { + Some((m, n)) if !m.contains(',') => (m, n.parse::().ok()?), + _ => (s, LOAD_SLOTS as u8), + }; + let mix: [u8; 3] = match mix_s { + "v2" => return Some(LoadClass::V2), + "w4" => [100, 0, 0], + "w16" => [0, 100, 0], + "w64" => [0, 0, 100], + m => { + let v: Vec = m.split(',').map(|x| x.trim().parse::().ok()).collect::>>()?; + if v.len() != 3 || v.iter().map(|&x| x as u32).sum::() != 100 { + return None; + } + [v[0], v[1], v[2]] + } + }; + if slots == 0 || slots as usize >= INSTR_COUNT { + return None; + } + Some(LoadClass { mix, load_slots: slots, scratch: None }) + } + + /// Scratch read-modify-writes per program (0 without a scratch). + pub fn scratch_slots(&self) -> usize { + self.scratch.unwrap_or(0) as usize + } + + pub fn is_v2(&self) -> bool { + *self == LoadClass::V2 + } + + /// "v2", "w4", "w16", "w64", "w64x4", "mix50-35-15", "mix25-50-25x8", "scr4". + pub fn name(&self) -> String { + if self.is_v2() { + return "v2".to_string(); + } + if let Some(k) = self.scratch { + return format!("scr{k}"); + } + let base = match self.mix { + [100, 0, 0] => "w4".to_string(), + [0, 100, 0] => "w16".to_string(), + [0, 0, 100] => "w64".to_string(), + [a, b, c] => format!("mix{a}-{b}-{c}"), + }; + if self.load_slots as usize == LOAD_SLOTS { + base + } else { + format!("{base}x{}", self.load_slots) + } + } + + /// The width in words of a load whose width roll (0..99) is `roll`: the first entry of the mix whose cumulative + /// weight exceeds the roll. + pub fn width_for_roll(&self, roll: u64) -> u8 { + let mut acc = 0u64; + for (i, &m) in self.mix.iter().enumerate() { + acc += m as u64; + if roll < acc { + return WIDTH_WORDS[i]; + } + } + WIDTH_WORDS[2] + } + + /// Expected dataset bytes read per hash: dataset loads per hash times the mean width (scratch traffic apart). + pub fn expected_bytes_per_hash(&self) -> f64 { + let mean = self.mix.iter().zip(WIDTH_WORDS.iter()).map(|(&m, &w)| m as f64 / 100.0 * w as f64 * 4.0).sum::(); + (self.load_slots as usize - self.scratch_slots()) as f64 * ITERATIONS as f64 * mean + } +} + +impl Default for LoadClass { + fn default() -> Self { + LoadClass::V2 + } +} + impl Program { pub fn loads_per_hash(&self) -> usize { self.instrs.iter().filter(|i| i.op.is_load()).count() * ITERATIONS @@ -152,6 +304,27 @@ impl Program { pub fn has_wide(&self) -> bool { self.instrs.iter().any(|i| i.op == Op::WLoad) } + /// Dataset bytes read per hash: 4 per one-word load, 16 and 64 for the wider loads of the experiment. + pub fn bytes_per_hash(&self) -> usize { + self.instrs.iter().filter(|i| i.op == Op::Load).map(|i| i.width as usize * 4).sum::() * ITERATIONS + } + /// Scratch read-modify-writes per hash (variant 5): each reads 16 bytes and writes 16 bytes. + pub fn scratch_ops_per_hash(&self) -> usize { + self.instrs.iter().filter(|i| i.op == Op::Scratch).count() * ITERATIONS + } + pub fn has_scratch(&self) -> bool { + self.class.scratch.is_some() + } + /// Width histogram of the loads, in words: (1, 4, 16) counts. + pub fn width_counts(&self) -> [usize; 3] { + let mut c = [0usize; 3]; + for i in self.instrs.iter().filter(|i| i.op == Op::Load) { + if let Some(k) = WIDTH_WORDS.iter().position(|&w| w == i.width) { + c[k] += 1; + } + } + c + } /// Distinct dataset items a 32-lane warp touches per hash: 32 per plain load, 2 per wide load. pub fn items_per_warp(&self) -> usize { (self.loads_per_hash() - self.wide_loads_per_hash()) * 32 + self.wide_loads_per_hash() * 2 @@ -177,7 +350,11 @@ impl Program { /// || attempt_le32`. Written into every pack so a version 1 program, or another attempt of the same seed, /// can never be mistaken for this one. pub fn program_id(&self) -> u64 { - program_id(self.generator, &self.seed, self.attempt) + if self.class.is_v2() { + program_id(self.generator, &self.seed, self.attempt) + } else { + program_id_class(self.generator, &self.seed, self.attempt, &self.class) + } } } @@ -192,6 +369,28 @@ pub fn program_id(generator: u32, seed: &[u32; 8], attempt: u32) -> u64 { fnv1a64(&b) } +/// Domain tag of the program id of a read-width class (never collides with [`PROGRAM_ID_TAG`]). +pub const PROGRAM_ID_TAG_RW: &[u8] = b"igneum-program-rw/"; + +/// The program id of a non-default class: the tag, then the same fields as [`program_id`], then the three mix +/// percentages and the slot count as bytes. +pub fn program_id_class(generator: u32, seed: &[u32; 8], attempt: u32, class: &LoadClass) -> u64 { + let mut b = Vec::with_capacity(PROGRAM_ID_TAG_RW.len() + 4 + 32 + 4 + 4); + b.extend_from_slice(PROGRAM_ID_TAG_RW); + b.extend_from_slice(&generator.to_le_bytes()); + for w in seed { + b.extend_from_slice(&w.to_le_bytes()); + } + b.extend_from_slice(&attempt.to_le_bytes()); + b.extend_from_slice(&class.mix); + b.push(class.load_slots); + if let Some(k) = class.scratch { + b.extend_from_slice(b"scratch/"); + b.push(k); + } + fnv1a64(&b) +} + /// Weights of the ten non-load families under version 2, in draw order. Sum 75. The load family has no /// weight: its count is fixed by [`LOAD_SLOTS`]. pub const NONLOAD_WEIGHTS: [(Op, u64); 10] = [ @@ -237,20 +436,39 @@ pub fn attempt_words(seed_bytes: &[u8], attempt: u32) -> [u32; 8] { /// One version 2 candidate from its seed words, before the acceptance rule. Spec 01 section 1.4.3: 16 slot draws, /// then nine draws per instruction, 592 per program. pub fn candidate_from_words(seed_string: &str, seed_bytes: &[u8], seed: [u32; 8], attempt: u32) -> Program { + candidate_from_words_class(seed_string, seed_bytes, seed, attempt, LoadClass::V2) +} + +/// [`candidate_from_words`] for a load class. For [`LoadClass::V2`] this is the version 2 draw stream exactly; +/// for any other class the slot count is the class's and every instruction takes a tenth draw, `below(100)`, +/// the width roll (used only on a load slot, drawn on every slot so the stream stays uniform). +pub fn candidate_from_words_class( + seed_string: &str, + seed_bytes: &[u8], + seed: [u32; 8], + attempt: u32, + class: LoadClass, +) -> Program { let mut rng = program_rng(&seed); - // (1) Load slots: a uniform 16-subset of 1..63 by partial Fisher-Yates. Instruction 0 is never a load. + let slots = class.load_slots as usize; + // (1) Load slots: a uniform subset of 1..63 by partial Fisher-Yates. Instruction 0 is never a load. let mut p: [u8; INSTR_COUNT - 1] = [0; INSTR_COUNT - 1]; for (i, slot) in p.iter_mut().enumerate() { *slot = (i + 1) as u8; } - for i in 0..LOAD_SLOTS { + for i in 0..slots { let j = i + rng.below((INSTR_COUNT - 1 - i) as u64) as usize; p.swap(i, j); } let mut is_load = [false; INSTR_COUNT]; - for &slot in &p[..LOAD_SLOTS] { + for &slot in &p[..slots] { is_load[slot as usize] = true; } + // Variant 5: the first k drawn load slots (a uniform k-subset, the draw order is random) are scratch ops. + let mut is_scratch = [false; INSTR_COUNT]; + for &slot in &p[..class.scratch_slots()] { + is_scratch[slot as usize] = true; + } // (2) The instructions. `fresh[r]`: r was written by an earlier instruction and no load has read it since. let mut fresh = [false; 8]; let mut instrs = Vec::with_capacity(INSTR_COUNT); @@ -265,10 +483,10 @@ pub fn candidate_from_words(seed_string: &str, seed_bytes: &[u8], seed: [u32; 8] roll -= w; } if is_load[k] { - op = Op::Load; + op = if is_scratch[k] { Op::Scratch } else { Op::Load }; } let dst = rng.below(8); - let src = if op == Op::Load { + let src = if op.is_load() { let mut eligible = [0u64; 8]; let mut n = 0usize; for r in 0..8u64 { @@ -301,11 +519,13 @@ pub fn candidate_from_words(seed_string: &str, seed_bytes: &[u8], seed: [u32; 8] let rot = 1 + rng.below(31) as u32; let bit = rng.below(32); let mask = 1u8 << rng.below(5); - if op == Op::Load { + let width = if class.is_v2() { 1 } else { class.width_for_roll(rng.below(100)) }; + let width = if op == Op::Load { width } else { 1 }; + if op.is_load() { fresh[src as usize] = false; } fresh[dst as usize] = true; - instrs.push(Instr { op, dst: dst as u8, src: src as u8, src2: b as u8, imm, imm2, rot, bit: bit as u8, mask }); + instrs.push(Instr { op, dst: dst as u8, src: src as u8, src2: b as u8, imm, imm2, rot, bit: bit as u8, mask, width }); } Program { seed_string: seed_string.to_string(), @@ -313,6 +533,7 @@ pub fn candidate_from_words(seed_string: &str, seed_bytes: &[u8], seed: [u32; 8] seed, generator: GENERATOR_VERSION, attempt, + class, instrs, } } @@ -322,6 +543,11 @@ pub fn candidate(seed_string: &str, seed_bytes: &[u8], attempt: u32) -> Program candidate_from_words(seed_string, seed_bytes, attempt_words(seed_bytes, attempt), attempt) } +/// [`candidate`] for a load class. +pub fn candidate_class(seed_string: &str, seed_bytes: &[u8], attempt: u32, class: LoadClass) -> Program { + candidate_from_words_class(seed_string, seed_bytes, attempt_words(seed_bytes, attempt), attempt, class) +} + /// Why no program could be derived from a seed. #[derive(Clone, Debug, PartialEq, Eq)] pub struct Exhausted { @@ -342,9 +568,14 @@ impl std::error::Error for Exhausted {} /// This is what the chain calls (`Epoch::from_seed_bytes`) with the 32-byte epoch seed, and what the packs call /// with the UTF-8 of a seed string. pub fn try_generate_from_seed_bytes(seed_string: &str, seed_bytes: &[u8]) -> Result { + try_generate_class(seed_string, seed_bytes, LoadClass::V2) +} + +/// [`try_generate_from_seed_bytes`] for a load class. +pub fn try_generate_class(seed_string: &str, seed_bytes: &[u8], class: LoadClass) -> Result { let mut last = None; for attempt in 0..MAX_ATTEMPTS { - let p = candidate(seed_string, seed_bytes, attempt); + let p = candidate_class(seed_string, seed_bytes, attempt, class); match check(&p) { Ok(_) => return Ok(p), Err(r) => last = Some(r), @@ -358,16 +589,31 @@ pub fn generate_from_seed_bytes(seed_string: &str, seed_bytes: &[u8]) -> Program try_generate_from_seed_bytes(seed_string, seed_bytes).unwrap_or_else(|e| panic!("{e}")) } +/// [`generate_from_seed_bytes`] for a load class. +pub fn generate_from_seed_bytes_class(seed_string: &str, seed_bytes: &[u8], class: LoadClass) -> Program { + try_generate_class(seed_string, seed_bytes, class).unwrap_or_else(|e| panic!("{e}")) +} + /// The program of a seed string (its UTF-8 bytes are the program seed). pub fn generate(seed_string: &str) -> Program { generate_from_seed_bytes(seed_string, seed_string.as_bytes()) } +/// [`generate`] for a load class. +pub fn generate_class(seed_string: &str, class: LoadClass) -> Program { + generate_from_seed_bytes_class(seed_string, seed_string.as_bytes(), class) +} + /// Every candidate of a seed up to and including the accepted one, with each rejection. For reports and tests. pub fn attempts(seed_string: &str, seed_bytes: &[u8]) -> Vec<(Program, Result<(), Reject>)> { + attempts_class(seed_string, seed_bytes, LoadClass::V2) +} + +/// [`attempts`] for a load class. +pub fn attempts_class(seed_string: &str, seed_bytes: &[u8], class: LoadClass) -> Vec<(Program, Result<(), Reject>)> { let mut out = Vec::new(); for attempt in 0..MAX_ATTEMPTS { - let p = candidate(seed_string, seed_bytes, attempt); + let p = candidate_class(seed_string, seed_bytes, attempt, class); let verdict = check(&p).map(|_| ()); let accepted = verdict.is_ok(); out.push((p, verdict)); @@ -455,7 +701,7 @@ pub fn generate_v1_from_words(seed_string: &str, seed: [u32; 8], cfg: &Generator if op == Op::Load && bit * 100 < cfg.wide_frac * 32 { op = Op::WLoad; } - instrs.push(Instr { op, dst: dst as u8, src: a as u8, src2: b as u8, imm, imm2, rot, bit: bit as u8, mask }); + instrs.push(Instr { op, dst: dst as u8, src: a as u8, src2: b as u8, imm, imm2, rot, bit: bit as u8, mask, width: 1 }); } Program { seed_string: seed_string.to_string(), @@ -463,6 +709,7 @@ pub fn generate_v1_from_words(seed_string: &str, seed: [u32; 8], cfg: &Generator seed, generator: 1, attempt: 0, + class: LoadClass::V2, instrs, } } @@ -565,6 +812,81 @@ mod tests { assert_ne!(program_id(2, &p.seed, 0), program_id(2, &p.seed, 1)); } + /// The read-width classes (5 October 2026): the default class is the version 2 stream exactly; a class + /// program has its slot count, widths from its mix only, and an id that separates it from version 2 and from + /// the other classes. + #[test] + fn load_classes() { + let v2 = candidate("igneum-genesis", b"igneum-genesis", 0); + let same = candidate_class("igneum-genesis", b"igneum-genesis", 0, LoadClass::V2); + assert_eq!(v2, same); + assert!(v2.instrs.iter().all(|i| i.width == 1)); + assert_eq!(v2.bytes_per_hash(), 512); + assert_eq!(LoadClass::parse("w16"), Some(LoadClass::fixed(4, 16))); + assert_eq!(LoadClass::parse("w64x4"), Some(LoadClass::fixed(16, 4))); + assert_eq!(LoadClass::parse("50,35,15"), Some(LoadClass::mixed([50, 35, 15]))); + assert_eq!(LoadClass::parse("v2"), Some(LoadClass::V2)); + assert_eq!(LoadClass::parse("50,35,10"), None); + assert_eq!(LoadClass::fixed(16, 4).name(), "w64x4"); + assert_eq!(LoadClass::mixed([25, 50, 25]).name(), "mix25-50-25"); + // W = 4 with 16 slots IS the lottery hash: the w4 name parses to the default class + assert_eq!(LoadClass::fixed(1, 16), LoadClass::V2); + assert_eq!(LoadClass::parse("w4"), Some(LoadClass::V2)); + assert_eq!(LoadClass::fixed(1, 16).name(), "v2"); + assert_eq!(LoadClass::fixed(1, 8).name(), "w4x8"); + assert!((LoadClass::mixed([50, 35, 15]).expected_bytes_per_hash() - 2201.6).abs() < 1e-6); + assert!((LoadClass::fixed(16, 4).expected_bytes_per_hash() - 2048.0).abs() < 1e-9); + let mut ids = std::collections::HashSet::new(); + ids.insert(v2.program_id()); + for (name, slots, widths) in [("w4x8", 8, vec![1u8]), ("w16", 16, vec![4]), ("w64", 16, vec![16]), ("w64x4", 4, vec![16]), ("50,35,15", 16, vec![1, 4, 16]), ("25,50,25", 16, vec![1, 4, 16])] { + let c = LoadClass::parse(name).unwrap(); + let p = candidate_class("igneum-genesis", b"igneum-genesis", 0, c); + assert_eq!(p.class, c); + assert_eq!(p.instrs.len(), INSTR_COUNT); + assert_eq!(p.loads_per_hash(), 8 * slots); + assert_ne!(p.instrs[0].op, Op::Load); + for ins in &p.instrs { + if ins.op == Op::Load { + assert!(widths.contains(&ins.width), "{name}: width {}", ins.width); + } else { + assert_eq!(ins.width, 1); + } + } + assert!(ids.insert(p.program_id()), "{name}: program id collides"); + } + // variant 5: k scratch ops among the 16 memory operations, the rest one-word loads + for k in [0u8, 2, 4, 8] { + let c = LoadClass::parse(&format!("scr{k}")).unwrap(); + assert_eq!(c, LoadClass::scratch(k)); + assert_eq!(c.name(), format!("scr{k}")); + let p = candidate_class("igneum-genesis", b"igneum-genesis", 0, c); + assert_eq!(p.loads_per_hash(), 128); + assert_eq!(p.scratch_ops_per_hash(), 8 * k as usize); + assert_eq!(p.bytes_per_hash(), (16 - k as usize) * 8 * 4); + assert!(p.instrs.iter().all(|i| i.width == 1)); + assert!(ids.insert(p.program_id()), "scr{k}: program id collides"); + } + assert!(!LoadClass::scratch(0).is_v2()); + // a class with the version 2 widths but another slot count takes the extra roll: a different stream + let w4x8 = candidate_class("igneum-genesis", b"igneum-genesis", 0, LoadClass::fixed(1, 8)); + assert_ne!(w4x8.instrs, v2.instrs); + // the mix draws every width over a population + let mut counts = [0usize; 3]; + for i in 0..200u32 { + let s = format!("igneum-rw-mix/{i}"); + let p = candidate_class(&s, s.as_bytes(), 0, LoadClass::mixed([50, 35, 15])); + let c = p.width_counts(); + for k in 0..3 { + counts[k] += c[k]; + } + } + let total = (counts[0] + counts[1] + counts[2]) as f64; + assert_eq!(total as usize, 200 * 16); + assert!((counts[0] as f64 / total - 0.50).abs() < 0.05, "{counts:?}"); + assert!((counts[1] as f64 / total - 0.35).abs() < 0.05, "{counts:?}"); + assert!((counts[2] as f64 / total - 0.15).abs() < 0.05, "{counts:?}"); + } + #[test] fn generate_returns_an_accepted_program() { let p = generate("igneum-genesis"); diff --git a/igneum-pow/src/main.rs b/igneum-pow/src/main.rs index f09697fa6..1ece86032 100644 --- a/igneum-pow/src/main.rs +++ b/igneum-pow/src/main.rs @@ -7,8 +7,12 @@ //! [--epoch-hex <64 hex> --day-hex ] byte seeds instead of strings (Epoch::from_seed_bytes) //! igneum-pow accept --seed [--epoch-hex <64 hex>] every candidate of the seed with its verdict (spec 01 section 1.4.6) //! igneum-pow show --seed [--epoch-hex <64 hex>] the accepted program, one instruction per line +//! +//! Read-width experiment (5 October 2026, docs/plans/read-width.md): `--class v2|w4|w16|w64|w64x4|p4,p16,p64[xN]` +//! on every command selects the load class (default v2, the lottery hash). Nothing in a v2 run changes. use igneum_pow::emit::export_pack; +use igneum_pow::generator::LoadClass; use igneum_pow::memhard::Cache; use igneum_pow::seed::day_key; use igneum_pow::verify::{DatasetMode, Epoch, DEFAULT_DATASET_LOG2}; @@ -26,6 +30,7 @@ struct Args { prehash: String, epoch_hex: Option, day_hex: Option, + class: LoadClass, } fn usage() -> ! { @@ -36,7 +41,8 @@ fn usage() -> ! { \x20 hash --nonce print the 64-bit hash of one nonce (pack form, init words = seed words)\n\ \x20 hash-bound --prehash <64 hex> --nonce print the header-bound hash (bind.rs) of one 64-bit nonce\n\ \x20 accept every candidate of the seed (or --epoch-hex) with its acceptance verdict\n\ - \x20 show the accepted program, one instruction per line" + \x20 show the accepted program, one instruction per line\n\ + \x20 --class C load class (read-width experiment): v2 (default), w4, w16, w64, w64x4, or p4,p16,p64[xN]" ); std::process::exit(2) } @@ -54,6 +60,7 @@ fn parse() -> Args { prehash: "00".repeat(32), epoch_hex: None, day_hex: None, + class: LoadClass::V2, }; let mut it = std::env::args().skip(1); a.cmd = it.next().unwrap_or_else(|| usage()); @@ -70,6 +77,7 @@ fn parse() -> Args { "--prehash" => a.prehash = val(), "--epoch-hex" => a.epoch_hex = Some(val()), "--day-hex" => a.day_hex = Some(val()), + "--class" => a.class = LoadClass::parse(&val()).unwrap_or_else(|| usage()), _ => usage(), } } @@ -85,7 +93,7 @@ fn main() { "accept" => accept(&a), "show" => show(&a), "hash" => { - let e = Epoch::new(&a.seed, &a.day, mode, a.dataset_log2); + let e = Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class); println!("{:016x}", e.hash(a.nonce as u32)); } "hash-bound" => { @@ -96,9 +104,9 @@ fn main() { (Some(eh), Some(dh)) => { let eb = igneum_pow::bind::unhex(eh).unwrap_or_else(|| usage()); let db = igneum_pow::bind::unhex(dh).unwrap_or_else(|| usage()); - Epoch::from_seed_bytes(&eb, &db, "cli") + Epoch::from_seed_bytes_class(&eb, &db, "cli", a.class) } - _ => Epoch::new(&a.seed, &a.day, mode, a.dataset_log2), + _ => Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class), }; let init = igneum_pow::bind::block_init_words(&prehash, a.nonce); println!("init words {}", init.iter().map(|w| format!("{w:08x}")).collect::>().join(" ")); @@ -125,11 +133,14 @@ fn bench(a: &Args, mode: DatasetMode) { drop(c); } let t0 = Instant::now(); - let e = Epoch::new(&a.seed, &a.day, mode, a.dataset_log2); + let e = Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class); let build_ms = t0.elapsed().as_secs_f64() * 1e3; println!( - "program: {} loads/hash, {} items/warp, op mix {}; epoch built in {build_ms:.1} ms", + "program: class {}, {} loads/hash, {} bytes/hash, widths (1,4,16 words) {:?}, {} items/warp, op mix {}; epoch built in {build_ms:.1} ms", + e.program.class.name(), e.program.loads_per_hash(), + e.program.bytes_per_hash(), + e.program.width_counts(), e.program.items_per_warp(), e.program.op_mix() ); @@ -162,9 +173,9 @@ fn export(a: &Args, mode: DatasetMode) { (Some(eh), Some(dh)) => { let eb = igneum_pow::bind::unhex(eh).unwrap_or_else(|| usage()); let db = igneum_pow::bind::unhex(dh).unwrap_or_else(|| usage()); - (Epoch::from_seed_bytes(&eb, &db, &format!("igneum-epoch/{eh}/day/{dh}")), format!("bytes:{dh}")) + (Epoch::from_seed_bytes_class(&eb, &db, &format!("igneum-epoch/{eh}/day/{dh}"), a.class), format!("bytes:{dh}")) } - _ => (Epoch::new(&a.seed, &a.day, mode, a.dataset_log2), a.day.clone()), + _ => (Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class), a.day.clone()), }; let build_ms = t0.elapsed().as_secs_f64() * 1e3; println!("igneum-pow export {out}"); @@ -179,7 +190,7 @@ fn export(a: &Args, mode: DatasetMode) { e.program.program_id(), e.program.loads_per_hash() ); - println!("op mix: {}", e.program.op_mix()); + println!("op mix: {}; class {}, {} bytes/hash, widths (1,4,16 words) {:?}", e.program.op_mix(), e.program.class.name(), e.program.bytes_per_hash(), e.program.width_counts()); let source = format!("igneum-pow (Rust) CPU interpreter, generator v{}, {} dataset", e.program.generator, e.dataset.mode().name()); let pack = export_pack(&e, &day_label, &source); let dir = std::path::Path::new(&out); @@ -209,7 +220,7 @@ fn seed_bytes_of(a: &Args) -> (String, Vec) { fn accept(a: &Args) { let (label, bytes) = seed_bytes_of(a); let t0 = Instant::now(); - let tries = igneum_pow::generator::attempts(&label, &bytes); + let tries = igneum_pow::generator::attempts_class(&label, &bytes, a.class); let ms = t0.elapsed().as_secs_f64() * 1e3; for (p, verdict) in &tries { match verdict { @@ -233,19 +244,20 @@ fn accept(a: &Args) { fn show(a: &Args) { let (label, bytes) = seed_bytes_of(a); - let p = igneum_pow::generator::generate_from_seed_bytes(&label, &bytes); + let p = igneum_pow::generator::generate_from_seed_bytes_class(&label, &bytes, a.class); println!( - "seed \"{}\" generator v{} attempt {} program id {:016x} seed words {}", + "seed \"{}\" generator v{} class {} attempt {} program id {:016x} seed words {}", p.seed_string, p.generator, + p.class.name(), p.attempt, p.program_id(), p.seed.iter().map(|w| format!("{w:08x}")).collect::>().join(" ") ); - println!("op mix {} loads/hash {}", p.op_mix(), p.loads_per_hash()); + println!("op mix {} loads/hash {} bytes/hash {}", p.op_mix(), p.loads_per_hash(), p.bytes_per_hash()); for (k, i) in p.instrs.iter().enumerate() { println!( - "{k:2}: {:5} dst={} src={} src2={} imm={:#010x} imm2={:#010x} rot={} bit={} mask={}", + "{k:2}: {:5} dst={} src={} src2={} imm={:#010x} imm2={:#010x} rot={} bit={} mask={}{}", i.op.name(), i.dst, i.src, @@ -254,7 +266,8 @@ fn show(a: &Args) { i.imm2, i.rot, i.bit, - i.mask + i.mask, + if i.op == igneum_pow::generator::Op::Load && i.width > 1 { format!(" width={}B", i.width as u32 * 4) } else { String::new() } ); } } diff --git a/igneum-pow/src/memhard.rs b/igneum-pow/src/memhard.rs index 3141f7d04..c6f0727fd 100644 --- a/igneum-pow/src/memhard.rs +++ b/igneum-pow/src/memhard.rs @@ -269,6 +269,34 @@ impl MemhardCpu { } u } + /// `out[k][j] = dataset[base[k] + j]` for `j < width` (read-width experiment): `base[k]` is aligned to `width` + /// words, so every lane's words lie in one item, derived once per distinct item. Returns the distinct items. + pub fn fetch_wide(&self, base: &[u32], width: usize, out: &mut [[u32; 16]]) -> usize { + let n = base.len(); + assert!(n <= FETCH_MAX && out.len() >= n && width <= 16); + let mut uniq = [0u32; FETCH_MAX]; + let mut slot = [0u8; FETCH_MAX]; + let mut u = 0usize; + for k in 0..n { + let t = base[k] >> 4; + let j = match uniq[..u].iter().position(|&x| x == t) { + Some(j) => j, + None => { + uniq[u] = t; + u += 1; + u - 1 + } + }; + slot[k] = j as u8; + } + let mut items = [[0u32; 16]; FETCH_MAX]; + derive_items(&uniq[..u], &self.params, &self.cache, &mut items); + for k in 0..n { + let o = (base[k] & 15) as usize; + out[k][..width].copy_from_slice(&items[slot[k] as usize][o..o + width]); + } + u + } } #[cfg(test)] diff --git a/igneum-pow/src/verify.rs b/igneum-pow/src/verify.rs index 3472165ef..45808be94 100644 --- a/igneum-pow/src/verify.rs +++ b/igneum-pow/src/verify.rs @@ -1,10 +1,91 @@ //! The CPU reference interpreter for one 32-lane warp (`cpuWarpTraced` in the Swift) and the API the node //! calls. Dataset words come from the memory-hard cache (default) or from the closed form (old packs). -use crate::generator::{generate, Instr, Op, Program, ITERATIONS, LANES}; +use crate::generator::{ + generate, generate_class, Instr, LoadClass, Op, Program, ITERATIONS, LANES, SCRATCH_SLOTS, SCRATCH_SLOT_MASK, +}; use crate::memhard::MemhardCpu; use crate::seed::day_key; +/// Read-width experiment (5 October 2026): a `load` of `W` words folds every word into `dst`: +/// `x = dst XOR w[0]; for j in 1..W: x = (rotl(x, FOLD_ROT) * FOLD_MUL) XOR w[j]; dst = x`. For `W = 1` this is the +/// lottery hash's `dst XOR dataset[...]`. The fold is state-dependent (the rotate-multiply sits between the words), +/// so no function of the line alone replaces it: two different lines give two different maps of `dst`, and a +/// dataset of folded lines cannot be stored in place of the dataset (see `docs/plans/read-width.md`). +pub const FOLD_ROT: u32 = 11; +pub const FOLD_MUL: u32 = 0x9E3779B1; + +/// The fold of `words` into `dst` (at least one word). +#[inline(always)] +pub fn fold_words(dst: u32, words: &[u32]) -> u32 { + let mut x = dst ^ words[0]; + for &w in &words[1..] { + x = x.rotate_left(FOLD_ROT).wrapping_mul(FOLD_MUL) ^ w; + } + x +} + +/// Variant 5 (scratch): the fill value of word `j` (0..2) of slot `slot` of lane `lane` of the unit at base nonce +/// `base`, under program seed words `seed`. The scratch of a unit starts as these values; a slot written during +/// the unit's hash holds what was written. Mirrored as `scr_fill` in every emitted kernel. +#[inline(always)] +pub fn scratch_fill(seed: &[u32; 8], base: u32, lane: u32, slot: u32, j: u32) -> u32 { + splitmix32( + (base.wrapping_add(lane) ^ seed[j as usize]) + .wrapping_add(slot.wrapping_mul(0x9E3779B1)) + .wrapping_add((j + 1).wrapping_mul(0x85EBCA77)), + ) +} + +/// Variant 5: the 16-byte slot after a read-modify-write that read `w` and folded to `x`: `(x ^ w1, rotl(x, 7) ^ w2, +/// x + w0)` behind the slot's tag. +#[inline(always)] +pub fn scratch_rewrite(x: u32, w: &[u32; 3]) -> [u32; 3] { + [x ^ w[1], x.rotate_left(7) ^ w[2], x.wrapping_add(w[0])] +} + +/// The CPU model of one unit's scratch (variant 5): per lane, the written slots and their words. Unwritten slots +/// read as [`scratch_fill`]. A unit touches at most `scratch ops x 32` slots, so the model is small whatever the +/// nominal 1 MiB; a GPU keeps the real 1 MiB per resident warp with a per-unit tag per slot. +pub struct ScratchModel { + written: Vec, + data: Vec<[u32; 3]>, + pub reads: usize, + pub writes: usize, +} + +impl ScratchModel { + pub fn new() -> Self { + Self { written: vec![false; LANES * SCRATCH_SLOTS], data: vec![[0; 3]; LANES * SCRATCH_SLOTS], reads: 0, writes: 0 } + } + /// Read slot `slot` of `lane`, then rewrite it from the fold result `x`. Returns the three words read. + #[inline] + pub fn rmw(&mut self, seed: &[u32; 8], base: u32, lane: usize, slot: u32, dst: u32) -> u32 { + let i = lane * SCRATCH_SLOTS + slot as usize; + let w = if self.written[i] { + self.data[i] + } else { + [ + scratch_fill(seed, base, lane as u32, slot, 0), + scratch_fill(seed, base, lane as u32, slot, 1), + scratch_fill(seed, base, lane as u32, slot, 2), + ] + }; + let x = fold_words(dst, &w); + self.data[i] = scratch_rewrite(x, &w); + self.written[i] = true; + self.reads += 1; + self.writes += 1; + x + } +} + +impl Default for ScratchModel { + fn default() -> Self { + Self::new() + } +} + /// Dataset element, closed form of (day words, index). The original prototype's six-operation element. #[inline(always)] pub fn dataset_elem(i: u32, d0: u32, d1: u32) -> u32 { @@ -121,6 +202,23 @@ impl DatasetSource { Dataset::MemoryHard(m) => m.fetch(idx, out), } } + + /// `out[k][j] = dataset[base[k] + j]` for `j < width`; bases are masked and aligned to `width` words + /// (`width` 4 or 16, so a lane's words lie in one item). Returns items derived (0 for the closed form). + #[inline] + fn fetch_wide(&self, base: &[u32; LANES], width: usize, out: &mut [[u32; 16]; LANES]) -> usize { + match &self.dataset { + Dataset::ClosedForm { d0, d1 } => { + for k in 0..LANES { + for j in 0..width { + out[k][j] = dataset_elem(base[k] + j as u32, *d0, *d1); + } + } + 0 + } + Dataset::MemoryHard(m) => m.fetch_wide(base, width, out), + } + } } /// The result of interpreting one warp. @@ -160,10 +258,19 @@ pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32, let mut items_derived = 0usize; let mut idx = [0u32; LANES]; let mut val = [0u32; LANES]; + let mut scratch = if program.has_scratch() { Some(ScratchModel::new()) } else { None }; for _ in 0..ITERATIONS { let sel = r[0]; for ins in &program.instrs { step(ins, &mut r, &sel, mask, ds, &mut idx, &mut val, &mut items_derived); + if ins.op == Op::Scratch { + let m = scratch.as_mut().expect("a scratch op needs a scratch class"); + let (d, a) = (ins.dst as usize, ins.src as usize); + for lane in 0..LANES { + let slot = r[a][lane] & SCRATCH_SLOT_MASK; + r[d][lane] = m.rmw(&program.seed, base_nonce, lane, slot, r[d][lane]); + } + } } } let mut hashes = [0u64; LANES]; @@ -255,7 +362,7 @@ fn step( r[d][lane] ^= src[lane ^ m]; } } - Op::Load => { + Op::Load if ins.width == 1 => { for lane in 0..LANES { idx[lane] = r[a][lane] & mask; } @@ -264,6 +371,22 @@ fn step( r[d][lane] ^= val[lane]; } } + Op::Load => { + // Read-width experiment: `width` words from the aligned address, every word folded into dst. + let width = ins.width as usize; + let align = !(ins.width as u32 - 1); + for lane in 0..LANES { + idx[lane] = (r[a][lane] & mask) & align; + } + let mut vals = [[0u32; 16]; LANES]; + *items_derived += ds.fetch_wide(idx, width, &mut vals); + for lane in 0..LANES { + r[d][lane] = fold_words(r[d][lane], &vals[lane][..width]); + } + } + Op::Scratch => { + // handled by the caller (interpret_warp_init), which owns the unit's scratch model + } Op::WLoad => { // Lane 0's register, masked, aligned down to 32 words; lane l reads word base + l. let base = (r[a][0] & mask) & !31; @@ -299,6 +422,11 @@ impl Epoch { Self { program: generate(seed), dataset: DatasetSource::new(day, mode, dataset_log2) } } + /// [`Epoch::new`] with a load class (read-width experiment). + pub fn new_class(seed: &str, day: &str, mode: DatasetMode, dataset_log2: u32, class: LoadClass) -> Self { + Self { program: generate_class(seed, class), dataset: DatasetSource::new(day, mode, dataset_log2) } + } + /// The production shape: memory-hard, 1 GiB dataset. pub fn memory_hard(seed: &str, day: &str) -> Self { Self::new(seed, day, DatasetMode::MemoryHard, DEFAULT_DATASET_LOG2) @@ -309,7 +437,12 @@ impl Epoch { /// `seed_words_from_bytes(day_bytes)` (`bind::day_bytes`). Memory-hard, 1 GiB dataset. `label` is only /// recorded in emitted packs. pub fn from_seed_bytes(epoch_seed: &[u8], day_bytes: &[u8], label: &str) -> Self { - let program = crate::generator::generate_from_seed_bytes(label, epoch_seed); + Self::from_seed_bytes_class(epoch_seed, day_bytes, label, LoadClass::V2) + } + + /// [`Epoch::from_seed_bytes`] with a load class (read-width experiment). + pub fn from_seed_bytes_class(epoch_seed: &[u8], day_bytes: &[u8], label: &str, class: LoadClass) -> Self { + let program = crate::generator::generate_from_seed_bytes_class(label, epoch_seed, class); let key = crate::seed::seed_words_from_bytes(day_bytes); let mut dataset = DatasetSource::from_key(key, DatasetMode::MemoryHard, DEFAULT_DATASET_LOG2); dataset.key_bytes = day_bytes.to_vec(); @@ -356,6 +489,80 @@ mod tests { assert_eq!(ds.word(0x0fffffff), 0xf78c84a4); } + /// Read-width experiment: the fold for one word is a plain xor; a wide fetch hands each lane the words the + /// scalar path would; two distinct lines give two distinct maps of dst (one point suffices as a smoke check). + #[test] + fn fold_and_wide_fetch() { + assert_eq!(fold_words(0x1234_5678, &[0xdead_beef]), 0x1234_5678 ^ 0xdead_beef); + let w = [1u32, 2, 3, 4]; + let x = fold_words(7, &w); + let mut y: u32 = 7 ^ 1; + for &v in &w[1..] { + y = y.rotate_left(FOLD_ROT).wrapping_mul(FOLD_MUL) ^ v; + } + assert_eq!(x, y); + assert_ne!(fold_words(7, &[1, 2, 3, 4]), fold_words(7, &[1, 2, 3, 5])); + let ds = DatasetSource::new("2026-10-03", DatasetMode::MemoryHard, 20); + let mut base = [0u32; LANES]; + for (k, b) in base.iter_mut().enumerate() { + *b = ((k as u32 * 0x9E37_79B1) & ds.mask) & !15; + } + let mut out = [[0u32; 16]; LANES]; + let items = ds.fetch_wide(&base, 16, &mut out); + assert!(items >= 1 && items <= LANES); + for k in 0..LANES { + for j in 0..16 { + assert_eq!(out[k][j], ds.word(base[k] + j as u32), "lane {k} word {j}"); + } + } + let mut base4 = base; + for b in base4.iter_mut() { + *b += 8; + } + let items4 = ds.fetch_wide(&base4, 4, &mut out); + assert_eq!(items4, items); + for k in 0..LANES { + for j in 0..4 { + assert_eq!(out[k][j], ds.word(base4[k] + j as u32)); + } + } + } + + /// A wide-load program interprets identically on the closed form and through the memory-hard path's fold + /// (the same fold code), and a mixed-class epoch builds and hashes. + /// Variant 5: a fill word is deterministic, a rewrite changes the slot, and a second read of a written slot + /// returns the rewrite, not the fill. + #[test] + fn scratch_model() { + let seed = [1u32, 2, 3, 4, 5, 6, 7, 8]; + assert_eq!(scratch_fill(&seed, 32, 3, 100, 1), scratch_fill(&seed, 32, 3, 100, 1)); + assert_ne!(scratch_fill(&seed, 32, 3, 100, 1), scratch_fill(&seed, 32, 3, 100, 2)); + assert_ne!(scratch_fill(&seed, 32, 3, 100, 1), scratch_fill(&seed, 64, 3, 100, 1)); + let mut m = ScratchModel::new(); + let w = [scratch_fill(&seed, 32, 3, 100, 0), scratch_fill(&seed, 32, 3, 100, 1), scratch_fill(&seed, 32, 3, 100, 2)]; + let x = m.rmw(&seed, 32, 3, 100, 0xabcd); + assert_eq!(x, fold_words(0xabcd, &w)); + let x2 = m.rmw(&seed, 32, 3, 100, 0xabcd); + assert_eq!(x2, fold_words(0xabcd, &scratch_rewrite(x, &w))); + assert_eq!(m.reads, 2); + let e = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::ClosedForm, 20, LoadClass::scratch(4)); + assert_eq!(e.program.scratch_ops_per_hash(), 32); + assert_eq!(e.hash_warp(0), e.hash_warp(0)); + } + + #[test] + fn wide_class_epochs_hash() { + for name in ["w16", "w64x4", "50,35,15"] { + let c = LoadClass::parse(name).unwrap(); + let e = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::ClosedForm, 20, c); + assert_eq!(e.program.class, c); + let a = e.hash_warp(0); + let b = e.hash_warp(0); + assert_eq!(a, b); + assert_ne!(a[0], a[1]); + } + } + #[test] fn closed_form_genesis_vector_lane0() { // Generator v2 vectors (4 October 2026), proto-cuda/packs/igneum-genesis/vectors.json. diff --git a/proto-cuda/emu/cuda_runtime.h b/proto-cuda/emu/cuda_runtime.h index 995d5c509..d4c31b40a 100644 --- a/proto-cuda/emu/cuda_runtime.h +++ b/proto-cuda/emu/cuda_runtime.h @@ -20,6 +20,9 @@ #define __forceinline__ inline struct uint3 { unsigned x, y, z; }; +// uint4 and make_uint4 (read-width experiment, 5 October 2026): the wide loads and the scratch slots are 16-byte vectors. +struct uint4 { unsigned x, y, z, w; }; +static inline uint4 make_uint4(unsigned x, unsigned y, unsigned z, unsigned w) { uint4 v; v.x = x; v.y = y; v.z = z; v.w = w; return v; } struct dim3 { unsigned x, y, z; dim3(unsigned x_ = 1, unsigned y_ = 1, unsigned z_ = 1) : x(x_), y(y_), z(z_) {} }; extern thread_local uint3 threadIdx; extern thread_local uint3 blockIdx; diff --git a/proto-cuda/packs-readwidth/mixA-0/kernel.cl b/proto-cuda/packs-readwidth/mixA-0/kernel.cl new file mode 100644 index 000000000..fac9a117e --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-0/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/0". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0xaa5a3f6eu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x5c0410c3u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x5c0410c3u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x9d994375u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x9d994375u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0xd8d53386u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0xd8d53386u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x2c956ee2u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x2c956ee2u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xe313c2f9u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xe313c2f9u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x495ace1bu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x495ace1bu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x8238bb25u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x8238bb25u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0xaa5a3f6eu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r2 + r7 + ((((sel >> 30u) & 1u) != 0u) ? 0xb788d5f8u : 0x92d00116u); // 0 add + r3 = mul_hi(r3, r2); // 1 mulhi + r6 = r6 ^ ds[r3 & mask]; // 2 load + r0 = r0 ^ ds[r6 & mask]; // 3 load + r5 = r5 * r2; // 4 mul + r7 = r7 + r2 + ((((sel >> 20u) & 1u) != 0u) ? 0x9a773a55u : 0xd5f6f389u); // 5 add + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 6 load + r5 = r5 ^ ds[r6 & mask]; // 7 load + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 8 load + r4 = r4 * r5; // 9 mul + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 10 load + r3 = r3 ^ ds[r2 & mask]; // 11 load + r4 = r4 ^ r0; // 12 xor + r7 = r7 ^ ds[r3 & mask]; // 13 load + r7 = r7 ^ r1; // 14 xor + r4 = r1 * r4 + r4; // 15 mad + r3 = r0 * r4 + r3; // 16 mad + r4 = r6 * r2 + r4; // 17 mad + r0 = rotr_var(r0, r2); // 18 rotr + r5 = rotr_var(r5, r6); // 19 rotr + r4 = r4 ^ ds[r1 & mask]; // 20 load + r2 = r2 ^ ds[r4 & mask]; // 21 load + r1 = rotl_imm(r1, 14u); // 22 rotl + r7 = r7 - r1; // 23 sub + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r2 = r2 ^ t_; } // 24 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r1 = r1 ^ t_; } // 25 shfl + r6 = r6 ^ ds[r0 & mask]; // 26 load + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r0 = r0 ^ t_; } // 27 shfl + r0 = r0 + r2 + ((((sel >> 25u) & 1u) != 0u) ? 0xc6311db1u : 0x1a3cecfbu); // 28 add + r5 = mul_hi(r5, r7); // 29 mulhi + r0 = r1 * r7 + r0; // 30 mad + r5 = rotl_imm(r5, 23u); // 31 rotl + r0 = r0 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0xf1282553u : 0x7aa05e39u); // 32 add + r1 = r1 ^ r6; // 33 xor + r0 = r0 - r1; // 34 sub + r1 = r1 ^ ds[r6 & mask]; // 35 load + r1 = rotl_imm(r1, 7u); // 36 rotl + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r0 = r0 ^ t_; } // 37 shfl + r3 = r3 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x70173cfdu : 0x4a9bf1a8u); // 38 add + r3 = mul_hi(r3, r4); // 39 mulhi + r3 = r3 + r6 + ((((sel >> 8u) & 1u) != 0u) ? 0x2596fd35u : 0x26b3e1a7u); // 40 add + r3 = r3 - r7; // 41 sub + r0 = r0 - r4; // 42 sub + r5 = rotr_var(r5, r7); // 43 rotr + r3 = r4 * r1 + r3; // 44 mad + r4 = r4 ^ ds[r2 & mask]; // 45 load + r2 = mul_hi(r2, r5); // 46 mulhi + r3 = r3 ^ ds[r5 & mask]; // 47 load + r5 = mul_hi(r5, r6); // 48 mulhi + r5 = r5 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x724f77b9u : 0x9f31d23cu); // 49 add + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 1u); r2 = r2 ^ t_; } // 50 shfl + r1 = r1 * r5; // 51 mul + r1 = r1 * r6; // 52 mul + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 53 load + r7 = r7 ^ r2; // 54 xor + r2 = r2 - r1; // 55 sub + r7 = r7 ^ r3; // 56 xor + r3 = rotl_imm(r3, 2u); // 57 rotl + r6 = mul_hi(r6, r4); // 58 mulhi + r4 = rotr_var(r4, r5); // 59 rotr + r3 = mul_hi(r3, r1); // 60 mulhi + r3 = r3 ^ r6; // 61 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 16u); r5 = r5 ^ t_; } // 62 shfl + r1 = r1 ^ ds[r6 & mask]; // 63 load + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/mixA-0/kernel.cu b/proto-cuda/packs-readwidth/mixA-0/kernel.cu new file mode 100644 index 000000000..3a56ea1e1 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-0/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/0". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0xaa5a3f6eu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x5c0410c3u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x5c0410c3u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x9d994375u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0x9d994375u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0xd8d53386u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0xd8d53386u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x2c956ee2u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0x2c956ee2u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xe313c2f9u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xe313c2f9u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x495ace1bu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x495ace1bu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x8238bb25u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x8238bb25u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0xaa5a3f6eu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r2 + r7 + ((((sel >> 30u) & 1u) != 0u) ? 0xb788d5f8u : 0x92d00116u); // 0 add + r3 = __umulhi(r3, r2); // 1 mulhi + r6 = r6 ^ ds[r3 & mask]; // 2 load + r0 = r0 ^ ds[r6 & mask]; // 3 load + r5 = r5 * r2; // 4 mul + r7 = r7 + r2 + ((((sel >> 20u) & 1u) != 0u) ? 0x9a773a55u : 0xd5f6f389u); // 5 add + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 6 load + r5 = r5 ^ ds[r6 & mask]; // 7 load + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 8 load + r4 = r4 * r5; // 9 mul + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 10 load + r3 = r3 ^ ds[r2 & mask]; // 11 load + r4 = r4 ^ r0; // 12 xor + r7 = r7 ^ ds[r3 & mask]; // 13 load + r7 = r7 ^ r1; // 14 xor + r4 = r1 * r4 + r4; // 15 mad + r3 = r0 * r4 + r3; // 16 mad + r4 = r6 * r2 + r4; // 17 mad + r0 = rotr_var(r0, r2); // 18 rotr + r5 = rotr_var(r5, r6); // 19 rotr + r4 = r4 ^ ds[r1 & mask]; // 20 load + r2 = r2 ^ ds[r4 & mask]; // 21 load + r1 = rotl_imm(r1, 14u); // 22 rotl + r7 = r7 - r1; // 23 sub + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 24 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 25 shfl + r6 = r6 ^ ds[r0 & mask]; // 26 load + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 27 shfl + r0 = r0 + r2 + ((((sel >> 25u) & 1u) != 0u) ? 0xc6311db1u : 0x1a3cecfbu); // 28 add + r5 = __umulhi(r5, r7); // 29 mulhi + r0 = r1 * r7 + r0; // 30 mad + r5 = rotl_imm(r5, 23u); // 31 rotl + r0 = r0 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0xf1282553u : 0x7aa05e39u); // 32 add + r1 = r1 ^ r6; // 33 xor + r0 = r0 - r1; // 34 sub + r1 = r1 ^ ds[r6 & mask]; // 35 load + r1 = rotl_imm(r1, 7u); // 36 rotl + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 37 shfl + r3 = r3 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x70173cfdu : 0x4a9bf1a8u); // 38 add + r3 = __umulhi(r3, r4); // 39 mulhi + r3 = r3 + r6 + ((((sel >> 8u) & 1u) != 0u) ? 0x2596fd35u : 0x26b3e1a7u); // 40 add + r3 = r3 - r7; // 41 sub + r0 = r0 - r4; // 42 sub + r5 = rotr_var(r5, r7); // 43 rotr + r3 = r4 * r1 + r3; // 44 mad + r4 = r4 ^ ds[r2 & mask]; // 45 load + r2 = __umulhi(r2, r5); // 46 mulhi + r3 = r3 ^ ds[r5 & mask]; // 47 load + r5 = __umulhi(r5, r6); // 48 mulhi + r5 = r5 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x724f77b9u : 0x9f31d23cu); // 49 add + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r7, 1); // 50 shfl + r1 = r1 * r5; // 51 mul + r1 = r1 * r6; // 52 mul + { uint32_t b_ = (r1 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 53 load + r7 = r7 ^ r2; // 54 xor + r2 = r2 - r1; // 55 sub + r7 = r7 ^ r3; // 56 xor + r3 = rotl_imm(r3, 2u); // 57 rotl + r6 = __umulhi(r6, r4); // 58 mulhi + r4 = rotr_var(r4, r5); // 59 rotr + r3 = __umulhi(r3, r1); // 60 mulhi + r3 = r3 ^ r6; // 61 xor + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r1, 16); // 62 shfl + r1 = r1 ^ ds[r6 & mask]; // 63 load + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixA-0/kernel_bound.cl b/proto-cuda/packs-readwidth/mixA-0/kernel_bound.cl new file mode 100644 index 000000000..7ed518363 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-0/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/0". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0xaa5a3f6eu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x5c0410c3u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x5c0410c3u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x9d994375u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x9d994375u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0xd8d53386u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0xd8d53386u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x2c956ee2u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x2c956ee2u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xe313c2f9u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xe313c2f9u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x495ace1bu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x495ace1bu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x8238bb25u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x8238bb25u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0xaa5a3f6eu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r2 + r7 + ((((sel >> 30u) & 1u) != 0u) ? 0xb788d5f8u : 0x92d00116u); // 0 add + r3 = mul_hi(r3, r2); // 1 mulhi + r6 = r6 ^ ds[r3 & mask]; // 2 load + r0 = r0 ^ ds[r6 & mask]; // 3 load + r5 = r5 * r2; // 4 mul + r7 = r7 + r2 + ((((sel >> 20u) & 1u) != 0u) ? 0x9a773a55u : 0xd5f6f389u); // 5 add + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 6 load + r5 = r5 ^ ds[r6 & mask]; // 7 load + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 8 load + r4 = r4 * r5; // 9 mul + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 10 load + r3 = r3 ^ ds[r2 & mask]; // 11 load + r4 = r4 ^ r0; // 12 xor + r7 = r7 ^ ds[r3 & mask]; // 13 load + r7 = r7 ^ r1; // 14 xor + r4 = r1 * r4 + r4; // 15 mad + r3 = r0 * r4 + r3; // 16 mad + r4 = r6 * r2 + r4; // 17 mad + r0 = rotr_var(r0, r2); // 18 rotr + r5 = rotr_var(r5, r6); // 19 rotr + r4 = r4 ^ ds[r1 & mask]; // 20 load + r2 = r2 ^ ds[r4 & mask]; // 21 load + r1 = rotl_imm(r1, 14u); // 22 rotl + r7 = r7 - r1; // 23 sub + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r2 = r2 ^ t_; } // 24 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r1 = r1 ^ t_; } // 25 shfl + r6 = r6 ^ ds[r0 & mask]; // 26 load + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r0 = r0 ^ t_; } // 27 shfl + r0 = r0 + r2 + ((((sel >> 25u) & 1u) != 0u) ? 0xc6311db1u : 0x1a3cecfbu); // 28 add + r5 = mul_hi(r5, r7); // 29 mulhi + r0 = r1 * r7 + r0; // 30 mad + r5 = rotl_imm(r5, 23u); // 31 rotl + r0 = r0 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0xf1282553u : 0x7aa05e39u); // 32 add + r1 = r1 ^ r6; // 33 xor + r0 = r0 - r1; // 34 sub + r1 = r1 ^ ds[r6 & mask]; // 35 load + r1 = rotl_imm(r1, 7u); // 36 rotl + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r0 = r0 ^ t_; } // 37 shfl + r3 = r3 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x70173cfdu : 0x4a9bf1a8u); // 38 add + r3 = mul_hi(r3, r4); // 39 mulhi + r3 = r3 + r6 + ((((sel >> 8u) & 1u) != 0u) ? 0x2596fd35u : 0x26b3e1a7u); // 40 add + r3 = r3 - r7; // 41 sub + r0 = r0 - r4; // 42 sub + r5 = rotr_var(r5, r7); // 43 rotr + r3 = r4 * r1 + r3; // 44 mad + r4 = r4 ^ ds[r2 & mask]; // 45 load + r2 = mul_hi(r2, r5); // 46 mulhi + r3 = r3 ^ ds[r5 & mask]; // 47 load + r5 = mul_hi(r5, r6); // 48 mulhi + r5 = r5 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x724f77b9u : 0x9f31d23cu); // 49 add + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 1u); r2 = r2 ^ t_; } // 50 shfl + r1 = r1 * r5; // 51 mul + r1 = r1 * r6; // 52 mul + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 53 load + r7 = r7 ^ r2; // 54 xor + r2 = r2 - r1; // 55 sub + r7 = r7 ^ r3; // 56 xor + r3 = rotl_imm(r3, 2u); // 57 rotl + r6 = mul_hi(r6, r4); // 58 mulhi + r4 = rotr_var(r4, r5); // 59 rotr + r3 = mul_hi(r3, r1); // 60 mulhi + r3 = r3 ^ r6; // 61 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 16u); r5 = r5 ^ t_; } // 62 shfl + r1 = r1 ^ ds[r6 & mask]; // 63 load + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r2 + r7 + ((((sel >> 30u) & 1u) != 0u) ? 0xb788d5f8u : 0x92d00116u); // 0 add + r3 = mul_hi(r3, r2); // 1 mulhi + r6 = r6 ^ ds[r3 & mask]; // 2 load + r0 = r0 ^ ds[r6 & mask]; // 3 load + r5 = r5 * r2; // 4 mul + r7 = r7 + r2 + ((((sel >> 20u) & 1u) != 0u) ? 0x9a773a55u : 0xd5f6f389u); // 5 add + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 6 load + r5 = r5 ^ ds[r6 & mask]; // 7 load + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 8 load + r4 = r4 * r5; // 9 mul + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 10 load + r3 = r3 ^ ds[r2 & mask]; // 11 load + r4 = r4 ^ r0; // 12 xor + r7 = r7 ^ ds[r3 & mask]; // 13 load + r7 = r7 ^ r1; // 14 xor + r4 = r1 * r4 + r4; // 15 mad + r3 = r0 * r4 + r3; // 16 mad + r4 = r6 * r2 + r4; // 17 mad + r0 = rotr_var(r0, r2); // 18 rotr + r5 = rotr_var(r5, r6); // 19 rotr + r4 = r4 ^ ds[r1 & mask]; // 20 load + r2 = r2 ^ ds[r4 & mask]; // 21 load + r1 = rotl_imm(r1, 14u); // 22 rotl + r7 = r7 - r1; // 23 sub + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r2 = r2 ^ t_; } // 24 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r1 = r1 ^ t_; } // 25 shfl + r6 = r6 ^ ds[r0 & mask]; // 26 load + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r0 = r0 ^ t_; } // 27 shfl + r0 = r0 + r2 + ((((sel >> 25u) & 1u) != 0u) ? 0xc6311db1u : 0x1a3cecfbu); // 28 add + r5 = mul_hi(r5, r7); // 29 mulhi + r0 = r1 * r7 + r0; // 30 mad + r5 = rotl_imm(r5, 23u); // 31 rotl + r0 = r0 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0xf1282553u : 0x7aa05e39u); // 32 add + r1 = r1 ^ r6; // 33 xor + r0 = r0 - r1; // 34 sub + r1 = r1 ^ ds[r6 & mask]; // 35 load + r1 = rotl_imm(r1, 7u); // 36 rotl + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r0 = r0 ^ t_; } // 37 shfl + r3 = r3 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x70173cfdu : 0x4a9bf1a8u); // 38 add + r3 = mul_hi(r3, r4); // 39 mulhi + r3 = r3 + r6 + ((((sel >> 8u) & 1u) != 0u) ? 0x2596fd35u : 0x26b3e1a7u); // 40 add + r3 = r3 - r7; // 41 sub + r0 = r0 - r4; // 42 sub + r5 = rotr_var(r5, r7); // 43 rotr + r3 = r4 * r1 + r3; // 44 mad + r4 = r4 ^ ds[r2 & mask]; // 45 load + r2 = mul_hi(r2, r5); // 46 mulhi + r3 = r3 ^ ds[r5 & mask]; // 47 load + r5 = mul_hi(r5, r6); // 48 mulhi + r5 = r5 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x724f77b9u : 0x9f31d23cu); // 49 add + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 1u); r2 = r2 ^ t_; } // 50 shfl + r1 = r1 * r5; // 51 mul + r1 = r1 * r6; // 52 mul + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 53 load + r7 = r7 ^ r2; // 54 xor + r2 = r2 - r1; // 55 sub + r7 = r7 ^ r3; // 56 xor + r3 = rotl_imm(r3, 2u); // 57 rotl + r6 = mul_hi(r6, r4); // 58 mulhi + r4 = rotr_var(r4, r5); // 59 rotr + r3 = mul_hi(r3, r1); // 60 mulhi + r3 = r3 ^ r6; // 61 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 16u); r5 = r5 ^ t_; } // 62 shfl + r1 = r1 ^ ds[r6 & mask]; // 63 load + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-0/kernel_bound.cu b/proto-cuda/packs-readwidth/mixA-0/kernel_bound.cu new file mode 100644 index 000000000..390d69235 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-0/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/0". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r2 + r7 + ((((sel >> 30u) & 1u) != 0u) ? 0xb788d5f8u : 0x92d00116u); // 0 add + r3 = __umulhi(r3, r2); // 1 mulhi + r6 = r6 ^ ds[r3 & mask]; // 2 load + r0 = r0 ^ ds[r6 & mask]; // 3 load + r5 = r5 * r2; // 4 mul + r7 = r7 + r2 + ((((sel >> 20u) & 1u) != 0u) ? 0x9a773a55u : 0xd5f6f389u); // 5 add + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 6 load + r5 = r5 ^ ds[r6 & mask]; // 7 load + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 8 load + r4 = r4 * r5; // 9 mul + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 10 load + r3 = r3 ^ ds[r2 & mask]; // 11 load + r4 = r4 ^ r0; // 12 xor + r7 = r7 ^ ds[r3 & mask]; // 13 load + r7 = r7 ^ r1; // 14 xor + r4 = r1 * r4 + r4; // 15 mad + r3 = r0 * r4 + r3; // 16 mad + r4 = r6 * r2 + r4; // 17 mad + r0 = rotr_var(r0, r2); // 18 rotr + r5 = rotr_var(r5, r6); // 19 rotr + r4 = r4 ^ ds[r1 & mask]; // 20 load + r2 = r2 ^ ds[r4 & mask]; // 21 load + r1 = rotl_imm(r1, 14u); // 22 rotl + r7 = r7 - r1; // 23 sub + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 24 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 25 shfl + r6 = r6 ^ ds[r0 & mask]; // 26 load + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 27 shfl + r0 = r0 + r2 + ((((sel >> 25u) & 1u) != 0u) ? 0xc6311db1u : 0x1a3cecfbu); // 28 add + r5 = __umulhi(r5, r7); // 29 mulhi + r0 = r1 * r7 + r0; // 30 mad + r5 = rotl_imm(r5, 23u); // 31 rotl + r0 = r0 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0xf1282553u : 0x7aa05e39u); // 32 add + r1 = r1 ^ r6; // 33 xor + r0 = r0 - r1; // 34 sub + r1 = r1 ^ ds[r6 & mask]; // 35 load + r1 = rotl_imm(r1, 7u); // 36 rotl + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 37 shfl + r3 = r3 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x70173cfdu : 0x4a9bf1a8u); // 38 add + r3 = __umulhi(r3, r4); // 39 mulhi + r3 = r3 + r6 + ((((sel >> 8u) & 1u) != 0u) ? 0x2596fd35u : 0x26b3e1a7u); // 40 add + r3 = r3 - r7; // 41 sub + r0 = r0 - r4; // 42 sub + r5 = rotr_var(r5, r7); // 43 rotr + r3 = r4 * r1 + r3; // 44 mad + r4 = r4 ^ ds[r2 & mask]; // 45 load + r2 = __umulhi(r2, r5); // 46 mulhi + r3 = r3 ^ ds[r5 & mask]; // 47 load + r5 = __umulhi(r5, r6); // 48 mulhi + r5 = r5 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x724f77b9u : 0x9f31d23cu); // 49 add + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r7, 1); // 50 shfl + r1 = r1 * r5; // 51 mul + r1 = r1 * r6; // 52 mul + { uint32_t b_ = (r1 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 53 load + r7 = r7 ^ r2; // 54 xor + r2 = r2 - r1; // 55 sub + r7 = r7 ^ r3; // 56 xor + r3 = rotl_imm(r3, 2u); // 57 rotl + r6 = __umulhi(r6, r4); // 58 mulhi + r4 = rotr_var(r4, r5); // 59 rotr + r3 = __umulhi(r3, r1); // 60 mulhi + r3 = r3 ^ r6; // 61 xor + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r1, 16); // 62 shfl + r1 = r1 ^ ds[r6 & mask]; // 63 load + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixA-0/memhard.h b/proto-cuda/packs-readwidth/mixA-0/memhard.h new file mode 100644 index 000000000..1eeed9bbc --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-0/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/0". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/mixA-0/memhard.metal b/proto-cuda/packs-readwidth/mixA-0/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-0/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/mixA-0/program.h b/proto-cuda/packs-readwidth/mixA-0/program.h new file mode 100644 index 000000000..b88c24aee --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-0/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/0". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-readwidth/A/0" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d7265616477696474682f412f30" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0xd30baa94fa82322eull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=7 mulhi=7 shfl=6 xor=6 mad=5 sub=5 mul=4 rotl=4 rotr=4" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "mix50-35-15" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 50, 35, 15 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 12, 2, 2 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 1664 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0xaa5a3f6eu, 0x5c0410c3u, 0x9d994375u, 0xd8d53386u, 0x2c956ee2u, 0xe313c2f9u, 0x495ace1bu, 0x8238bb25u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/mixA-0/program.json b/proto-cuda/packs-readwidth/mixA-0/program.json new file mode 100644 index 000000000..c38e841da --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-0/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0xd30baa94fa82322e", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-readwidth/A/0", + "seed_bytes": "69676e65756d2d7265616477696474682f412f30", + "seed_words": ["0xaa5a3f6e", "0x5c0410c3", "0x9d994375", "0xd8d53386", "0x2c956ee2", "0xe313c2f9", "0x495ace1b", "0x8238bb25"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "mix50-35-15", + "load_slots": 16, + "load_mix_percent_4_16_64": [50, 35, 15], + "load_width_counts_4_16_64": [12, 2, 2], + "bytes_per_hash": 1664, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "add": 7, "mulhi": 7, "shfl": 6, "xor": 6, "mad": 5, "sub": 5, "mul": 4, "rotl": 4, "rotr": 4}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "add", "dst": 2, "src": 7, "src2": 2, "imm": "0x92d00116", "imm2": "0xb788d5f8", "rot": 18, "bit": 30, "mask": 16, "width": 1}, + {"i": 1, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0xda7cdbf6", "imm2": "0xb722360d", "rot": 28, "bit": 18, "mask": 2, "width": 1}, + {"i": 2, "op": "load", "dst": 6, "src": 3, "src2": 3, "imm": "0x4c76905d", "imm2": "0x8700b920", "rot": 20, "bit": 19, "mask": 16, "width": 1}, + {"i": 3, "op": "load", "dst": 0, "src": 6, "src2": 5, "imm": "0xc88bd196", "imm2": "0x0bec95ee", "rot": 27, "bit": 2, "mask": 1, "width": 1}, + {"i": 4, "op": "mul", "dst": 5, "src": 2, "src2": 0, "imm": "0x7ed20556", "imm2": "0x3cb2440a", "rot": 19, "bit": 8, "mask": 8, "width": 1}, + {"i": 5, "op": "add", "dst": 7, "src": 2, "src2": 5, "imm": "0xd5f6f389", "imm2": "0x9a773a55", "rot": 29, "bit": 20, "mask": 2, "width": 1}, + {"i": 6, "op": "load", "dst": 6, "src": 5, "src2": 7, "imm": "0xd9894991", "imm2": "0xf41ba497", "rot": 14, "bit": 9, "mask": 4, "width": 16}, + {"i": 7, "op": "load", "dst": 5, "src": 6, "src2": 0, "imm": "0xab148320", "imm2": "0x83097251", "rot": 31, "bit": 6, "mask": 16, "width": 1}, + {"i": 8, "op": "load", "dst": 1, "src": 2, "src2": 0, "imm": "0x9f83b9b8", "imm2": "0x5b9a1ccc", "rot": 17, "bit": 28, "mask": 1, "width": 4}, + {"i": 9, "op": "mul", "dst": 4, "src": 5, "src2": 1, "imm": "0x70d7287a", "imm2": "0x803900fa", "rot": 26, "bit": 24, "mask": 4, "width": 1}, + {"i": 10, "op": "load", "dst": 2, "src": 0, "src2": 1, "imm": "0xf15303dd", "imm2": "0x4725246d", "rot": 4, "bit": 10, "mask": 2, "width": 4}, + {"i": 11, "op": "load", "dst": 3, "src": 2, "src2": 6, "imm": "0xd6d526e4", "imm2": "0x0af56645", "rot": 19, "bit": 18, "mask": 4, "width": 1}, + {"i": 12, "op": "xor", "dst": 4, "src": 0, "src2": 5, "imm": "0xa08eca24", "imm2": "0x21fe64a3", "rot": 29, "bit": 29, "mask": 8, "width": 1}, + {"i": 13, "op": "load", "dst": 7, "src": 3, "src2": 2, "imm": "0x0f32a107", "imm2": "0x92e66063", "rot": 24, "bit": 22, "mask": 8, "width": 1}, + {"i": 14, "op": "xor", "dst": 7, "src": 1, "src2": 2, "imm": "0x9eec6af9", "imm2": "0x21dc3f6f", "rot": 28, "bit": 18, "mask": 4, "width": 1}, + {"i": 15, "op": "mad", "dst": 4, "src": 1, "src2": 4, "imm": "0xfa6381a2", "imm2": "0x2409fa4f", "rot": 18, "bit": 31, "mask": 4, "width": 1}, + {"i": 16, "op": "mad", "dst": 3, "src": 0, "src2": 4, "imm": "0x52a02e74", "imm2": "0xcdd43c2d", "rot": 9, "bit": 12, "mask": 2, "width": 1}, + {"i": 17, "op": "mad", "dst": 4, "src": 6, "src2": 2, "imm": "0x60e472f6", "imm2": "0xe8623e17", "rot": 5, "bit": 30, "mask": 4, "width": 1}, + {"i": 18, "op": "rotr", "dst": 0, "src": 2, "src2": 1, "imm": "0x08cbdd3e", "imm2": "0x3ab8f3a3", "rot": 27, "bit": 2, "mask": 8, "width": 1}, + {"i": 19, "op": "rotr", "dst": 5, "src": 6, "src2": 6, "imm": "0x09d41964", "imm2": "0x98709424", "rot": 28, "bit": 2, "mask": 2, "width": 1}, + {"i": 20, "op": "load", "dst": 4, "src": 1, "src2": 3, "imm": "0x3831d4c0", "imm2": "0xb095fd4b", "rot": 28, "bit": 18, "mask": 2, "width": 1}, + {"i": 21, "op": "load", "dst": 2, "src": 4, "src2": 0, "imm": "0x61f00f86", "imm2": "0x2bd27add", "rot": 1, "bit": 5, "mask": 2, "width": 1}, + {"i": 22, "op": "rotl", "dst": 1, "src": 5, "src2": 6, "imm": "0xe8785677", "imm2": "0xafeefe07", "rot": 14, "bit": 20, "mask": 2, "width": 1}, + {"i": 23, "op": "sub", "dst": 7, "src": 1, "src2": 3, "imm": "0x06c22ebe", "imm2": "0x486a522e", "rot": 19, "bit": 17, "mask": 2, "width": 1}, + {"i": 24, "op": "shfl", "dst": 2, "src": 6, "src2": 5, "imm": "0xf684e748", "imm2": "0xd67fa827", "rot": 31, "bit": 24, "mask": 2, "width": 1}, + {"i": 25, "op": "shfl", "dst": 1, "src": 4, "src2": 1, "imm": "0x712a551f", "imm2": "0xf6b8ff0b", "rot": 7, "bit": 7, "mask": 16, "width": 1}, + {"i": 26, "op": "load", "dst": 6, "src": 0, "src2": 6, "imm": "0xcc4fcce5", "imm2": "0xd121fe63", "rot": 31, "bit": 20, "mask": 1, "width": 1}, + {"i": 27, "op": "shfl", "dst": 0, "src": 6, "src2": 6, "imm": "0x1a31bcef", "imm2": "0x0152fc58", "rot": 19, "bit": 31, "mask": 1, "width": 1}, + {"i": 28, "op": "add", "dst": 0, "src": 2, "src2": 6, "imm": "0x1a3cecfb", "imm2": "0xc6311db1", "rot": 25, "bit": 25, "mask": 16, "width": 1}, + {"i": 29, "op": "mulhi", "dst": 5, "src": 7, "src2": 7, "imm": "0x7eda9d10", "imm2": "0x74eacbd7", "rot": 8, "bit": 20, "mask": 1, "width": 1}, + {"i": 30, "op": "mad", "dst": 0, "src": 1, "src2": 7, "imm": "0xf1992097", "imm2": "0xcc9bf87e", "rot": 14, "bit": 0, "mask": 16, "width": 1}, + {"i": 31, "op": "rotl", "dst": 5, "src": 1, "src2": 6, "imm": "0xf7adfcf5", "imm2": "0x16693ac6", "rot": 23, "bit": 10, "mask": 4, "width": 1}, + {"i": 32, "op": "add", "dst": 0, "src": 2, "src2": 7, "imm": "0x7aa05e39", "imm2": "0xf1282553", "rot": 21, "bit": 28, "mask": 4, "width": 1}, + {"i": 33, "op": "xor", "dst": 1, "src": 6, "src2": 4, "imm": "0xa773d935", "imm2": "0x38d34547", "rot": 8, "bit": 16, "mask": 8, "width": 1}, + {"i": 34, "op": "sub", "dst": 0, "src": 1, "src2": 4, "imm": "0x20af644a", "imm2": "0x28e75a4d", "rot": 7, "bit": 28, "mask": 8, "width": 1}, + {"i": 35, "op": "load", "dst": 1, "src": 6, "src2": 2, "imm": "0x6fb47122", "imm2": "0x562f2b23", "rot": 9, "bit": 23, "mask": 8, "width": 1}, + {"i": 36, "op": "rotl", "dst": 1, "src": 6, "src2": 3, "imm": "0xafd08d8b", "imm2": "0xa7f3e93d", "rot": 7, "bit": 30, "mask": 16, "width": 1}, + {"i": 37, "op": "shfl", "dst": 0, "src": 4, "src2": 0, "imm": "0x84fda071", "imm2": "0x3ec80cba", "rot": 25, "bit": 11, "mask": 8, "width": 1}, + {"i": 38, "op": "add", "dst": 3, "src": 5, "src2": 1, "imm": "0x4a9bf1a8", "imm2": "0x70173cfd", "rot": 15, "bit": 13, "mask": 2, "width": 1}, + {"i": 39, "op": "mulhi", "dst": 3, "src": 4, "src2": 3, "imm": "0x2ea4e0bd", "imm2": "0x1431b05f", "rot": 6, "bit": 22, "mask": 1, "width": 1}, + {"i": 40, "op": "add", "dst": 3, "src": 6, "src2": 3, "imm": "0x26b3e1a7", "imm2": "0x2596fd35", "rot": 4, "bit": 8, "mask": 1, "width": 1}, + {"i": 41, "op": "sub", "dst": 3, "src": 7, "src2": 5, "imm": "0x50df0069", "imm2": "0xba092804", "rot": 28, "bit": 31, "mask": 8, "width": 1}, + {"i": 42, "op": "sub", "dst": 0, "src": 4, "src2": 1, "imm": "0x22b4b65e", "imm2": "0x567989de", "rot": 3, "bit": 2, "mask": 16, "width": 1}, + {"i": 43, "op": "rotr", "dst": 5, "src": 7, "src2": 6, "imm": "0xeebb6e92", "imm2": "0xdb303272", "rot": 7, "bit": 26, "mask": 16, "width": 1}, + {"i": 44, "op": "mad", "dst": 3, "src": 4, "src2": 1, "imm": "0x17907041", "imm2": "0x2665fc2a", "rot": 24, "bit": 10, "mask": 8, "width": 1}, + {"i": 45, "op": "load", "dst": 4, "src": 2, "src2": 0, "imm": "0xf85e047a", "imm2": "0x514751b2", "rot": 21, "bit": 3, "mask": 16, "width": 1}, + {"i": 46, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x79ee74b0", "imm2": "0x71d32fd1", "rot": 26, "bit": 5, "mask": 2, "width": 1}, + {"i": 47, "op": "load", "dst": 3, "src": 5, "src2": 6, "imm": "0x73fa2964", "imm2": "0x25d08887", "rot": 22, "bit": 22, "mask": 4, "width": 1}, + {"i": 48, "op": "mulhi", "dst": 5, "src": 6, "src2": 3, "imm": "0xb686c1c8", "imm2": "0x31b980fa", "rot": 26, "bit": 20, "mask": 8, "width": 1}, + {"i": 49, "op": "add", "dst": 5, "src": 7, "src2": 0, "imm": "0x9f31d23c", "imm2": "0x724f77b9", "rot": 9, "bit": 22, "mask": 4, "width": 1}, + {"i": 50, "op": "shfl", "dst": 2, "src": 7, "src2": 6, "imm": "0x0793dd9f", "imm2": "0xbed3bbbb", "rot": 31, "bit": 23, "mask": 1, "width": 1}, + {"i": 51, "op": "mul", "dst": 1, "src": 5, "src2": 7, "imm": "0x713cb7a1", "imm2": "0x9e06cef8", "rot": 20, "bit": 2, "mask": 1, "width": 1}, + {"i": 52, "op": "mul", "dst": 1, "src": 6, "src2": 5, "imm": "0x57dcbb02", "imm2": "0xf262ce05", "rot": 11, "bit": 17, "mask": 4, "width": 1}, + {"i": 53, "op": "load", "dst": 6, "src": 1, "src2": 2, "imm": "0x6a4caa97", "imm2": "0x26d85624", "rot": 22, "bit": 28, "mask": 2, "width": 16}, + {"i": 54, "op": "xor", "dst": 7, "src": 2, "src2": 5, "imm": "0x2cb21c3f", "imm2": "0x53fa6b79", "rot": 2, "bit": 17, "mask": 4, "width": 1}, + {"i": 55, "op": "sub", "dst": 2, "src": 1, "src2": 5, "imm": "0x19c7d010", "imm2": "0xd7f64a4b", "rot": 28, "bit": 30, "mask": 4, "width": 1}, + {"i": 56, "op": "xor", "dst": 7, "src": 3, "src2": 2, "imm": "0xecabf384", "imm2": "0x0587918d", "rot": 25, "bit": 5, "mask": 8, "width": 1}, + {"i": 57, "op": "rotl", "dst": 3, "src": 6, "src2": 0, "imm": "0x5dba32f7", "imm2": "0x3228574b", "rot": 2, "bit": 20, "mask": 4, "width": 1}, + {"i": 58, "op": "mulhi", "dst": 6, "src": 4, "src2": 5, "imm": "0x558cb969", "imm2": "0xb0531a17", "rot": 2, "bit": 17, "mask": 2, "width": 1}, + {"i": 59, "op": "rotr", "dst": 4, "src": 5, "src2": 2, "imm": "0x8384814a", "imm2": "0xad7ee308", "rot": 7, "bit": 2, "mask": 16, "width": 1}, + {"i": 60, "op": "mulhi", "dst": 3, "src": 1, "src2": 7, "imm": "0x86193717", "imm2": "0x3daece51", "rot": 10, "bit": 12, "mask": 4, "width": 1}, + {"i": 61, "op": "xor", "dst": 3, "src": 6, "src2": 7, "imm": "0x83dada1b", "imm2": "0xc9f04198", "rot": 24, "bit": 22, "mask": 2, "width": 1}, + {"i": 62, "op": "shfl", "dst": 5, "src": 1, "src2": 0, "imm": "0x21dc46bb", "imm2": "0xe2eb5f1e", "rot": 19, "bit": 19, "mask": 16, "width": 1}, + {"i": 63, "op": "load", "dst": 1, "src": 6, "src2": 5, "imm": "0x5a37a588", "imm2": "0xcb1e1122", "rot": 13, "bit": 1, "mask": 1, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/mixA-0/program.metal b/proto-cuda/packs-readwidth/mixA-0/program.metal new file mode 100644 index 000000000..0ceb64870 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-0/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0xaa5a3f6eu, 0x5c0410c3u, 0x9d994375u, 0xd8d53386u, 0x2c956ee2u, 0xe313c2f9u, 0x495ace1bu, 0x8238bb25u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r2 + r7 + select(0x92d00116u, 0xb788d5f8u, ((sel >> 30u) & 1u) != 0u); // 0 + r3 = mulhi(r3, r2); // 1 + r6 = r6 ^ dataset[r3 & MASK]; // 2 + r0 = r0 ^ dataset[r6 & MASK]; // 3 + r5 = r5 * r2; // 4 + r7 = r7 + r2 + select(0xd5f6f389u, 0x9a773a55u, ((sel >> 20u) & 1u) != 0u); // 5 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 6 + r5 = r5 ^ dataset[r6 & MASK]; // 7 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 8 + r4 = r4 * r5; // 9 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 10 + r3 = r3 ^ dataset[r2 & MASK]; // 11 + r4 = r4 ^ r0; // 12 + r7 = r7 ^ dataset[r3 & MASK]; // 13 + r7 = r7 ^ r1; // 14 + r4 = r1 * r4 + r4; // 15 + r3 = r0 * r4 + r3; // 16 + r4 = r6 * r2 + r4; // 17 + r0 = rotr_var(r0, r2); // 18 + r5 = rotr_var(r5, r6); // 19 + r4 = r4 ^ dataset[r1 & MASK]; // 20 + r2 = r2 ^ dataset[r4 & MASK]; // 21 + r1 = rotl_imm(r1, 14u); // 22 + r7 = r7 - r1; // 23 + r2 = r2 ^ simd_shuffle_xor(r6, (ushort)2); // 24 + r1 = r1 ^ simd_shuffle_xor(r4, (ushort)16); // 25 + r6 = r6 ^ dataset[r0 & MASK]; // 26 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)1); // 27 + r0 = r0 + r2 + select(0x1a3cecfbu, 0xc6311db1u, ((sel >> 25u) & 1u) != 0u); // 28 + r5 = mulhi(r5, r7); // 29 + r0 = r1 * r7 + r0; // 30 + r5 = rotl_imm(r5, 23u); // 31 + r0 = r0 + r2 + select(0x7aa05e39u, 0xf1282553u, ((sel >> 28u) & 1u) != 0u); // 32 + r1 = r1 ^ r6; // 33 + r0 = r0 - r1; // 34 + r1 = r1 ^ dataset[r6 & MASK]; // 35 + r1 = rotl_imm(r1, 7u); // 36 + r0 = r0 ^ simd_shuffle_xor(r4, (ushort)8); // 37 + r3 = r3 + r5 + select(0x4a9bf1a8u, 0x70173cfdu, ((sel >> 13u) & 1u) != 0u); // 38 + r3 = mulhi(r3, r4); // 39 + r3 = r3 + r6 + select(0x26b3e1a7u, 0x2596fd35u, ((sel >> 8u) & 1u) != 0u); // 40 + r3 = r3 - r7; // 41 + r0 = r0 - r4; // 42 + r5 = rotr_var(r5, r7); // 43 + r3 = r4 * r1 + r3; // 44 + r4 = r4 ^ dataset[r2 & MASK]; // 45 + r2 = mulhi(r2, r5); // 46 + r3 = r3 ^ dataset[r5 & MASK]; // 47 + r5 = mulhi(r5, r6); // 48 + r5 = r5 + r7 + select(0x9f31d23cu, 0x724f77b9u, ((sel >> 22u) & 1u) != 0u); // 49 + r2 = r2 ^ simd_shuffle_xor(r7, (ushort)1); // 50 + r1 = r1 * r5; // 51 + r1 = r1 * r6; // 52 + { uint b_ = (r1 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 53 + r7 = r7 ^ r2; // 54 + r2 = r2 - r1; // 55 + r7 = r7 ^ r3; // 56 + r3 = rotl_imm(r3, 2u); // 57 + r6 = mulhi(r6, r4); // 58 + r4 = rotr_var(r4, r5); // 59 + r3 = mulhi(r3, r1); // 60 + r3 = r3 ^ r6; // 61 + r5 = r5 ^ simd_shuffle_xor(r1, (ushort)16); // 62 + r1 = r1 ^ dataset[r6 & MASK]; // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-0/program_bound.metal b/proto-cuda/packs-readwidth/mixA-0/program_bound.metal new file mode 100644 index 000000000..6013bad34 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-0/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0xaa5a3f6eu, 0x5c0410c3u, 0x9d994375u, 0xd8d53386u, 0x2c956ee2u, 0xe313c2f9u, 0x495ace1bu, 0x8238bb25u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r2 + r7 + select(0x92d00116u, 0xb788d5f8u, ((sel >> 30u) & 1u) != 0u); // 0 + r3 = mulhi(r3, r2); // 1 + r6 = r6 ^ dataset[r3 & MASK]; // 2 + r0 = r0 ^ dataset[r6 & MASK]; // 3 + r5 = r5 * r2; // 4 + r7 = r7 + r2 + select(0xd5f6f389u, 0x9a773a55u, ((sel >> 20u) & 1u) != 0u); // 5 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 6 + r5 = r5 ^ dataset[r6 & MASK]; // 7 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 8 + r4 = r4 * r5; // 9 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 10 + r3 = r3 ^ dataset[r2 & MASK]; // 11 + r4 = r4 ^ r0; // 12 + r7 = r7 ^ dataset[r3 & MASK]; // 13 + r7 = r7 ^ r1; // 14 + r4 = r1 * r4 + r4; // 15 + r3 = r0 * r4 + r3; // 16 + r4 = r6 * r2 + r4; // 17 + r0 = rotr_var(r0, r2); // 18 + r5 = rotr_var(r5, r6); // 19 + r4 = r4 ^ dataset[r1 & MASK]; // 20 + r2 = r2 ^ dataset[r4 & MASK]; // 21 + r1 = rotl_imm(r1, 14u); // 22 + r7 = r7 - r1; // 23 + r2 = r2 ^ simd_shuffle_xor(r6, (ushort)2); // 24 + r1 = r1 ^ simd_shuffle_xor(r4, (ushort)16); // 25 + r6 = r6 ^ dataset[r0 & MASK]; // 26 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)1); // 27 + r0 = r0 + r2 + select(0x1a3cecfbu, 0xc6311db1u, ((sel >> 25u) & 1u) != 0u); // 28 + r5 = mulhi(r5, r7); // 29 + r0 = r1 * r7 + r0; // 30 + r5 = rotl_imm(r5, 23u); // 31 + r0 = r0 + r2 + select(0x7aa05e39u, 0xf1282553u, ((sel >> 28u) & 1u) != 0u); // 32 + r1 = r1 ^ r6; // 33 + r0 = r0 - r1; // 34 + r1 = r1 ^ dataset[r6 & MASK]; // 35 + r1 = rotl_imm(r1, 7u); // 36 + r0 = r0 ^ simd_shuffle_xor(r4, (ushort)8); // 37 + r3 = r3 + r5 + select(0x4a9bf1a8u, 0x70173cfdu, ((sel >> 13u) & 1u) != 0u); // 38 + r3 = mulhi(r3, r4); // 39 + r3 = r3 + r6 + select(0x26b3e1a7u, 0x2596fd35u, ((sel >> 8u) & 1u) != 0u); // 40 + r3 = r3 - r7; // 41 + r0 = r0 - r4; // 42 + r5 = rotr_var(r5, r7); // 43 + r3 = r4 * r1 + r3; // 44 + r4 = r4 ^ dataset[r2 & MASK]; // 45 + r2 = mulhi(r2, r5); // 46 + r3 = r3 ^ dataset[r5 & MASK]; // 47 + r5 = mulhi(r5, r6); // 48 + r5 = r5 + r7 + select(0x9f31d23cu, 0x724f77b9u, ((sel >> 22u) & 1u) != 0u); // 49 + r2 = r2 ^ simd_shuffle_xor(r7, (ushort)1); // 50 + r1 = r1 * r5; // 51 + r1 = r1 * r6; // 52 + { uint b_ = (r1 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 53 + r7 = r7 ^ r2; // 54 + r2 = r2 - r1; // 55 + r7 = r7 ^ r3; // 56 + r3 = rotl_imm(r3, 2u); // 57 + r6 = mulhi(r6, r4); // 58 + r4 = rotr_var(r4, r5); // 59 + r3 = mulhi(r3, r1); // 60 + r3 = r3 ^ r6; // 61 + r5 = r5 ^ simd_shuffle_xor(r1, (ushort)16); // 62 + r1 = r1 ^ dataset[r6 & MASK]; // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-0/vectors.h b/proto-cuda/packs-readwidth/mixA-0/vectors.h new file mode 100644 index 000000000..1fa5c100e --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-0/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/0". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0xc02af7d21c32b6e2ull, 0xf9676c7b81e42590ull, 0x3b5b2f433e291756ull, 0x2307064009985ab7ull, 0xf1f2d84aa86f1f5aull, 0xa7ed219296c204beull, 0xf8105c9692ac08e7ull, 0x04377ca8373b8e1full, + 0x2f76672aa0a6104dull, 0xac85110c18a8edc8ull, 0x3b4c20d7d3c22e52ull, 0x365c221488fa37e4ull, 0x989dab8f749cd976ull, 0x37f6690cc11dd5a7ull, 0x6159da4241e04ae2ull, 0x3c19bfdcd7821055ull, + 0x5b48a58c99800f4full, 0x991018ff7b067a45ull, 0xf281167cb9c28ef2ull, 0xe9497f5c0207892dull, 0xac1ccf9060f7c78bull, 0xf6f4282e68ceff58ull, 0xfa2694f87cf04ef6ull, 0x62d01632e98c6a11ull, + 0x63a875eb903ee9afull, 0x680ad026ac9e5e1full, 0x63d80718c2e2912bull, 0x377a3c901dadb9f5ull, 0xb2c39a01420d0806ull, 0x5ce44109e5c38357ull, 0x8532912b848b1f5bull, 0x76141c17aa59c5f6ull + }, + { // base nonce 4096 + 0x30301278872d1766ull, 0xfd009372bb888e00ull, 0xb93e45944b9cb923ull, 0xe74cb4f099ac3ea9ull, 0x3cd9b2381d1cde3cull, 0xb69bddb5df734350ull, 0xbccc2d57615b8465ull, 0xa30bbff25bf3d723ull, + 0x66ca1a4c82d500d0ull, 0xbb60ed823f1bee84ull, 0x6312dd68c839352full, 0xf15a1a4e9b902225ull, 0xcf315f5496b73b41ull, 0x5e6aa39724dbe6caull, 0x3a4be831dca92fd7ull, 0xbd5f1ae9b2892322ull, + 0x69582160f1bdf99cull, 0xcc41242a8f4776a2ull, 0x377a337ca9a5f44aull, 0x12e453c8c831a462ull, 0xfe278d9da6dcaf29ull, 0xfddd6db8998a7f00ull, 0x6f9d4db889ef7e98ull, 0x278294842317978dull, + 0xe4bab477a44177b0ull, 0x52c1ea0d495c960dull, 0xf1504ce7e9c0d818ull, 0x15673b7d791269d7ull, 0x18828b1a6fd35583ull, 0xa6d769abd44e3827ull, 0x82b838f1172afc66ull, 0x5b18a4842a119defull + }, + { // base nonce 1000000 + 0xfd55a3c4be548e4eull, 0x00f110e15155ab33ull, 0x374d6799a6b45a35ull, 0x27f9962d27ae0515ull, 0xd0ffb5d4c0f1dd83ull, 0x718d962e2ec9d7e7ull, 0xe07c572ca8705952ull, 0x99db60352bae63b7ull, + 0x065808eaebc7f7c8ull, 0x0072533c9ddc9f30ull, 0x7b11d0becd7297edull, 0xcf4b94c52549288aull, 0x0e8d98977c965c26ull, 0x890c6664b01fb638ull, 0x4d0a4b30a326dca6ull, 0x948a02536dca9925ull, + 0x1d9b3442fd4c3e0aull, 0x8b555804a68adabbull, 0x4942826f941b2089ull, 0x967b94a90c51196eull, 0xea8cf4b416d1878cull, 0xbefa6e39bfc1860aull, 0x53cbf72fe366ef94ull, 0xbfe204292b3f8b03ull, + 0xe269bfe9953bcf03ull, 0x24a758d6174e0da3ull, 0x25f57ab508d87895ull, 0xe45ec6208eb56f7dull, 0x47c891dda3f102c6ull, 0x088698d191e8dc33ull, 0xc1c4cd437ae11388ull, 0xa17c529d1a26c64dull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/mixA-0/vectors.json b/proto-cuda/packs-readwidth/mixA-0/vectors.json new file mode 100644 index 000000000..8f159a321 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-0/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-readwidth/A/0", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0xc02af7d21c32b6e2", "0xf9676c7b81e42590", "0x3b5b2f433e291756", "0x2307064009985ab7", "0xf1f2d84aa86f1f5a", "0xa7ed219296c204be", "0xf8105c9692ac08e7", "0x04377ca8373b8e1f", + "0x2f76672aa0a6104d", "0xac85110c18a8edc8", "0x3b4c20d7d3c22e52", "0x365c221488fa37e4", "0x989dab8f749cd976", "0x37f6690cc11dd5a7", "0x6159da4241e04ae2", "0x3c19bfdcd7821055", + "0x5b48a58c99800f4f", "0x991018ff7b067a45", "0xf281167cb9c28ef2", "0xe9497f5c0207892d", "0xac1ccf9060f7c78b", "0xf6f4282e68ceff58", "0xfa2694f87cf04ef6", "0x62d01632e98c6a11", + "0x63a875eb903ee9af", "0x680ad026ac9e5e1f", "0x63d80718c2e2912b", "0x377a3c901dadb9f5", "0xb2c39a01420d0806", "0x5ce44109e5c38357", "0x8532912b848b1f5b", "0x76141c17aa59c5f6" + ]}, + {"base_nonce": 4096, "expected": [ + "0x30301278872d1766", "0xfd009372bb888e00", "0xb93e45944b9cb923", "0xe74cb4f099ac3ea9", "0x3cd9b2381d1cde3c", "0xb69bddb5df734350", "0xbccc2d57615b8465", "0xa30bbff25bf3d723", + "0x66ca1a4c82d500d0", "0xbb60ed823f1bee84", "0x6312dd68c839352f", "0xf15a1a4e9b902225", "0xcf315f5496b73b41", "0x5e6aa39724dbe6ca", "0x3a4be831dca92fd7", "0xbd5f1ae9b2892322", + "0x69582160f1bdf99c", "0xcc41242a8f4776a2", "0x377a337ca9a5f44a", "0x12e453c8c831a462", "0xfe278d9da6dcaf29", "0xfddd6db8998a7f00", "0x6f9d4db889ef7e98", "0x278294842317978d", + "0xe4bab477a44177b0", "0x52c1ea0d495c960d", "0xf1504ce7e9c0d818", "0x15673b7d791269d7", "0x18828b1a6fd35583", "0xa6d769abd44e3827", "0x82b838f1172afc66", "0x5b18a4842a119def" + ]}, + {"base_nonce": 1000000, "expected": [ + "0xfd55a3c4be548e4e", "0x00f110e15155ab33", "0x374d6799a6b45a35", "0x27f9962d27ae0515", "0xd0ffb5d4c0f1dd83", "0x718d962e2ec9d7e7", "0xe07c572ca8705952", "0x99db60352bae63b7", + "0x065808eaebc7f7c8", "0x0072533c9ddc9f30", "0x7b11d0becd7297ed", "0xcf4b94c52549288a", "0x0e8d98977c965c26", "0x890c6664b01fb638", "0x4d0a4b30a326dca6", "0x948a02536dca9925", + "0x1d9b3442fd4c3e0a", "0x8b555804a68adabb", "0x4942826f941b2089", "0x967b94a90c51196e", "0xea8cf4b416d1878c", "0xbefa6e39bfc1860a", "0x53cbf72fe366ef94", "0xbfe204292b3f8b03", + "0xe269bfe9953bcf03", "0x24a758d6174e0da3", "0x25f57ab508d87895", "0xe45ec6208eb56f7d", "0x47c891dda3f102c6", "0x088698d191e8dc33", "0xc1c4cd437ae11388", "0xa17c529d1a26c64d" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/mixA-1/kernel.cl b/proto-cuda/packs-readwidth/mixA-1/kernel.cl new file mode 100644 index 000000000..58bd56c5e --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-1/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/1". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0xa7198abfu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x6cd0f8cfu; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x6cd0f8cfu; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xe4ef8ebfu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xe4ef8ebfu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x03ee1e65u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x03ee1e65u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xcfcfc5c0u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xcfcfc5c0u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x82e9e19bu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x82e9e19bu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x5d7a8a2fu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x5d7a8a2fu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xfafccd93u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0xfafccd93u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0xa7198abfu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r4 = r4 + r3 + ((((sel >> 8u) & 1u) != 0u) ? 0x5b2d5fe3u : 0x4a4c6caau); // 0 add + r7 = r7 | r6; // 1 or + r2 = r2 * r7; // 2 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 2u); r5 = r5 ^ t_; } // 3 shfl + r4 = r4 ^ r6; // 4 xor + r0 = r0 ^ ds[r4 & mask]; // 5 load + r3 = rotl_imm(r3, 17u); // 6 rotl + r1 = r1 + r4 + ((((sel >> 27u) & 1u) != 0u) ? 0x83aa2c52u : 0xaeb38cc5u); // 7 add + r1 = rotl_imm(r1, 31u); // 8 rotl + r4 = r4 ^ r7; // 9 xor + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 10 load + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 11 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 1u); r5 = r5 ^ t_; } // 12 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r1 = r1 ^ t_; } // 13 shfl + r5 = r5 * r7; // 14 mul + r3 = r3 ^ r0; // 15 xor + r5 = r5 * r4; // 16 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r2 = r2 ^ t_; } // 17 shfl + r4 = r4 + r1 + ((((sel >> 31u) & 1u) != 0u) ? 0xfb36bddau : 0x87d3a998u); // 18 add + r0 = r0 + r1 + ((((sel >> 22u) & 1u) != 0u) ? 0xf2ef7076u : 0x88921092u); // 19 add + r6 = r6 + r1 + ((((sel >> 3u) & 1u) != 0u) ? 0xe42e69c6u : 0x35e06b74u); // 20 add + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r2 = r2 ^ t_; } // 21 shfl + r3 = r3 + r6 + ((((sel >> 3u) & 1u) != 0u) ? 0xb45d9671u : 0x74df5734u); // 22 add + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 23 shfl + r4 = r4 ^ ds[r2 & mask]; // 24 load + r2 = r2 ^ r1; // 25 xor + r7 = r7 * r6; // 26 mul + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 27 load + r3 = rotl_imm(r3, 29u); // 28 rotl + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 8u); r3 = r3 ^ t_; } // 29 shfl + r1 = r1 ^ ds[r4 & mask]; // 30 load + r4 = r4 - r6; // 31 sub + r2 = r2 * r4; // 32 mul + r4 = r3 * r3 + r4; // 33 mad + r0 = r0 + r6 + ((((sel >> 15u) & 1u) != 0u) ? 0x97d2762du : 0x19318d72u); // 34 add + r7 = r7 + r1 + ((((sel >> 19u) & 1u) != 0u) ? 0x79830eadu : 0x26b63296u); // 35 add + r7 = r7 ^ ds[r2 & mask]; // 36 load + r0 = rotr_var(r0, r7); // 37 rotr + r2 = r2 + r7 + ((((sel >> 28u) & 1u) != 0u) ? 0x23c8dec4u : 0xb228dc81u); // 38 add + r0 = r0 ^ ds[r5 & mask]; // 39 load + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 2u); r7 = r7 ^ t_; } // 40 shfl + r5 = r0 * r1 + r5; // 41 mad + r2 = r2 ^ r3; // 42 xor + r7 = r7 ^ ds[r0 & mask]; // 43 load + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 44 load + r2 = r2 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x5a3a7fe1u : 0xcc64df8eu); // 45 add + r0 = rotr_var(r0, r5); // 46 rotr + r3 = rotr_var(r3, r7); // 47 rotr + r2 = r7 * r6 + r2; // 48 mad + r6 = r3 * r5 + r6; // 49 mad + r1 = r1 ^ ds[r6 & mask]; // 50 load + r7 = r7 ^ ds[r1 & mask]; // 51 load + r3 = r3 + r4 + ((((sel >> 6u) & 1u) != 0u) ? 0xcba22643u : 0x7a646d78u); // 52 add + r4 = r4 ^ r0; // 53 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r3 = r3 ^ t_; } // 54 shfl + r0 = rotr_var(r0, r1); // 55 rotr + r5 = r5 ^ r3; // 56 xor + r1 = r1 + r0 + ((((sel >> 8u) & 1u) != 0u) ? 0x1834af00u : 0x6f26b909u); // 57 add + r6 = r6 ^ ds[r4 & mask]; // 58 load + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r1 = r1 ^ ds[r0 & mask]; // 60 load + r4 = r4 ^ r7; // 61 xor + r6 = r2 * r1 + r6; // 62 mad + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 63 load + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/mixA-1/kernel.cu b/proto-cuda/packs-readwidth/mixA-1/kernel.cu new file mode 100644 index 000000000..be52afc4d --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-1/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/1". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0xa7198abfu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x6cd0f8cfu; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x6cd0f8cfu; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xe4ef8ebfu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xe4ef8ebfu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x03ee1e65u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x03ee1e65u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xcfcfc5c0u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xcfcfc5c0u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x82e9e19bu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0x82e9e19bu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x5d7a8a2fu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x5d7a8a2fu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xfafccd93u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0xfafccd93u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0xa7198abfu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r4 = r4 + r3 + ((((sel >> 8u) & 1u) != 0u) ? 0x5b2d5fe3u : 0x4a4c6caau); // 0 add + r7 = r7 | r6; // 1 or + r2 = r2 * r7; // 2 mul + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r1, 2); // 3 shfl + r4 = r4 ^ r6; // 4 xor + r0 = r0 ^ ds[r4 & mask]; // 5 load + r3 = rotl_imm(r3, 17u); // 6 rotl + r1 = r1 + r4 + ((((sel >> 27u) & 1u) != 0u) ? 0x83aa2c52u : 0xaeb38cc5u); // 7 add + r1 = rotl_imm(r1, 31u); // 8 rotl + r4 = r4 ^ r7; // 9 xor + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 10 load + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 11 load + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 1); // 12 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 13 shfl + r5 = r5 * r7; // 14 mul + r3 = r3 ^ r0; // 15 xor + r5 = r5 * r4; // 16 mul + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 17 shfl + r4 = r4 + r1 + ((((sel >> 31u) & 1u) != 0u) ? 0xfb36bddau : 0x87d3a998u); // 18 add + r0 = r0 + r1 + ((((sel >> 22u) & 1u) != 0u) ? 0xf2ef7076u : 0x88921092u); // 19 add + r6 = r6 + r1 + ((((sel >> 3u) & 1u) != 0u) ? 0xe42e69c6u : 0x35e06b74u); // 20 add + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 21 shfl + r3 = r3 + r6 + ((((sel >> 3u) & 1u) != 0u) ? 0xb45d9671u : 0x74df5734u); // 22 add + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 23 shfl + r4 = r4 ^ ds[r2 & mask]; // 24 load + r2 = r2 ^ r1; // 25 xor + r7 = r7 * r6; // 26 mul + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 27 load + r3 = rotl_imm(r3, 29u); // 28 rotl + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r0, 8); // 29 shfl + r1 = r1 ^ ds[r4 & mask]; // 30 load + r4 = r4 - r6; // 31 sub + r2 = r2 * r4; // 32 mul + r4 = r3 * r3 + r4; // 33 mad + r0 = r0 + r6 + ((((sel >> 15u) & 1u) != 0u) ? 0x97d2762du : 0x19318d72u); // 34 add + r7 = r7 + r1 + ((((sel >> 19u) & 1u) != 0u) ? 0x79830eadu : 0x26b63296u); // 35 add + r7 = r7 ^ ds[r2 & mask]; // 36 load + r0 = rotr_var(r0, r7); // 37 rotr + r2 = r2 + r7 + ((((sel >> 28u) & 1u) != 0u) ? 0x23c8dec4u : 0xb228dc81u); // 38 add + r0 = r0 ^ ds[r5 & mask]; // 39 load + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r5, 2); // 40 shfl + r5 = r0 * r1 + r5; // 41 mad + r2 = r2 ^ r3; // 42 xor + r7 = r7 ^ ds[r0 & mask]; // 43 load + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 44 load + r2 = r2 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x5a3a7fe1u : 0xcc64df8eu); // 45 add + r0 = rotr_var(r0, r5); // 46 rotr + r3 = rotr_var(r3, r7); // 47 rotr + r2 = r7 * r6 + r2; // 48 mad + r6 = r3 * r5 + r6; // 49 mad + r1 = r1 ^ ds[r6 & mask]; // 50 load + r7 = r7 ^ ds[r1 & mask]; // 51 load + r3 = r3 + r4 + ((((sel >> 6u) & 1u) != 0u) ? 0xcba22643u : 0x7a646d78u); // 52 add + r4 = r4 ^ r0; // 53 xor + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 54 shfl + r0 = rotr_var(r0, r1); // 55 rotr + r5 = r5 ^ r3; // 56 xor + r1 = r1 + r0 + ((((sel >> 8u) & 1u) != 0u) ? 0x1834af00u : 0x6f26b909u); // 57 add + r6 = r6 ^ ds[r4 & mask]; // 58 load + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r1 = r1 ^ ds[r0 & mask]; // 60 load + r4 = r4 ^ r7; // 61 xor + r6 = r2 * r1 + r6; // 62 mad + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 63 load + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixA-1/kernel_bound.cl b/proto-cuda/packs-readwidth/mixA-1/kernel_bound.cl new file mode 100644 index 000000000..b7f29f2f7 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-1/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/1". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0xa7198abfu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x6cd0f8cfu; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x6cd0f8cfu; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xe4ef8ebfu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xe4ef8ebfu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x03ee1e65u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x03ee1e65u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xcfcfc5c0u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xcfcfc5c0u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x82e9e19bu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x82e9e19bu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x5d7a8a2fu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x5d7a8a2fu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xfafccd93u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0xfafccd93u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0xa7198abfu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r4 = r4 + r3 + ((((sel >> 8u) & 1u) != 0u) ? 0x5b2d5fe3u : 0x4a4c6caau); // 0 add + r7 = r7 | r6; // 1 or + r2 = r2 * r7; // 2 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 2u); r5 = r5 ^ t_; } // 3 shfl + r4 = r4 ^ r6; // 4 xor + r0 = r0 ^ ds[r4 & mask]; // 5 load + r3 = rotl_imm(r3, 17u); // 6 rotl + r1 = r1 + r4 + ((((sel >> 27u) & 1u) != 0u) ? 0x83aa2c52u : 0xaeb38cc5u); // 7 add + r1 = rotl_imm(r1, 31u); // 8 rotl + r4 = r4 ^ r7; // 9 xor + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 10 load + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 11 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 1u); r5 = r5 ^ t_; } // 12 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r1 = r1 ^ t_; } // 13 shfl + r5 = r5 * r7; // 14 mul + r3 = r3 ^ r0; // 15 xor + r5 = r5 * r4; // 16 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r2 = r2 ^ t_; } // 17 shfl + r4 = r4 + r1 + ((((sel >> 31u) & 1u) != 0u) ? 0xfb36bddau : 0x87d3a998u); // 18 add + r0 = r0 + r1 + ((((sel >> 22u) & 1u) != 0u) ? 0xf2ef7076u : 0x88921092u); // 19 add + r6 = r6 + r1 + ((((sel >> 3u) & 1u) != 0u) ? 0xe42e69c6u : 0x35e06b74u); // 20 add + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r2 = r2 ^ t_; } // 21 shfl + r3 = r3 + r6 + ((((sel >> 3u) & 1u) != 0u) ? 0xb45d9671u : 0x74df5734u); // 22 add + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 23 shfl + r4 = r4 ^ ds[r2 & mask]; // 24 load + r2 = r2 ^ r1; // 25 xor + r7 = r7 * r6; // 26 mul + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 27 load + r3 = rotl_imm(r3, 29u); // 28 rotl + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 8u); r3 = r3 ^ t_; } // 29 shfl + r1 = r1 ^ ds[r4 & mask]; // 30 load + r4 = r4 - r6; // 31 sub + r2 = r2 * r4; // 32 mul + r4 = r3 * r3 + r4; // 33 mad + r0 = r0 + r6 + ((((sel >> 15u) & 1u) != 0u) ? 0x97d2762du : 0x19318d72u); // 34 add + r7 = r7 + r1 + ((((sel >> 19u) & 1u) != 0u) ? 0x79830eadu : 0x26b63296u); // 35 add + r7 = r7 ^ ds[r2 & mask]; // 36 load + r0 = rotr_var(r0, r7); // 37 rotr + r2 = r2 + r7 + ((((sel >> 28u) & 1u) != 0u) ? 0x23c8dec4u : 0xb228dc81u); // 38 add + r0 = r0 ^ ds[r5 & mask]; // 39 load + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 2u); r7 = r7 ^ t_; } // 40 shfl + r5 = r0 * r1 + r5; // 41 mad + r2 = r2 ^ r3; // 42 xor + r7 = r7 ^ ds[r0 & mask]; // 43 load + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 44 load + r2 = r2 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x5a3a7fe1u : 0xcc64df8eu); // 45 add + r0 = rotr_var(r0, r5); // 46 rotr + r3 = rotr_var(r3, r7); // 47 rotr + r2 = r7 * r6 + r2; // 48 mad + r6 = r3 * r5 + r6; // 49 mad + r1 = r1 ^ ds[r6 & mask]; // 50 load + r7 = r7 ^ ds[r1 & mask]; // 51 load + r3 = r3 + r4 + ((((sel >> 6u) & 1u) != 0u) ? 0xcba22643u : 0x7a646d78u); // 52 add + r4 = r4 ^ r0; // 53 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r3 = r3 ^ t_; } // 54 shfl + r0 = rotr_var(r0, r1); // 55 rotr + r5 = r5 ^ r3; // 56 xor + r1 = r1 + r0 + ((((sel >> 8u) & 1u) != 0u) ? 0x1834af00u : 0x6f26b909u); // 57 add + r6 = r6 ^ ds[r4 & mask]; // 58 load + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r1 = r1 ^ ds[r0 & mask]; // 60 load + r4 = r4 ^ r7; // 61 xor + r6 = r2 * r1 + r6; // 62 mad + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 63 load + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r4 = r4 + r3 + ((((sel >> 8u) & 1u) != 0u) ? 0x5b2d5fe3u : 0x4a4c6caau); // 0 add + r7 = r7 | r6; // 1 or + r2 = r2 * r7; // 2 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 2u); r5 = r5 ^ t_; } // 3 shfl + r4 = r4 ^ r6; // 4 xor + r0 = r0 ^ ds[r4 & mask]; // 5 load + r3 = rotl_imm(r3, 17u); // 6 rotl + r1 = r1 + r4 + ((((sel >> 27u) & 1u) != 0u) ? 0x83aa2c52u : 0xaeb38cc5u); // 7 add + r1 = rotl_imm(r1, 31u); // 8 rotl + r4 = r4 ^ r7; // 9 xor + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 10 load + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 11 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 1u); r5 = r5 ^ t_; } // 12 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r1 = r1 ^ t_; } // 13 shfl + r5 = r5 * r7; // 14 mul + r3 = r3 ^ r0; // 15 xor + r5 = r5 * r4; // 16 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r2 = r2 ^ t_; } // 17 shfl + r4 = r4 + r1 + ((((sel >> 31u) & 1u) != 0u) ? 0xfb36bddau : 0x87d3a998u); // 18 add + r0 = r0 + r1 + ((((sel >> 22u) & 1u) != 0u) ? 0xf2ef7076u : 0x88921092u); // 19 add + r6 = r6 + r1 + ((((sel >> 3u) & 1u) != 0u) ? 0xe42e69c6u : 0x35e06b74u); // 20 add + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r2 = r2 ^ t_; } // 21 shfl + r3 = r3 + r6 + ((((sel >> 3u) & 1u) != 0u) ? 0xb45d9671u : 0x74df5734u); // 22 add + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 23 shfl + r4 = r4 ^ ds[r2 & mask]; // 24 load + r2 = r2 ^ r1; // 25 xor + r7 = r7 * r6; // 26 mul + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 27 load + r3 = rotl_imm(r3, 29u); // 28 rotl + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 8u); r3 = r3 ^ t_; } // 29 shfl + r1 = r1 ^ ds[r4 & mask]; // 30 load + r4 = r4 - r6; // 31 sub + r2 = r2 * r4; // 32 mul + r4 = r3 * r3 + r4; // 33 mad + r0 = r0 + r6 + ((((sel >> 15u) & 1u) != 0u) ? 0x97d2762du : 0x19318d72u); // 34 add + r7 = r7 + r1 + ((((sel >> 19u) & 1u) != 0u) ? 0x79830eadu : 0x26b63296u); // 35 add + r7 = r7 ^ ds[r2 & mask]; // 36 load + r0 = rotr_var(r0, r7); // 37 rotr + r2 = r2 + r7 + ((((sel >> 28u) & 1u) != 0u) ? 0x23c8dec4u : 0xb228dc81u); // 38 add + r0 = r0 ^ ds[r5 & mask]; // 39 load + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 2u); r7 = r7 ^ t_; } // 40 shfl + r5 = r0 * r1 + r5; // 41 mad + r2 = r2 ^ r3; // 42 xor + r7 = r7 ^ ds[r0 & mask]; // 43 load + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 44 load + r2 = r2 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x5a3a7fe1u : 0xcc64df8eu); // 45 add + r0 = rotr_var(r0, r5); // 46 rotr + r3 = rotr_var(r3, r7); // 47 rotr + r2 = r7 * r6 + r2; // 48 mad + r6 = r3 * r5 + r6; // 49 mad + r1 = r1 ^ ds[r6 & mask]; // 50 load + r7 = r7 ^ ds[r1 & mask]; // 51 load + r3 = r3 + r4 + ((((sel >> 6u) & 1u) != 0u) ? 0xcba22643u : 0x7a646d78u); // 52 add + r4 = r4 ^ r0; // 53 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r3 = r3 ^ t_; } // 54 shfl + r0 = rotr_var(r0, r1); // 55 rotr + r5 = r5 ^ r3; // 56 xor + r1 = r1 + r0 + ((((sel >> 8u) & 1u) != 0u) ? 0x1834af00u : 0x6f26b909u); // 57 add + r6 = r6 ^ ds[r4 & mask]; // 58 load + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r1 = r1 ^ ds[r0 & mask]; // 60 load + r4 = r4 ^ r7; // 61 xor + r6 = r2 * r1 + r6; // 62 mad + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 63 load + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-1/kernel_bound.cu b/proto-cuda/packs-readwidth/mixA-1/kernel_bound.cu new file mode 100644 index 000000000..3c97ecde0 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-1/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/1". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r4 = r4 + r3 + ((((sel >> 8u) & 1u) != 0u) ? 0x5b2d5fe3u : 0x4a4c6caau); // 0 add + r7 = r7 | r6; // 1 or + r2 = r2 * r7; // 2 mul + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r1, 2); // 3 shfl + r4 = r4 ^ r6; // 4 xor + r0 = r0 ^ ds[r4 & mask]; // 5 load + r3 = rotl_imm(r3, 17u); // 6 rotl + r1 = r1 + r4 + ((((sel >> 27u) & 1u) != 0u) ? 0x83aa2c52u : 0xaeb38cc5u); // 7 add + r1 = rotl_imm(r1, 31u); // 8 rotl + r4 = r4 ^ r7; // 9 xor + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 10 load + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 11 load + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 1); // 12 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 13 shfl + r5 = r5 * r7; // 14 mul + r3 = r3 ^ r0; // 15 xor + r5 = r5 * r4; // 16 mul + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 17 shfl + r4 = r4 + r1 + ((((sel >> 31u) & 1u) != 0u) ? 0xfb36bddau : 0x87d3a998u); // 18 add + r0 = r0 + r1 + ((((sel >> 22u) & 1u) != 0u) ? 0xf2ef7076u : 0x88921092u); // 19 add + r6 = r6 + r1 + ((((sel >> 3u) & 1u) != 0u) ? 0xe42e69c6u : 0x35e06b74u); // 20 add + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 21 shfl + r3 = r3 + r6 + ((((sel >> 3u) & 1u) != 0u) ? 0xb45d9671u : 0x74df5734u); // 22 add + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 23 shfl + r4 = r4 ^ ds[r2 & mask]; // 24 load + r2 = r2 ^ r1; // 25 xor + r7 = r7 * r6; // 26 mul + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 27 load + r3 = rotl_imm(r3, 29u); // 28 rotl + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r0, 8); // 29 shfl + r1 = r1 ^ ds[r4 & mask]; // 30 load + r4 = r4 - r6; // 31 sub + r2 = r2 * r4; // 32 mul + r4 = r3 * r3 + r4; // 33 mad + r0 = r0 + r6 + ((((sel >> 15u) & 1u) != 0u) ? 0x97d2762du : 0x19318d72u); // 34 add + r7 = r7 + r1 + ((((sel >> 19u) & 1u) != 0u) ? 0x79830eadu : 0x26b63296u); // 35 add + r7 = r7 ^ ds[r2 & mask]; // 36 load + r0 = rotr_var(r0, r7); // 37 rotr + r2 = r2 + r7 + ((((sel >> 28u) & 1u) != 0u) ? 0x23c8dec4u : 0xb228dc81u); // 38 add + r0 = r0 ^ ds[r5 & mask]; // 39 load + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r5, 2); // 40 shfl + r5 = r0 * r1 + r5; // 41 mad + r2 = r2 ^ r3; // 42 xor + r7 = r7 ^ ds[r0 & mask]; // 43 load + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 44 load + r2 = r2 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x5a3a7fe1u : 0xcc64df8eu); // 45 add + r0 = rotr_var(r0, r5); // 46 rotr + r3 = rotr_var(r3, r7); // 47 rotr + r2 = r7 * r6 + r2; // 48 mad + r6 = r3 * r5 + r6; // 49 mad + r1 = r1 ^ ds[r6 & mask]; // 50 load + r7 = r7 ^ ds[r1 & mask]; // 51 load + r3 = r3 + r4 + ((((sel >> 6u) & 1u) != 0u) ? 0xcba22643u : 0x7a646d78u); // 52 add + r4 = r4 ^ r0; // 53 xor + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 54 shfl + r0 = rotr_var(r0, r1); // 55 rotr + r5 = r5 ^ r3; // 56 xor + r1 = r1 + r0 + ((((sel >> 8u) & 1u) != 0u) ? 0x1834af00u : 0x6f26b909u); // 57 add + r6 = r6 ^ ds[r4 & mask]; // 58 load + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r1 = r1 ^ ds[r0 & mask]; // 60 load + r4 = r4 ^ r7; // 61 xor + r6 = r2 * r1 + r6; // 62 mad + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 63 load + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixA-1/memhard.h b/proto-cuda/packs-readwidth/mixA-1/memhard.h new file mode 100644 index 000000000..b813a93c4 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-1/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/1". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/mixA-1/memhard.metal b/proto-cuda/packs-readwidth/mixA-1/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-1/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/mixA-1/program.h b/proto-cuda/packs-readwidth/mixA-1/program.h new file mode 100644 index 000000000..16c67a1bd --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-1/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/1". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-readwidth/A/1" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d7265616477696474682f412f31" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x6408f2e28fdd344cull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=12 shfl=9 xor=8 mad=5 mul=5 rotr=4 rotl=3 or=1 sub=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "mix50-35-15" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 50, 35, 15 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 10, 4, 2 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 1856 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0xa7198abfu, 0x6cd0f8cfu, 0xe4ef8ebfu, 0x03ee1e65u, 0xcfcfc5c0u, 0x82e9e19bu, 0x5d7a8a2fu, 0xfafccd93u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/mixA-1/program.json b/proto-cuda/packs-readwidth/mixA-1/program.json new file mode 100644 index 000000000..279a53367 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-1/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x6408f2e28fdd344c", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-readwidth/A/1", + "seed_bytes": "69676e65756d2d7265616477696474682f412f31", + "seed_words": ["0xa7198abf", "0x6cd0f8cf", "0xe4ef8ebf", "0x03ee1e65", "0xcfcfc5c0", "0x82e9e19b", "0x5d7a8a2f", "0xfafccd93"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "mix50-35-15", + "load_slots": 16, + "load_mix_percent_4_16_64": [50, 35, 15], + "load_width_counts_4_16_64": [10, 4, 2], + "bytes_per_hash": 1856, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "add": 12, "shfl": 9, "xor": 8, "mad": 5, "mul": 5, "rotr": 4, "rotl": 3, "or": 1, "sub": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "add", "dst": 4, "src": 3, "src2": 4, "imm": "0x4a4c6caa", "imm2": "0x5b2d5fe3", "rot": 28, "bit": 8, "mask": 1, "width": 1}, + {"i": 1, "op": "or", "dst": 7, "src": 6, "src2": 0, "imm": "0x8c4c9f7c", "imm2": "0x3b47c42a", "rot": 3, "bit": 13, "mask": 8, "width": 1}, + {"i": 2, "op": "mul", "dst": 2, "src": 7, "src2": 1, "imm": "0xe7cf827b", "imm2": "0x786c6890", "rot": 24, "bit": 25, "mask": 8, "width": 1}, + {"i": 3, "op": "shfl", "dst": 5, "src": 1, "src2": 6, "imm": "0xa686f6b7", "imm2": "0x48920bc0", "rot": 1, "bit": 5, "mask": 2, "width": 1}, + {"i": 4, "op": "xor", "dst": 4, "src": 6, "src2": 1, "imm": "0x58489525", "imm2": "0x1c56b23f", "rot": 25, "bit": 21, "mask": 1, "width": 1}, + {"i": 5, "op": "load", "dst": 0, "src": 4, "src2": 6, "imm": "0xacf5203a", "imm2": "0x76e3d8a1", "rot": 22, "bit": 20, "mask": 1, "width": 1}, + {"i": 6, "op": "rotl", "dst": 3, "src": 1, "src2": 6, "imm": "0x2507e906", "imm2": "0xf9c5c1f1", "rot": 17, "bit": 11, "mask": 4, "width": 1}, + {"i": 7, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xaeb38cc5", "imm2": "0x83aa2c52", "rot": 10, "bit": 27, "mask": 4, "width": 1}, + {"i": 8, "op": "rotl", "dst": 1, "src": 2, "src2": 4, "imm": "0x261b9557", "imm2": "0xe0df5769", "rot": 31, "bit": 30, "mask": 8, "width": 1}, + {"i": 9, "op": "xor", "dst": 4, "src": 7, "src2": 1, "imm": "0x8f733ec9", "imm2": "0x468afd76", "rot": 3, "bit": 1, "mask": 2, "width": 1}, + {"i": 10, "op": "load", "dst": 6, "src": 0, "src2": 1, "imm": "0x53eef18d", "imm2": "0x763b1f0a", "rot": 20, "bit": 20, "mask": 4, "width": 4}, + {"i": 11, "op": "load", "dst": 0, "src": 6, "src2": 4, "imm": "0xa2f22f71", "imm2": "0xe0a8243e", "rot": 20, "bit": 30, "mask": 1, "width": 4}, + {"i": 12, "op": "shfl", "dst": 5, "src": 2, "src2": 6, "imm": "0x88827407", "imm2": "0xe9f1befa", "rot": 10, "bit": 11, "mask": 1, "width": 1}, + {"i": 13, "op": "shfl", "dst": 1, "src": 4, "src2": 7, "imm": "0x88867e60", "imm2": "0x2726b465", "rot": 9, "bit": 6, "mask": 2, "width": 1}, + {"i": 14, "op": "mul", "dst": 5, "src": 7, "src2": 2, "imm": "0x41530f45", "imm2": "0xf0badd76", "rot": 1, "bit": 13, "mask": 1, "width": 1}, + {"i": 15, "op": "xor", "dst": 3, "src": 0, "src2": 6, "imm": "0x29e41472", "imm2": "0xd2c1de0c", "rot": 22, "bit": 5, "mask": 16, "width": 1}, + {"i": 16, "op": "mul", "dst": 5, "src": 4, "src2": 4, "imm": "0xafa3d0bd", "imm2": "0x1568e384", "rot": 19, "bit": 20, "mask": 2, "width": 1}, + {"i": 17, "op": "shfl", "dst": 2, "src": 0, "src2": 6, "imm": "0xffccd357", "imm2": "0x4b21f8cf", "rot": 3, "bit": 6, "mask": 4, "width": 1}, + {"i": 18, "op": "add", "dst": 4, "src": 1, "src2": 3, "imm": "0x87d3a998", "imm2": "0xfb36bdda", "rot": 29, "bit": 31, "mask": 4, "width": 1}, + {"i": 19, "op": "add", "dst": 0, "src": 1, "src2": 0, "imm": "0x88921092", "imm2": "0xf2ef7076", "rot": 20, "bit": 22, "mask": 2, "width": 1}, + {"i": 20, "op": "add", "dst": 6, "src": 1, "src2": 4, "imm": "0x35e06b74", "imm2": "0xe42e69c6", "rot": 3, "bit": 3, "mask": 2, "width": 1}, + {"i": 21, "op": "shfl", "dst": 2, "src": 5, "src2": 0, "imm": "0xa7323c7a", "imm2": "0x81d43278", "rot": 16, "bit": 10, "mask": 8, "width": 1}, + {"i": 22, "op": "add", "dst": 3, "src": 6, "src2": 2, "imm": "0x74df5734", "imm2": "0xb45d9671", "rot": 26, "bit": 3, "mask": 2, "width": 1}, + {"i": 23, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x1e0660fe", "imm2": "0x13c7e83f", "rot": 13, "bit": 15, "mask": 8, "width": 1}, + {"i": 24, "op": "load", "dst": 4, "src": 2, "src2": 2, "imm": "0x68bfe9f0", "imm2": "0x2067c339", "rot": 5, "bit": 11, "mask": 1, "width": 1}, + {"i": 25, "op": "xor", "dst": 2, "src": 1, "src2": 3, "imm": "0x383e819c", "imm2": "0xef4f43f5", "rot": 2, "bit": 25, "mask": 2, "width": 1}, + {"i": 26, "op": "mul", "dst": 7, "src": 6, "src2": 4, "imm": "0x45892e5e", "imm2": "0x5c358e4b", "rot": 13, "bit": 9, "mask": 1, "width": 1}, + {"i": 27, "op": "load", "dst": 2, "src": 3, "src2": 2, "imm": "0x8fe1d831", "imm2": "0xe2f0a7df", "rot": 18, "bit": 5, "mask": 16, "width": 16}, + {"i": 28, "op": "rotl", "dst": 3, "src": 4, "src2": 4, "imm": "0x78e2bf43", "imm2": "0x723070c4", "rot": 29, "bit": 6, "mask": 2, "width": 1}, + {"i": 29, "op": "shfl", "dst": 3, "src": 0, "src2": 3, "imm": "0x6a40f0ce", "imm2": "0x8f751dd0", "rot": 26, "bit": 12, "mask": 8, "width": 1}, + {"i": 30, "op": "load", "dst": 1, "src": 4, "src2": 6, "imm": "0x3bcdf182", "imm2": "0x5ec397eb", "rot": 5, "bit": 24, "mask": 16, "width": 1}, + {"i": 31, "op": "sub", "dst": 4, "src": 6, "src2": 5, "imm": "0x28631d4e", "imm2": "0x1c8b6e91", "rot": 2, "bit": 13, "mask": 4, "width": 1}, + {"i": 32, "op": "mul", "dst": 2, "src": 4, "src2": 4, "imm": "0xf8b83d3f", "imm2": "0x6bbab3d3", "rot": 3, "bit": 19, "mask": 1, "width": 1}, + {"i": 33, "op": "mad", "dst": 4, "src": 3, "src2": 3, "imm": "0xdf39e98a", "imm2": "0x1c4dbde8", "rot": 8, "bit": 17, "mask": 2, "width": 1}, + {"i": 34, "op": "add", "dst": 0, "src": 6, "src2": 4, "imm": "0x19318d72", "imm2": "0x97d2762d", "rot": 14, "bit": 15, "mask": 16, "width": 1}, + {"i": 35, "op": "add", "dst": 7, "src": 1, "src2": 1, "imm": "0x26b63296", "imm2": "0x79830ead", "rot": 26, "bit": 19, "mask": 16, "width": 1}, + {"i": 36, "op": "load", "dst": 7, "src": 2, "src2": 2, "imm": "0xf2fe58f7", "imm2": "0x3813073a", "rot": 19, "bit": 7, "mask": 2, "width": 1}, + {"i": 37, "op": "rotr", "dst": 0, "src": 7, "src2": 7, "imm": "0x268c32f9", "imm2": "0x1f02f30a", "rot": 20, "bit": 10, "mask": 1, "width": 1}, + {"i": 38, "op": "add", "dst": 2, "src": 7, "src2": 1, "imm": "0xb228dc81", "imm2": "0x23c8dec4", "rot": 14, "bit": 28, "mask": 4, "width": 1}, + {"i": 39, "op": "load", "dst": 0, "src": 5, "src2": 6, "imm": "0x08c2fe36", "imm2": "0xa232fa5b", "rot": 22, "bit": 13, "mask": 16, "width": 1}, + {"i": 40, "op": "shfl", "dst": 7, "src": 5, "src2": 4, "imm": "0x8601d298", "imm2": "0x04a37e40", "rot": 27, "bit": 21, "mask": 2, "width": 1}, + {"i": 41, "op": "mad", "dst": 5, "src": 0, "src2": 1, "imm": "0xc0112c85", "imm2": "0xcae01ef3", "rot": 16, "bit": 5, "mask": 4, "width": 1}, + {"i": 42, "op": "xor", "dst": 2, "src": 3, "src2": 1, "imm": "0x19f77b37", "imm2": "0x0952e878", "rot": 28, "bit": 23, "mask": 4, "width": 1}, + {"i": 43, "op": "load", "dst": 7, "src": 0, "src2": 6, "imm": "0x020804cf", "imm2": "0x9eb73dda", "rot": 30, "bit": 2, "mask": 1, "width": 1}, + {"i": 44, "op": "load", "dst": 0, "src": 5, "src2": 6, "imm": "0x485a8a15", "imm2": "0xdeac555a", "rot": 13, "bit": 21, "mask": 8, "width": 4}, + {"i": 45, "op": "add", "dst": 2, "src": 7, "src2": 0, "imm": "0xcc64df8e", "imm2": "0x5a3a7fe1", "rot": 18, "bit": 22, "mask": 8, "width": 1}, + {"i": 46, "op": "rotr", "dst": 0, "src": 5, "src2": 5, "imm": "0xe6e78bd9", "imm2": "0x5c932006", "rot": 31, "bit": 20, "mask": 2, "width": 1}, + {"i": 47, "op": "rotr", "dst": 3, "src": 7, "src2": 4, "imm": "0xa77d4cb5", "imm2": "0x32f2b544", "rot": 15, "bit": 16, "mask": 8, "width": 1}, + {"i": 48, "op": "mad", "dst": 2, "src": 7, "src2": 6, "imm": "0x197fc3bf", "imm2": "0x02536a02", "rot": 3, "bit": 5, "mask": 16, "width": 1}, + {"i": 49, "op": "mad", "dst": 6, "src": 3, "src2": 5, "imm": "0x6aa6ad47", "imm2": "0x0f1c2695", "rot": 6, "bit": 31, "mask": 8, "width": 1}, + {"i": 50, "op": "load", "dst": 1, "src": 6, "src2": 1, "imm": "0x106ee155", "imm2": "0x6e409245", "rot": 29, "bit": 22, "mask": 2, "width": 1}, + {"i": 51, "op": "load", "dst": 7, "src": 1, "src2": 0, "imm": "0xc604b17f", "imm2": "0x7ba442ad", "rot": 8, "bit": 27, "mask": 1, "width": 1}, + {"i": 52, "op": "add", "dst": 3, "src": 4, "src2": 3, "imm": "0x7a646d78", "imm2": "0xcba22643", "rot": 1, "bit": 6, "mask": 2, "width": 1}, + {"i": 53, "op": "xor", "dst": 4, "src": 0, "src2": 1, "imm": "0x0c951e55", "imm2": "0xe4242078", "rot": 10, "bit": 14, "mask": 4, "width": 1}, + {"i": 54, "op": "shfl", "dst": 3, "src": 7, "src2": 5, "imm": "0x1e8e8986", "imm2": "0xcc7c23e9", "rot": 26, "bit": 11, "mask": 4, "width": 1}, + {"i": 55, "op": "rotr", "dst": 0, "src": 1, "src2": 1, "imm": "0x384aa1de", "imm2": "0x03bd71f3", "rot": 7, "bit": 27, "mask": 16, "width": 1}, + {"i": 56, "op": "xor", "dst": 5, "src": 3, "src2": 6, "imm": "0x0f3bbfb3", "imm2": "0x829b8b8d", "rot": 15, "bit": 3, "mask": 4, "width": 1}, + {"i": 57, "op": "add", "dst": 1, "src": 0, "src2": 0, "imm": "0x6f26b909", "imm2": "0x1834af00", "rot": 28, "bit": 8, "mask": 2, "width": 1}, + {"i": 58, "op": "load", "dst": 6, "src": 4, "src2": 2, "imm": "0x6bcbd615", "imm2": "0xadbaf0ef", "rot": 3, "bit": 4, "mask": 8, "width": 1}, + {"i": 59, "op": "load", "dst": 1, "src": 7, "src2": 1, "imm": "0x79dc0a3f", "imm2": "0x436e360d", "rot": 30, "bit": 13, "mask": 8, "width": 16}, + {"i": 60, "op": "load", "dst": 1, "src": 0, "src2": 2, "imm": "0x30cd5237", "imm2": "0x8bd46148", "rot": 22, "bit": 29, "mask": 1, "width": 1}, + {"i": 61, "op": "xor", "dst": 4, "src": 7, "src2": 4, "imm": "0x74419ddb", "imm2": "0x70d9f3d5", "rot": 3, "bit": 26, "mask": 16, "width": 1}, + {"i": 62, "op": "mad", "dst": 6, "src": 2, "src2": 1, "imm": "0x938645fc", "imm2": "0xa9e49728", "rot": 20, "bit": 27, "mask": 8, "width": 1}, + {"i": 63, "op": "load", "dst": 3, "src": 2, "src2": 2, "imm": "0x2089409e", "imm2": "0x0437ba4e", "rot": 3, "bit": 31, "mask": 16, "width": 4} + ] +} diff --git a/proto-cuda/packs-readwidth/mixA-1/program.metal b/proto-cuda/packs-readwidth/mixA-1/program.metal new file mode 100644 index 000000000..efbde9916 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-1/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0xa7198abfu, 0x6cd0f8cfu, 0xe4ef8ebfu, 0x03ee1e65u, 0xcfcfc5c0u, 0x82e9e19bu, 0x5d7a8a2fu, 0xfafccd93u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r4 = r4 + r3 + select(0x4a4c6caau, 0x5b2d5fe3u, ((sel >> 8u) & 1u) != 0u); // 0 + r7 = r7 | r6; // 1 + r2 = r2 * r7; // 2 + r5 = r5 ^ simd_shuffle_xor(r1, (ushort)2); // 3 + r4 = r4 ^ r6; // 4 + r0 = r0 ^ dataset[r4 & MASK]; // 5 + r3 = rotl_imm(r3, 17u); // 6 + r1 = r1 + r4 + select(0xaeb38cc5u, 0x83aa2c52u, ((sel >> 27u) & 1u) != 0u); // 7 + r1 = rotl_imm(r1, 31u); // 8 + r4 = r4 ^ r7; // 9 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 10 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 11 + r5 = r5 ^ simd_shuffle_xor(r2, (ushort)1); // 12 + r1 = r1 ^ simd_shuffle_xor(r4, (ushort)2); // 13 + r5 = r5 * r7; // 14 + r3 = r3 ^ r0; // 15 + r5 = r5 * r4; // 16 + r2 = r2 ^ simd_shuffle_xor(r0, (ushort)4); // 17 + r4 = r4 + r1 + select(0x87d3a998u, 0xfb36bddau, ((sel >> 31u) & 1u) != 0u); // 18 + r0 = r0 + r1 + select(0x88921092u, 0xf2ef7076u, ((sel >> 22u) & 1u) != 0u); // 19 + r6 = r6 + r1 + select(0x35e06b74u, 0xe42e69c6u, ((sel >> 3u) & 1u) != 0u); // 20 + r2 = r2 ^ simd_shuffle_xor(r5, (ushort)8); // 21 + r3 = r3 + r6 + select(0x74df5734u, 0xb45d9671u, ((sel >> 3u) & 1u) != 0u); // 22 + r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 23 + r4 = r4 ^ dataset[r2 & MASK]; // 24 + r2 = r2 ^ r1; // 25 + r7 = r7 * r6; // 26 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 27 + r3 = rotl_imm(r3, 29u); // 28 + r3 = r3 ^ simd_shuffle_xor(r0, (ushort)8); // 29 + r1 = r1 ^ dataset[r4 & MASK]; // 30 + r4 = r4 - r6; // 31 + r2 = r2 * r4; // 32 + r4 = r3 * r3 + r4; // 33 + r0 = r0 + r6 + select(0x19318d72u, 0x97d2762du, ((sel >> 15u) & 1u) != 0u); // 34 + r7 = r7 + r1 + select(0x26b63296u, 0x79830eadu, ((sel >> 19u) & 1u) != 0u); // 35 + r7 = r7 ^ dataset[r2 & MASK]; // 36 + r0 = rotr_var(r0, r7); // 37 + r2 = r2 + r7 + select(0xb228dc81u, 0x23c8dec4u, ((sel >> 28u) & 1u) != 0u); // 38 + r0 = r0 ^ dataset[r5 & MASK]; // 39 + r7 = r7 ^ simd_shuffle_xor(r5, (ushort)2); // 40 + r5 = r0 * r1 + r5; // 41 + r2 = r2 ^ r3; // 42 + r7 = r7 ^ dataset[r0 & MASK]; // 43 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 44 + r2 = r2 + r7 + select(0xcc64df8eu, 0x5a3a7fe1u, ((sel >> 22u) & 1u) != 0u); // 45 + r0 = rotr_var(r0, r5); // 46 + r3 = rotr_var(r3, r7); // 47 + r2 = r7 * r6 + r2; // 48 + r6 = r3 * r5 + r6; // 49 + r1 = r1 ^ dataset[r6 & MASK]; // 50 + r7 = r7 ^ dataset[r1 & MASK]; // 51 + r3 = r3 + r4 + select(0x7a646d78u, 0xcba22643u, ((sel >> 6u) & 1u) != 0u); // 52 + r4 = r4 ^ r0; // 53 + r3 = r3 ^ simd_shuffle_xor(r7, (ushort)4); // 54 + r0 = rotr_var(r0, r1); // 55 + r5 = r5 ^ r3; // 56 + r1 = r1 + r0 + select(0x6f26b909u, 0x1834af00u, ((sel >> 8u) & 1u) != 0u); // 57 + r6 = r6 ^ dataset[r4 & MASK]; // 58 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 + r1 = r1 ^ dataset[r0 & MASK]; // 60 + r4 = r4 ^ r7; // 61 + r6 = r2 * r1 + r6; // 62 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-1/program_bound.metal b/proto-cuda/packs-readwidth/mixA-1/program_bound.metal new file mode 100644 index 000000000..86b7966d1 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-1/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0xa7198abfu, 0x6cd0f8cfu, 0xe4ef8ebfu, 0x03ee1e65u, 0xcfcfc5c0u, 0x82e9e19bu, 0x5d7a8a2fu, 0xfafccd93u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r4 = r4 + r3 + select(0x4a4c6caau, 0x5b2d5fe3u, ((sel >> 8u) & 1u) != 0u); // 0 + r7 = r7 | r6; // 1 + r2 = r2 * r7; // 2 + r5 = r5 ^ simd_shuffle_xor(r1, (ushort)2); // 3 + r4 = r4 ^ r6; // 4 + r0 = r0 ^ dataset[r4 & MASK]; // 5 + r3 = rotl_imm(r3, 17u); // 6 + r1 = r1 + r4 + select(0xaeb38cc5u, 0x83aa2c52u, ((sel >> 27u) & 1u) != 0u); // 7 + r1 = rotl_imm(r1, 31u); // 8 + r4 = r4 ^ r7; // 9 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 10 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 11 + r5 = r5 ^ simd_shuffle_xor(r2, (ushort)1); // 12 + r1 = r1 ^ simd_shuffle_xor(r4, (ushort)2); // 13 + r5 = r5 * r7; // 14 + r3 = r3 ^ r0; // 15 + r5 = r5 * r4; // 16 + r2 = r2 ^ simd_shuffle_xor(r0, (ushort)4); // 17 + r4 = r4 + r1 + select(0x87d3a998u, 0xfb36bddau, ((sel >> 31u) & 1u) != 0u); // 18 + r0 = r0 + r1 + select(0x88921092u, 0xf2ef7076u, ((sel >> 22u) & 1u) != 0u); // 19 + r6 = r6 + r1 + select(0x35e06b74u, 0xe42e69c6u, ((sel >> 3u) & 1u) != 0u); // 20 + r2 = r2 ^ simd_shuffle_xor(r5, (ushort)8); // 21 + r3 = r3 + r6 + select(0x74df5734u, 0xb45d9671u, ((sel >> 3u) & 1u) != 0u); // 22 + r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 23 + r4 = r4 ^ dataset[r2 & MASK]; // 24 + r2 = r2 ^ r1; // 25 + r7 = r7 * r6; // 26 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 27 + r3 = rotl_imm(r3, 29u); // 28 + r3 = r3 ^ simd_shuffle_xor(r0, (ushort)8); // 29 + r1 = r1 ^ dataset[r4 & MASK]; // 30 + r4 = r4 - r6; // 31 + r2 = r2 * r4; // 32 + r4 = r3 * r3 + r4; // 33 + r0 = r0 + r6 + select(0x19318d72u, 0x97d2762du, ((sel >> 15u) & 1u) != 0u); // 34 + r7 = r7 + r1 + select(0x26b63296u, 0x79830eadu, ((sel >> 19u) & 1u) != 0u); // 35 + r7 = r7 ^ dataset[r2 & MASK]; // 36 + r0 = rotr_var(r0, r7); // 37 + r2 = r2 + r7 + select(0xb228dc81u, 0x23c8dec4u, ((sel >> 28u) & 1u) != 0u); // 38 + r0 = r0 ^ dataset[r5 & MASK]; // 39 + r7 = r7 ^ simd_shuffle_xor(r5, (ushort)2); // 40 + r5 = r0 * r1 + r5; // 41 + r2 = r2 ^ r3; // 42 + r7 = r7 ^ dataset[r0 & MASK]; // 43 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 44 + r2 = r2 + r7 + select(0xcc64df8eu, 0x5a3a7fe1u, ((sel >> 22u) & 1u) != 0u); // 45 + r0 = rotr_var(r0, r5); // 46 + r3 = rotr_var(r3, r7); // 47 + r2 = r7 * r6 + r2; // 48 + r6 = r3 * r5 + r6; // 49 + r1 = r1 ^ dataset[r6 & MASK]; // 50 + r7 = r7 ^ dataset[r1 & MASK]; // 51 + r3 = r3 + r4 + select(0x7a646d78u, 0xcba22643u, ((sel >> 6u) & 1u) != 0u); // 52 + r4 = r4 ^ r0; // 53 + r3 = r3 ^ simd_shuffle_xor(r7, (ushort)4); // 54 + r0 = rotr_var(r0, r1); // 55 + r5 = r5 ^ r3; // 56 + r1 = r1 + r0 + select(0x6f26b909u, 0x1834af00u, ((sel >> 8u) & 1u) != 0u); // 57 + r6 = r6 ^ dataset[r4 & MASK]; // 58 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 + r1 = r1 ^ dataset[r0 & MASK]; // 60 + r4 = r4 ^ r7; // 61 + r6 = r2 * r1 + r6; // 62 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-1/vectors.h b/proto-cuda/packs-readwidth/mixA-1/vectors.h new file mode 100644 index 000000000..467e35d77 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-1/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/1". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x9a39e4c3b724e139ull, 0x0b2f9aefe9b5a7d9ull, 0x489f25154d36b837ull, 0x2bacca9851bedf63ull, 0x23b9e9d6f6f024a2ull, 0xa297adc388367ac7ull, 0x5f77305cf55428e2ull, 0x13f19dd950363e8cull, + 0xd3cc9c7923f85ae7ull, 0x5fdfc76e2559aaebull, 0x24bbc07c2b73b227ull, 0x5c128b8728383521ull, 0x95536c4dd0071840ull, 0xc2337e88f9eba324ull, 0x8a0d0d5db0b8c0c0ull, 0x4aa918c941e517f8ull, + 0x00b15821dc1a563dull, 0xa13888b391d5c4ffull, 0x2545a17fbecd2233ull, 0xf340d8b211921866ull, 0x827e215990856eccull, 0xa524b3dff5acb2f2ull, 0x495b540800df7c9eull, 0x9e274c9e9ed7edc5ull, + 0x5b8c358f08d10146ull, 0x7745916245fd6bd6ull, 0x20ab3090ecbf86e7ull, 0x3c2429a761881d58ull, 0xb14ad7a2b75a3e5bull, 0xd040113da7d8929aull, 0x0aac7bad8ef1785full, 0x69ef296bba30963dull + }, + { // base nonce 4096 + 0x380980b425973372ull, 0xb039f8c8f2085f2eull, 0xdb9093fb95d8bffbull, 0x87fb86cda3374fcfull, 0x4138cea25d6a4d5full, 0xbe0a9b404a88c69full, 0xc696db4aa95e6473ull, 0x6d7fd33fb2be2a34ull, + 0x46acefa0c5e54e1full, 0xe898e7d08e946d6cull, 0xea998c0f7965842aull, 0xf6b60065df70f28dull, 0x24622e14557fd974ull, 0x5730af77da1186ffull, 0xf776630a6ff0a022ull, 0xa69c82d11c376933ull, + 0x93fb4e2c9c12456full, 0xa7a30a3c776b32cfull, 0x1669db65ca0f2724ull, 0x491df6828bfc7ad0ull, 0xeb5bf506947713c2ull, 0xe83fd62f415cfe9eull, 0xbc4415e97b359dd0ull, 0x2ab09721b97d003full, + 0xaf04c23a63a7bb7full, 0x6c4986d838379e4aull, 0x0dcecd1729525d44ull, 0x92ba240b16950bbaull, 0xed6df8cc108a7e50ull, 0x60252d3be009c7eaull, 0x300d103ffc4e8d35ull, 0x73522acf3d2f56d4ull + }, + { // base nonce 1000000 + 0x254a52e3824ffcb7ull, 0x35f6dfb86ffa9f8eull, 0x27838da0f63ce460ull, 0xd25ec8a499e26256ull, 0xfe3b4f012071b8b2ull, 0xde760f5a3e5378cdull, 0x0810bd3a5f232ac3ull, 0x7ee980b9098f2df2ull, + 0xbae54f58f6daab8eull, 0x6f89d253a6b184f1ull, 0xb9c87bb5f8914f6full, 0xd264d719aa5d7bd5ull, 0x23e89f82b32a5a44ull, 0x080a72f8a360f627ull, 0xe28b507b1cba42beull, 0x25b77fdc149f8eb7ull, + 0x1a34f036f09596f5ull, 0x3e67e706d4719f7eull, 0x33764e53c58097a6ull, 0xbcf62943f155349aull, 0x8eb89841fbcb8ee5ull, 0x386b0433a1567ee7ull, 0x0c5775cb8fce2a91ull, 0x64518cc52f393a59ull, + 0x8f3cdb40e81a1cddull, 0x6af5a66a0a2436c1ull, 0xb76a5386454bef44ull, 0xfdd12f4ff2f53256ull, 0xf9ff7b79b577274full, 0xe7ed16e8ad89f703ull, 0x5a64c33fdfec335eull, 0xda50f54c2c547ac1ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/mixA-1/vectors.json b/proto-cuda/packs-readwidth/mixA-1/vectors.json new file mode 100644 index 000000000..35abcdfd3 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-1/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-readwidth/A/1", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x9a39e4c3b724e139", "0x0b2f9aefe9b5a7d9", "0x489f25154d36b837", "0x2bacca9851bedf63", "0x23b9e9d6f6f024a2", "0xa297adc388367ac7", "0x5f77305cf55428e2", "0x13f19dd950363e8c", + "0xd3cc9c7923f85ae7", "0x5fdfc76e2559aaeb", "0x24bbc07c2b73b227", "0x5c128b8728383521", "0x95536c4dd0071840", "0xc2337e88f9eba324", "0x8a0d0d5db0b8c0c0", "0x4aa918c941e517f8", + "0x00b15821dc1a563d", "0xa13888b391d5c4ff", "0x2545a17fbecd2233", "0xf340d8b211921866", "0x827e215990856ecc", "0xa524b3dff5acb2f2", "0x495b540800df7c9e", "0x9e274c9e9ed7edc5", + "0x5b8c358f08d10146", "0x7745916245fd6bd6", "0x20ab3090ecbf86e7", "0x3c2429a761881d58", "0xb14ad7a2b75a3e5b", "0xd040113da7d8929a", "0x0aac7bad8ef1785f", "0x69ef296bba30963d" + ]}, + {"base_nonce": 4096, "expected": [ + "0x380980b425973372", "0xb039f8c8f2085f2e", "0xdb9093fb95d8bffb", "0x87fb86cda3374fcf", "0x4138cea25d6a4d5f", "0xbe0a9b404a88c69f", "0xc696db4aa95e6473", "0x6d7fd33fb2be2a34", + "0x46acefa0c5e54e1f", "0xe898e7d08e946d6c", "0xea998c0f7965842a", "0xf6b60065df70f28d", "0x24622e14557fd974", "0x5730af77da1186ff", "0xf776630a6ff0a022", "0xa69c82d11c376933", + "0x93fb4e2c9c12456f", "0xa7a30a3c776b32cf", "0x1669db65ca0f2724", "0x491df6828bfc7ad0", "0xeb5bf506947713c2", "0xe83fd62f415cfe9e", "0xbc4415e97b359dd0", "0x2ab09721b97d003f", + "0xaf04c23a63a7bb7f", "0x6c4986d838379e4a", "0x0dcecd1729525d44", "0x92ba240b16950bba", "0xed6df8cc108a7e50", "0x60252d3be009c7ea", "0x300d103ffc4e8d35", "0x73522acf3d2f56d4" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x254a52e3824ffcb7", "0x35f6dfb86ffa9f8e", "0x27838da0f63ce460", "0xd25ec8a499e26256", "0xfe3b4f012071b8b2", "0xde760f5a3e5378cd", "0x0810bd3a5f232ac3", "0x7ee980b9098f2df2", + "0xbae54f58f6daab8e", "0x6f89d253a6b184f1", "0xb9c87bb5f8914f6f", "0xd264d719aa5d7bd5", "0x23e89f82b32a5a44", "0x080a72f8a360f627", "0xe28b507b1cba42be", "0x25b77fdc149f8eb7", + "0x1a34f036f09596f5", "0x3e67e706d4719f7e", "0x33764e53c58097a6", "0xbcf62943f155349a", "0x8eb89841fbcb8ee5", "0x386b0433a1567ee7", "0x0c5775cb8fce2a91", "0x64518cc52f393a59", + "0x8f3cdb40e81a1cdd", "0x6af5a66a0a2436c1", "0xb76a5386454bef44", "0xfdd12f4ff2f53256", "0xf9ff7b79b577274f", "0xe7ed16e8ad89f703", "0x5a64c33fdfec335e", "0xda50f54c2c547ac1" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/mixA-2/kernel.cl b/proto-cuda/packs-readwidth/mixA-2/kernel.cl new file mode 100644 index 000000000..9db6d5af2 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-2/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/2". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x774fd410u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x520f91b2u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x520f91b2u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x9357786fu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x9357786fu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x8bbe44d7u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x8bbe44d7u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x23db18feu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x23db18feu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x52022aecu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x52022aecu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x653ea608u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x653ea608u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x57788e47u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x57788e47u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x774fd410u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r2 + r7 + ((((sel >> 7u) & 1u) != 0u) ? 0xd8952471u : 0x17c8e8eeu); // 0 add + r0 = rotr_var(r0, r5); // 1 rotr + r5 = r5 + r0 + ((((sel >> 21u) & 1u) != 0u) ? 0x4325cc6au : 0x37d9560au); // 2 add + r7 = r7 ^ r5; // 3 xor + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 4 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 5 load + r0 = mul_hi(r0, r5); // 6 mulhi + r4 = r4 * r7; // 7 mul + r0 = r5 * r6 + r0; // 8 mad + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 9 load + r0 = r0 ^ r6; // 10 xor + r5 = r5 + r3 + ((((sel >> 22u) & 1u) != 0u) ? 0x81619a4cu : 0xbc4eca12u); // 11 add + r4 = mul_hi(r4, r2); // 12 mulhi + r2 = r2 ^ ds[r0 & mask]; // 13 load + r0 = r0 + r2 + ((((sel >> 3u) & 1u) != 0u) ? 0xa557fd2bu : 0xcf9919eeu); // 14 add + r4 = r4 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x33857f70u : 0x36ec1d20u); // 15 add + r7 = r7 + r5 + ((((sel >> 6u) & 1u) != 0u) ? 0x40495714u : 0x42f5db15u); // 16 add + r7 = r7 ^ ds[r5 & mask]; // 17 load + r3 = r3 ^ r6; // 18 xor + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 19 load + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 20 load + r1 = r1 ^ ds[r0 & mask]; // 21 load + r1 = mul_hi(r1, r3); // 22 mulhi + r1 = rotr_var(r1, r5); // 23 rotr + r2 = mul_hi(r2, r0); // 24 mulhi + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 load + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 26 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r5 = r5 ^ t_; } // 27 shfl + r5 = r5 + r4 + ((((sel >> 12u) & 1u) != 0u) ? 0x58a2eb85u : 0x22029cb9u); // 28 add + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 29 load + r6 = r6 * r5; // 30 mul + r6 = r6 ^ ds[r2 & mask]; // 31 load + r7 = mul_hi(r7, r5); // 32 mulhi + r0 = r0 ^ ds[r7 & mask]; // 33 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 1u); r7 = r7 ^ t_; } // 34 shfl + r7 = r7 * r0; // 35 mul + r5 = r5 ^ ds[r7 & mask]; // 36 load + r3 = rotr_var(r3, r2); // 37 rotr + r6 = rotl_imm(r6, 1u); // 38 rotl + r3 = r3 * r0; // 39 mul + r3 = mul_hi(r3, r7); // 40 mulhi + r5 = r5 ^ r2; // 41 xor + r4 = r4 * r0; // 42 mul + r3 = r3 + r0 + ((((sel >> 4u) & 1u) != 0u) ? 0xfaaf2d1eu : 0x831bfab3u); // 43 add + r0 = r1 * r6 + r0; // 44 mad + r6 = rotl_imm(r6, 26u); // 45 rotl + r2 = r2 ^ r1; // 46 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 2u); r0 = r0 ^ t_; } // 47 shfl + r1 = r1 ^ r2; // 48 xor + r3 = r3 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0xc5759120u : 0x31df1a86u); // 49 add + r1 = r4 * r1 + r1; // 50 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r1 = r1 ^ t_; } // 51 shfl + r5 = r5 + r4 + ((((sel >> 16u) & 1u) != 0u) ? 0x4e4a2759u : 0x553b85d1u); // 52 add + r1 = r1 ^ ds[r3 & mask]; // 53 load + r6 = rotl_imm(r6, 21u); // 54 rotl + r1 = mul_hi(r1, r4); // 55 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 1u); r0 = r0 ^ t_; } // 56 shfl + r1 = r1 ^ r2; // 57 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 16u); r3 = r3 ^ t_; } // 58 shfl + r5 = r5 * r0; // 59 mul + r4 = r4 ^ ds[r2 & mask]; // 60 load + r0 = rotr_var(r0, r4); // 61 rotr + r6 = r6 * r7; // 62 mul + r6 = rotl_imm(r6, 2u); // 63 rotl + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/mixA-2/kernel.cu b/proto-cuda/packs-readwidth/mixA-2/kernel.cu new file mode 100644 index 000000000..451e4b652 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-2/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/2". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x774fd410u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x520f91b2u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x520f91b2u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x9357786fu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0x9357786fu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x8bbe44d7u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x8bbe44d7u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x23db18feu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0x23db18feu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x52022aecu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0x52022aecu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x653ea608u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x653ea608u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x57788e47u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x57788e47u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x774fd410u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r2 + r7 + ((((sel >> 7u) & 1u) != 0u) ? 0xd8952471u : 0x17c8e8eeu); // 0 add + r0 = rotr_var(r0, r5); // 1 rotr + r5 = r5 + r0 + ((((sel >> 21u) & 1u) != 0u) ? 0x4325cc6au : 0x37d9560au); // 2 add + r7 = r7 ^ r5; // 3 xor + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 4 load + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 5 load + r0 = __umulhi(r0, r5); // 6 mulhi + r4 = r4 * r7; // 7 mul + r0 = r5 * r6 + r0; // 8 mad + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 9 load + r0 = r0 ^ r6; // 10 xor + r5 = r5 + r3 + ((((sel >> 22u) & 1u) != 0u) ? 0x81619a4cu : 0xbc4eca12u); // 11 add + r4 = __umulhi(r4, r2); // 12 mulhi + r2 = r2 ^ ds[r0 & mask]; // 13 load + r0 = r0 + r2 + ((((sel >> 3u) & 1u) != 0u) ? 0xa557fd2bu : 0xcf9919eeu); // 14 add + r4 = r4 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x33857f70u : 0x36ec1d20u); // 15 add + r7 = r7 + r5 + ((((sel >> 6u) & 1u) != 0u) ? 0x40495714u : 0x42f5db15u); // 16 add + r7 = r7 ^ ds[r5 & mask]; // 17 load + r3 = r3 ^ r6; // 18 xor + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 19 load + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 20 load + r1 = r1 ^ ds[r0 & mask]; // 21 load + r1 = __umulhi(r1, r3); // 22 mulhi + r1 = rotr_var(r1, r5); // 23 rotr + r2 = __umulhi(r2, r0); // 24 mulhi + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 load + { uint32_t b_ = (r1 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 26 load + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 2); // 27 shfl + r5 = r5 + r4 + ((((sel >> 12u) & 1u) != 0u) ? 0x58a2eb85u : 0x22029cb9u); // 28 add + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 29 load + r6 = r6 * r5; // 30 mul + r6 = r6 ^ ds[r2 & mask]; // 31 load + r7 = __umulhi(r7, r5); // 32 mulhi + r0 = r0 ^ ds[r7 & mask]; // 33 load + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r2, 1); // 34 shfl + r7 = r7 * r0; // 35 mul + r5 = r5 ^ ds[r7 & mask]; // 36 load + r3 = rotr_var(r3, r2); // 37 rotr + r6 = rotl_imm(r6, 1u); // 38 rotl + r3 = r3 * r0; // 39 mul + r3 = __umulhi(r3, r7); // 40 mulhi + r5 = r5 ^ r2; // 41 xor + r4 = r4 * r0; // 42 mul + r3 = r3 + r0 + ((((sel >> 4u) & 1u) != 0u) ? 0xfaaf2d1eu : 0x831bfab3u); // 43 add + r0 = r1 * r6 + r0; // 44 mad + r6 = rotl_imm(r6, 26u); // 45 rotl + r2 = r2 ^ r1; // 46 xor + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 2); // 47 shfl + r1 = r1 ^ r2; // 48 xor + r3 = r3 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0xc5759120u : 0x31df1a86u); // 49 add + r1 = r4 * r1 + r1; // 50 mad + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 51 shfl + r5 = r5 + r4 + ((((sel >> 16u) & 1u) != 0u) ? 0x4e4a2759u : 0x553b85d1u); // 52 add + r1 = r1 ^ ds[r3 & mask]; // 53 load + r6 = rotl_imm(r6, 21u); // 54 rotl + r1 = __umulhi(r1, r4); // 55 mulhi + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r3, 1); // 56 shfl + r1 = r1 ^ r2; // 57 xor + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 16); // 58 shfl + r5 = r5 * r0; // 59 mul + r4 = r4 ^ ds[r2 & mask]; // 60 load + r0 = rotr_var(r0, r4); // 61 rotr + r6 = r6 * r7; // 62 mul + r6 = rotl_imm(r6, 2u); // 63 rotl + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixA-2/kernel_bound.cl b/proto-cuda/packs-readwidth/mixA-2/kernel_bound.cl new file mode 100644 index 000000000..a5325a547 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-2/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/2". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x774fd410u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x520f91b2u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x520f91b2u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x9357786fu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x9357786fu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x8bbe44d7u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x8bbe44d7u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x23db18feu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x23db18feu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x52022aecu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x52022aecu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x653ea608u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x653ea608u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x57788e47u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x57788e47u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x774fd410u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r2 + r7 + ((((sel >> 7u) & 1u) != 0u) ? 0xd8952471u : 0x17c8e8eeu); // 0 add + r0 = rotr_var(r0, r5); // 1 rotr + r5 = r5 + r0 + ((((sel >> 21u) & 1u) != 0u) ? 0x4325cc6au : 0x37d9560au); // 2 add + r7 = r7 ^ r5; // 3 xor + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 4 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 5 load + r0 = mul_hi(r0, r5); // 6 mulhi + r4 = r4 * r7; // 7 mul + r0 = r5 * r6 + r0; // 8 mad + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 9 load + r0 = r0 ^ r6; // 10 xor + r5 = r5 + r3 + ((((sel >> 22u) & 1u) != 0u) ? 0x81619a4cu : 0xbc4eca12u); // 11 add + r4 = mul_hi(r4, r2); // 12 mulhi + r2 = r2 ^ ds[r0 & mask]; // 13 load + r0 = r0 + r2 + ((((sel >> 3u) & 1u) != 0u) ? 0xa557fd2bu : 0xcf9919eeu); // 14 add + r4 = r4 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x33857f70u : 0x36ec1d20u); // 15 add + r7 = r7 + r5 + ((((sel >> 6u) & 1u) != 0u) ? 0x40495714u : 0x42f5db15u); // 16 add + r7 = r7 ^ ds[r5 & mask]; // 17 load + r3 = r3 ^ r6; // 18 xor + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 19 load + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 20 load + r1 = r1 ^ ds[r0 & mask]; // 21 load + r1 = mul_hi(r1, r3); // 22 mulhi + r1 = rotr_var(r1, r5); // 23 rotr + r2 = mul_hi(r2, r0); // 24 mulhi + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 load + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 26 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r5 = r5 ^ t_; } // 27 shfl + r5 = r5 + r4 + ((((sel >> 12u) & 1u) != 0u) ? 0x58a2eb85u : 0x22029cb9u); // 28 add + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 29 load + r6 = r6 * r5; // 30 mul + r6 = r6 ^ ds[r2 & mask]; // 31 load + r7 = mul_hi(r7, r5); // 32 mulhi + r0 = r0 ^ ds[r7 & mask]; // 33 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 1u); r7 = r7 ^ t_; } // 34 shfl + r7 = r7 * r0; // 35 mul + r5 = r5 ^ ds[r7 & mask]; // 36 load + r3 = rotr_var(r3, r2); // 37 rotr + r6 = rotl_imm(r6, 1u); // 38 rotl + r3 = r3 * r0; // 39 mul + r3 = mul_hi(r3, r7); // 40 mulhi + r5 = r5 ^ r2; // 41 xor + r4 = r4 * r0; // 42 mul + r3 = r3 + r0 + ((((sel >> 4u) & 1u) != 0u) ? 0xfaaf2d1eu : 0x831bfab3u); // 43 add + r0 = r1 * r6 + r0; // 44 mad + r6 = rotl_imm(r6, 26u); // 45 rotl + r2 = r2 ^ r1; // 46 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 2u); r0 = r0 ^ t_; } // 47 shfl + r1 = r1 ^ r2; // 48 xor + r3 = r3 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0xc5759120u : 0x31df1a86u); // 49 add + r1 = r4 * r1 + r1; // 50 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r1 = r1 ^ t_; } // 51 shfl + r5 = r5 + r4 + ((((sel >> 16u) & 1u) != 0u) ? 0x4e4a2759u : 0x553b85d1u); // 52 add + r1 = r1 ^ ds[r3 & mask]; // 53 load + r6 = rotl_imm(r6, 21u); // 54 rotl + r1 = mul_hi(r1, r4); // 55 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 1u); r0 = r0 ^ t_; } // 56 shfl + r1 = r1 ^ r2; // 57 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 16u); r3 = r3 ^ t_; } // 58 shfl + r5 = r5 * r0; // 59 mul + r4 = r4 ^ ds[r2 & mask]; // 60 load + r0 = rotr_var(r0, r4); // 61 rotr + r6 = r6 * r7; // 62 mul + r6 = rotl_imm(r6, 2u); // 63 rotl + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r2 + r7 + ((((sel >> 7u) & 1u) != 0u) ? 0xd8952471u : 0x17c8e8eeu); // 0 add + r0 = rotr_var(r0, r5); // 1 rotr + r5 = r5 + r0 + ((((sel >> 21u) & 1u) != 0u) ? 0x4325cc6au : 0x37d9560au); // 2 add + r7 = r7 ^ r5; // 3 xor + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 4 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 5 load + r0 = mul_hi(r0, r5); // 6 mulhi + r4 = r4 * r7; // 7 mul + r0 = r5 * r6 + r0; // 8 mad + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 9 load + r0 = r0 ^ r6; // 10 xor + r5 = r5 + r3 + ((((sel >> 22u) & 1u) != 0u) ? 0x81619a4cu : 0xbc4eca12u); // 11 add + r4 = mul_hi(r4, r2); // 12 mulhi + r2 = r2 ^ ds[r0 & mask]; // 13 load + r0 = r0 + r2 + ((((sel >> 3u) & 1u) != 0u) ? 0xa557fd2bu : 0xcf9919eeu); // 14 add + r4 = r4 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x33857f70u : 0x36ec1d20u); // 15 add + r7 = r7 + r5 + ((((sel >> 6u) & 1u) != 0u) ? 0x40495714u : 0x42f5db15u); // 16 add + r7 = r7 ^ ds[r5 & mask]; // 17 load + r3 = r3 ^ r6; // 18 xor + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 19 load + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 20 load + r1 = r1 ^ ds[r0 & mask]; // 21 load + r1 = mul_hi(r1, r3); // 22 mulhi + r1 = rotr_var(r1, r5); // 23 rotr + r2 = mul_hi(r2, r0); // 24 mulhi + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 load + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 26 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r5 = r5 ^ t_; } // 27 shfl + r5 = r5 + r4 + ((((sel >> 12u) & 1u) != 0u) ? 0x58a2eb85u : 0x22029cb9u); // 28 add + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 29 load + r6 = r6 * r5; // 30 mul + r6 = r6 ^ ds[r2 & mask]; // 31 load + r7 = mul_hi(r7, r5); // 32 mulhi + r0 = r0 ^ ds[r7 & mask]; // 33 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 1u); r7 = r7 ^ t_; } // 34 shfl + r7 = r7 * r0; // 35 mul + r5 = r5 ^ ds[r7 & mask]; // 36 load + r3 = rotr_var(r3, r2); // 37 rotr + r6 = rotl_imm(r6, 1u); // 38 rotl + r3 = r3 * r0; // 39 mul + r3 = mul_hi(r3, r7); // 40 mulhi + r5 = r5 ^ r2; // 41 xor + r4 = r4 * r0; // 42 mul + r3 = r3 + r0 + ((((sel >> 4u) & 1u) != 0u) ? 0xfaaf2d1eu : 0x831bfab3u); // 43 add + r0 = r1 * r6 + r0; // 44 mad + r6 = rotl_imm(r6, 26u); // 45 rotl + r2 = r2 ^ r1; // 46 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 2u); r0 = r0 ^ t_; } // 47 shfl + r1 = r1 ^ r2; // 48 xor + r3 = r3 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0xc5759120u : 0x31df1a86u); // 49 add + r1 = r4 * r1 + r1; // 50 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r1 = r1 ^ t_; } // 51 shfl + r5 = r5 + r4 + ((((sel >> 16u) & 1u) != 0u) ? 0x4e4a2759u : 0x553b85d1u); // 52 add + r1 = r1 ^ ds[r3 & mask]; // 53 load + r6 = rotl_imm(r6, 21u); // 54 rotl + r1 = mul_hi(r1, r4); // 55 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 1u); r0 = r0 ^ t_; } // 56 shfl + r1 = r1 ^ r2; // 57 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 16u); r3 = r3 ^ t_; } // 58 shfl + r5 = r5 * r0; // 59 mul + r4 = r4 ^ ds[r2 & mask]; // 60 load + r0 = rotr_var(r0, r4); // 61 rotr + r6 = r6 * r7; // 62 mul + r6 = rotl_imm(r6, 2u); // 63 rotl + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-2/kernel_bound.cu b/proto-cuda/packs-readwidth/mixA-2/kernel_bound.cu new file mode 100644 index 000000000..24f185a86 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-2/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/2". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r2 + r7 + ((((sel >> 7u) & 1u) != 0u) ? 0xd8952471u : 0x17c8e8eeu); // 0 add + r0 = rotr_var(r0, r5); // 1 rotr + r5 = r5 + r0 + ((((sel >> 21u) & 1u) != 0u) ? 0x4325cc6au : 0x37d9560au); // 2 add + r7 = r7 ^ r5; // 3 xor + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 4 load + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 5 load + r0 = __umulhi(r0, r5); // 6 mulhi + r4 = r4 * r7; // 7 mul + r0 = r5 * r6 + r0; // 8 mad + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 9 load + r0 = r0 ^ r6; // 10 xor + r5 = r5 + r3 + ((((sel >> 22u) & 1u) != 0u) ? 0x81619a4cu : 0xbc4eca12u); // 11 add + r4 = __umulhi(r4, r2); // 12 mulhi + r2 = r2 ^ ds[r0 & mask]; // 13 load + r0 = r0 + r2 + ((((sel >> 3u) & 1u) != 0u) ? 0xa557fd2bu : 0xcf9919eeu); // 14 add + r4 = r4 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x33857f70u : 0x36ec1d20u); // 15 add + r7 = r7 + r5 + ((((sel >> 6u) & 1u) != 0u) ? 0x40495714u : 0x42f5db15u); // 16 add + r7 = r7 ^ ds[r5 & mask]; // 17 load + r3 = r3 ^ r6; // 18 xor + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 19 load + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 20 load + r1 = r1 ^ ds[r0 & mask]; // 21 load + r1 = __umulhi(r1, r3); // 22 mulhi + r1 = rotr_var(r1, r5); // 23 rotr + r2 = __umulhi(r2, r0); // 24 mulhi + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 load + { uint32_t b_ = (r1 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 26 load + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 2); // 27 shfl + r5 = r5 + r4 + ((((sel >> 12u) & 1u) != 0u) ? 0x58a2eb85u : 0x22029cb9u); // 28 add + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 29 load + r6 = r6 * r5; // 30 mul + r6 = r6 ^ ds[r2 & mask]; // 31 load + r7 = __umulhi(r7, r5); // 32 mulhi + r0 = r0 ^ ds[r7 & mask]; // 33 load + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r2, 1); // 34 shfl + r7 = r7 * r0; // 35 mul + r5 = r5 ^ ds[r7 & mask]; // 36 load + r3 = rotr_var(r3, r2); // 37 rotr + r6 = rotl_imm(r6, 1u); // 38 rotl + r3 = r3 * r0; // 39 mul + r3 = __umulhi(r3, r7); // 40 mulhi + r5 = r5 ^ r2; // 41 xor + r4 = r4 * r0; // 42 mul + r3 = r3 + r0 + ((((sel >> 4u) & 1u) != 0u) ? 0xfaaf2d1eu : 0x831bfab3u); // 43 add + r0 = r1 * r6 + r0; // 44 mad + r6 = rotl_imm(r6, 26u); // 45 rotl + r2 = r2 ^ r1; // 46 xor + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 2); // 47 shfl + r1 = r1 ^ r2; // 48 xor + r3 = r3 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0xc5759120u : 0x31df1a86u); // 49 add + r1 = r4 * r1 + r1; // 50 mad + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 51 shfl + r5 = r5 + r4 + ((((sel >> 16u) & 1u) != 0u) ? 0x4e4a2759u : 0x553b85d1u); // 52 add + r1 = r1 ^ ds[r3 & mask]; // 53 load + r6 = rotl_imm(r6, 21u); // 54 rotl + r1 = __umulhi(r1, r4); // 55 mulhi + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r3, 1); // 56 shfl + r1 = r1 ^ r2; // 57 xor + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 16); // 58 shfl + r5 = r5 * r0; // 59 mul + r4 = r4 ^ ds[r2 & mask]; // 60 load + r0 = rotr_var(r0, r4); // 61 rotr + r6 = r6 * r7; // 62 mul + r6 = rotl_imm(r6, 2u); // 63 rotl + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixA-2/memhard.h b/proto-cuda/packs-readwidth/mixA-2/memhard.h new file mode 100644 index 000000000..3a5185b98 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-2/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/2". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/mixA-2/memhard.metal b/proto-cuda/packs-readwidth/mixA-2/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-2/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/mixA-2/program.h b/proto-cuda/packs-readwidth/mixA-2/program.h new file mode 100644 index 000000000..6ae69b935 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-2/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/2". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-readwidth/A/2" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d7265616477696474682f412f32" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x11d47bbfc868120eull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=10 mul=7 mulhi=7 xor=7 shfl=6 rotl=4 rotr=4 mad=3" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "mix50-35-15" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 50, 35, 15 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 8, 5, 3 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 2432 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x774fd410u, 0x520f91b2u, 0x9357786fu, 0x8bbe44d7u, 0x23db18feu, 0x52022aecu, 0x653ea608u, 0x57788e47u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/mixA-2/program.json b/proto-cuda/packs-readwidth/mixA-2/program.json new file mode 100644 index 000000000..7356f3b11 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-2/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x11d47bbfc868120e", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-readwidth/A/2", + "seed_bytes": "69676e65756d2d7265616477696474682f412f32", + "seed_words": ["0x774fd410", "0x520f91b2", "0x9357786f", "0x8bbe44d7", "0x23db18fe", "0x52022aec", "0x653ea608", "0x57788e47"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "mix50-35-15", + "load_slots": 16, + "load_mix_percent_4_16_64": [50, 35, 15], + "load_width_counts_4_16_64": [8, 5, 3], + "bytes_per_hash": 2432, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "add": 10, "mul": 7, "mulhi": 7, "xor": 7, "shfl": 6, "rotl": 4, "rotr": 4, "mad": 3}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "add", "dst": 2, "src": 7, "src2": 4, "imm": "0x17c8e8ee", "imm2": "0xd8952471", "rot": 31, "bit": 7, "mask": 4, "width": 1}, + {"i": 1, "op": "rotr", "dst": 0, "src": 5, "src2": 2, "imm": "0xcd7dcc5c", "imm2": "0xa1188ae7", "rot": 15, "bit": 23, "mask": 16, "width": 1}, + {"i": 2, "op": "add", "dst": 5, "src": 0, "src2": 1, "imm": "0x37d9560a", "imm2": "0x4325cc6a", "rot": 4, "bit": 21, "mask": 2, "width": 1}, + {"i": 3, "op": "xor", "dst": 7, "src": 5, "src2": 3, "imm": "0x6e598036", "imm2": "0x3370fa87", "rot": 22, "bit": 10, "mask": 4, "width": 1}, + {"i": 4, "op": "load", "dst": 3, "src": 0, "src2": 4, "imm": "0x072213d2", "imm2": "0xa024aa08", "rot": 9, "bit": 1, "mask": 16, "width": 16}, + {"i": 5, "op": "load", "dst": 5, "src": 7, "src2": 4, "imm": "0x3f74f8d4", "imm2": "0x29345708", "rot": 6, "bit": 9, "mask": 8, "width": 4}, + {"i": 6, "op": "mulhi", "dst": 0, "src": 5, "src2": 2, "imm": "0x0b8090fa", "imm2": "0x6677fdf7", "rot": 15, "bit": 24, "mask": 8, "width": 1}, + {"i": 7, "op": "mul", "dst": 4, "src": 7, "src2": 4, "imm": "0xf40f7deb", "imm2": "0x93c66180", "rot": 23, "bit": 1, "mask": 1, "width": 1}, + {"i": 8, "op": "mad", "dst": 0, "src": 5, "src2": 6, "imm": "0xfdce2834", "imm2": "0x61107ebf", "rot": 1, "bit": 10, "mask": 8, "width": 1}, + {"i": 9, "op": "load", "dst": 6, "src": 2, "src2": 6, "imm": "0x99c257d5", "imm2": "0x7b8a8224", "rot": 10, "bit": 29, "mask": 2, "width": 4}, + {"i": 10, "op": "xor", "dst": 0, "src": 6, "src2": 6, "imm": "0x1d598c2f", "imm2": "0x989541c3", "rot": 5, "bit": 7, "mask": 4, "width": 1}, + {"i": 11, "op": "add", "dst": 5, "src": 3, "src2": 3, "imm": "0xbc4eca12", "imm2": "0x81619a4c", "rot": 9, "bit": 22, "mask": 16, "width": 1}, + {"i": 12, "op": "mulhi", "dst": 4, "src": 2, "src2": 0, "imm": "0xdd65e362", "imm2": "0x3928839a", "rot": 14, "bit": 28, "mask": 4, "width": 1}, + {"i": 13, "op": "load", "dst": 2, "src": 0, "src2": 6, "imm": "0x059fcec1", "imm2": "0x1e541b07", "rot": 4, "bit": 21, "mask": 1, "width": 1}, + {"i": 14, "op": "add", "dst": 0, "src": 2, "src2": 6, "imm": "0xcf9919ee", "imm2": "0xa557fd2b", "rot": 6, "bit": 3, "mask": 1, "width": 1}, + {"i": 15, "op": "add", "dst": 4, "src": 5, "src2": 0, "imm": "0x36ec1d20", "imm2": "0x33857f70", "rot": 21, "bit": 7, "mask": 8, "width": 1}, + {"i": 16, "op": "add", "dst": 7, "src": 5, "src2": 1, "imm": "0x42f5db15", "imm2": "0x40495714", "rot": 18, "bit": 6, "mask": 8, "width": 1}, + {"i": 17, "op": "load", "dst": 7, "src": 5, "src2": 1, "imm": "0xf649fe2d", "imm2": "0x4477242b", "rot": 27, "bit": 3, "mask": 8, "width": 1}, + {"i": 18, "op": "xor", "dst": 3, "src": 6, "src2": 7, "imm": "0xae781ff8", "imm2": "0xbe6ecc6d", "rot": 30, "bit": 18, "mask": 1, "width": 1}, + {"i": 19, "op": "load", "dst": 7, "src": 2, "src2": 4, "imm": "0x68e68ea0", "imm2": "0x4dc840da", "rot": 19, "bit": 5, "mask": 1, "width": 16}, + {"i": 20, "op": "load", "dst": 5, "src": 3, "src2": 3, "imm": "0x59d3176d", "imm2": "0xdb77debd", "rot": 28, "bit": 5, "mask": 16, "width": 4}, + {"i": 21, "op": "load", "dst": 1, "src": 0, "src2": 5, "imm": "0xbdc0a385", "imm2": "0x464f9bc3", "rot": 17, "bit": 12, "mask": 8, "width": 1}, + {"i": 22, "op": "mulhi", "dst": 1, "src": 3, "src2": 2, "imm": "0xa654255b", "imm2": "0xea4103de", "rot": 27, "bit": 31, "mask": 2, "width": 1}, + {"i": 23, "op": "rotr", "dst": 1, "src": 5, "src2": 7, "imm": "0xa3991a21", "imm2": "0x4a407fc9", "rot": 22, "bit": 9, "mask": 16, "width": 1}, + {"i": 24, "op": "mulhi", "dst": 2, "src": 0, "src2": 7, "imm": "0x06fa7a88", "imm2": "0x06f12fe9", "rot": 3, "bit": 30, "mask": 16, "width": 1}, + {"i": 25, "op": "load", "dst": 0, "src": 4, "src2": 5, "imm": "0xf9075dec", "imm2": "0xe902c837", "rot": 16, "bit": 22, "mask": 2, "width": 16}, + {"i": 26, "op": "load", "dst": 0, "src": 1, "src2": 2, "imm": "0xd363edf4", "imm2": "0x8445583d", "rot": 10, "bit": 21, "mask": 1, "width": 4}, + {"i": 27, "op": "shfl", "dst": 5, "src": 2, "src2": 7, "imm": "0x8327fa2f", "imm2": "0x3a6092b0", "rot": 1, "bit": 23, "mask": 2, "width": 1}, + {"i": 28, "op": "add", "dst": 5, "src": 4, "src2": 3, "imm": "0x22029cb9", "imm2": "0x58a2eb85", "rot": 20, "bit": 12, "mask": 8, "width": 1}, + {"i": 29, "op": "load", "dst": 2, "src": 5, "src2": 0, "imm": "0x508d6c6e", "imm2": "0x51e70669", "rot": 23, "bit": 5, "mask": 4, "width": 4}, + {"i": 30, "op": "mul", "dst": 6, "src": 5, "src2": 0, "imm": "0xbbd10d76", "imm2": "0xcc8e1885", "rot": 29, "bit": 31, "mask": 1, "width": 1}, + {"i": 31, "op": "load", "dst": 6, "src": 2, "src2": 7, "imm": "0x032a9acc", "imm2": "0x0e202d4e", "rot": 9, "bit": 31, "mask": 16, "width": 1}, + {"i": 32, "op": "mulhi", "dst": 7, "src": 5, "src2": 7, "imm": "0x050c9aca", "imm2": "0x68f23dcc", "rot": 8, "bit": 8, "mask": 8, "width": 1}, + {"i": 33, "op": "load", "dst": 0, "src": 7, "src2": 6, "imm": "0x170bc6d7", "imm2": "0x8a08c2e4", "rot": 2, "bit": 15, "mask": 1, "width": 1}, + {"i": 34, "op": "shfl", "dst": 7, "src": 2, "src2": 6, "imm": "0x691b40d0", "imm2": "0x44b77bc4", "rot": 11, "bit": 16, "mask": 1, "width": 1}, + {"i": 35, "op": "mul", "dst": 7, "src": 0, "src2": 0, "imm": "0x708adb53", "imm2": "0x8101af24", "rot": 12, "bit": 19, "mask": 1, "width": 1}, + {"i": 36, "op": "load", "dst": 5, "src": 7, "src2": 3, "imm": "0x0c731099", "imm2": "0x8f29d683", "rot": 3, "bit": 9, "mask": 4, "width": 1}, + {"i": 37, "op": "rotr", "dst": 3, "src": 2, "src2": 0, "imm": "0x08ea65a2", "imm2": "0xc22ef0f2", "rot": 26, "bit": 7, "mask": 8, "width": 1}, + {"i": 38, "op": "rotl", "dst": 6, "src": 7, "src2": 6, "imm": "0xa5e2e082", "imm2": "0x3e9f9d96", "rot": 1, "bit": 22, "mask": 16, "width": 1}, + {"i": 39, "op": "mul", "dst": 3, "src": 0, "src2": 0, "imm": "0xb08ca537", "imm2": "0x98bd2f28", "rot": 15, "bit": 8, "mask": 2, "width": 1}, + {"i": 40, "op": "mulhi", "dst": 3, "src": 7, "src2": 0, "imm": "0x6b4f575a", "imm2": "0x0de8bd13", "rot": 14, "bit": 27, "mask": 2, "width": 1}, + {"i": 41, "op": "xor", "dst": 5, "src": 2, "src2": 1, "imm": "0xb651de8e", "imm2": "0x74c00b0c", "rot": 24, "bit": 21, "mask": 16, "width": 1}, + {"i": 42, "op": "mul", "dst": 4, "src": 0, "src2": 6, "imm": "0xd604a95a", "imm2": "0x93ad566c", "rot": 20, "bit": 16, "mask": 8, "width": 1}, + {"i": 43, "op": "add", "dst": 3, "src": 0, "src2": 1, "imm": "0x831bfab3", "imm2": "0xfaaf2d1e", "rot": 2, "bit": 4, "mask": 16, "width": 1}, + {"i": 44, "op": "mad", "dst": 0, "src": 1, "src2": 6, "imm": "0x40201b32", "imm2": "0xde9bd02f", "rot": 8, "bit": 6, "mask": 4, "width": 1}, + {"i": 45, "op": "rotl", "dst": 6, "src": 1, "src2": 5, "imm": "0x4bfbff1c", "imm2": "0x02d87bfc", "rot": 26, "bit": 8, "mask": 4, "width": 1}, + {"i": 46, "op": "xor", "dst": 2, "src": 1, "src2": 2, "imm": "0xc069583c", "imm2": "0x88265351", "rot": 26, "bit": 15, "mask": 2, "width": 1}, + {"i": 47, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0xffe42b00", "imm2": "0x8c0acfed", "rot": 21, "bit": 7, "mask": 2, "width": 1}, + {"i": 48, "op": "xor", "dst": 1, "src": 2, "src2": 1, "imm": "0x2f594ed1", "imm2": "0x7d51b0fb", "rot": 12, "bit": 19, "mask": 4, "width": 1}, + {"i": 49, "op": "add", "dst": 3, "src": 5, "src2": 4, "imm": "0x31df1a86", "imm2": "0xc5759120", "rot": 23, "bit": 7, "mask": 1, "width": 1}, + {"i": 50, "op": "mad", "dst": 1, "src": 4, "src2": 1, "imm": "0x1a81eb4c", "imm2": "0xb0108e0b", "rot": 27, "bit": 7, "mask": 1, "width": 1}, + {"i": 51, "op": "shfl", "dst": 1, "src": 4, "src2": 7, "imm": "0x1a207d4c", "imm2": "0x44993b80", "rot": 27, "bit": 27, "mask": 2, "width": 1}, + {"i": 52, "op": "add", "dst": 5, "src": 4, "src2": 4, "imm": "0x553b85d1", "imm2": "0x4e4a2759", "rot": 8, "bit": 16, "mask": 16, "width": 1}, + {"i": 53, "op": "load", "dst": 1, "src": 3, "src2": 4, "imm": "0x0c0b1b74", "imm2": "0x39083c8c", "rot": 4, "bit": 11, "mask": 1, "width": 1}, + {"i": 54, "op": "rotl", "dst": 6, "src": 4, "src2": 0, "imm": "0x1da26beb", "imm2": "0x0ddad11c", "rot": 21, "bit": 27, "mask": 2, "width": 1}, + {"i": 55, "op": "mulhi", "dst": 1, "src": 4, "src2": 7, "imm": "0x1fe9ea90", "imm2": "0x9457576e", "rot": 2, "bit": 2, "mask": 1, "width": 1}, + {"i": 56, "op": "shfl", "dst": 0, "src": 3, "src2": 4, "imm": "0x15776f9f", "imm2": "0xe087579f", "rot": 14, "bit": 11, "mask": 1, "width": 1}, + {"i": 57, "op": "xor", "dst": 1, "src": 2, "src2": 5, "imm": "0x1c5bb8e2", "imm2": "0x977e10b3", "rot": 18, "bit": 27, "mask": 16, "width": 1}, + {"i": 58, "op": "shfl", "dst": 3, "src": 6, "src2": 3, "imm": "0x37ecda35", "imm2": "0x21944e21", "rot": 3, "bit": 12, "mask": 16, "width": 1}, + {"i": 59, "op": "mul", "dst": 5, "src": 0, "src2": 6, "imm": "0xf3685385", "imm2": "0x86144005", "rot": 21, "bit": 30, "mask": 4, "width": 1}, + {"i": 60, "op": "load", "dst": 4, "src": 2, "src2": 0, "imm": "0xfdcc24f2", "imm2": "0x9bc020dc", "rot": 7, "bit": 18, "mask": 2, "width": 1}, + {"i": 61, "op": "rotr", "dst": 0, "src": 4, "src2": 2, "imm": "0x9a85b0f3", "imm2": "0xc04760bd", "rot": 26, "bit": 4, "mask": 16, "width": 1}, + {"i": 62, "op": "mul", "dst": 6, "src": 7, "src2": 7, "imm": "0x261783fc", "imm2": "0x7e75e2f9", "rot": 9, "bit": 0, "mask": 16, "width": 1}, + {"i": 63, "op": "rotl", "dst": 6, "src": 0, "src2": 4, "imm": "0x3569c3f0", "imm2": "0xbe2e19c5", "rot": 2, "bit": 18, "mask": 8, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/mixA-2/program.metal b/proto-cuda/packs-readwidth/mixA-2/program.metal new file mode 100644 index 000000000..229fd7e3a --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-2/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x774fd410u, 0x520f91b2u, 0x9357786fu, 0x8bbe44d7u, 0x23db18feu, 0x52022aecu, 0x653ea608u, 0x57788e47u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r2 + r7 + select(0x17c8e8eeu, 0xd8952471u, ((sel >> 7u) & 1u) != 0u); // 0 + r0 = rotr_var(r0, r5); // 1 + r5 = r5 + r0 + select(0x37d9560au, 0x4325cc6au, ((sel >> 21u) & 1u) != 0u); // 2 + r7 = r7 ^ r5; // 3 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 4 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 5 + r0 = mulhi(r0, r5); // 6 + r4 = r4 * r7; // 7 + r0 = r5 * r6 + r0; // 8 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 9 + r0 = r0 ^ r6; // 10 + r5 = r5 + r3 + select(0xbc4eca12u, 0x81619a4cu, ((sel >> 22u) & 1u) != 0u); // 11 + r4 = mulhi(r4, r2); // 12 + r2 = r2 ^ dataset[r0 & MASK]; // 13 + r0 = r0 + r2 + select(0xcf9919eeu, 0xa557fd2bu, ((sel >> 3u) & 1u) != 0u); // 14 + r4 = r4 + r5 + select(0x36ec1d20u, 0x33857f70u, ((sel >> 7u) & 1u) != 0u); // 15 + r7 = r7 + r5 + select(0x42f5db15u, 0x40495714u, ((sel >> 6u) & 1u) != 0u); // 16 + r7 = r7 ^ dataset[r5 & MASK]; // 17 + r3 = r3 ^ r6; // 18 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 19 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 20 + r1 = r1 ^ dataset[r0 & MASK]; // 21 + r1 = mulhi(r1, r3); // 22 + r1 = rotr_var(r1, r5); // 23 + r2 = mulhi(r2, r0); // 24 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 + { uint b_ = (r1 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 26 + r5 = r5 ^ simd_shuffle_xor(r2, (ushort)2); // 27 + r5 = r5 + r4 + select(0x22029cb9u, 0x58a2eb85u, ((sel >> 12u) & 1u) != 0u); // 28 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 29 + r6 = r6 * r5; // 30 + r6 = r6 ^ dataset[r2 & MASK]; // 31 + r7 = mulhi(r7, r5); // 32 + r0 = r0 ^ dataset[r7 & MASK]; // 33 + r7 = r7 ^ simd_shuffle_xor(r2, (ushort)1); // 34 + r7 = r7 * r0; // 35 + r5 = r5 ^ dataset[r7 & MASK]; // 36 + r3 = rotr_var(r3, r2); // 37 + r6 = rotl_imm(r6, 1u); // 38 + r3 = r3 * r0; // 39 + r3 = mulhi(r3, r7); // 40 + r5 = r5 ^ r2; // 41 + r4 = r4 * r0; // 42 + r3 = r3 + r0 + select(0x831bfab3u, 0xfaaf2d1eu, ((sel >> 4u) & 1u) != 0u); // 43 + r0 = r1 * r6 + r0; // 44 + r6 = rotl_imm(r6, 26u); // 45 + r2 = r2 ^ r1; // 46 + r0 = r0 ^ simd_shuffle_xor(r7, (ushort)2); // 47 + r1 = r1 ^ r2; // 48 + r3 = r3 + r5 + select(0x31df1a86u, 0xc5759120u, ((sel >> 7u) & 1u) != 0u); // 49 + r1 = r4 * r1 + r1; // 50 + r1 = r1 ^ simd_shuffle_xor(r4, (ushort)2); // 51 + r5 = r5 + r4 + select(0x553b85d1u, 0x4e4a2759u, ((sel >> 16u) & 1u) != 0u); // 52 + r1 = r1 ^ dataset[r3 & MASK]; // 53 + r6 = rotl_imm(r6, 21u); // 54 + r1 = mulhi(r1, r4); // 55 + r0 = r0 ^ simd_shuffle_xor(r3, (ushort)1); // 56 + r1 = r1 ^ r2; // 57 + r3 = r3 ^ simd_shuffle_xor(r6, (ushort)16); // 58 + r5 = r5 * r0; // 59 + r4 = r4 ^ dataset[r2 & MASK]; // 60 + r0 = rotr_var(r0, r4); // 61 + r6 = r6 * r7; // 62 + r6 = rotl_imm(r6, 2u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-2/program_bound.metal b/proto-cuda/packs-readwidth/mixA-2/program_bound.metal new file mode 100644 index 000000000..6136bf221 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-2/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x774fd410u, 0x520f91b2u, 0x9357786fu, 0x8bbe44d7u, 0x23db18feu, 0x52022aecu, 0x653ea608u, 0x57788e47u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r2 + r7 + select(0x17c8e8eeu, 0xd8952471u, ((sel >> 7u) & 1u) != 0u); // 0 + r0 = rotr_var(r0, r5); // 1 + r5 = r5 + r0 + select(0x37d9560au, 0x4325cc6au, ((sel >> 21u) & 1u) != 0u); // 2 + r7 = r7 ^ r5; // 3 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 4 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 5 + r0 = mulhi(r0, r5); // 6 + r4 = r4 * r7; // 7 + r0 = r5 * r6 + r0; // 8 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 9 + r0 = r0 ^ r6; // 10 + r5 = r5 + r3 + select(0xbc4eca12u, 0x81619a4cu, ((sel >> 22u) & 1u) != 0u); // 11 + r4 = mulhi(r4, r2); // 12 + r2 = r2 ^ dataset[r0 & MASK]; // 13 + r0 = r0 + r2 + select(0xcf9919eeu, 0xa557fd2bu, ((sel >> 3u) & 1u) != 0u); // 14 + r4 = r4 + r5 + select(0x36ec1d20u, 0x33857f70u, ((sel >> 7u) & 1u) != 0u); // 15 + r7 = r7 + r5 + select(0x42f5db15u, 0x40495714u, ((sel >> 6u) & 1u) != 0u); // 16 + r7 = r7 ^ dataset[r5 & MASK]; // 17 + r3 = r3 ^ r6; // 18 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 19 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 20 + r1 = r1 ^ dataset[r0 & MASK]; // 21 + r1 = mulhi(r1, r3); // 22 + r1 = rotr_var(r1, r5); // 23 + r2 = mulhi(r2, r0); // 24 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 + { uint b_ = (r1 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 26 + r5 = r5 ^ simd_shuffle_xor(r2, (ushort)2); // 27 + r5 = r5 + r4 + select(0x22029cb9u, 0x58a2eb85u, ((sel >> 12u) & 1u) != 0u); // 28 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 29 + r6 = r6 * r5; // 30 + r6 = r6 ^ dataset[r2 & MASK]; // 31 + r7 = mulhi(r7, r5); // 32 + r0 = r0 ^ dataset[r7 & MASK]; // 33 + r7 = r7 ^ simd_shuffle_xor(r2, (ushort)1); // 34 + r7 = r7 * r0; // 35 + r5 = r5 ^ dataset[r7 & MASK]; // 36 + r3 = rotr_var(r3, r2); // 37 + r6 = rotl_imm(r6, 1u); // 38 + r3 = r3 * r0; // 39 + r3 = mulhi(r3, r7); // 40 + r5 = r5 ^ r2; // 41 + r4 = r4 * r0; // 42 + r3 = r3 + r0 + select(0x831bfab3u, 0xfaaf2d1eu, ((sel >> 4u) & 1u) != 0u); // 43 + r0 = r1 * r6 + r0; // 44 + r6 = rotl_imm(r6, 26u); // 45 + r2 = r2 ^ r1; // 46 + r0 = r0 ^ simd_shuffle_xor(r7, (ushort)2); // 47 + r1 = r1 ^ r2; // 48 + r3 = r3 + r5 + select(0x31df1a86u, 0xc5759120u, ((sel >> 7u) & 1u) != 0u); // 49 + r1 = r4 * r1 + r1; // 50 + r1 = r1 ^ simd_shuffle_xor(r4, (ushort)2); // 51 + r5 = r5 + r4 + select(0x553b85d1u, 0x4e4a2759u, ((sel >> 16u) & 1u) != 0u); // 52 + r1 = r1 ^ dataset[r3 & MASK]; // 53 + r6 = rotl_imm(r6, 21u); // 54 + r1 = mulhi(r1, r4); // 55 + r0 = r0 ^ simd_shuffle_xor(r3, (ushort)1); // 56 + r1 = r1 ^ r2; // 57 + r3 = r3 ^ simd_shuffle_xor(r6, (ushort)16); // 58 + r5 = r5 * r0; // 59 + r4 = r4 ^ dataset[r2 & MASK]; // 60 + r0 = rotr_var(r0, r4); // 61 + r6 = r6 * r7; // 62 + r6 = rotl_imm(r6, 2u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-2/vectors.h b/proto-cuda/packs-readwidth/mixA-2/vectors.h new file mode 100644 index 000000000..c6b2875a0 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-2/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/2". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x1e5a818edb31f50eull, 0x39ab5ff405f05a3aull, 0x8ba99b65319aea26ull, 0x00e66c9a5d9edc85ull, 0x673a7e5b68954010ull, 0x7182b60aa5acefacull, 0x14242454aa3bdad1ull, 0x68c88b2b4ee2d9a4ull, + 0xd5a63f69306d1985ull, 0x7f1ed7474726b715ull, 0x10b2f2b157e45181ull, 0x267f18d74afdb191ull, 0xe6d713eba1ee4ec4ull, 0xd1112f5d2ca5a7e6ull, 0xd4407d80e22e732dull, 0xf0ddeb92365545aeull, + 0xc877c1ad94896995ull, 0x9ecf17f1b4b9be25ull, 0x1f529607bcdc4065ull, 0xe7cc5c40d4e4d57bull, 0xb29bf3beadf6f5d7ull, 0xbc674c35c773fcacull, 0x35ffa42b4c965803ull, 0xaf0c5c09e68ba314ull, + 0x6a313419bd31d9a6ull, 0xae5dd12365e96057ull, 0xd36b68108f13bf22ull, 0xb24bce1326403870ull, 0x4e79d454a075b862ull, 0x6301a9ec32ab72a7ull, 0x2be3716f8db33d01ull, 0x0259670d3010002eull + }, + { // base nonce 4096 + 0x8caaf6166eb165bbull, 0xa589ad41c360235full, 0x221d65887c34c329ull, 0xb167f8e67bddb1c8ull, 0x3368dcad623b40adull, 0x0143881c07493949ull, 0x0c0d44b06195e623ull, 0xa23ac0e6640ff95full, + 0x60e159dc557915a2ull, 0xea2b4ca8b06d5887ull, 0x5ff87ec91d1267e9ull, 0x49999f98b08dc991ull, 0x4a27b1ae7e9d2044ull, 0xad16f2167074cbceull, 0x2fe4de42eac4c102ull, 0x02b2e4945eb0f25full, + 0x00ec90afcc4c7be1ull, 0xedd7998f09f6202bull, 0xe871ff53c81ba442ull, 0xea2cbe8c86632596ull, 0x8c4ef1bbc35954f2ull, 0x2fe50dcc0129f200ull, 0x45089308785c4e5aull, 0xecdb829e3482a126ull, + 0x6092081b5ac666a3ull, 0xf20050f6cf7c36d0ull, 0x577b199c944cff9aull, 0xe567ab065574693cull, 0xd38fa9a00813f5bbull, 0x3f3b602229a4e685ull, 0x4bc63c32b3eee422ull, 0x0267ae18d258e96bull + }, + { // base nonce 1000000 + 0xbc346fa7af395ccaull, 0xf6ab289812882ea5ull, 0x1af5ae5da7af7f73ull, 0x3f1c7f8f7c1a0dd2ull, 0x01f938d5c871fedcull, 0x375a820f2322f11aull, 0x33be6d6c9d5e59e1ull, 0xc401cdcff783f93bull, + 0x5e41cc8a873b6cdeull, 0x17b7410058197847ull, 0x9ea7c3ca06afb2d1ull, 0x601ddfbee56db4c5ull, 0x69e68ca477d75cb8ull, 0x7c226802f7f403a9ull, 0xcdb9bceade734442ull, 0x0a0a9e8e2524ceebull, + 0x52a88bf50c886dbfull, 0x7a395acb2fc42cdfull, 0x060fe78ab0c1382bull, 0xeb448b87bdfd6af3ull, 0x8f9832144201dba5ull, 0x65a8bd919a4a5230ull, 0xea3ec06d53c77fa3ull, 0xab2ea3b4dcb53ae0ull, + 0xbacf785210da5784ull, 0x01ccffb87836c39bull, 0x95eff222af8774d5ull, 0xeb12cd87ca802c98ull, 0x500473c6eb87f917ull, 0xf63aebcada9c2a14ull, 0xc3416375ce92b971ull, 0xe23abae6636093f9ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/mixA-2/vectors.json b/proto-cuda/packs-readwidth/mixA-2/vectors.json new file mode 100644 index 000000000..4bf05dec3 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-2/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-readwidth/A/2", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x1e5a818edb31f50e", "0x39ab5ff405f05a3a", "0x8ba99b65319aea26", "0x00e66c9a5d9edc85", "0x673a7e5b68954010", "0x7182b60aa5acefac", "0x14242454aa3bdad1", "0x68c88b2b4ee2d9a4", + "0xd5a63f69306d1985", "0x7f1ed7474726b715", "0x10b2f2b157e45181", "0x267f18d74afdb191", "0xe6d713eba1ee4ec4", "0xd1112f5d2ca5a7e6", "0xd4407d80e22e732d", "0xf0ddeb92365545ae", + "0xc877c1ad94896995", "0x9ecf17f1b4b9be25", "0x1f529607bcdc4065", "0xe7cc5c40d4e4d57b", "0xb29bf3beadf6f5d7", "0xbc674c35c773fcac", "0x35ffa42b4c965803", "0xaf0c5c09e68ba314", + "0x6a313419bd31d9a6", "0xae5dd12365e96057", "0xd36b68108f13bf22", "0xb24bce1326403870", "0x4e79d454a075b862", "0x6301a9ec32ab72a7", "0x2be3716f8db33d01", "0x0259670d3010002e" + ]}, + {"base_nonce": 4096, "expected": [ + "0x8caaf6166eb165bb", "0xa589ad41c360235f", "0x221d65887c34c329", "0xb167f8e67bddb1c8", "0x3368dcad623b40ad", "0x0143881c07493949", "0x0c0d44b06195e623", "0xa23ac0e6640ff95f", + "0x60e159dc557915a2", "0xea2b4ca8b06d5887", "0x5ff87ec91d1267e9", "0x49999f98b08dc991", "0x4a27b1ae7e9d2044", "0xad16f2167074cbce", "0x2fe4de42eac4c102", "0x02b2e4945eb0f25f", + "0x00ec90afcc4c7be1", "0xedd7998f09f6202b", "0xe871ff53c81ba442", "0xea2cbe8c86632596", "0x8c4ef1bbc35954f2", "0x2fe50dcc0129f200", "0x45089308785c4e5a", "0xecdb829e3482a126", + "0x6092081b5ac666a3", "0xf20050f6cf7c36d0", "0x577b199c944cff9a", "0xe567ab065574693c", "0xd38fa9a00813f5bb", "0x3f3b602229a4e685", "0x4bc63c32b3eee422", "0x0267ae18d258e96b" + ]}, + {"base_nonce": 1000000, "expected": [ + "0xbc346fa7af395cca", "0xf6ab289812882ea5", "0x1af5ae5da7af7f73", "0x3f1c7f8f7c1a0dd2", "0x01f938d5c871fedc", "0x375a820f2322f11a", "0x33be6d6c9d5e59e1", "0xc401cdcff783f93b", + "0x5e41cc8a873b6cde", "0x17b7410058197847", "0x9ea7c3ca06afb2d1", "0x601ddfbee56db4c5", "0x69e68ca477d75cb8", "0x7c226802f7f403a9", "0xcdb9bceade734442", "0x0a0a9e8e2524ceeb", + "0x52a88bf50c886dbf", "0x7a395acb2fc42cdf", "0x060fe78ab0c1382b", "0xeb448b87bdfd6af3", "0x8f9832144201dba5", "0x65a8bd919a4a5230", "0xea3ec06d53c77fa3", "0xab2ea3b4dcb53ae0", + "0xbacf785210da5784", "0x01ccffb87836c39b", "0x95eff222af8774d5", "0xeb12cd87ca802c98", "0x500473c6eb87f917", "0xf63aebcada9c2a14", "0xc3416375ce92b971", "0xe23abae6636093f9" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/mixA-3/kernel.cl b/proto-cuda/packs-readwidth/mixA-3/kernel.cl new file mode 100644 index 000000000..820ef37d1 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-3/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/3". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x30b957f0u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x374e2a95u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x374e2a95u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xf416345eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xf416345eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x7af15ccbu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x7af15ccbu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xf0bcabc5u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xf0bcabc5u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xb5b36f35u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xb5b36f35u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xf99641c7u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xf99641c7u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xd0312afeu; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0xd0312afeu; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x30b957f0u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = rotr_var(r3, r4); // 0 rotr + r0 = r0 ^ r7; // 1 xor + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 2 load + r5 = rotl_imm(r5, 15u); // 3 rotl + r2 = r2 + r3 + ((((sel >> 6u) & 1u) != 0u) ? 0xd35e575cu : 0xd7a264d9u); // 4 add + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 1u); r5 = r5 ^ t_; } // 5 shfl + r1 = r2 * r5 + r1; // 6 mad + r2 = r2 | r3; // 7 or + r3 = r3 ^ r5; // 8 xor + r0 = r0 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0xd98be6edu : 0x4e543e70u); // 9 add + r5 = rotr_var(r5, r1); // 10 rotr + r3 = r3 * r5; // 11 mul + r2 = r2 ^ r7; // 12 xor + r1 = r1 + r6 + ((((sel >> 30u) & 1u) != 0u) ? 0x97a9219cu : 0xe36f3c9fu); // 13 add + r2 = r2 | r1; // 14 or + r1 = r1 | r3; // 15 or + r7 = r3 * r7 + r7; // 16 mad + r3 = r3 + r2 + ((((sel >> 13u) & 1u) != 0u) ? 0xf799b153u : 0x0c0a7b49u); // 17 add + r7 = r7 ^ ds[r5 & mask]; // 18 load + r7 = r7 ^ ds[r2 & mask]; // 19 load + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 8u); r5 = r5 ^ t_; } // 20 shfl + r4 = r4 ^ ds[r1 & mask]; // 21 load + r5 = r5 ^ ds[r7 & mask]; // 22 load + r4 = r4 ^ r2; // 23 xor + r5 = r5 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x26eb8325u : 0x011f9670u); // 24 add + r0 = r0 * r1; // 25 mul + r4 = r4 ^ r7; // 26 xor + r7 = r7 + r5 + ((((sel >> 26u) & 1u) != 0u) ? 0xf12a4057u : 0x42177757u); // 27 add + r5 = r0 * r7 + r5; // 28 mad + r3 = r3 * r6; // 29 mul + r6 = rotl_imm(r6, 24u); // 30 rotl + r4 = rotr_var(r4, r0); // 31 rotr + r6 = r6 ^ ds[r4 & mask]; // 32 load + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 1u); r0 = r0 ^ t_; } // 33 shfl + r4 = r4 * r5; // 34 mul + r2 = rotl_imm(r2, 15u); // 35 rotl + r7 = r7 + r3 + ((((sel >> 5u) & 1u) != 0u) ? 0xc1b9573fu : 0xe0ebc725u); // 36 add + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 37 load + r7 = r0 * r2 + r7; // 38 mad + r0 = r0 + r7 + ((((sel >> 16u) & 1u) != 0u) ? 0xde4cd365u : 0xd0b844deu); // 39 add + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 40 load + r0 = r0 ^ r1; // 41 xor + r1 = rotl_imm(r1, 19u); // 42 rotl + r7 = r7 ^ r1; // 43 xor + r2 = r2 * r4; // 44 mul + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 45 load + r7 = r7 + r3 + ((((sel >> 28u) & 1u) != 0u) ? 0x5d12f1a2u : 0x7633c48cu); // 46 add + r3 = r3 ^ ds[r2 & mask]; // 47 load + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 48 load + r7 = r7 * r5; // 49 mul + r3 = r7 * r0 + r3; // 50 mad + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 51 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 52 load + r4 = mul_hi(r4, r1); // 53 mulhi + r7 = r7 + r4 + ((((sel >> 24u) & 1u) != 0u) ? 0xd8ed09bau : 0xf64a6e41u); // 54 add + r5 = mul_hi(r5, r3); // 55 mulhi + r5 = r5 * r7; // 56 mul + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 57 load + r5 = mul_hi(r5, r2); // 58 mulhi + r7 = r7 * r5; // 59 mul + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 load + r7 = r1 * r1 + r7; // 61 mad + r0 = r0 ^ ds[r7 & mask]; // 62 load + r5 = r5 * r3; // 63 mul + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/mixA-3/kernel.cu b/proto-cuda/packs-readwidth/mixA-3/kernel.cu new file mode 100644 index 000000000..a2270aeb5 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-3/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/3". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x30b957f0u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x374e2a95u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x374e2a95u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xf416345eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xf416345eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x7af15ccbu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x7af15ccbu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xf0bcabc5u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xf0bcabc5u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xb5b36f35u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xb5b36f35u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xf99641c7u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0xf99641c7u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xd0312afeu; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0xd0312afeu; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x30b957f0u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r3 = rotr_var(r3, r4); // 0 rotr + r0 = r0 ^ r7; // 1 xor + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 2 load + r5 = rotl_imm(r5, 15u); // 3 rotl + r2 = r2 + r3 + ((((sel >> 6u) & 1u) != 0u) ? 0xd35e575cu : 0xd7a264d9u); // 4 add + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 1); // 5 shfl + r1 = r2 * r5 + r1; // 6 mad + r2 = r2 | r3; // 7 or + r3 = r3 ^ r5; // 8 xor + r0 = r0 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0xd98be6edu : 0x4e543e70u); // 9 add + r5 = rotr_var(r5, r1); // 10 rotr + r3 = r3 * r5; // 11 mul + r2 = r2 ^ r7; // 12 xor + r1 = r1 + r6 + ((((sel >> 30u) & 1u) != 0u) ? 0x97a9219cu : 0xe36f3c9fu); // 13 add + r2 = r2 | r1; // 14 or + r1 = r1 | r3; // 15 or + r7 = r3 * r7 + r7; // 16 mad + r3 = r3 + r2 + ((((sel >> 13u) & 1u) != 0u) ? 0xf799b153u : 0x0c0a7b49u); // 17 add + r7 = r7 ^ ds[r5 & mask]; // 18 load + r7 = r7 ^ ds[r2 & mask]; // 19 load + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r0, 8); // 20 shfl + r4 = r4 ^ ds[r1 & mask]; // 21 load + r5 = r5 ^ ds[r7 & mask]; // 22 load + r4 = r4 ^ r2; // 23 xor + r5 = r5 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x26eb8325u : 0x011f9670u); // 24 add + r0 = r0 * r1; // 25 mul + r4 = r4 ^ r7; // 26 xor + r7 = r7 + r5 + ((((sel >> 26u) & 1u) != 0u) ? 0xf12a4057u : 0x42177757u); // 27 add + r5 = r0 * r7 + r5; // 28 mad + r3 = r3 * r6; // 29 mul + r6 = rotl_imm(r6, 24u); // 30 rotl + r4 = rotr_var(r4, r0); // 31 rotr + r6 = r6 ^ ds[r4 & mask]; // 32 load + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r1, 1); // 33 shfl + r4 = r4 * r5; // 34 mul + r2 = rotl_imm(r2, 15u); // 35 rotl + r7 = r7 + r3 + ((((sel >> 5u) & 1u) != 0u) ? 0xc1b9573fu : 0xe0ebc725u); // 36 add + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 37 load + r7 = r0 * r2 + r7; // 38 mad + r0 = r0 + r7 + ((((sel >> 16u) & 1u) != 0u) ? 0xde4cd365u : 0xd0b844deu); // 39 add + { uint32_t b_ = (r6 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 40 load + r0 = r0 ^ r1; // 41 xor + r1 = rotl_imm(r1, 19u); // 42 rotl + r7 = r7 ^ r1; // 43 xor + r2 = r2 * r4; // 44 mul + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 45 load + r7 = r7 + r3 + ((((sel >> 28u) & 1u) != 0u) ? 0x5d12f1a2u : 0x7633c48cu); // 46 add + r3 = r3 ^ ds[r2 & mask]; // 47 load + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 48 load + r7 = r7 * r5; // 49 mul + r3 = r7 * r0 + r3; // 50 mad + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 51 load + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 52 load + r4 = __umulhi(r4, r1); // 53 mulhi + r7 = r7 + r4 + ((((sel >> 24u) & 1u) != 0u) ? 0xd8ed09bau : 0xf64a6e41u); // 54 add + r5 = __umulhi(r5, r3); // 55 mulhi + r5 = r5 * r7; // 56 mul + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 57 load + r5 = __umulhi(r5, r2); // 58 mulhi + r7 = r7 * r5; // 59 mul + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 load + r7 = r1 * r1 + r7; // 61 mad + r0 = r0 ^ ds[r7 & mask]; // 62 load + r5 = r5 * r3; // 63 mul + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixA-3/kernel_bound.cl b/proto-cuda/packs-readwidth/mixA-3/kernel_bound.cl new file mode 100644 index 000000000..7188b50f3 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-3/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/3". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x30b957f0u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x374e2a95u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x374e2a95u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xf416345eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xf416345eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x7af15ccbu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x7af15ccbu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xf0bcabc5u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xf0bcabc5u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xb5b36f35u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xb5b36f35u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xf99641c7u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xf99641c7u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xd0312afeu; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0xd0312afeu; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x30b957f0u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = rotr_var(r3, r4); // 0 rotr + r0 = r0 ^ r7; // 1 xor + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 2 load + r5 = rotl_imm(r5, 15u); // 3 rotl + r2 = r2 + r3 + ((((sel >> 6u) & 1u) != 0u) ? 0xd35e575cu : 0xd7a264d9u); // 4 add + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 1u); r5 = r5 ^ t_; } // 5 shfl + r1 = r2 * r5 + r1; // 6 mad + r2 = r2 | r3; // 7 or + r3 = r3 ^ r5; // 8 xor + r0 = r0 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0xd98be6edu : 0x4e543e70u); // 9 add + r5 = rotr_var(r5, r1); // 10 rotr + r3 = r3 * r5; // 11 mul + r2 = r2 ^ r7; // 12 xor + r1 = r1 + r6 + ((((sel >> 30u) & 1u) != 0u) ? 0x97a9219cu : 0xe36f3c9fu); // 13 add + r2 = r2 | r1; // 14 or + r1 = r1 | r3; // 15 or + r7 = r3 * r7 + r7; // 16 mad + r3 = r3 + r2 + ((((sel >> 13u) & 1u) != 0u) ? 0xf799b153u : 0x0c0a7b49u); // 17 add + r7 = r7 ^ ds[r5 & mask]; // 18 load + r7 = r7 ^ ds[r2 & mask]; // 19 load + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 8u); r5 = r5 ^ t_; } // 20 shfl + r4 = r4 ^ ds[r1 & mask]; // 21 load + r5 = r5 ^ ds[r7 & mask]; // 22 load + r4 = r4 ^ r2; // 23 xor + r5 = r5 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x26eb8325u : 0x011f9670u); // 24 add + r0 = r0 * r1; // 25 mul + r4 = r4 ^ r7; // 26 xor + r7 = r7 + r5 + ((((sel >> 26u) & 1u) != 0u) ? 0xf12a4057u : 0x42177757u); // 27 add + r5 = r0 * r7 + r5; // 28 mad + r3 = r3 * r6; // 29 mul + r6 = rotl_imm(r6, 24u); // 30 rotl + r4 = rotr_var(r4, r0); // 31 rotr + r6 = r6 ^ ds[r4 & mask]; // 32 load + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 1u); r0 = r0 ^ t_; } // 33 shfl + r4 = r4 * r5; // 34 mul + r2 = rotl_imm(r2, 15u); // 35 rotl + r7 = r7 + r3 + ((((sel >> 5u) & 1u) != 0u) ? 0xc1b9573fu : 0xe0ebc725u); // 36 add + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 37 load + r7 = r0 * r2 + r7; // 38 mad + r0 = r0 + r7 + ((((sel >> 16u) & 1u) != 0u) ? 0xde4cd365u : 0xd0b844deu); // 39 add + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 40 load + r0 = r0 ^ r1; // 41 xor + r1 = rotl_imm(r1, 19u); // 42 rotl + r7 = r7 ^ r1; // 43 xor + r2 = r2 * r4; // 44 mul + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 45 load + r7 = r7 + r3 + ((((sel >> 28u) & 1u) != 0u) ? 0x5d12f1a2u : 0x7633c48cu); // 46 add + r3 = r3 ^ ds[r2 & mask]; // 47 load + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 48 load + r7 = r7 * r5; // 49 mul + r3 = r7 * r0 + r3; // 50 mad + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 51 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 52 load + r4 = mul_hi(r4, r1); // 53 mulhi + r7 = r7 + r4 + ((((sel >> 24u) & 1u) != 0u) ? 0xd8ed09bau : 0xf64a6e41u); // 54 add + r5 = mul_hi(r5, r3); // 55 mulhi + r5 = r5 * r7; // 56 mul + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 57 load + r5 = mul_hi(r5, r2); // 58 mulhi + r7 = r7 * r5; // 59 mul + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 load + r7 = r1 * r1 + r7; // 61 mad + r0 = r0 ^ ds[r7 & mask]; // 62 load + r5 = r5 * r3; // 63 mul + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = rotr_var(r3, r4); // 0 rotr + r0 = r0 ^ r7; // 1 xor + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 2 load + r5 = rotl_imm(r5, 15u); // 3 rotl + r2 = r2 + r3 + ((((sel >> 6u) & 1u) != 0u) ? 0xd35e575cu : 0xd7a264d9u); // 4 add + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 1u); r5 = r5 ^ t_; } // 5 shfl + r1 = r2 * r5 + r1; // 6 mad + r2 = r2 | r3; // 7 or + r3 = r3 ^ r5; // 8 xor + r0 = r0 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0xd98be6edu : 0x4e543e70u); // 9 add + r5 = rotr_var(r5, r1); // 10 rotr + r3 = r3 * r5; // 11 mul + r2 = r2 ^ r7; // 12 xor + r1 = r1 + r6 + ((((sel >> 30u) & 1u) != 0u) ? 0x97a9219cu : 0xe36f3c9fu); // 13 add + r2 = r2 | r1; // 14 or + r1 = r1 | r3; // 15 or + r7 = r3 * r7 + r7; // 16 mad + r3 = r3 + r2 + ((((sel >> 13u) & 1u) != 0u) ? 0xf799b153u : 0x0c0a7b49u); // 17 add + r7 = r7 ^ ds[r5 & mask]; // 18 load + r7 = r7 ^ ds[r2 & mask]; // 19 load + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 8u); r5 = r5 ^ t_; } // 20 shfl + r4 = r4 ^ ds[r1 & mask]; // 21 load + r5 = r5 ^ ds[r7 & mask]; // 22 load + r4 = r4 ^ r2; // 23 xor + r5 = r5 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x26eb8325u : 0x011f9670u); // 24 add + r0 = r0 * r1; // 25 mul + r4 = r4 ^ r7; // 26 xor + r7 = r7 + r5 + ((((sel >> 26u) & 1u) != 0u) ? 0xf12a4057u : 0x42177757u); // 27 add + r5 = r0 * r7 + r5; // 28 mad + r3 = r3 * r6; // 29 mul + r6 = rotl_imm(r6, 24u); // 30 rotl + r4 = rotr_var(r4, r0); // 31 rotr + r6 = r6 ^ ds[r4 & mask]; // 32 load + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 1u); r0 = r0 ^ t_; } // 33 shfl + r4 = r4 * r5; // 34 mul + r2 = rotl_imm(r2, 15u); // 35 rotl + r7 = r7 + r3 + ((((sel >> 5u) & 1u) != 0u) ? 0xc1b9573fu : 0xe0ebc725u); // 36 add + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 37 load + r7 = r0 * r2 + r7; // 38 mad + r0 = r0 + r7 + ((((sel >> 16u) & 1u) != 0u) ? 0xde4cd365u : 0xd0b844deu); // 39 add + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 40 load + r0 = r0 ^ r1; // 41 xor + r1 = rotl_imm(r1, 19u); // 42 rotl + r7 = r7 ^ r1; // 43 xor + r2 = r2 * r4; // 44 mul + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 45 load + r7 = r7 + r3 + ((((sel >> 28u) & 1u) != 0u) ? 0x5d12f1a2u : 0x7633c48cu); // 46 add + r3 = r3 ^ ds[r2 & mask]; // 47 load + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 48 load + r7 = r7 * r5; // 49 mul + r3 = r7 * r0 + r3; // 50 mad + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 51 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 52 load + r4 = mul_hi(r4, r1); // 53 mulhi + r7 = r7 + r4 + ((((sel >> 24u) & 1u) != 0u) ? 0xd8ed09bau : 0xf64a6e41u); // 54 add + r5 = mul_hi(r5, r3); // 55 mulhi + r5 = r5 * r7; // 56 mul + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 57 load + r5 = mul_hi(r5, r2); // 58 mulhi + r7 = r7 * r5; // 59 mul + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 load + r7 = r1 * r1 + r7; // 61 mad + r0 = r0 ^ ds[r7 & mask]; // 62 load + r5 = r5 * r3; // 63 mul + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-3/kernel_bound.cu b/proto-cuda/packs-readwidth/mixA-3/kernel_bound.cu new file mode 100644 index 000000000..1a2c93e2b --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-3/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/3". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r3 = rotr_var(r3, r4); // 0 rotr + r0 = r0 ^ r7; // 1 xor + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 2 load + r5 = rotl_imm(r5, 15u); // 3 rotl + r2 = r2 + r3 + ((((sel >> 6u) & 1u) != 0u) ? 0xd35e575cu : 0xd7a264d9u); // 4 add + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 1); // 5 shfl + r1 = r2 * r5 + r1; // 6 mad + r2 = r2 | r3; // 7 or + r3 = r3 ^ r5; // 8 xor + r0 = r0 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0xd98be6edu : 0x4e543e70u); // 9 add + r5 = rotr_var(r5, r1); // 10 rotr + r3 = r3 * r5; // 11 mul + r2 = r2 ^ r7; // 12 xor + r1 = r1 + r6 + ((((sel >> 30u) & 1u) != 0u) ? 0x97a9219cu : 0xe36f3c9fu); // 13 add + r2 = r2 | r1; // 14 or + r1 = r1 | r3; // 15 or + r7 = r3 * r7 + r7; // 16 mad + r3 = r3 + r2 + ((((sel >> 13u) & 1u) != 0u) ? 0xf799b153u : 0x0c0a7b49u); // 17 add + r7 = r7 ^ ds[r5 & mask]; // 18 load + r7 = r7 ^ ds[r2 & mask]; // 19 load + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r0, 8); // 20 shfl + r4 = r4 ^ ds[r1 & mask]; // 21 load + r5 = r5 ^ ds[r7 & mask]; // 22 load + r4 = r4 ^ r2; // 23 xor + r5 = r5 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x26eb8325u : 0x011f9670u); // 24 add + r0 = r0 * r1; // 25 mul + r4 = r4 ^ r7; // 26 xor + r7 = r7 + r5 + ((((sel >> 26u) & 1u) != 0u) ? 0xf12a4057u : 0x42177757u); // 27 add + r5 = r0 * r7 + r5; // 28 mad + r3 = r3 * r6; // 29 mul + r6 = rotl_imm(r6, 24u); // 30 rotl + r4 = rotr_var(r4, r0); // 31 rotr + r6 = r6 ^ ds[r4 & mask]; // 32 load + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r1, 1); // 33 shfl + r4 = r4 * r5; // 34 mul + r2 = rotl_imm(r2, 15u); // 35 rotl + r7 = r7 + r3 + ((((sel >> 5u) & 1u) != 0u) ? 0xc1b9573fu : 0xe0ebc725u); // 36 add + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 37 load + r7 = r0 * r2 + r7; // 38 mad + r0 = r0 + r7 + ((((sel >> 16u) & 1u) != 0u) ? 0xde4cd365u : 0xd0b844deu); // 39 add + { uint32_t b_ = (r6 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 40 load + r0 = r0 ^ r1; // 41 xor + r1 = rotl_imm(r1, 19u); // 42 rotl + r7 = r7 ^ r1; // 43 xor + r2 = r2 * r4; // 44 mul + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 45 load + r7 = r7 + r3 + ((((sel >> 28u) & 1u) != 0u) ? 0x5d12f1a2u : 0x7633c48cu); // 46 add + r3 = r3 ^ ds[r2 & mask]; // 47 load + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 48 load + r7 = r7 * r5; // 49 mul + r3 = r7 * r0 + r3; // 50 mad + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 51 load + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 52 load + r4 = __umulhi(r4, r1); // 53 mulhi + r7 = r7 + r4 + ((((sel >> 24u) & 1u) != 0u) ? 0xd8ed09bau : 0xf64a6e41u); // 54 add + r5 = __umulhi(r5, r3); // 55 mulhi + r5 = r5 * r7; // 56 mul + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 57 load + r5 = __umulhi(r5, r2); // 58 mulhi + r7 = r7 * r5; // 59 mul + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 load + r7 = r1 * r1 + r7; // 61 mad + r0 = r0 ^ ds[r7 & mask]; // 62 load + r5 = r5 * r3; // 63 mul + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixA-3/memhard.h b/proto-cuda/packs-readwidth/mixA-3/memhard.h new file mode 100644 index 000000000..d2f99f378 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-3/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/3". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/mixA-3/memhard.metal b/proto-cuda/packs-readwidth/mixA-3/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-3/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/mixA-3/program.h b/proto-cuda/packs-readwidth/mixA-3/program.h new file mode 100644 index 000000000..efe52f8d5 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-3/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/3". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-readwidth/A/3" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d7265616477696474682f412f33" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x67648c3193a33f6cull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=10 mul=9 xor=7 mad=6 rotl=4 mulhi=3 or=3 rotr=3 shfl=3" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "mix50-35-15" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 50, 35, 15 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 7, 3, 6 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 3680 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x30b957f0u, 0x374e2a95u, 0xf416345eu, 0x7af15ccbu, 0xf0bcabc5u, 0xb5b36f35u, 0xf99641c7u, 0xd0312afeu } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/mixA-3/program.json b/proto-cuda/packs-readwidth/mixA-3/program.json new file mode 100644 index 000000000..e7e024c88 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-3/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x67648c3193a33f6c", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-readwidth/A/3", + "seed_bytes": "69676e65756d2d7265616477696474682f412f33", + "seed_words": ["0x30b957f0", "0x374e2a95", "0xf416345e", "0x7af15ccb", "0xf0bcabc5", "0xb5b36f35", "0xf99641c7", "0xd0312afe"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "mix50-35-15", + "load_slots": 16, + "load_mix_percent_4_16_64": [50, 35, 15], + "load_width_counts_4_16_64": [7, 3, 6], + "bytes_per_hash": 3680, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "add": 10, "mul": 9, "xor": 7, "mad": 6, "rotl": 4, "mulhi": 3, "or": 3, "rotr": 3, "shfl": 3}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "rotr", "dst": 3, "src": 4, "src2": 3, "imm": "0x33f6fe7b", "imm2": "0xd46e1ac2", "rot": 21, "bit": 9, "mask": 8, "width": 1}, + {"i": 1, "op": "xor", "dst": 0, "src": 7, "src2": 4, "imm": "0x12e86486", "imm2": "0x1e848be1", "rot": 3, "bit": 7, "mask": 4, "width": 1}, + {"i": 2, "op": "load", "dst": 1, "src": 0, "src2": 7, "imm": "0x25bf1dfe", "imm2": "0x85cff09c", "rot": 27, "bit": 24, "mask": 4, "width": 16}, + {"i": 3, "op": "rotl", "dst": 5, "src": 3, "src2": 3, "imm": "0xbfde1d4f", "imm2": "0x098e8aab", "rot": 15, "bit": 30, "mask": 4, "width": 1}, + {"i": 4, "op": "add", "dst": 2, "src": 3, "src2": 6, "imm": "0xd7a264d9", "imm2": "0xd35e575c", "rot": 7, "bit": 6, "mask": 4, "width": 1}, + {"i": 5, "op": "shfl", "dst": 5, "src": 2, "src2": 4, "imm": "0x5d5210ad", "imm2": "0xd984eb95", "rot": 27, "bit": 13, "mask": 1, "width": 1}, + {"i": 6, "op": "mad", "dst": 1, "src": 2, "src2": 5, "imm": "0xb969487c", "imm2": "0x61ea012e", "rot": 26, "bit": 25, "mask": 1, "width": 1}, + {"i": 7, "op": "or", "dst": 2, "src": 3, "src2": 4, "imm": "0xc1c5d61e", "imm2": "0x50d4e85c", "rot": 19, "bit": 25, "mask": 4, "width": 1}, + {"i": 8, "op": "xor", "dst": 3, "src": 5, "src2": 2, "imm": "0x33a2e407", "imm2": "0x4c9951f2", "rot": 2, "bit": 24, "mask": 16, "width": 1}, + {"i": 9, "op": "add", "dst": 0, "src": 3, "src2": 0, "imm": "0x4e543e70", "imm2": "0xd98be6ed", "rot": 4, "bit": 29, "mask": 1, "width": 1}, + {"i": 10, "op": "rotr", "dst": 5, "src": 1, "src2": 0, "imm": "0x76afe302", "imm2": "0x9eca97e5", "rot": 3, "bit": 20, "mask": 8, "width": 1}, + {"i": 11, "op": "mul", "dst": 3, "src": 5, "src2": 3, "imm": "0xbfb4e238", "imm2": "0x2746818f", "rot": 9, "bit": 25, "mask": 8, "width": 1}, + {"i": 12, "op": "xor", "dst": 2, "src": 7, "src2": 3, "imm": "0x29d5f28f", "imm2": "0xab4d8139", "rot": 10, "bit": 4, "mask": 1, "width": 1}, + {"i": 13, "op": "add", "dst": 1, "src": 6, "src2": 1, "imm": "0xe36f3c9f", "imm2": "0x97a9219c", "rot": 30, "bit": 30, "mask": 2, "width": 1}, + {"i": 14, "op": "or", "dst": 2, "src": 1, "src2": 0, "imm": "0x3c03e7d3", "imm2": "0xd16859dd", "rot": 15, "bit": 28, "mask": 2, "width": 1}, + {"i": 15, "op": "or", "dst": 1, "src": 3, "src2": 0, "imm": "0x38a5677d", "imm2": "0xa67fd283", "rot": 29, "bit": 15, "mask": 2, "width": 1}, + {"i": 16, "op": "mad", "dst": 7, "src": 3, "src2": 7, "imm": "0x78f5e68e", "imm2": "0x9d9644ec", "rot": 25, "bit": 1, "mask": 2, "width": 1}, + {"i": 17, "op": "add", "dst": 3, "src": 2, "src2": 0, "imm": "0x0c0a7b49", "imm2": "0xf799b153", "rot": 22, "bit": 13, "mask": 8, "width": 1}, + {"i": 18, "op": "load", "dst": 7, "src": 5, "src2": 6, "imm": "0xf2026981", "imm2": "0x15d47d37", "rot": 28, "bit": 27, "mask": 2, "width": 1}, + {"i": 19, "op": "load", "dst": 7, "src": 2, "src2": 4, "imm": "0x9d237b38", "imm2": "0x527aa464", "rot": 5, "bit": 22, "mask": 2, "width": 1}, + {"i": 20, "op": "shfl", "dst": 5, "src": 0, "src2": 0, "imm": "0x102562b7", "imm2": "0xe66dad9b", "rot": 9, "bit": 22, "mask": 8, "width": 1}, + {"i": 21, "op": "load", "dst": 4, "src": 1, "src2": 5, "imm": "0xe8365586", "imm2": "0x6f13fc63", "rot": 26, "bit": 6, "mask": 8, "width": 1}, + {"i": 22, "op": "load", "dst": 5, "src": 7, "src2": 2, "imm": "0xacd9154f", "imm2": "0x785e81e4", "rot": 22, "bit": 29, "mask": 1, "width": 1}, + {"i": 23, "op": "xor", "dst": 4, "src": 2, "src2": 6, "imm": "0x03eb8f2d", "imm2": "0x2f686222", "rot": 30, "bit": 30, "mask": 2, "width": 1}, + {"i": 24, "op": "add", "dst": 5, "src": 3, "src2": 3, "imm": "0x011f9670", "imm2": "0x26eb8325", "rot": 15, "bit": 14, "mask": 8, "width": 1}, + {"i": 25, "op": "mul", "dst": 0, "src": 1, "src2": 0, "imm": "0x0cacc6a3", "imm2": "0x28bca959", "rot": 8, "bit": 19, "mask": 2, "width": 1}, + {"i": 26, "op": "xor", "dst": 4, "src": 7, "src2": 7, "imm": "0x60fc8d5f", "imm2": "0xb8453fd3", "rot": 22, "bit": 3, "mask": 8, "width": 1}, + {"i": 27, "op": "add", "dst": 7, "src": 5, "src2": 6, "imm": "0x42177757", "imm2": "0xf12a4057", "rot": 15, "bit": 26, "mask": 16, "width": 1}, + {"i": 28, "op": "mad", "dst": 5, "src": 0, "src2": 7, "imm": "0x0f21b242", "imm2": "0x2fa77c71", "rot": 30, "bit": 20, "mask": 1, "width": 1}, + {"i": 29, "op": "mul", "dst": 3, "src": 6, "src2": 0, "imm": "0x69e62f8d", "imm2": "0x0037ce73", "rot": 18, "bit": 12, "mask": 2, "width": 1}, + {"i": 30, "op": "rotl", "dst": 6, "src": 7, "src2": 6, "imm": "0xb259e02f", "imm2": "0x0a33cfdf", "rot": 24, "bit": 12, "mask": 4, "width": 1}, + {"i": 31, "op": "rotr", "dst": 4, "src": 0, "src2": 6, "imm": "0xa8839588", "imm2": "0xd5af17af", "rot": 18, "bit": 19, "mask": 1, "width": 1}, + {"i": 32, "op": "load", "dst": 6, "src": 4, "src2": 4, "imm": "0xd7e19199", "imm2": "0xf45eb79a", "rot": 24, "bit": 11, "mask": 2, "width": 1}, + {"i": 33, "op": "shfl", "dst": 0, "src": 1, "src2": 5, "imm": "0x3e9f25b5", "imm2": "0x405b0189", "rot": 9, "bit": 30, "mask": 1, "width": 1}, + {"i": 34, "op": "mul", "dst": 4, "src": 5, "src2": 2, "imm": "0x42ed1682", "imm2": "0xdd577015", "rot": 27, "bit": 2, "mask": 8, "width": 1}, + {"i": 35, "op": "rotl", "dst": 2, "src": 4, "src2": 0, "imm": "0x39954523", "imm2": "0x9d5b1079", "rot": 15, "bit": 13, "mask": 16, "width": 1}, + {"i": 36, "op": "add", "dst": 7, "src": 3, "src2": 0, "imm": "0xe0ebc725", "imm2": "0xc1b9573f", "rot": 14, "bit": 5, "mask": 2, "width": 1}, + {"i": 37, "op": "load", "dst": 0, "src": 2, "src2": 5, "imm": "0x921d347e", "imm2": "0x18b83d03", "rot": 27, "bit": 7, "mask": 4, "width": 4}, + {"i": 38, "op": "mad", "dst": 7, "src": 0, "src2": 2, "imm": "0x211c8e0d", "imm2": "0x963c2eec", "rot": 6, "bit": 23, "mask": 16, "width": 1}, + {"i": 39, "op": "add", "dst": 0, "src": 7, "src2": 2, "imm": "0xd0b844de", "imm2": "0xde4cd365", "rot": 27, "bit": 16, "mask": 4, "width": 1}, + {"i": 40, "op": "load", "dst": 7, "src": 6, "src2": 6, "imm": "0x3b8d9718", "imm2": "0x8ab1654a", "rot": 7, "bit": 16, "mask": 16, "width": 16}, + {"i": 41, "op": "xor", "dst": 0, "src": 1, "src2": 2, "imm": "0x014f5a59", "imm2": "0xc8398216", "rot": 2, "bit": 17, "mask": 2, "width": 1}, + {"i": 42, "op": "rotl", "dst": 1, "src": 5, "src2": 1, "imm": "0x3bee3d81", "imm2": "0x2ce34bcd", "rot": 19, "bit": 6, "mask": 1, "width": 1}, + {"i": 43, "op": "xor", "dst": 7, "src": 1, "src2": 1, "imm": "0x27aaf2ee", "imm2": "0xec39fef7", "rot": 9, "bit": 18, "mask": 8, "width": 1}, + {"i": 44, "op": "mul", "dst": 2, "src": 4, "src2": 3, "imm": "0x1b770975", "imm2": "0xa556f55d", "rot": 4, "bit": 23, "mask": 1, "width": 1}, + {"i": 45, "op": "load", "dst": 6, "src": 4, "src2": 5, "imm": "0xe6ebfbe4", "imm2": "0x0a9cc201", "rot": 2, "bit": 5, "mask": 4, "width": 16}, + {"i": 46, "op": "add", "dst": 7, "src": 3, "src2": 7, "imm": "0x7633c48c", "imm2": "0x5d12f1a2", "rot": 10, "bit": 28, "mask": 1, "width": 1}, + {"i": 47, "op": "load", "dst": 3, "src": 2, "src2": 5, "imm": "0xe92f250d", "imm2": "0x626fb82e", "rot": 27, "bit": 8, "mask": 1, "width": 1}, + {"i": 48, "op": "load", "dst": 2, "src": 5, "src2": 5, "imm": "0x32ec784c", "imm2": "0x5c40c45a", "rot": 23, "bit": 0, "mask": 4, "width": 16}, + {"i": 49, "op": "mul", "dst": 7, "src": 5, "src2": 0, "imm": "0x68a298b5", "imm2": "0xded99974", "rot": 31, "bit": 22, "mask": 8, "width": 1}, + {"i": 50, "op": "mad", "dst": 3, "src": 7, "src2": 0, "imm": "0xf89a028b", "imm2": "0x8ece86f6", "rot": 24, "bit": 20, "mask": 2, "width": 1}, + {"i": 51, "op": "load", "dst": 3, "src": 2, "src2": 7, "imm": "0xe1a0d923", "imm2": "0xdb8faa0b", "rot": 15, "bit": 23, "mask": 8, "width": 4}, + {"i": 52, "op": "load", "dst": 4, "src": 7, "src2": 6, "imm": "0x28e50e93", "imm2": "0x965bcc9c", "rot": 25, "bit": 28, "mask": 4, "width": 4}, + {"i": 53, "op": "mulhi", "dst": 4, "src": 1, "src2": 2, "imm": "0x0065ecaf", "imm2": "0xe870289e", "rot": 8, "bit": 18, "mask": 4, "width": 1}, + {"i": 54, "op": "add", "dst": 7, "src": 4, "src2": 3, "imm": "0xf64a6e41", "imm2": "0xd8ed09ba", "rot": 14, "bit": 24, "mask": 16, "width": 1}, + {"i": 55, "op": "mulhi", "dst": 5, "src": 3, "src2": 1, "imm": "0xd12951e6", "imm2": "0xe3ae79fe", "rot": 2, "bit": 0, "mask": 2, "width": 1}, + {"i": 56, "op": "mul", "dst": 5, "src": 7, "src2": 4, "imm": "0xac8f65d4", "imm2": "0x11401bd1", "rot": 22, "bit": 6, "mask": 2, "width": 1}, + {"i": 57, "op": "load", "dst": 3, "src": 0, "src2": 1, "imm": "0x2815f83f", "imm2": "0x34b3e25b", "rot": 7, "bit": 27, "mask": 2, "width": 16}, + {"i": 58, "op": "mulhi", "dst": 5, "src": 2, "src2": 7, "imm": "0xbc696afd", "imm2": "0x488bfc7d", "rot": 24, "bit": 31, "mask": 4, "width": 1}, + {"i": 59, "op": "mul", "dst": 7, "src": 5, "src2": 7, "imm": "0x0f1bd7bd", "imm2": "0xf3a4c1c8", "rot": 25, "bit": 10, "mask": 4, "width": 1}, + {"i": 60, "op": "load", "dst": 6, "src": 3, "src2": 1, "imm": "0x2ebad44e", "imm2": "0x6b6fcf95", "rot": 4, "bit": 16, "mask": 8, "width": 16}, + {"i": 61, "op": "mad", "dst": 7, "src": 1, "src2": 1, "imm": "0xd9791f15", "imm2": "0xdd0cbcb2", "rot": 8, "bit": 14, "mask": 16, "width": 1}, + {"i": 62, "op": "load", "dst": 0, "src": 7, "src2": 1, "imm": "0x057ced89", "imm2": "0x08fdfbc1", "rot": 20, "bit": 15, "mask": 16, "width": 1}, + {"i": 63, "op": "mul", "dst": 5, "src": 3, "src2": 4, "imm": "0xbd9f5fdd", "imm2": "0xac802ec9", "rot": 27, "bit": 11, "mask": 2, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/mixA-3/program.metal b/proto-cuda/packs-readwidth/mixA-3/program.metal new file mode 100644 index 000000000..399061383 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-3/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x30b957f0u, 0x374e2a95u, 0xf416345eu, 0x7af15ccbu, 0xf0bcabc5u, 0xb5b36f35u, 0xf99641c7u, 0xd0312afeu }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = rotr_var(r3, r4); // 0 + r0 = r0 ^ r7; // 1 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 2 + r5 = rotl_imm(r5, 15u); // 3 + r2 = r2 + r3 + select(0xd7a264d9u, 0xd35e575cu, ((sel >> 6u) & 1u) != 0u); // 4 + r5 = r5 ^ simd_shuffle_xor(r2, (ushort)1); // 5 + r1 = r2 * r5 + r1; // 6 + r2 = r2 | r3; // 7 + r3 = r3 ^ r5; // 8 + r0 = r0 + r3 + select(0x4e543e70u, 0xd98be6edu, ((sel >> 29u) & 1u) != 0u); // 9 + r5 = rotr_var(r5, r1); // 10 + r3 = r3 * r5; // 11 + r2 = r2 ^ r7; // 12 + r1 = r1 + r6 + select(0xe36f3c9fu, 0x97a9219cu, ((sel >> 30u) & 1u) != 0u); // 13 + r2 = r2 | r1; // 14 + r1 = r1 | r3; // 15 + r7 = r3 * r7 + r7; // 16 + r3 = r3 + r2 + select(0x0c0a7b49u, 0xf799b153u, ((sel >> 13u) & 1u) != 0u); // 17 + r7 = r7 ^ dataset[r5 & MASK]; // 18 + r7 = r7 ^ dataset[r2 & MASK]; // 19 + r5 = r5 ^ simd_shuffle_xor(r0, (ushort)8); // 20 + r4 = r4 ^ dataset[r1 & MASK]; // 21 + r5 = r5 ^ dataset[r7 & MASK]; // 22 + r4 = r4 ^ r2; // 23 + r5 = r5 + r3 + select(0x011f9670u, 0x26eb8325u, ((sel >> 14u) & 1u) != 0u); // 24 + r0 = r0 * r1; // 25 + r4 = r4 ^ r7; // 26 + r7 = r7 + r5 + select(0x42177757u, 0xf12a4057u, ((sel >> 26u) & 1u) != 0u); // 27 + r5 = r0 * r7 + r5; // 28 + r3 = r3 * r6; // 29 + r6 = rotl_imm(r6, 24u); // 30 + r4 = rotr_var(r4, r0); // 31 + r6 = r6 ^ dataset[r4 & MASK]; // 32 + r0 = r0 ^ simd_shuffle_xor(r1, (ushort)1); // 33 + r4 = r4 * r5; // 34 + r2 = rotl_imm(r2, 15u); // 35 + r7 = r7 + r3 + select(0xe0ebc725u, 0xc1b9573fu, ((sel >> 5u) & 1u) != 0u); // 36 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 37 + r7 = r0 * r2 + r7; // 38 + r0 = r0 + r7 + select(0xd0b844deu, 0xde4cd365u, ((sel >> 16u) & 1u) != 0u); // 39 + { uint b_ = (r6 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 40 + r0 = r0 ^ r1; // 41 + r1 = rotl_imm(r1, 19u); // 42 + r7 = r7 ^ r1; // 43 + r2 = r2 * r4; // 44 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 45 + r7 = r7 + r3 + select(0x7633c48cu, 0x5d12f1a2u, ((sel >> 28u) & 1u) != 0u); // 46 + r3 = r3 ^ dataset[r2 & MASK]; // 47 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 48 + r7 = r7 * r5; // 49 + r3 = r7 * r0 + r3; // 50 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 51 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 52 + r4 = mulhi(r4, r1); // 53 + r7 = r7 + r4 + select(0xf64a6e41u, 0xd8ed09bau, ((sel >> 24u) & 1u) != 0u); // 54 + r5 = mulhi(r5, r3); // 55 + r5 = r5 * r7; // 56 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 57 + r5 = mulhi(r5, r2); // 58 + r7 = r7 * r5; // 59 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 + r7 = r1 * r1 + r7; // 61 + r0 = r0 ^ dataset[r7 & MASK]; // 62 + r5 = r5 * r3; // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-3/program_bound.metal b/proto-cuda/packs-readwidth/mixA-3/program_bound.metal new file mode 100644 index 000000000..58c1f01dc --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-3/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x30b957f0u, 0x374e2a95u, 0xf416345eu, 0x7af15ccbu, 0xf0bcabc5u, 0xb5b36f35u, 0xf99641c7u, 0xd0312afeu }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = rotr_var(r3, r4); // 0 + r0 = r0 ^ r7; // 1 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 2 + r5 = rotl_imm(r5, 15u); // 3 + r2 = r2 + r3 + select(0xd7a264d9u, 0xd35e575cu, ((sel >> 6u) & 1u) != 0u); // 4 + r5 = r5 ^ simd_shuffle_xor(r2, (ushort)1); // 5 + r1 = r2 * r5 + r1; // 6 + r2 = r2 | r3; // 7 + r3 = r3 ^ r5; // 8 + r0 = r0 + r3 + select(0x4e543e70u, 0xd98be6edu, ((sel >> 29u) & 1u) != 0u); // 9 + r5 = rotr_var(r5, r1); // 10 + r3 = r3 * r5; // 11 + r2 = r2 ^ r7; // 12 + r1 = r1 + r6 + select(0xe36f3c9fu, 0x97a9219cu, ((sel >> 30u) & 1u) != 0u); // 13 + r2 = r2 | r1; // 14 + r1 = r1 | r3; // 15 + r7 = r3 * r7 + r7; // 16 + r3 = r3 + r2 + select(0x0c0a7b49u, 0xf799b153u, ((sel >> 13u) & 1u) != 0u); // 17 + r7 = r7 ^ dataset[r5 & MASK]; // 18 + r7 = r7 ^ dataset[r2 & MASK]; // 19 + r5 = r5 ^ simd_shuffle_xor(r0, (ushort)8); // 20 + r4 = r4 ^ dataset[r1 & MASK]; // 21 + r5 = r5 ^ dataset[r7 & MASK]; // 22 + r4 = r4 ^ r2; // 23 + r5 = r5 + r3 + select(0x011f9670u, 0x26eb8325u, ((sel >> 14u) & 1u) != 0u); // 24 + r0 = r0 * r1; // 25 + r4 = r4 ^ r7; // 26 + r7 = r7 + r5 + select(0x42177757u, 0xf12a4057u, ((sel >> 26u) & 1u) != 0u); // 27 + r5 = r0 * r7 + r5; // 28 + r3 = r3 * r6; // 29 + r6 = rotl_imm(r6, 24u); // 30 + r4 = rotr_var(r4, r0); // 31 + r6 = r6 ^ dataset[r4 & MASK]; // 32 + r0 = r0 ^ simd_shuffle_xor(r1, (ushort)1); // 33 + r4 = r4 * r5; // 34 + r2 = rotl_imm(r2, 15u); // 35 + r7 = r7 + r3 + select(0xe0ebc725u, 0xc1b9573fu, ((sel >> 5u) & 1u) != 0u); // 36 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 37 + r7 = r0 * r2 + r7; // 38 + r0 = r0 + r7 + select(0xd0b844deu, 0xde4cd365u, ((sel >> 16u) & 1u) != 0u); // 39 + { uint b_ = (r6 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 40 + r0 = r0 ^ r1; // 41 + r1 = rotl_imm(r1, 19u); // 42 + r7 = r7 ^ r1; // 43 + r2 = r2 * r4; // 44 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 45 + r7 = r7 + r3 + select(0x7633c48cu, 0x5d12f1a2u, ((sel >> 28u) & 1u) != 0u); // 46 + r3 = r3 ^ dataset[r2 & MASK]; // 47 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 48 + r7 = r7 * r5; // 49 + r3 = r7 * r0 + r3; // 50 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 51 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 52 + r4 = mulhi(r4, r1); // 53 + r7 = r7 + r4 + select(0xf64a6e41u, 0xd8ed09bau, ((sel >> 24u) & 1u) != 0u); // 54 + r5 = mulhi(r5, r3); // 55 + r5 = r5 * r7; // 56 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 57 + r5 = mulhi(r5, r2); // 58 + r7 = r7 * r5; // 59 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 + r7 = r1 * r1 + r7; // 61 + r0 = r0 ^ dataset[r7 & MASK]; // 62 + r5 = r5 * r3; // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-3/vectors.h b/proto-cuda/packs-readwidth/mixA-3/vectors.h new file mode 100644 index 000000000..09162d6d3 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-3/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/3". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x828587efbc9cfc52ull, 0xd16fc546d5ec261cull, 0xb8af925838b9e1acull, 0x1fc15054e4659c04ull, 0x93ebf4629d0269bbull, 0x52d88b5a306a4737ull, 0x0831861028fa46deull, 0xa4d101fceb303faaull, + 0x96e9ddfb6a3af702ull, 0x2684793f4309a5baull, 0x12af0be41555dc8cull, 0xa5a7e5ff81a320abull, 0xa3cdfe85fb3696bfull, 0x2771918e4de403a4ull, 0xd46c364e94b850faull, 0x41f55fd8b9e13547ull, + 0x31922377a1f64cf8ull, 0x13b1dedfb55abdb7ull, 0xd4d88bf9b5ad332dull, 0xda146417bd74c676ull, 0x360aed886b80ce04ull, 0x82340715236f4d67ull, 0x830d5072d63ffc31ull, 0xd08271eaa0cab2dcull, + 0x9abe271aec0bd6cfull, 0xc12cb8813051e021ull, 0xc0e2a5765296703aull, 0x40afd7e8bcd7d766ull, 0x8d086fb3a523d08full, 0x2693f50209f3451aull, 0x7efb655d14f1ebeeull, 0x99ee7b1b6c7d4bc1ull + }, + { // base nonce 4096 + 0xbb4af8028dd1238full, 0xb1d96223fb465f3aull, 0xf6177d3717c5902dull, 0x7eb185cfde7ef8cdull, 0xed4d5f3af5e2455cull, 0x0b27846cbbb73cd6ull, 0x237cfab99f0d5ac0ull, 0xbb1e762688a5fa1full, + 0xbd0462583cfd0143ull, 0xaa071a547999c287ull, 0x42f373a7cc6b30d0ull, 0xab2f623965cd83c7ull, 0x0ffe1417daac3f15ull, 0x3517ce71179553f4ull, 0xd6a3417bf3cd74e1ull, 0xcb3d43858bc48ff3ull, + 0xaee5c7ff929e9c42ull, 0x1306f6d212e43830ull, 0x14c859a79eb2dad0ull, 0x8cfab5f11caec5e2ull, 0x2de1093748836118ull, 0xc1fa999644710490ull, 0x7c1afa846083ca13ull, 0x14a3ace55059a23dull, + 0x0a5624ae684434deull, 0xbad4e3a866d28c00ull, 0x0708afb27b44085aull, 0xea27cfd72bcf1862ull, 0x5c9538461c82d517ull, 0x95e0033bf5d07d57ull, 0x49b334ff88786acfull, 0xade49578b94c594bull + }, + { // base nonce 1000000 + 0x448bdf6ce862af09ull, 0xb137b3a92e26de3bull, 0xf57d56bf9cdb1ba1ull, 0x60a48fcbd0a34564ull, 0xd4578282342e482aull, 0x09d7d7bd7e7166e1ull, 0xa13ea5faa65a969bull, 0x1c8f45874270eb1bull, + 0x6c6b1c5799bceaf4ull, 0x11505938be6c01c3ull, 0x46d70f2787f3deaaull, 0xe7546fa4fdf67d02ull, 0x27780d83c95eb515ull, 0x5206a7adfba9f15dull, 0x485504ae37c128e0ull, 0xd442920b668e97dbull, + 0x56fdc1e95fb2d170ull, 0x09e38ff904a5b56aull, 0xb2ae26db1222f910ull, 0xa19864d4801daf4dull, 0x06a933927c9d1911ull, 0x3dab78408de69705ull, 0x17f568a8f4bf8d07ull, 0x3bc2d52c807a0290ull, + 0x58626fdb2dac1a6eull, 0x5bf21bfea7756f58ull, 0xd40dc38290489e78ull, 0xa0851f74c4783040ull, 0x9f521fa799d21d1eull, 0x363b3f454f9feef6ull, 0xb383f1a11e51190full, 0x4c83f6eb054d58a8ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/mixA-3/vectors.json b/proto-cuda/packs-readwidth/mixA-3/vectors.json new file mode 100644 index 000000000..d9b7365b7 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-3/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-readwidth/A/3", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x828587efbc9cfc52", "0xd16fc546d5ec261c", "0xb8af925838b9e1ac", "0x1fc15054e4659c04", "0x93ebf4629d0269bb", "0x52d88b5a306a4737", "0x0831861028fa46de", "0xa4d101fceb303faa", + "0x96e9ddfb6a3af702", "0x2684793f4309a5ba", "0x12af0be41555dc8c", "0xa5a7e5ff81a320ab", "0xa3cdfe85fb3696bf", "0x2771918e4de403a4", "0xd46c364e94b850fa", "0x41f55fd8b9e13547", + "0x31922377a1f64cf8", "0x13b1dedfb55abdb7", "0xd4d88bf9b5ad332d", "0xda146417bd74c676", "0x360aed886b80ce04", "0x82340715236f4d67", "0x830d5072d63ffc31", "0xd08271eaa0cab2dc", + "0x9abe271aec0bd6cf", "0xc12cb8813051e021", "0xc0e2a5765296703a", "0x40afd7e8bcd7d766", "0x8d086fb3a523d08f", "0x2693f50209f3451a", "0x7efb655d14f1ebee", "0x99ee7b1b6c7d4bc1" + ]}, + {"base_nonce": 4096, "expected": [ + "0xbb4af8028dd1238f", "0xb1d96223fb465f3a", "0xf6177d3717c5902d", "0x7eb185cfde7ef8cd", "0xed4d5f3af5e2455c", "0x0b27846cbbb73cd6", "0x237cfab99f0d5ac0", "0xbb1e762688a5fa1f", + "0xbd0462583cfd0143", "0xaa071a547999c287", "0x42f373a7cc6b30d0", "0xab2f623965cd83c7", "0x0ffe1417daac3f15", "0x3517ce71179553f4", "0xd6a3417bf3cd74e1", "0xcb3d43858bc48ff3", + "0xaee5c7ff929e9c42", "0x1306f6d212e43830", "0x14c859a79eb2dad0", "0x8cfab5f11caec5e2", "0x2de1093748836118", "0xc1fa999644710490", "0x7c1afa846083ca13", "0x14a3ace55059a23d", + "0x0a5624ae684434de", "0xbad4e3a866d28c00", "0x0708afb27b44085a", "0xea27cfd72bcf1862", "0x5c9538461c82d517", "0x95e0033bf5d07d57", "0x49b334ff88786acf", "0xade49578b94c594b" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x448bdf6ce862af09", "0xb137b3a92e26de3b", "0xf57d56bf9cdb1ba1", "0x60a48fcbd0a34564", "0xd4578282342e482a", "0x09d7d7bd7e7166e1", "0xa13ea5faa65a969b", "0x1c8f45874270eb1b", + "0x6c6b1c5799bceaf4", "0x11505938be6c01c3", "0x46d70f2787f3deaa", "0xe7546fa4fdf67d02", "0x27780d83c95eb515", "0x5206a7adfba9f15d", "0x485504ae37c128e0", "0xd442920b668e97db", + "0x56fdc1e95fb2d170", "0x09e38ff904a5b56a", "0xb2ae26db1222f910", "0xa19864d4801daf4d", "0x06a933927c9d1911", "0x3dab78408de69705", "0x17f568a8f4bf8d07", "0x3bc2d52c807a0290", + "0x58626fdb2dac1a6e", "0x5bf21bfea7756f58", "0xd40dc38290489e78", "0xa0851f74c4783040", "0x9f521fa799d21d1e", "0x363b3f454f9feef6", "0xb383f1a11e51190f", "0x4c83f6eb054d58a8" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/mixA-4/kernel.cl b/proto-cuda/packs-readwidth/mixA-4/kernel.cl new file mode 100644 index 000000000..b903aedca --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-4/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/5". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0xc255b2bfu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0xdba9d396u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0xdba9d396u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x6ea527e4u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x6ea527e4u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x05f1866au; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x05f1866au; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x2947e02eu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x2947e02eu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x0ecc5c6fu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x0ecc5c6fu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xc3b7068du; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xc3b7068du; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x282d1c2eu; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x282d1c2eu; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0xc255b2bfu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = rotl_imm(r3, 26u); // 0 rotl + r2 = rotr_var(r2, r0); // 1 rotr + r0 = r0 ^ r7; // 2 xor + r4 = r4 ^ r7; // 3 xor + r2 = r2 ^ ds[r4 & mask]; // 4 load + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 6 shfl + r6 = rotl_imm(r6, 31u); // 7 rotl + r1 = r4 * r2 + r1; // 8 mad + r1 = r7 * r0 + r1; // 9 mad + r4 = r4 ^ ds[r1 & mask]; // 10 load + r6 = r7 * r0 + r6; // 11 mad + r2 = r2 * r6; // 12 mul + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 13 load + r2 = r2 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x4a3db5a5u : 0x5df3957du); // 14 add + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 1u); r5 = r5 ^ t_; } // 15 shfl + r3 = r5 * r6 + r3; // 16 mad + r2 = r2 * r5; // 17 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r1 = r1 ^ t_; } // 18 shfl + r0 = r0 - r1; // 19 sub + r4 = r6 * r4 + r4; // 20 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 16u); r3 = r3 ^ t_; } // 21 shfl + r1 = rotl_imm(r1, 24u); // 22 rotl + r1 = r1 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0xfc07c54cu : 0x5c141117u); // 23 add + r2 = r2 ^ ds[r0 & mask]; // 24 load + r0 = r0 ^ r1; // 25 xor + r3 = r3 ^ r7; // 26 xor + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 load + r0 = r0 ^ r6; // 28 xor + r4 = r4 ^ r1; // 29 xor + r6 = r6 * r3; // 30 mul + r3 = rotl_imm(r3, 23u); // 31 rotl + r7 = rotr_var(r7, r1); // 32 rotr + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r6 = r6 ^ t_; } // 33 shfl + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 load + r6 = r6 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0x0b74657bu : 0xfb55c58du); // 35 add + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 36 load + r1 = r1 ^ ds[r4 & mask]; // 37 load + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 38 load + r4 = r4 ^ ds[r7 & mask]; // 39 load + r7 = mul_hi(r7, r3); // 40 mulhi + r0 = r0 | r6; // 41 or + r0 = r3 * r0 + r0; // 42 mad + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 43 load + r2 = r4 * r0 + r2; // 44 mad + r2 = r4 * r2 + r2; // 45 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r7 = r7 ^ t_; } // 46 shfl + r6 = r1 * r4 + r6; // 47 mad + r7 = r7 ^ ds[r1 & mask]; // 48 load + r6 = mul_hi(r6, r5); // 49 mulhi + r5 = mul_hi(r5, r0); // 50 mulhi + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 51 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 1u); r7 = r7 ^ t_; } // 52 shfl + r2 = r2 ^ r5; // 53 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r0 = r0 ^ t_; } // 54 shfl + r2 = r2 * r5; // 55 mul + r5 = r5 * r6; // 56 mul + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 57 load + r4 = r7 * r7 + r4; // 58 mad + r2 = r2 ^ ds[r1 & mask]; // 59 load + r1 = r1 * r4; // 60 mul + r3 = r3 * r4; // 61 mul + r6 = r6 ^ r4; // 62 xor + r3 = rotr_var(r3, r1); // 63 rotr + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/mixA-4/kernel.cu b/proto-cuda/packs-readwidth/mixA-4/kernel.cu new file mode 100644 index 000000000..0c357e499 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-4/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/5". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0xc255b2bfu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0xdba9d396u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0xdba9d396u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x6ea527e4u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0x6ea527e4u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x05f1866au; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x05f1866au; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x2947e02eu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0x2947e02eu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x0ecc5c6fu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0x0ecc5c6fu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xc3b7068du; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0xc3b7068du; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x282d1c2eu; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x282d1c2eu; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0xc255b2bfu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r3 = rotl_imm(r3, 26u); // 0 rotl + r2 = rotr_var(r2, r0); // 1 rotr + r0 = r0 ^ r7; // 2 xor + r4 = r4 ^ r7; // 3 xor + r2 = r2 ^ ds[r4 & mask]; // 4 load + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 5 load + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 6 shfl + r6 = rotl_imm(r6, 31u); // 7 rotl + r1 = r4 * r2 + r1; // 8 mad + r1 = r7 * r0 + r1; // 9 mad + r4 = r4 ^ ds[r1 & mask]; // 10 load + r6 = r7 * r0 + r6; // 11 mad + r2 = r2 * r6; // 12 mul + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 13 load + r2 = r2 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x4a3db5a5u : 0x5df3957du); // 14 add + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r7, 1); // 15 shfl + r3 = r5 * r6 + r3; // 16 mad + r2 = r2 * r5; // 17 mul + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 18 shfl + r0 = r0 - r1; // 19 sub + r4 = r6 * r4 + r4; // 20 mad + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r0, 16); // 21 shfl + r1 = rotl_imm(r1, 24u); // 22 rotl + r1 = r1 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0xfc07c54cu : 0x5c141117u); // 23 add + r2 = r2 ^ ds[r0 & mask]; // 24 load + r0 = r0 ^ r1; // 25 xor + r3 = r3 ^ r7; // 26 xor + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 load + r0 = r0 ^ r6; // 28 xor + r4 = r4 ^ r1; // 29 xor + r6 = r6 * r3; // 30 mul + r3 = rotl_imm(r3, 23u); // 31 rotl + r7 = rotr_var(r7, r1); // 32 rotr + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r2, 2); // 33 shfl + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 load + r6 = r6 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0x0b74657bu : 0xfb55c58du); // 35 add + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 36 load + r1 = r1 ^ ds[r4 & mask]; // 37 load + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 38 load + r4 = r4 ^ ds[r7 & mask]; // 39 load + r7 = __umulhi(r7, r3); // 40 mulhi + r0 = r0 | r6; // 41 or + r0 = r3 * r0 + r0; // 42 mad + { uint32_t b_ = (r6 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 43 load + r2 = r4 * r0 + r2; // 44 mad + r2 = r4 * r2 + r2; // 45 mad + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 46 shfl + r6 = r1 * r4 + r6; // 47 mad + r7 = r7 ^ ds[r1 & mask]; // 48 load + r6 = __umulhi(r6, r5); // 49 mulhi + r5 = __umulhi(r5, r0); // 50 mulhi + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 51 load + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 1); // 52 shfl + r2 = r2 ^ r5; // 53 xor + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 4); // 54 shfl + r2 = r2 * r5; // 55 mul + r5 = r5 * r6; // 56 mul + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 57 load + r4 = r7 * r7 + r4; // 58 mad + r2 = r2 ^ ds[r1 & mask]; // 59 load + r1 = r1 * r4; // 60 mul + r3 = r3 * r4; // 61 mul + r6 = r6 ^ r4; // 62 xor + r3 = rotr_var(r3, r1); // 63 rotr + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixA-4/kernel_bound.cl b/proto-cuda/packs-readwidth/mixA-4/kernel_bound.cl new file mode 100644 index 000000000..27ea423ff --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-4/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/5". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0xc255b2bfu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0xdba9d396u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0xdba9d396u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x6ea527e4u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x6ea527e4u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x05f1866au; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x05f1866au; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x2947e02eu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x2947e02eu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x0ecc5c6fu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x0ecc5c6fu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xc3b7068du; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xc3b7068du; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x282d1c2eu; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x282d1c2eu; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0xc255b2bfu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = rotl_imm(r3, 26u); // 0 rotl + r2 = rotr_var(r2, r0); // 1 rotr + r0 = r0 ^ r7; // 2 xor + r4 = r4 ^ r7; // 3 xor + r2 = r2 ^ ds[r4 & mask]; // 4 load + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 6 shfl + r6 = rotl_imm(r6, 31u); // 7 rotl + r1 = r4 * r2 + r1; // 8 mad + r1 = r7 * r0 + r1; // 9 mad + r4 = r4 ^ ds[r1 & mask]; // 10 load + r6 = r7 * r0 + r6; // 11 mad + r2 = r2 * r6; // 12 mul + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 13 load + r2 = r2 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x4a3db5a5u : 0x5df3957du); // 14 add + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 1u); r5 = r5 ^ t_; } // 15 shfl + r3 = r5 * r6 + r3; // 16 mad + r2 = r2 * r5; // 17 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r1 = r1 ^ t_; } // 18 shfl + r0 = r0 - r1; // 19 sub + r4 = r6 * r4 + r4; // 20 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 16u); r3 = r3 ^ t_; } // 21 shfl + r1 = rotl_imm(r1, 24u); // 22 rotl + r1 = r1 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0xfc07c54cu : 0x5c141117u); // 23 add + r2 = r2 ^ ds[r0 & mask]; // 24 load + r0 = r0 ^ r1; // 25 xor + r3 = r3 ^ r7; // 26 xor + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 load + r0 = r0 ^ r6; // 28 xor + r4 = r4 ^ r1; // 29 xor + r6 = r6 * r3; // 30 mul + r3 = rotl_imm(r3, 23u); // 31 rotl + r7 = rotr_var(r7, r1); // 32 rotr + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r6 = r6 ^ t_; } // 33 shfl + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 load + r6 = r6 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0x0b74657bu : 0xfb55c58du); // 35 add + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 36 load + r1 = r1 ^ ds[r4 & mask]; // 37 load + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 38 load + r4 = r4 ^ ds[r7 & mask]; // 39 load + r7 = mul_hi(r7, r3); // 40 mulhi + r0 = r0 | r6; // 41 or + r0 = r3 * r0 + r0; // 42 mad + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 43 load + r2 = r4 * r0 + r2; // 44 mad + r2 = r4 * r2 + r2; // 45 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r7 = r7 ^ t_; } // 46 shfl + r6 = r1 * r4 + r6; // 47 mad + r7 = r7 ^ ds[r1 & mask]; // 48 load + r6 = mul_hi(r6, r5); // 49 mulhi + r5 = mul_hi(r5, r0); // 50 mulhi + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 51 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 1u); r7 = r7 ^ t_; } // 52 shfl + r2 = r2 ^ r5; // 53 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r0 = r0 ^ t_; } // 54 shfl + r2 = r2 * r5; // 55 mul + r5 = r5 * r6; // 56 mul + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 57 load + r4 = r7 * r7 + r4; // 58 mad + r2 = r2 ^ ds[r1 & mask]; // 59 load + r1 = r1 * r4; // 60 mul + r3 = r3 * r4; // 61 mul + r6 = r6 ^ r4; // 62 xor + r3 = rotr_var(r3, r1); // 63 rotr + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = rotl_imm(r3, 26u); // 0 rotl + r2 = rotr_var(r2, r0); // 1 rotr + r0 = r0 ^ r7; // 2 xor + r4 = r4 ^ r7; // 3 xor + r2 = r2 ^ ds[r4 & mask]; // 4 load + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 6 shfl + r6 = rotl_imm(r6, 31u); // 7 rotl + r1 = r4 * r2 + r1; // 8 mad + r1 = r7 * r0 + r1; // 9 mad + r4 = r4 ^ ds[r1 & mask]; // 10 load + r6 = r7 * r0 + r6; // 11 mad + r2 = r2 * r6; // 12 mul + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 13 load + r2 = r2 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x4a3db5a5u : 0x5df3957du); // 14 add + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 1u); r5 = r5 ^ t_; } // 15 shfl + r3 = r5 * r6 + r3; // 16 mad + r2 = r2 * r5; // 17 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r1 = r1 ^ t_; } // 18 shfl + r0 = r0 - r1; // 19 sub + r4 = r6 * r4 + r4; // 20 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 16u); r3 = r3 ^ t_; } // 21 shfl + r1 = rotl_imm(r1, 24u); // 22 rotl + r1 = r1 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0xfc07c54cu : 0x5c141117u); // 23 add + r2 = r2 ^ ds[r0 & mask]; // 24 load + r0 = r0 ^ r1; // 25 xor + r3 = r3 ^ r7; // 26 xor + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 load + r0 = r0 ^ r6; // 28 xor + r4 = r4 ^ r1; // 29 xor + r6 = r6 * r3; // 30 mul + r3 = rotl_imm(r3, 23u); // 31 rotl + r7 = rotr_var(r7, r1); // 32 rotr + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r6 = r6 ^ t_; } // 33 shfl + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 load + r6 = r6 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0x0b74657bu : 0xfb55c58du); // 35 add + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 36 load + r1 = r1 ^ ds[r4 & mask]; // 37 load + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 38 load + r4 = r4 ^ ds[r7 & mask]; // 39 load + r7 = mul_hi(r7, r3); // 40 mulhi + r0 = r0 | r6; // 41 or + r0 = r3 * r0 + r0; // 42 mad + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 43 load + r2 = r4 * r0 + r2; // 44 mad + r2 = r4 * r2 + r2; // 45 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r7 = r7 ^ t_; } // 46 shfl + r6 = r1 * r4 + r6; // 47 mad + r7 = r7 ^ ds[r1 & mask]; // 48 load + r6 = mul_hi(r6, r5); // 49 mulhi + r5 = mul_hi(r5, r0); // 50 mulhi + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 51 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 1u); r7 = r7 ^ t_; } // 52 shfl + r2 = r2 ^ r5; // 53 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r0 = r0 ^ t_; } // 54 shfl + r2 = r2 * r5; // 55 mul + r5 = r5 * r6; // 56 mul + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 57 load + r4 = r7 * r7 + r4; // 58 mad + r2 = r2 ^ ds[r1 & mask]; // 59 load + r1 = r1 * r4; // 60 mul + r3 = r3 * r4; // 61 mul + r6 = r6 ^ r4; // 62 xor + r3 = rotr_var(r3, r1); // 63 rotr + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-4/kernel_bound.cu b/proto-cuda/packs-readwidth/mixA-4/kernel_bound.cu new file mode 100644 index 000000000..d2a3e6222 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-4/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/5". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r3 = rotl_imm(r3, 26u); // 0 rotl + r2 = rotr_var(r2, r0); // 1 rotr + r0 = r0 ^ r7; // 2 xor + r4 = r4 ^ r7; // 3 xor + r2 = r2 ^ ds[r4 & mask]; // 4 load + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 5 load + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 6 shfl + r6 = rotl_imm(r6, 31u); // 7 rotl + r1 = r4 * r2 + r1; // 8 mad + r1 = r7 * r0 + r1; // 9 mad + r4 = r4 ^ ds[r1 & mask]; // 10 load + r6 = r7 * r0 + r6; // 11 mad + r2 = r2 * r6; // 12 mul + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 13 load + r2 = r2 + r7 + ((((sel >> 22u) & 1u) != 0u) ? 0x4a3db5a5u : 0x5df3957du); // 14 add + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r7, 1); // 15 shfl + r3 = r5 * r6 + r3; // 16 mad + r2 = r2 * r5; // 17 mul + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 18 shfl + r0 = r0 - r1; // 19 sub + r4 = r6 * r4 + r4; // 20 mad + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r0, 16); // 21 shfl + r1 = rotl_imm(r1, 24u); // 22 rotl + r1 = r1 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0xfc07c54cu : 0x5c141117u); // 23 add + r2 = r2 ^ ds[r0 & mask]; // 24 load + r0 = r0 ^ r1; // 25 xor + r3 = r3 ^ r7; // 26 xor + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 load + r0 = r0 ^ r6; // 28 xor + r4 = r4 ^ r1; // 29 xor + r6 = r6 * r3; // 30 mul + r3 = rotl_imm(r3, 23u); // 31 rotl + r7 = rotr_var(r7, r1); // 32 rotr + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r2, 2); // 33 shfl + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 load + r6 = r6 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0x0b74657bu : 0xfb55c58du); // 35 add + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 36 load + r1 = r1 ^ ds[r4 & mask]; // 37 load + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 38 load + r4 = r4 ^ ds[r7 & mask]; // 39 load + r7 = __umulhi(r7, r3); // 40 mulhi + r0 = r0 | r6; // 41 or + r0 = r3 * r0 + r0; // 42 mad + { uint32_t b_ = (r6 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 43 load + r2 = r4 * r0 + r2; // 44 mad + r2 = r4 * r2 + r2; // 45 mad + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 46 shfl + r6 = r1 * r4 + r6; // 47 mad + r7 = r7 ^ ds[r1 & mask]; // 48 load + r6 = __umulhi(r6, r5); // 49 mulhi + r5 = __umulhi(r5, r0); // 50 mulhi + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 51 load + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 1); // 52 shfl + r2 = r2 ^ r5; // 53 xor + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 4); // 54 shfl + r2 = r2 * r5; // 55 mul + r5 = r5 * r6; // 56 mul + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 57 load + r4 = r7 * r7 + r4; // 58 mad + r2 = r2 ^ ds[r1 & mask]; // 59 load + r1 = r1 * r4; // 60 mul + r3 = r3 * r4; // 61 mul + r6 = r6 ^ r4; // 62 xor + r3 = rotr_var(r3, r1); // 63 rotr + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixA-4/memhard.h b/proto-cuda/packs-readwidth/mixA-4/memhard.h new file mode 100644 index 000000000..288047a67 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-4/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/5". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/mixA-4/memhard.metal b/proto-cuda/packs-readwidth/mixA-4/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-4/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/mixA-4/program.h b/proto-cuda/packs-readwidth/mixA-4/program.h new file mode 100644 index 000000000..ce2bfa3cf --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-4/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/5". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-readwidth/A/5" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d7265616477696474682f412f35" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x8c7f626ae00b8a7cull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 mad=10 shfl=8 xor=8 mul=7 rotl=4 add=3 mulhi=3 rotr=3 or=1 sub=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "mix50-35-15" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 50, 35, 15 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 7, 6, 3 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 2528 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0xc255b2bfu, 0xdba9d396u, 0x6ea527e4u, 0x05f1866au, 0x2947e02eu, 0x0ecc5c6fu, 0xc3b7068du, 0x282d1c2eu } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/mixA-4/program.json b/proto-cuda/packs-readwidth/mixA-4/program.json new file mode 100644 index 000000000..156ae14b4 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-4/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x8c7f626ae00b8a7c", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-readwidth/A/5", + "seed_bytes": "69676e65756d2d7265616477696474682f412f35", + "seed_words": ["0xc255b2bf", "0xdba9d396", "0x6ea527e4", "0x05f1866a", "0x2947e02e", "0x0ecc5c6f", "0xc3b7068d", "0x282d1c2e"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "mix50-35-15", + "load_slots": 16, + "load_mix_percent_4_16_64": [50, 35, 15], + "load_width_counts_4_16_64": [7, 6, 3], + "bytes_per_hash": 2528, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "mad": 10, "shfl": 8, "xor": 8, "mul": 7, "rotl": 4, "add": 3, "mulhi": 3, "rotr": 3, "or": 1, "sub": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "rotl", "dst": 3, "src": 5, "src2": 6, "imm": "0x7e513467", "imm2": "0x0fed3e98", "rot": 26, "bit": 29, "mask": 8, "width": 1}, + {"i": 1, "op": "rotr", "dst": 2, "src": 0, "src2": 1, "imm": "0x944f466f", "imm2": "0xbbb1d985", "rot": 17, "bit": 20, "mask": 1, "width": 1}, + {"i": 2, "op": "xor", "dst": 0, "src": 7, "src2": 4, "imm": "0x42d344ab", "imm2": "0xc4fa2084", "rot": 29, "bit": 10, "mask": 4, "width": 1}, + {"i": 3, "op": "xor", "dst": 4, "src": 7, "src2": 2, "imm": "0x6a558d6d", "imm2": "0x21feb6e9", "rot": 31, "bit": 0, "mask": 8, "width": 1}, + {"i": 4, "op": "load", "dst": 2, "src": 4, "src2": 4, "imm": "0xe6c6b0d6", "imm2": "0xf7ff31ee", "rot": 25, "bit": 21, "mask": 4, "width": 1}, + {"i": 5, "op": "load", "dst": 3, "src": 2, "src2": 0, "imm": "0x8c1b23b2", "imm2": "0x7db0b535", "rot": 16, "bit": 1, "mask": 1, "width": 4}, + {"i": 6, "op": "shfl", "dst": 0, "src": 7, "src2": 4, "imm": "0x9d7e34c0", "imm2": "0xf15d4662", "rot": 4, "bit": 16, "mask": 4, "width": 1}, + {"i": 7, "op": "rotl", "dst": 6, "src": 5, "src2": 1, "imm": "0x8981adb9", "imm2": "0x4bb830b4", "rot": 31, "bit": 24, "mask": 4, "width": 1}, + {"i": 8, "op": "mad", "dst": 1, "src": 4, "src2": 2, "imm": "0xcb805ebd", "imm2": "0xde933f7d", "rot": 29, "bit": 27, "mask": 8, "width": 1}, + {"i": 9, "op": "mad", "dst": 1, "src": 7, "src2": 0, "imm": "0x66835a74", "imm2": "0xde14a90d", "rot": 30, "bit": 29, "mask": 4, "width": 1}, + {"i": 10, "op": "load", "dst": 4, "src": 1, "src2": 3, "imm": "0x0e0ff63d", "imm2": "0x848393fd", "rot": 20, "bit": 13, "mask": 4, "width": 1}, + {"i": 11, "op": "mad", "dst": 6, "src": 7, "src2": 0, "imm": "0x54b20c82", "imm2": "0x6be278b9", "rot": 27, "bit": 25, "mask": 4, "width": 1}, + {"i": 12, "op": "mul", "dst": 2, "src": 6, "src2": 4, "imm": "0x152bfc86", "imm2": "0x403f65d5", "rot": 31, "bit": 12, "mask": 8, "width": 1}, + {"i": 13, "op": "load", "dst": 0, "src": 2, "src2": 5, "imm": "0x0c3b44ce", "imm2": "0x03bbe6be", "rot": 13, "bit": 27, "mask": 4, "width": 4}, + {"i": 14, "op": "add", "dst": 2, "src": 7, "src2": 2, "imm": "0x5df3957d", "imm2": "0x4a3db5a5", "rot": 19, "bit": 22, "mask": 16, "width": 1}, + {"i": 15, "op": "shfl", "dst": 5, "src": 7, "src2": 2, "imm": "0x62602275", "imm2": "0x8dd64d3c", "rot": 14, "bit": 1, "mask": 1, "width": 1}, + {"i": 16, "op": "mad", "dst": 3, "src": 5, "src2": 6, "imm": "0xdcbd85e9", "imm2": "0x6aa7c473", "rot": 3, "bit": 30, "mask": 8, "width": 1}, + {"i": 17, "op": "mul", "dst": 2, "src": 5, "src2": 0, "imm": "0xbc049a58", "imm2": "0x65f78542", "rot": 22, "bit": 28, "mask": 8, "width": 1}, + {"i": 18, "op": "shfl", "dst": 1, "src": 5, "src2": 3, "imm": "0x1381d2fa", "imm2": "0xb37a6855", "rot": 22, "bit": 29, "mask": 1, "width": 1}, + {"i": 19, "op": "sub", "dst": 0, "src": 1, "src2": 3, "imm": "0x0975c87d", "imm2": "0xcf931617", "rot": 10, "bit": 22, "mask": 8, "width": 1}, + {"i": 20, "op": "mad", "dst": 4, "src": 6, "src2": 4, "imm": "0x10759307", "imm2": "0x3eae5571", "rot": 15, "bit": 7, "mask": 8, "width": 1}, + {"i": 21, "op": "shfl", "dst": 3, "src": 0, "src2": 0, "imm": "0x047b824d", "imm2": "0xf79c23a1", "rot": 28, "bit": 5, "mask": 16, "width": 1}, + {"i": 22, "op": "rotl", "dst": 1, "src": 3, "src2": 3, "imm": "0x6ebf5dd2", "imm2": "0xcc6928a6", "rot": 24, "bit": 0, "mask": 1, "width": 1}, + {"i": 23, "op": "add", "dst": 1, "src": 0, "src2": 2, "imm": "0x5c141117", "imm2": "0xfc07c54c", "rot": 6, "bit": 22, "mask": 4, "width": 1}, + {"i": 24, "op": "load", "dst": 2, "src": 0, "src2": 0, "imm": "0x4ca5dd93", "imm2": "0x50f7ea7a", "rot": 15, "bit": 21, "mask": 16, "width": 1}, + {"i": 25, "op": "xor", "dst": 0, "src": 1, "src2": 0, "imm": "0xb02c8523", "imm2": "0xc20cbd93", "rot": 8, "bit": 14, "mask": 4, "width": 1}, + {"i": 26, "op": "xor", "dst": 3, "src": 7, "src2": 6, "imm": "0x4ac5aa82", "imm2": "0xd0a3b8ca", "rot": 1, "bit": 22, "mask": 16, "width": 1}, + {"i": 27, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xc7cb167e", "imm2": "0xe604db1b", "rot": 3, "bit": 1, "mask": 16, "width": 4}, + {"i": 28, "op": "xor", "dst": 0, "src": 6, "src2": 5, "imm": "0x68319ac4", "imm2": "0x1d0fc6ef", "rot": 31, "bit": 13, "mask": 8, "width": 1}, + {"i": 29, "op": "xor", "dst": 4, "src": 1, "src2": 4, "imm": "0x46e6b295", "imm2": "0x279187a0", "rot": 22, "bit": 6, "mask": 4, "width": 1}, + {"i": 30, "op": "mul", "dst": 6, "src": 3, "src2": 6, "imm": "0x087009f4", "imm2": "0x69bd5f06", "rot": 17, "bit": 13, "mask": 1, "width": 1}, + {"i": 31, "op": "rotl", "dst": 3, "src": 0, "src2": 3, "imm": "0x96261bd7", "imm2": "0x2304582f", "rot": 23, "bit": 31, "mask": 4, "width": 1}, + {"i": 32, "op": "rotr", "dst": 7, "src": 1, "src2": 1, "imm": "0x975e5e73", "imm2": "0x7382921b", "rot": 11, "bit": 21, "mask": 2, "width": 1}, + {"i": 33, "op": "shfl", "dst": 6, "src": 2, "src2": 2, "imm": "0xfc64a106", "imm2": "0xfa7eb90c", "rot": 1, "bit": 16, "mask": 2, "width": 1}, + {"i": 34, "op": "load", "dst": 5, "src": 0, "src2": 7, "imm": "0xe536f810", "imm2": "0xe44d2045", "rot": 4, "bit": 9, "mask": 8, "width": 16}, + {"i": 35, "op": "add", "dst": 6, "src": 3, "src2": 6, "imm": "0xfb55c58d", "imm2": "0x0b74657b", "rot": 25, "bit": 29, "mask": 8, "width": 1}, + {"i": 36, "op": "load", "dst": 6, "src": 5, "src2": 1, "imm": "0x85c5e518", "imm2": "0xdf4f27ed", "rot": 10, "bit": 6, "mask": 2, "width": 16}, + {"i": 37, "op": "load", "dst": 1, "src": 4, "src2": 6, "imm": "0xd47ed295", "imm2": "0x062102da", "rot": 1, "bit": 15, "mask": 16, "width": 1}, + {"i": 38, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x07c3080a", "imm2": "0x2f4b3801", "rot": 8, "bit": 10, "mask": 4, "width": 4}, + {"i": 39, "op": "load", "dst": 4, "src": 7, "src2": 5, "imm": "0x1beb83e7", "imm2": "0x90d4963d", "rot": 16, "bit": 31, "mask": 4, "width": 1}, + {"i": 40, "op": "mulhi", "dst": 7, "src": 3, "src2": 7, "imm": "0x42b300c8", "imm2": "0x05b01fcb", "rot": 5, "bit": 17, "mask": 1, "width": 1}, + {"i": 41, "op": "or", "dst": 0, "src": 6, "src2": 5, "imm": "0xd6d93d15", "imm2": "0xcef78983", "rot": 31, "bit": 6, "mask": 1, "width": 1}, + {"i": 42, "op": "mad", "dst": 0, "src": 3, "src2": 0, "imm": "0xd14e02da", "imm2": "0x72d866a4", "rot": 9, "bit": 29, "mask": 4, "width": 1}, + {"i": 43, "op": "load", "dst": 4, "src": 6, "src2": 6, "imm": "0x84684f5c", "imm2": "0x396d612e", "rot": 2, "bit": 7, "mask": 2, "width": 16}, + {"i": 44, "op": "mad", "dst": 2, "src": 4, "src2": 0, "imm": "0x0a129ff8", "imm2": "0x8dbb8864", "rot": 20, "bit": 29, "mask": 8, "width": 1}, + {"i": 45, "op": "mad", "dst": 2, "src": 4, "src2": 2, "imm": "0x8e8de36b", "imm2": "0x5f1ad032", "rot": 11, "bit": 7, "mask": 2, "width": 1}, + {"i": 46, "op": "shfl", "dst": 7, "src": 5, "src2": 0, "imm": "0x395fb2f0", "imm2": "0x1f8120ba", "rot": 13, "bit": 25, "mask": 16, "width": 1}, + {"i": 47, "op": "mad", "dst": 6, "src": 1, "src2": 4, "imm": "0x46dd77ef", "imm2": "0x3344daf3", "rot": 24, "bit": 26, "mask": 4, "width": 1}, + {"i": 48, "op": "load", "dst": 7, "src": 1, "src2": 5, "imm": "0x72e62a51", "imm2": "0x6887cd34", "rot": 28, "bit": 6, "mask": 4, "width": 1}, + {"i": 49, "op": "mulhi", "dst": 6, "src": 5, "src2": 6, "imm": "0x15785a34", "imm2": "0x4cd4815d", "rot": 18, "bit": 22, "mask": 1, "width": 1}, + {"i": 50, "op": "mulhi", "dst": 5, "src": 0, "src2": 3, "imm": "0xf4adde36", "imm2": "0x3a596180", "rot": 28, "bit": 11, "mask": 2, "width": 1}, + {"i": 51, "op": "load", "dst": 1, "src": 6, "src2": 0, "imm": "0xbe0a1ce3", "imm2": "0xa67b98b2", "rot": 19, "bit": 12, "mask": 8, "width": 4}, + {"i": 52, "op": "shfl", "dst": 7, "src": 3, "src2": 2, "imm": "0xde71a6c7", "imm2": "0x089678f0", "rot": 11, "bit": 22, "mask": 1, "width": 1}, + {"i": 53, "op": "xor", "dst": 2, "src": 5, "src2": 6, "imm": "0x25e0b714", "imm2": "0xc498694f", "rot": 31, "bit": 20, "mask": 2, "width": 1}, + {"i": 54, "op": "shfl", "dst": 0, "src": 6, "src2": 5, "imm": "0x4f08890f", "imm2": "0x9214703e", "rot": 1, "bit": 23, "mask": 4, "width": 1}, + {"i": 55, "op": "mul", "dst": 2, "src": 5, "src2": 3, "imm": "0x3c2cb502", "imm2": "0xd2b3be81", "rot": 31, "bit": 18, "mask": 1, "width": 1}, + {"i": 56, "op": "mul", "dst": 5, "src": 6, "src2": 3, "imm": "0x856a28d5", "imm2": "0xd0a711af", "rot": 24, "bit": 0, "mask": 1, "width": 1}, + {"i": 57, "op": "load", "dst": 6, "src": 4, "src2": 7, "imm": "0xf0c439cf", "imm2": "0x36507133", "rot": 14, "bit": 2, "mask": 4, "width": 4}, + {"i": 58, "op": "mad", "dst": 4, "src": 7, "src2": 7, "imm": "0x1df01e69", "imm2": "0xae8c4b8e", "rot": 28, "bit": 17, "mask": 2, "width": 1}, + {"i": 59, "op": "load", "dst": 2, "src": 1, "src2": 6, "imm": "0x3796e56b", "imm2": "0x4b519b72", "rot": 30, "bit": 31, "mask": 2, "width": 1}, + {"i": 60, "op": "mul", "dst": 1, "src": 4, "src2": 2, "imm": "0x4c165ca5", "imm2": "0xd4c97f2a", "rot": 18, "bit": 20, "mask": 1, "width": 1}, + {"i": 61, "op": "mul", "dst": 3, "src": 4, "src2": 1, "imm": "0x5d55a322", "imm2": "0xdde0b168", "rot": 1, "bit": 27, "mask": 4, "width": 1}, + {"i": 62, "op": "xor", "dst": 6, "src": 4, "src2": 5, "imm": "0x82cf3c97", "imm2": "0x7c368d65", "rot": 26, "bit": 17, "mask": 4, "width": 1}, + {"i": 63, "op": "rotr", "dst": 3, "src": 1, "src2": 2, "imm": "0x53bbae08", "imm2": "0x223ccdc5", "rot": 28, "bit": 16, "mask": 4, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/mixA-4/program.metal b/proto-cuda/packs-readwidth/mixA-4/program.metal new file mode 100644 index 000000000..a10aefab8 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-4/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0xc255b2bfu, 0xdba9d396u, 0x6ea527e4u, 0x05f1866au, 0x2947e02eu, 0x0ecc5c6fu, 0xc3b7068du, 0x282d1c2eu }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = rotl_imm(r3, 26u); // 0 + r2 = rotr_var(r2, r0); // 1 + r0 = r0 ^ r7; // 2 + r4 = r4 ^ r7; // 3 + r2 = r2 ^ dataset[r4 & MASK]; // 4 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 5 + r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 6 + r6 = rotl_imm(r6, 31u); // 7 + r1 = r4 * r2 + r1; // 8 + r1 = r7 * r0 + r1; // 9 + r4 = r4 ^ dataset[r1 & MASK]; // 10 + r6 = r7 * r0 + r6; // 11 + r2 = r2 * r6; // 12 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 13 + r2 = r2 + r7 + select(0x5df3957du, 0x4a3db5a5u, ((sel >> 22u) & 1u) != 0u); // 14 + r5 = r5 ^ simd_shuffle_xor(r7, (ushort)1); // 15 + r3 = r5 * r6 + r3; // 16 + r2 = r2 * r5; // 17 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)1); // 18 + r0 = r0 - r1; // 19 + r4 = r6 * r4 + r4; // 20 + r3 = r3 ^ simd_shuffle_xor(r0, (ushort)16); // 21 + r1 = rotl_imm(r1, 24u); // 22 + r1 = r1 + r0 + select(0x5c141117u, 0xfc07c54cu, ((sel >> 22u) & 1u) != 0u); // 23 + r2 = r2 ^ dataset[r0 & MASK]; // 24 + r0 = r0 ^ r1; // 25 + r3 = r3 ^ r7; // 26 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 + r0 = r0 ^ r6; // 28 + r4 = r4 ^ r1; // 29 + r6 = r6 * r3; // 30 + r3 = rotl_imm(r3, 23u); // 31 + r7 = rotr_var(r7, r1); // 32 + r6 = r6 ^ simd_shuffle_xor(r2, (ushort)2); // 33 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 + r6 = r6 + r3 + select(0xfb55c58du, 0x0b74657bu, ((sel >> 29u) & 1u) != 0u); // 35 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 36 + r1 = r1 ^ dataset[r4 & MASK]; // 37 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 38 + r4 = r4 ^ dataset[r7 & MASK]; // 39 + r7 = mulhi(r7, r3); // 40 + r0 = r0 | r6; // 41 + r0 = r3 * r0 + r0; // 42 + { uint b_ = (r6 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 43 + r2 = r4 * r0 + r2; // 44 + r2 = r4 * r2 + r2; // 45 + r7 = r7 ^ simd_shuffle_xor(r5, (ushort)16); // 46 + r6 = r1 * r4 + r6; // 47 + r7 = r7 ^ dataset[r1 & MASK]; // 48 + r6 = mulhi(r6, r5); // 49 + r5 = mulhi(r5, r0); // 50 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 51 + r7 = r7 ^ simd_shuffle_xor(r3, (ushort)1); // 52 + r2 = r2 ^ r5; // 53 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)4); // 54 + r2 = r2 * r5; // 55 + r5 = r5 * r6; // 56 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 57 + r4 = r7 * r7 + r4; // 58 + r2 = r2 ^ dataset[r1 & MASK]; // 59 + r1 = r1 * r4; // 60 + r3 = r3 * r4; // 61 + r6 = r6 ^ r4; // 62 + r3 = rotr_var(r3, r1); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-4/program_bound.metal b/proto-cuda/packs-readwidth/mixA-4/program_bound.metal new file mode 100644 index 000000000..a625a972b --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-4/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0xc255b2bfu, 0xdba9d396u, 0x6ea527e4u, 0x05f1866au, 0x2947e02eu, 0x0ecc5c6fu, 0xc3b7068du, 0x282d1c2eu }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = rotl_imm(r3, 26u); // 0 + r2 = rotr_var(r2, r0); // 1 + r0 = r0 ^ r7; // 2 + r4 = r4 ^ r7; // 3 + r2 = r2 ^ dataset[r4 & MASK]; // 4 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 5 + r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 6 + r6 = rotl_imm(r6, 31u); // 7 + r1 = r4 * r2 + r1; // 8 + r1 = r7 * r0 + r1; // 9 + r4 = r4 ^ dataset[r1 & MASK]; // 10 + r6 = r7 * r0 + r6; // 11 + r2 = r2 * r6; // 12 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 13 + r2 = r2 + r7 + select(0x5df3957du, 0x4a3db5a5u, ((sel >> 22u) & 1u) != 0u); // 14 + r5 = r5 ^ simd_shuffle_xor(r7, (ushort)1); // 15 + r3 = r5 * r6 + r3; // 16 + r2 = r2 * r5; // 17 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)1); // 18 + r0 = r0 - r1; // 19 + r4 = r6 * r4 + r4; // 20 + r3 = r3 ^ simd_shuffle_xor(r0, (ushort)16); // 21 + r1 = rotl_imm(r1, 24u); // 22 + r1 = r1 + r0 + select(0x5c141117u, 0xfc07c54cu, ((sel >> 22u) & 1u) != 0u); // 23 + r2 = r2 ^ dataset[r0 & MASK]; // 24 + r0 = r0 ^ r1; // 25 + r3 = r3 ^ r7; // 26 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 + r0 = r0 ^ r6; // 28 + r4 = r4 ^ r1; // 29 + r6 = r6 * r3; // 30 + r3 = rotl_imm(r3, 23u); // 31 + r7 = rotr_var(r7, r1); // 32 + r6 = r6 ^ simd_shuffle_xor(r2, (ushort)2); // 33 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 + r6 = r6 + r3 + select(0xfb55c58du, 0x0b74657bu, ((sel >> 29u) & 1u) != 0u); // 35 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 36 + r1 = r1 ^ dataset[r4 & MASK]; // 37 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 38 + r4 = r4 ^ dataset[r7 & MASK]; // 39 + r7 = mulhi(r7, r3); // 40 + r0 = r0 | r6; // 41 + r0 = r3 * r0 + r0; // 42 + { uint b_ = (r6 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 43 + r2 = r4 * r0 + r2; // 44 + r2 = r4 * r2 + r2; // 45 + r7 = r7 ^ simd_shuffle_xor(r5, (ushort)16); // 46 + r6 = r1 * r4 + r6; // 47 + r7 = r7 ^ dataset[r1 & MASK]; // 48 + r6 = mulhi(r6, r5); // 49 + r5 = mulhi(r5, r0); // 50 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 51 + r7 = r7 ^ simd_shuffle_xor(r3, (ushort)1); // 52 + r2 = r2 ^ r5; // 53 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)4); // 54 + r2 = r2 * r5; // 55 + r5 = r5 * r6; // 56 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 57 + r4 = r7 * r7 + r4; // 58 + r2 = r2 ^ dataset[r1 & MASK]; // 59 + r1 = r1 * r4; // 60 + r3 = r3 * r4; // 61 + r6 = r6 ^ r4; // 62 + r3 = rotr_var(r3, r1); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-4/vectors.h b/proto-cuda/packs-readwidth/mixA-4/vectors.h new file mode 100644 index 000000000..2d1f11191 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-4/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/5". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0xeaee437d9c4bba42ull, 0xa89fc64fb516d23bull, 0xf50c45c711918300ull, 0x7f02b33440a1ee24ull, 0x2010bc3f907d5040ull, 0x74d0198a4fbd6269ull, 0xe6248af74a39174eull, 0x1eb569508ef6c62bull, + 0xa33d857ed5fe79beull, 0x373c29561fbc6793ull, 0xf085a877af1bb6e8ull, 0x642c45f3214b4d6dull, 0xe1ad911edc598629ull, 0xff3e0523b7834ed3ull, 0x7c8752e2db521931ull, 0xe14314ca7f3637bcull, + 0x325b4f284dcc8f53ull, 0x4d2d6b6e3dd0acc0ull, 0x435670a68d505a3dull, 0x719f3f336a843deeull, 0x101662b22c52db66ull, 0x4467ce7e0020e5deull, 0xef2031f7841e1b86ull, 0xccac3386fc48adbdull, + 0x0634a7aa079c7fcfull, 0x5139e023bceb13d3ull, 0xc114f6be02fd1befull, 0x3401895eb20ff261ull, 0xf1e3dcf94fa94b59ull, 0xa4af930f43a9a6e7ull, 0xcf2a8f7bc60fd349ull, 0x3cf48986c94c816cull + }, + { // base nonce 4096 + 0x81fc83f13ff4be40ull, 0x7c42d114d8e674fbull, 0x816154766b564cccull, 0xb7bf2a4a3d544d2eull, 0x7de35c931f18d8c9ull, 0x1d3829de6e647586ull, 0x76a8d92d75a87fc4ull, 0x2308825663bba9b7ull, + 0xe6d226a37785ffa1ull, 0xdeef0eed55374e0full, 0xde5cfde59a1a53ecull, 0xc4799ac83c5a0602ull, 0xaab7dada14aa7c47ull, 0x7e5eacebfb6ec92cull, 0x2c99a69019e48570ull, 0xae065fa5a54b586aull, + 0x23cb704f5a33e2caull, 0xf0bdea48e2c03edbull, 0xe150785d233211ceull, 0xb45388e11497f6bbull, 0xf862683cde4568bcull, 0x2e4e81f9fafef5c2ull, 0x8e6aef8b82861379ull, 0x7883eea61e6b6cbdull, + 0xcc89c8ea28a7a715ull, 0x4adb824fc1302fc7ull, 0x0b99a18825cf092cull, 0xac0f0a2a901bc59aull, 0xef5e0f0bf1c9b8aeull, 0x2220f11e1b69bc0aull, 0x3e9695b77c69892cull, 0xbfe252bff26b14fcull + }, + { // base nonce 1000000 + 0xb2cc4dc245d1522dull, 0x3f89483adcc726bcull, 0x3df0f900a0027a6cull, 0xa47b69fe239ee7dcull, 0xff539ee62e5ffa98ull, 0x622a019529ba3cc4ull, 0xf92f75b66deb9f43ull, 0x5bdb07bf6d151978ull, + 0xc6b6ee7fb2350d3aull, 0x0d47d76e37088cc7ull, 0xc662723c083ae26dull, 0x68ddb8f6849ffa3cull, 0x5439d7ce29a28ac0ull, 0x59ffe6543a643d07ull, 0xec4286570ff0b04cull, 0xef042881c63fdda4ull, + 0xf8185da0349a2894ull, 0xdc4e87d0650fe7a1ull, 0x7df286ee2777a93full, 0x4c2b743817f1a427ull, 0x1b7aef4be8cf3bb3ull, 0xfb64afb73c0fbab8ull, 0x365b83541877755dull, 0x59872f36106c4827ull, + 0x452e18f12aa47c59ull, 0xed37e59b230c5ad7ull, 0x383768d545312f62ull, 0x30c548bee141f5f5ull, 0x073494619104a4c4ull, 0x0ba09ddfc99ee224ull, 0x0a371f34ea9b52a4ull, 0xe3c763c14c3e2f38ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/mixA-4/vectors.json b/proto-cuda/packs-readwidth/mixA-4/vectors.json new file mode 100644 index 000000000..366cb1f0c --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-4/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-readwidth/A/5", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0xeaee437d9c4bba42", "0xa89fc64fb516d23b", "0xf50c45c711918300", "0x7f02b33440a1ee24", "0x2010bc3f907d5040", "0x74d0198a4fbd6269", "0xe6248af74a39174e", "0x1eb569508ef6c62b", + "0xa33d857ed5fe79be", "0x373c29561fbc6793", "0xf085a877af1bb6e8", "0x642c45f3214b4d6d", "0xe1ad911edc598629", "0xff3e0523b7834ed3", "0x7c8752e2db521931", "0xe14314ca7f3637bc", + "0x325b4f284dcc8f53", "0x4d2d6b6e3dd0acc0", "0x435670a68d505a3d", "0x719f3f336a843dee", "0x101662b22c52db66", "0x4467ce7e0020e5de", "0xef2031f7841e1b86", "0xccac3386fc48adbd", + "0x0634a7aa079c7fcf", "0x5139e023bceb13d3", "0xc114f6be02fd1bef", "0x3401895eb20ff261", "0xf1e3dcf94fa94b59", "0xa4af930f43a9a6e7", "0xcf2a8f7bc60fd349", "0x3cf48986c94c816c" + ]}, + {"base_nonce": 4096, "expected": [ + "0x81fc83f13ff4be40", "0x7c42d114d8e674fb", "0x816154766b564ccc", "0xb7bf2a4a3d544d2e", "0x7de35c931f18d8c9", "0x1d3829de6e647586", "0x76a8d92d75a87fc4", "0x2308825663bba9b7", + "0xe6d226a37785ffa1", "0xdeef0eed55374e0f", "0xde5cfde59a1a53ec", "0xc4799ac83c5a0602", "0xaab7dada14aa7c47", "0x7e5eacebfb6ec92c", "0x2c99a69019e48570", "0xae065fa5a54b586a", + "0x23cb704f5a33e2ca", "0xf0bdea48e2c03edb", "0xe150785d233211ce", "0xb45388e11497f6bb", "0xf862683cde4568bc", "0x2e4e81f9fafef5c2", "0x8e6aef8b82861379", "0x7883eea61e6b6cbd", + "0xcc89c8ea28a7a715", "0x4adb824fc1302fc7", "0x0b99a18825cf092c", "0xac0f0a2a901bc59a", "0xef5e0f0bf1c9b8ae", "0x2220f11e1b69bc0a", "0x3e9695b77c69892c", "0xbfe252bff26b14fc" + ]}, + {"base_nonce": 1000000, "expected": [ + "0xb2cc4dc245d1522d", "0x3f89483adcc726bc", "0x3df0f900a0027a6c", "0xa47b69fe239ee7dc", "0xff539ee62e5ffa98", "0x622a019529ba3cc4", "0xf92f75b66deb9f43", "0x5bdb07bf6d151978", + "0xc6b6ee7fb2350d3a", "0x0d47d76e37088cc7", "0xc662723c083ae26d", "0x68ddb8f6849ffa3c", "0x5439d7ce29a28ac0", "0x59ffe6543a643d07", "0xec4286570ff0b04c", "0xef042881c63fdda4", + "0xf8185da0349a2894", "0xdc4e87d0650fe7a1", "0x7df286ee2777a93f", "0x4c2b743817f1a427", "0x1b7aef4be8cf3bb3", "0xfb64afb73c0fbab8", "0x365b83541877755d", "0x59872f36106c4827", + "0x452e18f12aa47c59", "0xed37e59b230c5ad7", "0x383768d545312f62", "0x30c548bee141f5f5", "0x073494619104a4c4", "0x0ba09ddfc99ee224", "0x0a371f34ea9b52a4", "0xe3c763c14c3e2f38" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/mixA-5/kernel.cl b/proto-cuda/packs-readwidth/mixA-5/kernel.cl new file mode 100644 index 000000000..c74dfe203 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-5/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/6". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x2a53aad4u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x37fcb6e2u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x37fcb6e2u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x25d66a27u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x25d66a27u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0xf5249e5eu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0xf5249e5eu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x15f86a59u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x15f86a59u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x8144559au; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x8144559au; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xcad741a6u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xcad741a6u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xcb961373u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0xcb961373u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x2a53aad4u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = rotl_imm(r0, 10u); // 0 rotl + r5 = r3 * r1 + r5; // 1 mad + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 2 load + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 3 load + r5 = r2 * r2 + r5; // 4 mad + r4 = r4 + r3 + ((((sel >> 10u) & 1u) != 0u) ? 0x4edf355au : 0xbd587d10u); // 5 add + r3 = r3 ^ r2; // 6 xor + r5 = r5 ^ r3; // 7 xor + r7 = r6 * r2 + r7; // 8 mad + r7 = r6 * r3 + r7; // 9 mad + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 10 load + r1 = mul_hi(r1, r4); // 11 mulhi + r1 = r1 ^ r6; // 12 xor + r2 = mul_hi(r2, r1); // 13 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r6 = r6 ^ t_; } // 14 shfl + r4 = r6 * r4 + r4; // 15 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 8u); r1 = r1 ^ t_; } // 16 shfl + r5 = r5 ^ r4; // 17 xor + r0 = r0 * r2; // 18 mul + r6 = r6 ^ ds[r3 & mask]; // 19 load + r6 = mul_hi(r6, r2); // 20 mulhi + r7 = r2 * r6 + r7; // 21 mad + r4 = r4 * r5; // 22 mul + r3 = r3 * r5; // 23 mul + r6 = r6 * r4; // 24 mul + r3 = r3 ^ ds[r7 & mask]; // 25 load + r2 = rotr_var(r2, r5); // 26 rotr + r5 = r5 | r6; // 27 or + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 28 load + r6 = r6 | r0; // 29 or + r0 = r0 + r3 + ((((sel >> 2u) & 1u) != 0u) ? 0xdca976acu : 0x5faa0547u); // 30 add + r6 = mul_hi(r6, r0); // 31 mulhi + r3 = rotr_var(r3, r7); // 32 rotr + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 4u); r0 = r0 ^ t_; } // 33 shfl + r2 = rotl_imm(r2, 17u); // 34 rotl + r7 = mul_hi(r7, r6); // 35 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r3 = r3 ^ t_; } // 36 shfl + r6 = r6 ^ ds[r0 & mask]; // 37 load + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 38 load + r1 = r1 ^ r0; // 39 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 2u); r3 = r3 ^ t_; } // 40 shfl + r1 = r1 ^ ds[r5 & mask]; // 41 load + r3 = r3 * r2; // 42 mul + r2 = r2 ^ ds[r7 & mask]; // 43 load + r7 = r7 + r3 + ((((sel >> 7u) & 1u) != 0u) ? 0x610bd3c2u : 0xe28c457cu); // 44 add + r7 = mul_hi(r7, r6); // 45 mulhi + r3 = r3 ^ r4; // 46 xor + r3 = r3 ^ ds[r7 & mask]; // 47 load + r7 = mul_hi(r7, r2); // 48 mulhi + r6 = mul_hi(r6, r3); // 49 mulhi + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 50 load + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 51 load + r5 = r5 ^ r3; // 52 xor + r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xcca8fa7cu : 0xd6457f6eu); // 53 add + r7 = rotr_var(r7, r1); // 54 rotr + r6 = r6 + r7 + ((((sel >> 12u) & 1u) != 0u) ? 0x826cb755u : 0xf3e1301cu); // 55 add + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 56 load + r1 = rotr_var(r1, r5); // 57 rotr + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 58 load + r3 = r3 - r7; // 59 sub + r1 = r4 * r6 + r1; // 60 mad + r2 = r1 * r2 + r2; // 61 mad + r4 = r4 ^ r3; // 62 xor + r2 = r2 ^ ds[r4 & mask]; // 63 load + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/mixA-5/kernel.cu b/proto-cuda/packs-readwidth/mixA-5/kernel.cu new file mode 100644 index 000000000..930d2d554 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-5/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/6". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x2a53aad4u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x37fcb6e2u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x37fcb6e2u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x25d66a27u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0x25d66a27u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0xf5249e5eu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0xf5249e5eu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x15f86a59u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0x15f86a59u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x8144559au; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0x8144559au; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xcad741a6u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0xcad741a6u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xcb961373u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0xcb961373u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x2a53aad4u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r0 = rotl_imm(r0, 10u); // 0 rotl + r5 = r3 * r1 + r5; // 1 mad + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 2 load + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 3 load + r5 = r2 * r2 + r5; // 4 mad + r4 = r4 + r3 + ((((sel >> 10u) & 1u) != 0u) ? 0x4edf355au : 0xbd587d10u); // 5 add + r3 = r3 ^ r2; // 6 xor + r5 = r5 ^ r3; // 7 xor + r7 = r6 * r2 + r7; // 8 mad + r7 = r6 * r3 + r7; // 9 mad + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 10 load + r1 = __umulhi(r1, r4); // 11 mulhi + r1 = r1 ^ r6; // 12 xor + r2 = __umulhi(r2, r1); // 13 mulhi + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 14 shfl + r4 = r6 * r4 + r4; // 15 mad + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r0, 8); // 16 shfl + r5 = r5 ^ r4; // 17 xor + r0 = r0 * r2; // 18 mul + r6 = r6 ^ ds[r3 & mask]; // 19 load + r6 = __umulhi(r6, r2); // 20 mulhi + r7 = r2 * r6 + r7; // 21 mad + r4 = r4 * r5; // 22 mul + r3 = r3 * r5; // 23 mul + r6 = r6 * r4; // 24 mul + r3 = r3 ^ ds[r7 & mask]; // 25 load + r2 = rotr_var(r2, r5); // 26 rotr + r5 = r5 | r6; // 27 or + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 28 load + r6 = r6 | r0; // 29 or + r0 = r0 + r3 + ((((sel >> 2u) & 1u) != 0u) ? 0xdca976acu : 0x5faa0547u); // 30 add + r6 = __umulhi(r6, r0); // 31 mulhi + r3 = rotr_var(r3, r7); // 32 rotr + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r4, 4); // 33 shfl + r2 = rotl_imm(r2, 17u); // 34 rotl + r7 = __umulhi(r7, r6); // 35 mulhi + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r2, 2); // 36 shfl + r6 = r6 ^ ds[r0 & mask]; // 37 load + { uint32_t b_ = (r6 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 38 load + r1 = r1 ^ r0; // 39 xor + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r1, 2); // 40 shfl + r1 = r1 ^ ds[r5 & mask]; // 41 load + r3 = r3 * r2; // 42 mul + r2 = r2 ^ ds[r7 & mask]; // 43 load + r7 = r7 + r3 + ((((sel >> 7u) & 1u) != 0u) ? 0x610bd3c2u : 0xe28c457cu); // 44 add + r7 = __umulhi(r7, r6); // 45 mulhi + r3 = r3 ^ r4; // 46 xor + r3 = r3 ^ ds[r7 & mask]; // 47 load + r7 = __umulhi(r7, r2); // 48 mulhi + r6 = __umulhi(r6, r3); // 49 mulhi + { uint32_t b_ = (r1 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 50 load + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 51 load + r5 = r5 ^ r3; // 52 xor + r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xcca8fa7cu : 0xd6457f6eu); // 53 add + r7 = rotr_var(r7, r1); // 54 rotr + r6 = r6 + r7 + ((((sel >> 12u) & 1u) != 0u) ? 0x826cb755u : 0xf3e1301cu); // 55 add + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 56 load + r1 = rotr_var(r1, r5); // 57 rotr + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 58 load + r3 = r3 - r7; // 59 sub + r1 = r4 * r6 + r1; // 60 mad + r2 = r1 * r2 + r2; // 61 mad + r4 = r4 ^ r3; // 62 xor + r2 = r2 ^ ds[r4 & mask]; // 63 load + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixA-5/kernel_bound.cl b/proto-cuda/packs-readwidth/mixA-5/kernel_bound.cl new file mode 100644 index 000000000..78d6a9e24 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-5/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/6". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x2a53aad4u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x37fcb6e2u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x37fcb6e2u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x25d66a27u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x25d66a27u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0xf5249e5eu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0xf5249e5eu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x15f86a59u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x15f86a59u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x8144559au; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x8144559au; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xcad741a6u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xcad741a6u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xcb961373u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0xcb961373u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x2a53aad4u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = rotl_imm(r0, 10u); // 0 rotl + r5 = r3 * r1 + r5; // 1 mad + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 2 load + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 3 load + r5 = r2 * r2 + r5; // 4 mad + r4 = r4 + r3 + ((((sel >> 10u) & 1u) != 0u) ? 0x4edf355au : 0xbd587d10u); // 5 add + r3 = r3 ^ r2; // 6 xor + r5 = r5 ^ r3; // 7 xor + r7 = r6 * r2 + r7; // 8 mad + r7 = r6 * r3 + r7; // 9 mad + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 10 load + r1 = mul_hi(r1, r4); // 11 mulhi + r1 = r1 ^ r6; // 12 xor + r2 = mul_hi(r2, r1); // 13 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r6 = r6 ^ t_; } // 14 shfl + r4 = r6 * r4 + r4; // 15 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 8u); r1 = r1 ^ t_; } // 16 shfl + r5 = r5 ^ r4; // 17 xor + r0 = r0 * r2; // 18 mul + r6 = r6 ^ ds[r3 & mask]; // 19 load + r6 = mul_hi(r6, r2); // 20 mulhi + r7 = r2 * r6 + r7; // 21 mad + r4 = r4 * r5; // 22 mul + r3 = r3 * r5; // 23 mul + r6 = r6 * r4; // 24 mul + r3 = r3 ^ ds[r7 & mask]; // 25 load + r2 = rotr_var(r2, r5); // 26 rotr + r5 = r5 | r6; // 27 or + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 28 load + r6 = r6 | r0; // 29 or + r0 = r0 + r3 + ((((sel >> 2u) & 1u) != 0u) ? 0xdca976acu : 0x5faa0547u); // 30 add + r6 = mul_hi(r6, r0); // 31 mulhi + r3 = rotr_var(r3, r7); // 32 rotr + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 4u); r0 = r0 ^ t_; } // 33 shfl + r2 = rotl_imm(r2, 17u); // 34 rotl + r7 = mul_hi(r7, r6); // 35 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r3 = r3 ^ t_; } // 36 shfl + r6 = r6 ^ ds[r0 & mask]; // 37 load + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 38 load + r1 = r1 ^ r0; // 39 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 2u); r3 = r3 ^ t_; } // 40 shfl + r1 = r1 ^ ds[r5 & mask]; // 41 load + r3 = r3 * r2; // 42 mul + r2 = r2 ^ ds[r7 & mask]; // 43 load + r7 = r7 + r3 + ((((sel >> 7u) & 1u) != 0u) ? 0x610bd3c2u : 0xe28c457cu); // 44 add + r7 = mul_hi(r7, r6); // 45 mulhi + r3 = r3 ^ r4; // 46 xor + r3 = r3 ^ ds[r7 & mask]; // 47 load + r7 = mul_hi(r7, r2); // 48 mulhi + r6 = mul_hi(r6, r3); // 49 mulhi + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 50 load + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 51 load + r5 = r5 ^ r3; // 52 xor + r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xcca8fa7cu : 0xd6457f6eu); // 53 add + r7 = rotr_var(r7, r1); // 54 rotr + r6 = r6 + r7 + ((((sel >> 12u) & 1u) != 0u) ? 0x826cb755u : 0xf3e1301cu); // 55 add + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 56 load + r1 = rotr_var(r1, r5); // 57 rotr + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 58 load + r3 = r3 - r7; // 59 sub + r1 = r4 * r6 + r1; // 60 mad + r2 = r1 * r2 + r2; // 61 mad + r4 = r4 ^ r3; // 62 xor + r2 = r2 ^ ds[r4 & mask]; // 63 load + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = rotl_imm(r0, 10u); // 0 rotl + r5 = r3 * r1 + r5; // 1 mad + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 2 load + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 3 load + r5 = r2 * r2 + r5; // 4 mad + r4 = r4 + r3 + ((((sel >> 10u) & 1u) != 0u) ? 0x4edf355au : 0xbd587d10u); // 5 add + r3 = r3 ^ r2; // 6 xor + r5 = r5 ^ r3; // 7 xor + r7 = r6 * r2 + r7; // 8 mad + r7 = r6 * r3 + r7; // 9 mad + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 10 load + r1 = mul_hi(r1, r4); // 11 mulhi + r1 = r1 ^ r6; // 12 xor + r2 = mul_hi(r2, r1); // 13 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r6 = r6 ^ t_; } // 14 shfl + r4 = r6 * r4 + r4; // 15 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 8u); r1 = r1 ^ t_; } // 16 shfl + r5 = r5 ^ r4; // 17 xor + r0 = r0 * r2; // 18 mul + r6 = r6 ^ ds[r3 & mask]; // 19 load + r6 = mul_hi(r6, r2); // 20 mulhi + r7 = r2 * r6 + r7; // 21 mad + r4 = r4 * r5; // 22 mul + r3 = r3 * r5; // 23 mul + r6 = r6 * r4; // 24 mul + r3 = r3 ^ ds[r7 & mask]; // 25 load + r2 = rotr_var(r2, r5); // 26 rotr + r5 = r5 | r6; // 27 or + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 28 load + r6 = r6 | r0; // 29 or + r0 = r0 + r3 + ((((sel >> 2u) & 1u) != 0u) ? 0xdca976acu : 0x5faa0547u); // 30 add + r6 = mul_hi(r6, r0); // 31 mulhi + r3 = rotr_var(r3, r7); // 32 rotr + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 4u); r0 = r0 ^ t_; } // 33 shfl + r2 = rotl_imm(r2, 17u); // 34 rotl + r7 = mul_hi(r7, r6); // 35 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r3 = r3 ^ t_; } // 36 shfl + r6 = r6 ^ ds[r0 & mask]; // 37 load + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 38 load + r1 = r1 ^ r0; // 39 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 2u); r3 = r3 ^ t_; } // 40 shfl + r1 = r1 ^ ds[r5 & mask]; // 41 load + r3 = r3 * r2; // 42 mul + r2 = r2 ^ ds[r7 & mask]; // 43 load + r7 = r7 + r3 + ((((sel >> 7u) & 1u) != 0u) ? 0x610bd3c2u : 0xe28c457cu); // 44 add + r7 = mul_hi(r7, r6); // 45 mulhi + r3 = r3 ^ r4; // 46 xor + r3 = r3 ^ ds[r7 & mask]; // 47 load + r7 = mul_hi(r7, r2); // 48 mulhi + r6 = mul_hi(r6, r3); // 49 mulhi + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 50 load + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 51 load + r5 = r5 ^ r3; // 52 xor + r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xcca8fa7cu : 0xd6457f6eu); // 53 add + r7 = rotr_var(r7, r1); // 54 rotr + r6 = r6 + r7 + ((((sel >> 12u) & 1u) != 0u) ? 0x826cb755u : 0xf3e1301cu); // 55 add + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 56 load + r1 = rotr_var(r1, r5); // 57 rotr + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 58 load + r3 = r3 - r7; // 59 sub + r1 = r4 * r6 + r1; // 60 mad + r2 = r1 * r2 + r2; // 61 mad + r4 = r4 ^ r3; // 62 xor + r2 = r2 ^ ds[r4 & mask]; // 63 load + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-5/kernel_bound.cu b/proto-cuda/packs-readwidth/mixA-5/kernel_bound.cu new file mode 100644 index 000000000..96a8f86d7 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-5/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/6". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r0 = rotl_imm(r0, 10u); // 0 rotl + r5 = r3 * r1 + r5; // 1 mad + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 2 load + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 3 load + r5 = r2 * r2 + r5; // 4 mad + r4 = r4 + r3 + ((((sel >> 10u) & 1u) != 0u) ? 0x4edf355au : 0xbd587d10u); // 5 add + r3 = r3 ^ r2; // 6 xor + r5 = r5 ^ r3; // 7 xor + r7 = r6 * r2 + r7; // 8 mad + r7 = r6 * r3 + r7; // 9 mad + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 10 load + r1 = __umulhi(r1, r4); // 11 mulhi + r1 = r1 ^ r6; // 12 xor + r2 = __umulhi(r2, r1); // 13 mulhi + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 14 shfl + r4 = r6 * r4 + r4; // 15 mad + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r0, 8); // 16 shfl + r5 = r5 ^ r4; // 17 xor + r0 = r0 * r2; // 18 mul + r6 = r6 ^ ds[r3 & mask]; // 19 load + r6 = __umulhi(r6, r2); // 20 mulhi + r7 = r2 * r6 + r7; // 21 mad + r4 = r4 * r5; // 22 mul + r3 = r3 * r5; // 23 mul + r6 = r6 * r4; // 24 mul + r3 = r3 ^ ds[r7 & mask]; // 25 load + r2 = rotr_var(r2, r5); // 26 rotr + r5 = r5 | r6; // 27 or + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 28 load + r6 = r6 | r0; // 29 or + r0 = r0 + r3 + ((((sel >> 2u) & 1u) != 0u) ? 0xdca976acu : 0x5faa0547u); // 30 add + r6 = __umulhi(r6, r0); // 31 mulhi + r3 = rotr_var(r3, r7); // 32 rotr + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r4, 4); // 33 shfl + r2 = rotl_imm(r2, 17u); // 34 rotl + r7 = __umulhi(r7, r6); // 35 mulhi + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r2, 2); // 36 shfl + r6 = r6 ^ ds[r0 & mask]; // 37 load + { uint32_t b_ = (r6 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 38 load + r1 = r1 ^ r0; // 39 xor + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r1, 2); // 40 shfl + r1 = r1 ^ ds[r5 & mask]; // 41 load + r3 = r3 * r2; // 42 mul + r2 = r2 ^ ds[r7 & mask]; // 43 load + r7 = r7 + r3 + ((((sel >> 7u) & 1u) != 0u) ? 0x610bd3c2u : 0xe28c457cu); // 44 add + r7 = __umulhi(r7, r6); // 45 mulhi + r3 = r3 ^ r4; // 46 xor + r3 = r3 ^ ds[r7 & mask]; // 47 load + r7 = __umulhi(r7, r2); // 48 mulhi + r6 = __umulhi(r6, r3); // 49 mulhi + { uint32_t b_ = (r1 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 50 load + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 51 load + r5 = r5 ^ r3; // 52 xor + r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xcca8fa7cu : 0xd6457f6eu); // 53 add + r7 = rotr_var(r7, r1); // 54 rotr + r6 = r6 + r7 + ((((sel >> 12u) & 1u) != 0u) ? 0x826cb755u : 0xf3e1301cu); // 55 add + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 56 load + r1 = rotr_var(r1, r5); // 57 rotr + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 58 load + r3 = r3 - r7; // 59 sub + r1 = r4 * r6 + r1; // 60 mad + r2 = r1 * r2 + r2; // 61 mad + r4 = r4 ^ r3; // 62 xor + r2 = r2 ^ ds[r4 & mask]; // 63 load + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixA-5/memhard.h b/proto-cuda/packs-readwidth/mixA-5/memhard.h new file mode 100644 index 000000000..f30b14e83 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-5/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/6". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/mixA-5/memhard.metal b/proto-cuda/packs-readwidth/mixA-5/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-5/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/mixA-5/program.h b/proto-cuda/packs-readwidth/mixA-5/program.h new file mode 100644 index 000000000..a27b03489 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-5/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/6". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-readwidth/A/6" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d7265616477696474682f412f36" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0xd9bea03c18e5b8d6ull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 mad=8 mulhi=8 xor=8 add=5 mul=5 shfl=5 rotr=4 or=2 rotl=2 sub=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "mix50-35-15" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 50, 35, 15 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 7, 3, 6 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 3680 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x2a53aad4u, 0x37fcb6e2u, 0x25d66a27u, 0xf5249e5eu, 0x15f86a59u, 0x8144559au, 0xcad741a6u, 0xcb961373u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/mixA-5/program.json b/proto-cuda/packs-readwidth/mixA-5/program.json new file mode 100644 index 000000000..a6f93ce4b --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-5/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0xd9bea03c18e5b8d6", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-readwidth/A/6", + "seed_bytes": "69676e65756d2d7265616477696474682f412f36", + "seed_words": ["0x2a53aad4", "0x37fcb6e2", "0x25d66a27", "0xf5249e5e", "0x15f86a59", "0x8144559a", "0xcad741a6", "0xcb961373"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "mix50-35-15", + "load_slots": 16, + "load_mix_percent_4_16_64": [50, 35, 15], + "load_width_counts_4_16_64": [7, 3, 6], + "bytes_per_hash": 3680, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "mad": 8, "mulhi": 8, "xor": 8, "add": 5, "mul": 5, "shfl": 5, "rotr": 4, "or": 2, "rotl": 2, "sub": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "rotl", "dst": 0, "src": 7, "src2": 6, "imm": "0x2237bb81", "imm2": "0xb069b9f1", "rot": 10, "bit": 18, "mask": 2, "width": 1}, + {"i": 1, "op": "mad", "dst": 5, "src": 3, "src2": 1, "imm": "0x4fd258fa", "imm2": "0x918a4f37", "rot": 1, "bit": 23, "mask": 16, "width": 1}, + {"i": 2, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0xbf959e37", "imm2": "0x365cef9c", "rot": 27, "bit": 21, "mask": 16, "width": 16}, + {"i": 3, "op": "load", "dst": 6, "src": 5, "src2": 5, "imm": "0x48715ea5", "imm2": "0x70bc4e9a", "rot": 21, "bit": 7, "mask": 8, "width": 4}, + {"i": 4, "op": "mad", "dst": 5, "src": 2, "src2": 2, "imm": "0x84131e0c", "imm2": "0xd3744459", "rot": 24, "bit": 28, "mask": 2, "width": 1}, + {"i": 5, "op": "add", "dst": 4, "src": 3, "src2": 4, "imm": "0xbd587d10", "imm2": "0x4edf355a", "rot": 18, "bit": 10, "mask": 8, "width": 1}, + {"i": 6, "op": "xor", "dst": 3, "src": 2, "src2": 3, "imm": "0x41f40748", "imm2": "0x6cfe9987", "rot": 14, "bit": 6, "mask": 16, "width": 1}, + {"i": 7, "op": "xor", "dst": 5, "src": 3, "src2": 4, "imm": "0x2c14f9c9", "imm2": "0x4f37fd66", "rot": 3, "bit": 1, "mask": 1, "width": 1}, + {"i": 8, "op": "mad", "dst": 7, "src": 6, "src2": 2, "imm": "0xcb0b8c69", "imm2": "0x451450b3", "rot": 21, "bit": 7, "mask": 16, "width": 1}, + {"i": 9, "op": "mad", "dst": 7, "src": 6, "src2": 3, "imm": "0x5d258713", "imm2": "0x9822328d", "rot": 30, "bit": 31, "mask": 16, "width": 1}, + {"i": 10, "op": "load", "dst": 3, "src": 5, "src2": 4, "imm": "0x61b87fba", "imm2": "0x51702831", "rot": 20, "bit": 30, "mask": 8, "width": 16}, + {"i": 11, "op": "mulhi", "dst": 1, "src": 4, "src2": 7, "imm": "0xc9e85411", "imm2": "0xe89df827", "rot": 26, "bit": 23, "mask": 8, "width": 1}, + {"i": 12, "op": "xor", "dst": 1, "src": 6, "src2": 6, "imm": "0x1ceb672f", "imm2": "0xf15040b7", "rot": 27, "bit": 19, "mask": 16, "width": 1}, + {"i": 13, "op": "mulhi", "dst": 2, "src": 1, "src2": 3, "imm": "0xe2dcec46", "imm2": "0x6f243314", "rot": 1, "bit": 3, "mask": 4, "width": 1}, + {"i": 14, "op": "shfl", "dst": 6, "src": 7, "src2": 5, "imm": "0x8bac3c39", "imm2": "0x55595cd3", "rot": 1, "bit": 28, "mask": 4, "width": 1}, + {"i": 15, "op": "mad", "dst": 4, "src": 6, "src2": 4, "imm": "0x6d3cc858", "imm2": "0x3c5576ea", "rot": 2, "bit": 21, "mask": 1, "width": 1}, + {"i": 16, "op": "shfl", "dst": 1, "src": 0, "src2": 1, "imm": "0xdd571938", "imm2": "0x044350f3", "rot": 5, "bit": 5, "mask": 8, "width": 1}, + {"i": 17, "op": "xor", "dst": 5, "src": 4, "src2": 4, "imm": "0x2c2d15da", "imm2": "0xb15a5c88", "rot": 31, "bit": 19, "mask": 16, "width": 1}, + {"i": 18, "op": "mul", "dst": 0, "src": 2, "src2": 3, "imm": "0x033348a5", "imm2": "0xd44c3d67", "rot": 1, "bit": 10, "mask": 16, "width": 1}, + {"i": 19, "op": "load", "dst": 6, "src": 3, "src2": 1, "imm": "0x0dda22a1", "imm2": "0x1ae5993b", "rot": 8, "bit": 4, "mask": 2, "width": 1}, + {"i": 20, "op": "mulhi", "dst": 6, "src": 2, "src2": 7, "imm": "0xd48e8694", "imm2": "0x11ed9b7b", "rot": 24, "bit": 19, "mask": 16, "width": 1}, + {"i": 21, "op": "mad", "dst": 7, "src": 2, "src2": 6, "imm": "0x44beb2fe", "imm2": "0x044aa951", "rot": 28, "bit": 22, "mask": 1, "width": 1}, + {"i": 22, "op": "mul", "dst": 4, "src": 5, "src2": 0, "imm": "0xa0256d05", "imm2": "0xd61f8a1a", "rot": 1, "bit": 0, "mask": 16, "width": 1}, + {"i": 23, "op": "mul", "dst": 3, "src": 5, "src2": 3, "imm": "0x0b79eb26", "imm2": "0x7f479e48", "rot": 3, "bit": 0, "mask": 8, "width": 1}, + {"i": 24, "op": "mul", "dst": 6, "src": 4, "src2": 4, "imm": "0x91742ca1", "imm2": "0x9f369d2a", "rot": 19, "bit": 2, "mask": 1, "width": 1}, + {"i": 25, "op": "load", "dst": 3, "src": 7, "src2": 6, "imm": "0x820853df", "imm2": "0xb288d5cd", "rot": 3, "bit": 23, "mask": 8, "width": 1}, + {"i": 26, "op": "rotr", "dst": 2, "src": 5, "src2": 7, "imm": "0x65a24fb5", "imm2": "0xfaaf9fca", "rot": 19, "bit": 12, "mask": 8, "width": 1}, + {"i": 27, "op": "or", "dst": 5, "src": 6, "src2": 4, "imm": "0xb8dc6cbb", "imm2": "0xcac9d5a4", "rot": 8, "bit": 2, "mask": 8, "width": 1}, + {"i": 28, "op": "load", "dst": 4, "src": 2, "src2": 1, "imm": "0x98a93309", "imm2": "0x235dc454", "rot": 31, "bit": 16, "mask": 4, "width": 4}, + {"i": 29, "op": "or", "dst": 6, "src": 0, "src2": 0, "imm": "0xc9dc1885", "imm2": "0xade1c24e", "rot": 9, "bit": 20, "mask": 16, "width": 1}, + {"i": 30, "op": "add", "dst": 0, "src": 3, "src2": 4, "imm": "0x5faa0547", "imm2": "0xdca976ac", "rot": 2, "bit": 2, "mask": 8, "width": 1}, + {"i": 31, "op": "mulhi", "dst": 6, "src": 0, "src2": 7, "imm": "0xec8a588f", "imm2": "0x7e280eb3", "rot": 12, "bit": 19, "mask": 2, "width": 1}, + {"i": 32, "op": "rotr", "dst": 3, "src": 7, "src2": 2, "imm": "0x6deb6b1d", "imm2": "0x79c4c05a", "rot": 25, "bit": 22, "mask": 2, "width": 1}, + {"i": 33, "op": "shfl", "dst": 0, "src": 4, "src2": 4, "imm": "0x3e652de8", "imm2": "0x3a967830", "rot": 31, "bit": 5, "mask": 4, "width": 1}, + {"i": 34, "op": "rotl", "dst": 2, "src": 0, "src2": 7, "imm": "0xd524262b", "imm2": "0xb65f15f3", "rot": 17, "bit": 2, "mask": 4, "width": 1}, + {"i": 35, "op": "mulhi", "dst": 7, "src": 6, "src2": 5, "imm": "0x9c455c39", "imm2": "0x7241afc1", "rot": 9, "bit": 20, "mask": 8, "width": 1}, + {"i": 36, "op": "shfl", "dst": 3, "src": 2, "src2": 4, "imm": "0x78d91acc", "imm2": "0x4efaf581", "rot": 17, "bit": 6, "mask": 2, "width": 1}, + {"i": 37, "op": "load", "dst": 6, "src": 0, "src2": 1, "imm": "0x45f26b60", "imm2": "0xad134516", "rot": 19, "bit": 4, "mask": 4, "width": 1}, + {"i": 38, "op": "load", "dst": 0, "src": 6, "src2": 3, "imm": "0x1798a61a", "imm2": "0xc40c58c6", "rot": 21, "bit": 28, "mask": 4, "width": 16}, + {"i": 39, "op": "xor", "dst": 1, "src": 0, "src2": 4, "imm": "0x862680d2", "imm2": "0x92df77ec", "rot": 12, "bit": 9, "mask": 1, "width": 1}, + {"i": 40, "op": "shfl", "dst": 3, "src": 1, "src2": 4, "imm": "0x44316a4c", "imm2": "0x843d6213", "rot": 6, "bit": 24, "mask": 2, "width": 1}, + {"i": 41, "op": "load", "dst": 1, "src": 5, "src2": 1, "imm": "0xeaeb2a04", "imm2": "0xc4cba010", "rot": 7, "bit": 23, "mask": 16, "width": 1}, + {"i": 42, "op": "mul", "dst": 3, "src": 2, "src2": 5, "imm": "0xe108d4b1", "imm2": "0xf65c52e5", "rot": 3, "bit": 30, "mask": 16, "width": 1}, + {"i": 43, "op": "load", "dst": 2, "src": 7, "src2": 5, "imm": "0x940d6ee3", "imm2": "0x9d6e3a8f", "rot": 27, "bit": 28, "mask": 1, "width": 1}, + {"i": 44, "op": "add", "dst": 7, "src": 3, "src2": 0, "imm": "0xe28c457c", "imm2": "0x610bd3c2", "rot": 9, "bit": 7, "mask": 4, "width": 1}, + {"i": 45, "op": "mulhi", "dst": 7, "src": 6, "src2": 3, "imm": "0xec2679c5", "imm2": "0x522fb2f8", "rot": 4, "bit": 17, "mask": 16, "width": 1}, + {"i": 46, "op": "xor", "dst": 3, "src": 4, "src2": 5, "imm": "0xce5dce0b", "imm2": "0x5447f7da", "rot": 5, "bit": 28, "mask": 2, "width": 1}, + {"i": 47, "op": "load", "dst": 3, "src": 7, "src2": 4, "imm": "0x1be86e58", "imm2": "0xe9e1afb9", "rot": 4, "bit": 17, "mask": 2, "width": 1}, + {"i": 48, "op": "mulhi", "dst": 7, "src": 2, "src2": 1, "imm": "0x26aea9a4", "imm2": "0xa5fa94ce", "rot": 22, "bit": 22, "mask": 16, "width": 1}, + {"i": 49, "op": "mulhi", "dst": 6, "src": 3, "src2": 2, "imm": "0xc25a3528", "imm2": "0x7a974c09", "rot": 31, "bit": 31, "mask": 8, "width": 1}, + {"i": 50, "op": "load", "dst": 2, "src": 1, "src2": 0, "imm": "0x3abaf4ef", "imm2": "0x38d75b93", "rot": 26, "bit": 22, "mask": 16, "width": 16}, + {"i": 51, "op": "load", "dst": 6, "src": 3, "src2": 0, "imm": "0x8b41ee85", "imm2": "0x9171a667", "rot": 19, "bit": 18, "mask": 2, "width": 16}, + {"i": 52, "op": "xor", "dst": 5, "src": 3, "src2": 7, "imm": "0x9c7de4ed", "imm2": "0x72e1df1d", "rot": 18, "bit": 5, "mask": 2, "width": 1}, + {"i": 53, "op": "add", "dst": 5, "src": 7, "src2": 6, "imm": "0xd6457f6e", "imm2": "0xcca8fa7c", "rot": 11, "bit": 31, "mask": 4, "width": 1}, + {"i": 54, "op": "rotr", "dst": 7, "src": 1, "src2": 2, "imm": "0x5ad54c6e", "imm2": "0x0bf13a74", "rot": 2, "bit": 19, "mask": 4, "width": 1}, + {"i": 55, "op": "add", "dst": 6, "src": 7, "src2": 5, "imm": "0xf3e1301c", "imm2": "0x826cb755", "rot": 18, "bit": 12, "mask": 1, "width": 1}, + {"i": 56, "op": "load", "dst": 3, "src": 0, "src2": 6, "imm": "0x68c04498", "imm2": "0x09d73293", "rot": 31, "bit": 9, "mask": 16, "width": 16}, + {"i": 57, "op": "rotr", "dst": 1, "src": 5, "src2": 0, "imm": "0x6ae5f1c4", "imm2": "0x4acb3d16", "rot": 2, "bit": 25, "mask": 2, "width": 1}, + {"i": 58, "op": "load", "dst": 2, "src": 6, "src2": 3, "imm": "0x908dad61", "imm2": "0x3caa6b5c", "rot": 26, "bit": 23, "mask": 8, "width": 4}, + {"i": 59, "op": "sub", "dst": 3, "src": 7, "src2": 4, "imm": "0xf497693d", "imm2": "0x12a59627", "rot": 6, "bit": 15, "mask": 16, "width": 1}, + {"i": 60, "op": "mad", "dst": 1, "src": 4, "src2": 6, "imm": "0x87197b1e", "imm2": "0x9fe5c5a6", "rot": 5, "bit": 18, "mask": 1, "width": 1}, + {"i": 61, "op": "mad", "dst": 2, "src": 1, "src2": 2, "imm": "0x7ffc05da", "imm2": "0xdd4b88aa", "rot": 15, "bit": 8, "mask": 2, "width": 1}, + {"i": 62, "op": "xor", "dst": 4, "src": 3, "src2": 0, "imm": "0x660f3bb9", "imm2": "0x48d985a9", "rot": 2, "bit": 27, "mask": 1, "width": 1}, + {"i": 63, "op": "load", "dst": 2, "src": 4, "src2": 4, "imm": "0x45650a4c", "imm2": "0x4d838ea4", "rot": 15, "bit": 13, "mask": 1, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/mixA-5/program.metal b/proto-cuda/packs-readwidth/mixA-5/program.metal new file mode 100644 index 000000000..36a1e76c1 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-5/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x2a53aad4u, 0x37fcb6e2u, 0x25d66a27u, 0xf5249e5eu, 0x15f86a59u, 0x8144559au, 0xcad741a6u, 0xcb961373u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = rotl_imm(r0, 10u); // 0 + r5 = r3 * r1 + r5; // 1 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 2 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 3 + r5 = r2 * r2 + r5; // 4 + r4 = r4 + r3 + select(0xbd587d10u, 0x4edf355au, ((sel >> 10u) & 1u) != 0u); // 5 + r3 = r3 ^ r2; // 6 + r5 = r5 ^ r3; // 7 + r7 = r6 * r2 + r7; // 8 + r7 = r6 * r3 + r7; // 9 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 10 + r1 = mulhi(r1, r4); // 11 + r1 = r1 ^ r6; // 12 + r2 = mulhi(r2, r1); // 13 + r6 = r6 ^ simd_shuffle_xor(r7, (ushort)4); // 14 + r4 = r6 * r4 + r4; // 15 + r1 = r1 ^ simd_shuffle_xor(r0, (ushort)8); // 16 + r5 = r5 ^ r4; // 17 + r0 = r0 * r2; // 18 + r6 = r6 ^ dataset[r3 & MASK]; // 19 + r6 = mulhi(r6, r2); // 20 + r7 = r2 * r6 + r7; // 21 + r4 = r4 * r5; // 22 + r3 = r3 * r5; // 23 + r6 = r6 * r4; // 24 + r3 = r3 ^ dataset[r7 & MASK]; // 25 + r2 = rotr_var(r2, r5); // 26 + r5 = r5 | r6; // 27 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 28 + r6 = r6 | r0; // 29 + r0 = r0 + r3 + select(0x5faa0547u, 0xdca976acu, ((sel >> 2u) & 1u) != 0u); // 30 + r6 = mulhi(r6, r0); // 31 + r3 = rotr_var(r3, r7); // 32 + r0 = r0 ^ simd_shuffle_xor(r4, (ushort)4); // 33 + r2 = rotl_imm(r2, 17u); // 34 + r7 = mulhi(r7, r6); // 35 + r3 = r3 ^ simd_shuffle_xor(r2, (ushort)2); // 36 + r6 = r6 ^ dataset[r0 & MASK]; // 37 + { uint b_ = (r6 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 38 + r1 = r1 ^ r0; // 39 + r3 = r3 ^ simd_shuffle_xor(r1, (ushort)2); // 40 + r1 = r1 ^ dataset[r5 & MASK]; // 41 + r3 = r3 * r2; // 42 + r2 = r2 ^ dataset[r7 & MASK]; // 43 + r7 = r7 + r3 + select(0xe28c457cu, 0x610bd3c2u, ((sel >> 7u) & 1u) != 0u); // 44 + r7 = mulhi(r7, r6); // 45 + r3 = r3 ^ r4; // 46 + r3 = r3 ^ dataset[r7 & MASK]; // 47 + r7 = mulhi(r7, r2); // 48 + r6 = mulhi(r6, r3); // 49 + { uint b_ = (r1 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 50 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 51 + r5 = r5 ^ r3; // 52 + r5 = r5 + r7 + select(0xd6457f6eu, 0xcca8fa7cu, ((sel >> 31u) & 1u) != 0u); // 53 + r7 = rotr_var(r7, r1); // 54 + r6 = r6 + r7 + select(0xf3e1301cu, 0x826cb755u, ((sel >> 12u) & 1u) != 0u); // 55 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 56 + r1 = rotr_var(r1, r5); // 57 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 58 + r3 = r3 - r7; // 59 + r1 = r4 * r6 + r1; // 60 + r2 = r1 * r2 + r2; // 61 + r4 = r4 ^ r3; // 62 + r2 = r2 ^ dataset[r4 & MASK]; // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-5/program_bound.metal b/proto-cuda/packs-readwidth/mixA-5/program_bound.metal new file mode 100644 index 000000000..bd28f47d0 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-5/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x2a53aad4u, 0x37fcb6e2u, 0x25d66a27u, 0xf5249e5eu, 0x15f86a59u, 0x8144559au, 0xcad741a6u, 0xcb961373u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = rotl_imm(r0, 10u); // 0 + r5 = r3 * r1 + r5; // 1 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 2 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 3 + r5 = r2 * r2 + r5; // 4 + r4 = r4 + r3 + select(0xbd587d10u, 0x4edf355au, ((sel >> 10u) & 1u) != 0u); // 5 + r3 = r3 ^ r2; // 6 + r5 = r5 ^ r3; // 7 + r7 = r6 * r2 + r7; // 8 + r7 = r6 * r3 + r7; // 9 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 10 + r1 = mulhi(r1, r4); // 11 + r1 = r1 ^ r6; // 12 + r2 = mulhi(r2, r1); // 13 + r6 = r6 ^ simd_shuffle_xor(r7, (ushort)4); // 14 + r4 = r6 * r4 + r4; // 15 + r1 = r1 ^ simd_shuffle_xor(r0, (ushort)8); // 16 + r5 = r5 ^ r4; // 17 + r0 = r0 * r2; // 18 + r6 = r6 ^ dataset[r3 & MASK]; // 19 + r6 = mulhi(r6, r2); // 20 + r7 = r2 * r6 + r7; // 21 + r4 = r4 * r5; // 22 + r3 = r3 * r5; // 23 + r6 = r6 * r4; // 24 + r3 = r3 ^ dataset[r7 & MASK]; // 25 + r2 = rotr_var(r2, r5); // 26 + r5 = r5 | r6; // 27 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 28 + r6 = r6 | r0; // 29 + r0 = r0 + r3 + select(0x5faa0547u, 0xdca976acu, ((sel >> 2u) & 1u) != 0u); // 30 + r6 = mulhi(r6, r0); // 31 + r3 = rotr_var(r3, r7); // 32 + r0 = r0 ^ simd_shuffle_xor(r4, (ushort)4); // 33 + r2 = rotl_imm(r2, 17u); // 34 + r7 = mulhi(r7, r6); // 35 + r3 = r3 ^ simd_shuffle_xor(r2, (ushort)2); // 36 + r6 = r6 ^ dataset[r0 & MASK]; // 37 + { uint b_ = (r6 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 38 + r1 = r1 ^ r0; // 39 + r3 = r3 ^ simd_shuffle_xor(r1, (ushort)2); // 40 + r1 = r1 ^ dataset[r5 & MASK]; // 41 + r3 = r3 * r2; // 42 + r2 = r2 ^ dataset[r7 & MASK]; // 43 + r7 = r7 + r3 + select(0xe28c457cu, 0x610bd3c2u, ((sel >> 7u) & 1u) != 0u); // 44 + r7 = mulhi(r7, r6); // 45 + r3 = r3 ^ r4; // 46 + r3 = r3 ^ dataset[r7 & MASK]; // 47 + r7 = mulhi(r7, r2); // 48 + r6 = mulhi(r6, r3); // 49 + { uint b_ = (r1 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 50 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 51 + r5 = r5 ^ r3; // 52 + r5 = r5 + r7 + select(0xd6457f6eu, 0xcca8fa7cu, ((sel >> 31u) & 1u) != 0u); // 53 + r7 = rotr_var(r7, r1); // 54 + r6 = r6 + r7 + select(0xf3e1301cu, 0x826cb755u, ((sel >> 12u) & 1u) != 0u); // 55 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 56 + r1 = rotr_var(r1, r5); // 57 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 58 + r3 = r3 - r7; // 59 + r1 = r4 * r6 + r1; // 60 + r2 = r1 * r2 + r2; // 61 + r4 = r4 ^ r3; // 62 + r2 = r2 ^ dataset[r4 & MASK]; // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixA-5/vectors.h b/proto-cuda/packs-readwidth/mixA-5/vectors.h new file mode 100644 index 000000000..0c1d64c82 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-5/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/A/6". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x93e8388cac40c811ull, 0xa7fc8484877d2d42ull, 0x0e4fade6c54c3828ull, 0xd51024c9edc222c6ull, 0xf69126345bc6a789ull, 0xe4bf38ace928f612ull, 0xa3915889cbe42924ull, 0x338fb75c6a087e3dull, + 0xf27481a4ff979f3bull, 0xefe50d7bc95581abull, 0x7d2fa9d720b02ecaull, 0xfca2c211f788acf2ull, 0x2e66cc8095f3cb16ull, 0xae75758d3dc34127ull, 0x9553b782a362e16eull, 0xcb22382b49172a28ull, + 0x5fa9a18442efb5d8ull, 0x8bdd7aceb04694f1ull, 0xd7b19cf5b3dd0d2eull, 0x2f704fb0b2342d83ull, 0x8c3747e33ccefd96ull, 0x67c00a117dfc2e3cull, 0x3d548ce336a6c72full, 0xfca49cfec0f68f8full, + 0x1c253507e5d1a81cull, 0xa17ebb784aa17028ull, 0x511b1f8191bbf57full, 0xabfd11f3a4eac2c8ull, 0x729a42cfe7386d86ull, 0xfa1d04c46ce4f2c2ull, 0xf108ab5b9481c6b0ull, 0xe9fe2f9e682f40d8ull + }, + { // base nonce 4096 + 0xc5d7b8097741686bull, 0x7fae5d95c0939293ull, 0x9b115b4967fd0af2ull, 0xf347d09f0b570b6bull, 0xb611632e5e2cc724ull, 0x92a6046c22ce9826ull, 0x6a40216015aeabe7ull, 0x8c5b7bbf98ee8ad1ull, + 0x160da4c87f1ff350ull, 0xf898efb8f2ecbf70ull, 0x84ec67d175b97164ull, 0x5813809116708d53ull, 0xe5318deef160edcfull, 0x1d3a703990fe850eull, 0xf6a0b78ee2dc77adull, 0xfddb5e0da4c705d5ull, + 0x98af1bd6cbba216full, 0xd339b38e62f2f4a5ull, 0xa6ac444d1edabc78ull, 0xe37432779a991a79ull, 0x2e10f0ac4e7a00d6ull, 0xf948234a7b8a4b92ull, 0x44c7abbb3897dd06ull, 0x14c7287af3dcb661ull, + 0x9b6b3d1c5b2a53f8ull, 0x36f4b115baec85acull, 0x41192bfbaba0592full, 0x0a6a7e78a97065d0ull, 0x54a2f48437f43f94ull, 0x5c7b3a5411458527ull, 0x359ac92b685d331dull, 0x2746c6f53c998ac1ull + }, + { // base nonce 1000000 + 0x299e126aefa2a456ull, 0xd7cb35bb207a4744ull, 0x28755e512456e322ull, 0x7e12f9a852ed67b7ull, 0x54ff55e7b8fdf7f7ull, 0xb36a09f8c0fefd19ull, 0xb466f445db516c51ull, 0xd129d222de67ce0dull, + 0xc8a4d0e8591982a8ull, 0xb7566ad957bb1e1dull, 0xa56457ebd28d83efull, 0xd981708266e400c0ull, 0x73c629bbcedab2d9ull, 0x0d5fb3aa79870efbull, 0xf225b95724a24e46ull, 0xdf8d23b7472004bfull, + 0x4ca04ddf25cd4aceull, 0x80864b3b0740f888ull, 0xc114ed0d7042660aull, 0x8a3c7c3bcb5575c3ull, 0xa46c25ba69798388ull, 0x33a8aeab66d63875ull, 0x5bea09d563955e97ull, 0x81e81e3c8ea6d163ull, + 0x09f4dca8200b8220ull, 0x0227171d66f5197dull, 0x9d5ae8c2fe24c1abull, 0x572b88825c78884dull, 0x33be787f0b6d52fcull, 0x88689ec5df299a7full, 0xfad078a36872646eull, 0x8d7439bc62c19870ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/mixA-5/vectors.json b/proto-cuda/packs-readwidth/mixA-5/vectors.json new file mode 100644 index 000000000..cbaadbe40 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixA-5/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-readwidth/A/6", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x93e8388cac40c811", "0xa7fc8484877d2d42", "0x0e4fade6c54c3828", "0xd51024c9edc222c6", "0xf69126345bc6a789", "0xe4bf38ace928f612", "0xa3915889cbe42924", "0x338fb75c6a087e3d", + "0xf27481a4ff979f3b", "0xefe50d7bc95581ab", "0x7d2fa9d720b02eca", "0xfca2c211f788acf2", "0x2e66cc8095f3cb16", "0xae75758d3dc34127", "0x9553b782a362e16e", "0xcb22382b49172a28", + "0x5fa9a18442efb5d8", "0x8bdd7aceb04694f1", "0xd7b19cf5b3dd0d2e", "0x2f704fb0b2342d83", "0x8c3747e33ccefd96", "0x67c00a117dfc2e3c", "0x3d548ce336a6c72f", "0xfca49cfec0f68f8f", + "0x1c253507e5d1a81c", "0xa17ebb784aa17028", "0x511b1f8191bbf57f", "0xabfd11f3a4eac2c8", "0x729a42cfe7386d86", "0xfa1d04c46ce4f2c2", "0xf108ab5b9481c6b0", "0xe9fe2f9e682f40d8" + ]}, + {"base_nonce": 4096, "expected": [ + "0xc5d7b8097741686b", "0x7fae5d95c0939293", "0x9b115b4967fd0af2", "0xf347d09f0b570b6b", "0xb611632e5e2cc724", "0x92a6046c22ce9826", "0x6a40216015aeabe7", "0x8c5b7bbf98ee8ad1", + "0x160da4c87f1ff350", "0xf898efb8f2ecbf70", "0x84ec67d175b97164", "0x5813809116708d53", "0xe5318deef160edcf", "0x1d3a703990fe850e", "0xf6a0b78ee2dc77ad", "0xfddb5e0da4c705d5", + "0x98af1bd6cbba216f", "0xd339b38e62f2f4a5", "0xa6ac444d1edabc78", "0xe37432779a991a79", "0x2e10f0ac4e7a00d6", "0xf948234a7b8a4b92", "0x44c7abbb3897dd06", "0x14c7287af3dcb661", + "0x9b6b3d1c5b2a53f8", "0x36f4b115baec85ac", "0x41192bfbaba0592f", "0x0a6a7e78a97065d0", "0x54a2f48437f43f94", "0x5c7b3a5411458527", "0x359ac92b685d331d", "0x2746c6f53c998ac1" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x299e126aefa2a456", "0xd7cb35bb207a4744", "0x28755e512456e322", "0x7e12f9a852ed67b7", "0x54ff55e7b8fdf7f7", "0xb36a09f8c0fefd19", "0xb466f445db516c51", "0xd129d222de67ce0d", + "0xc8a4d0e8591982a8", "0xb7566ad957bb1e1d", "0xa56457ebd28d83ef", "0xd981708266e400c0", "0x73c629bbcedab2d9", "0x0d5fb3aa79870efb", "0xf225b95724a24e46", "0xdf8d23b7472004bf", + "0x4ca04ddf25cd4ace", "0x80864b3b0740f888", "0xc114ed0d7042660a", "0x8a3c7c3bcb5575c3", "0xa46c25ba69798388", "0x33a8aeab66d63875", "0x5bea09d563955e97", "0x81e81e3c8ea6d163", + "0x09f4dca8200b8220", "0x0227171d66f5197d", "0x9d5ae8c2fe24c1ab", "0x572b88825c78884d", "0x33be787f0b6d52fc", "0x88689ec5df299a7f", "0xfad078a36872646e", "0x8d7439bc62c19870" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/mixB-0/kernel.cl b/proto-cuda/packs-readwidth/mixB-0/kernel.cl new file mode 100644 index 000000000..bb1510fe3 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-0/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/0". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x3673211cu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0xaae550b4u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0xaae550b4u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x5a0ce2e0u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x5a0ce2e0u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x1d2471cfu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x1d2471cfu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xba944366u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xba944366u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xbdd4d1dbu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xbdd4d1dbu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xcb1984a9u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xcb1984a9u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x081ce12au; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x081ce12au; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x3673211cu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 1u); r0 = r0 ^ t_; } // 0 shfl + r2 = r2 * r0; // 1 mul + r0 = r0 + r6 + ((((sel >> 20u) & 1u) != 0u) ? 0x03fa29f3u : 0x7fbf4ae6u); // 2 add + r2 = r2 * r7; // 3 mul + r6 = r6 ^ ds[r2 & mask]; // 4 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r6 = r6 ^ t_; } // 6 shfl + r3 = r3 ^ r5; // 7 xor + r6 = r6 ^ ds[r7 & mask]; // 8 load + r0 = r0 + r7 + ((((sel >> 25u) & 1u) != 0u) ? 0x308c81bfu : 0xd59b21b0u); // 9 add + r5 = rotr_var(r5, r7); // 10 rotr + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 load + r2 = r2 * r6; // 12 mul + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 13 load + r2 = r5 * r2 + r2; // 14 mad + r4 = r4 + r0 + ((((sel >> 24u) & 1u) != 0u) ? 0x67081e7eu : 0x902dc661u); // 15 add + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 load + r3 = mul_hi(r3, r4); // 17 mulhi + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 18 load + r2 = r2 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0xd1e74db0u : 0x3ee3182cu); // 19 add + r7 = r7 ^ r6; // 20 xor + r5 = r5 ^ r3; // 21 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 16u); r6 = r6 ^ t_; } // 22 shfl + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 23 load + r2 = r2 + r5 + ((((sel >> 22u) & 1u) != 0u) ? 0x62ac9e52u : 0xd2d451c6u); // 24 add + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 25 load + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 16u); r4 = r4 ^ t_; } // 26 shfl + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 27 load + r1 = r1 + r5 + ((((sel >> 2u) & 1u) != 0u) ? 0x072cfabeu : 0x111ff813u); // 28 add + r4 = r1 * r2 + r4; // 29 mad + r2 = r2 * r0; // 30 mul + r0 = r0 ^ r5; // 31 xor + r1 = r1 ^ ds[r0 & mask]; // 32 load + r2 = r2 - r3; // 33 sub + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 2u); r2 = r2 ^ t_; } // 34 shfl + r0 = r0 - r6; // 35 sub + r4 = r4 * r7; // 36 mul + r5 = r5 ^ r6; // 37 xor + r0 = mul_hi(r0, r6); // 38 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r5 = r5 ^ t_; } // 39 shfl + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 40 load + r2 = r2 ^ ds[r4 & mask]; // 41 load + r5 = r5 ^ ds[r6 & mask]; // 42 load + r1 = rotr_var(r1, r4); // 43 rotr + r0 = r0 | r5; // 44 or + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r0 = r0 ^ t_; } // 45 shfl + r7 = r7 * r4; // 46 mul + r3 = r3 ^ r4; // 47 xor + r2 = mul_hi(r2, r6); // 48 mulhi + r5 = r5 ^ r4; // 49 xor + r1 = rotr_var(r1, r6); // 50 rotr + r7 = r7 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0x0fe79cecu : 0x68d3a5c0u); // 51 add + r5 = r5 + r4 + ((((sel >> 8u) & 1u) != 0u) ? 0x04309e6fu : 0xeb764d91u); // 52 add + r7 = r7 ^ r0; // 53 xor + r4 = r4 + r0 + ((((sel >> 6u) & 1u) != 0u) ? 0xba0b81c6u : 0x1390b188u); // 54 add + r0 = r0 + r5 + ((((sel >> 15u) & 1u) != 0u) ? 0x5ec96dd8u : 0x4c21a6cdu); // 55 add + r7 = r5 * r3 + r7; // 56 mad + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 57 load + r5 = r5 + r1 + ((((sel >> 1u) & 1u) != 0u) ? 0x6c0ac4ddu : 0x68297b06u); // 58 add + r6 = rotr_var(r6, r7); // 59 rotr + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 60 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 61 shfl + r4 = rotl_imm(r4, 6u); // 62 rotl + r1 = r2 * r1 + r1; // 63 mad + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/mixB-0/kernel.cu b/proto-cuda/packs-readwidth/mixB-0/kernel.cu new file mode 100644 index 000000000..84debc12d --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-0/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/0". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x3673211cu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0xaae550b4u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0xaae550b4u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x5a0ce2e0u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0x5a0ce2e0u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x1d2471cfu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x1d2471cfu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xba944366u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xba944366u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xbdd4d1dbu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xbdd4d1dbu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xcb1984a9u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0xcb1984a9u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x081ce12au; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x081ce12au; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x3673211cu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r2, 1); // 0 shfl + r2 = r2 * r0; // 1 mul + r0 = r0 + r6 + ((((sel >> 20u) & 1u) != 0u) ? 0x03fa29f3u : 0x7fbf4ae6u); // 2 add + r2 = r2 * r7; // 3 mul + r6 = r6 ^ ds[r2 & mask]; // 4 load + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 6 shfl + r3 = r3 ^ r5; // 7 xor + r6 = r6 ^ ds[r7 & mask]; // 8 load + r0 = r0 + r7 + ((((sel >> 25u) & 1u) != 0u) ? 0x308c81bfu : 0xd59b21b0u); // 9 add + r5 = rotr_var(r5, r7); // 10 rotr + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 load + r2 = r2 * r6; // 12 mul + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 13 load + r2 = r5 * r2 + r2; // 14 mad + r4 = r4 + r0 + ((((sel >> 24u) & 1u) != 0u) ? 0x67081e7eu : 0x902dc661u); // 15 add + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 load + r3 = __umulhi(r3, r4); // 17 mulhi + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 18 load + r2 = r2 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0xd1e74db0u : 0x3ee3182cu); // 19 add + r7 = r7 ^ r6; // 20 xor + r5 = r5 ^ r3; // 21 xor + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r2, 16); // 22 shfl + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 23 load + r2 = r2 + r5 + ((((sel >> 22u) & 1u) != 0u) ? 0x62ac9e52u : 0xd2d451c6u); // 24 add + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 25 load + r4 = r4 ^ __shfl_xor_sync(0xffffffffu, r1, 16); // 26 shfl + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 27 load + r1 = r1 + r5 + ((((sel >> 2u) & 1u) != 0u) ? 0x072cfabeu : 0x111ff813u); // 28 add + r4 = r1 * r2 + r4; // 29 mad + r2 = r2 * r0; // 30 mul + r0 = r0 ^ r5; // 31 xor + r1 = r1 ^ ds[r0 & mask]; // 32 load + r2 = r2 - r3; // 33 sub + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 2); // 34 shfl + r0 = r0 - r6; // 35 sub + r4 = r4 * r7; // 36 mul + r5 = r5 ^ r6; // 37 xor + r0 = __umulhi(r0, r6); // 38 mulhi + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 39 shfl + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 40 load + r2 = r2 ^ ds[r4 & mask]; // 41 load + r5 = r5 ^ ds[r6 & mask]; // 42 load + r1 = rotr_var(r1, r4); // 43 rotr + r0 = r0 | r5; // 44 or + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r2, 2); // 45 shfl + r7 = r7 * r4; // 46 mul + r3 = r3 ^ r4; // 47 xor + r2 = __umulhi(r2, r6); // 48 mulhi + r5 = r5 ^ r4; // 49 xor + r1 = rotr_var(r1, r6); // 50 rotr + r7 = r7 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0x0fe79cecu : 0x68d3a5c0u); // 51 add + r5 = r5 + r4 + ((((sel >> 8u) & 1u) != 0u) ? 0x04309e6fu : 0xeb764d91u); // 52 add + r7 = r7 ^ r0; // 53 xor + r4 = r4 + r0 + ((((sel >> 6u) & 1u) != 0u) ? 0xba0b81c6u : 0x1390b188u); // 54 add + r0 = r0 + r5 + ((((sel >> 15u) & 1u) != 0u) ? 0x5ec96dd8u : 0x4c21a6cdu); // 55 add + r7 = r5 * r3 + r7; // 56 mad + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 57 load + r5 = r5 + r1 + ((((sel >> 1u) & 1u) != 0u) ? 0x6c0ac4ddu : 0x68297b06u); // 58 add + r6 = rotr_var(r6, r7); // 59 rotr + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 60 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 61 shfl + r4 = rotl_imm(r4, 6u); // 62 rotl + r1 = r2 * r1 + r1; // 63 mad + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixB-0/kernel_bound.cl b/proto-cuda/packs-readwidth/mixB-0/kernel_bound.cl new file mode 100644 index 000000000..2c0425775 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-0/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/0". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x3673211cu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0xaae550b4u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0xaae550b4u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x5a0ce2e0u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x5a0ce2e0u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x1d2471cfu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x1d2471cfu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xba944366u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xba944366u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xbdd4d1dbu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xbdd4d1dbu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xcb1984a9u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xcb1984a9u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x081ce12au; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x081ce12au; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x3673211cu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 1u); r0 = r0 ^ t_; } // 0 shfl + r2 = r2 * r0; // 1 mul + r0 = r0 + r6 + ((((sel >> 20u) & 1u) != 0u) ? 0x03fa29f3u : 0x7fbf4ae6u); // 2 add + r2 = r2 * r7; // 3 mul + r6 = r6 ^ ds[r2 & mask]; // 4 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r6 = r6 ^ t_; } // 6 shfl + r3 = r3 ^ r5; // 7 xor + r6 = r6 ^ ds[r7 & mask]; // 8 load + r0 = r0 + r7 + ((((sel >> 25u) & 1u) != 0u) ? 0x308c81bfu : 0xd59b21b0u); // 9 add + r5 = rotr_var(r5, r7); // 10 rotr + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 load + r2 = r2 * r6; // 12 mul + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 13 load + r2 = r5 * r2 + r2; // 14 mad + r4 = r4 + r0 + ((((sel >> 24u) & 1u) != 0u) ? 0x67081e7eu : 0x902dc661u); // 15 add + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 load + r3 = mul_hi(r3, r4); // 17 mulhi + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 18 load + r2 = r2 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0xd1e74db0u : 0x3ee3182cu); // 19 add + r7 = r7 ^ r6; // 20 xor + r5 = r5 ^ r3; // 21 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 16u); r6 = r6 ^ t_; } // 22 shfl + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 23 load + r2 = r2 + r5 + ((((sel >> 22u) & 1u) != 0u) ? 0x62ac9e52u : 0xd2d451c6u); // 24 add + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 25 load + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 16u); r4 = r4 ^ t_; } // 26 shfl + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 27 load + r1 = r1 + r5 + ((((sel >> 2u) & 1u) != 0u) ? 0x072cfabeu : 0x111ff813u); // 28 add + r4 = r1 * r2 + r4; // 29 mad + r2 = r2 * r0; // 30 mul + r0 = r0 ^ r5; // 31 xor + r1 = r1 ^ ds[r0 & mask]; // 32 load + r2 = r2 - r3; // 33 sub + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 2u); r2 = r2 ^ t_; } // 34 shfl + r0 = r0 - r6; // 35 sub + r4 = r4 * r7; // 36 mul + r5 = r5 ^ r6; // 37 xor + r0 = mul_hi(r0, r6); // 38 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r5 = r5 ^ t_; } // 39 shfl + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 40 load + r2 = r2 ^ ds[r4 & mask]; // 41 load + r5 = r5 ^ ds[r6 & mask]; // 42 load + r1 = rotr_var(r1, r4); // 43 rotr + r0 = r0 | r5; // 44 or + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r0 = r0 ^ t_; } // 45 shfl + r7 = r7 * r4; // 46 mul + r3 = r3 ^ r4; // 47 xor + r2 = mul_hi(r2, r6); // 48 mulhi + r5 = r5 ^ r4; // 49 xor + r1 = rotr_var(r1, r6); // 50 rotr + r7 = r7 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0x0fe79cecu : 0x68d3a5c0u); // 51 add + r5 = r5 + r4 + ((((sel >> 8u) & 1u) != 0u) ? 0x04309e6fu : 0xeb764d91u); // 52 add + r7 = r7 ^ r0; // 53 xor + r4 = r4 + r0 + ((((sel >> 6u) & 1u) != 0u) ? 0xba0b81c6u : 0x1390b188u); // 54 add + r0 = r0 + r5 + ((((sel >> 15u) & 1u) != 0u) ? 0x5ec96dd8u : 0x4c21a6cdu); // 55 add + r7 = r5 * r3 + r7; // 56 mad + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 57 load + r5 = r5 + r1 + ((((sel >> 1u) & 1u) != 0u) ? 0x6c0ac4ddu : 0x68297b06u); // 58 add + r6 = rotr_var(r6, r7); // 59 rotr + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 60 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 61 shfl + r4 = rotl_imm(r4, 6u); // 62 rotl + r1 = r2 * r1 + r1; // 63 mad + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 1u); r0 = r0 ^ t_; } // 0 shfl + r2 = r2 * r0; // 1 mul + r0 = r0 + r6 + ((((sel >> 20u) & 1u) != 0u) ? 0x03fa29f3u : 0x7fbf4ae6u); // 2 add + r2 = r2 * r7; // 3 mul + r6 = r6 ^ ds[r2 & mask]; // 4 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r6 = r6 ^ t_; } // 6 shfl + r3 = r3 ^ r5; // 7 xor + r6 = r6 ^ ds[r7 & mask]; // 8 load + r0 = r0 + r7 + ((((sel >> 25u) & 1u) != 0u) ? 0x308c81bfu : 0xd59b21b0u); // 9 add + r5 = rotr_var(r5, r7); // 10 rotr + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 load + r2 = r2 * r6; // 12 mul + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 13 load + r2 = r5 * r2 + r2; // 14 mad + r4 = r4 + r0 + ((((sel >> 24u) & 1u) != 0u) ? 0x67081e7eu : 0x902dc661u); // 15 add + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 load + r3 = mul_hi(r3, r4); // 17 mulhi + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 18 load + r2 = r2 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0xd1e74db0u : 0x3ee3182cu); // 19 add + r7 = r7 ^ r6; // 20 xor + r5 = r5 ^ r3; // 21 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 16u); r6 = r6 ^ t_; } // 22 shfl + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 23 load + r2 = r2 + r5 + ((((sel >> 22u) & 1u) != 0u) ? 0x62ac9e52u : 0xd2d451c6u); // 24 add + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 25 load + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 16u); r4 = r4 ^ t_; } // 26 shfl + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 27 load + r1 = r1 + r5 + ((((sel >> 2u) & 1u) != 0u) ? 0x072cfabeu : 0x111ff813u); // 28 add + r4 = r1 * r2 + r4; // 29 mad + r2 = r2 * r0; // 30 mul + r0 = r0 ^ r5; // 31 xor + r1 = r1 ^ ds[r0 & mask]; // 32 load + r2 = r2 - r3; // 33 sub + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 2u); r2 = r2 ^ t_; } // 34 shfl + r0 = r0 - r6; // 35 sub + r4 = r4 * r7; // 36 mul + r5 = r5 ^ r6; // 37 xor + r0 = mul_hi(r0, r6); // 38 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r5 = r5 ^ t_; } // 39 shfl + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 40 load + r2 = r2 ^ ds[r4 & mask]; // 41 load + r5 = r5 ^ ds[r6 & mask]; // 42 load + r1 = rotr_var(r1, r4); // 43 rotr + r0 = r0 | r5; // 44 or + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r0 = r0 ^ t_; } // 45 shfl + r7 = r7 * r4; // 46 mul + r3 = r3 ^ r4; // 47 xor + r2 = mul_hi(r2, r6); // 48 mulhi + r5 = r5 ^ r4; // 49 xor + r1 = rotr_var(r1, r6); // 50 rotr + r7 = r7 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0x0fe79cecu : 0x68d3a5c0u); // 51 add + r5 = r5 + r4 + ((((sel >> 8u) & 1u) != 0u) ? 0x04309e6fu : 0xeb764d91u); // 52 add + r7 = r7 ^ r0; // 53 xor + r4 = r4 + r0 + ((((sel >> 6u) & 1u) != 0u) ? 0xba0b81c6u : 0x1390b188u); // 54 add + r0 = r0 + r5 + ((((sel >> 15u) & 1u) != 0u) ? 0x5ec96dd8u : 0x4c21a6cdu); // 55 add + r7 = r5 * r3 + r7; // 56 mad + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 57 load + r5 = r5 + r1 + ((((sel >> 1u) & 1u) != 0u) ? 0x6c0ac4ddu : 0x68297b06u); // 58 add + r6 = rotr_var(r6, r7); // 59 rotr + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 60 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 61 shfl + r4 = rotl_imm(r4, 6u); // 62 rotl + r1 = r2 * r1 + r1; // 63 mad + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-0/kernel_bound.cu b/proto-cuda/packs-readwidth/mixB-0/kernel_bound.cu new file mode 100644 index 000000000..e80cfbaa5 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-0/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/0". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r2, 1); // 0 shfl + r2 = r2 * r0; // 1 mul + r0 = r0 + r6 + ((((sel >> 20u) & 1u) != 0u) ? 0x03fa29f3u : 0x7fbf4ae6u); // 2 add + r2 = r2 * r7; // 3 mul + r6 = r6 ^ ds[r2 & mask]; // 4 load + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 6 shfl + r3 = r3 ^ r5; // 7 xor + r6 = r6 ^ ds[r7 & mask]; // 8 load + r0 = r0 + r7 + ((((sel >> 25u) & 1u) != 0u) ? 0x308c81bfu : 0xd59b21b0u); // 9 add + r5 = rotr_var(r5, r7); // 10 rotr + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 load + r2 = r2 * r6; // 12 mul + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 13 load + r2 = r5 * r2 + r2; // 14 mad + r4 = r4 + r0 + ((((sel >> 24u) & 1u) != 0u) ? 0x67081e7eu : 0x902dc661u); // 15 add + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 load + r3 = __umulhi(r3, r4); // 17 mulhi + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 18 load + r2 = r2 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0xd1e74db0u : 0x3ee3182cu); // 19 add + r7 = r7 ^ r6; // 20 xor + r5 = r5 ^ r3; // 21 xor + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r2, 16); // 22 shfl + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 23 load + r2 = r2 + r5 + ((((sel >> 22u) & 1u) != 0u) ? 0x62ac9e52u : 0xd2d451c6u); // 24 add + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 25 load + r4 = r4 ^ __shfl_xor_sync(0xffffffffu, r1, 16); // 26 shfl + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 27 load + r1 = r1 + r5 + ((((sel >> 2u) & 1u) != 0u) ? 0x072cfabeu : 0x111ff813u); // 28 add + r4 = r1 * r2 + r4; // 29 mad + r2 = r2 * r0; // 30 mul + r0 = r0 ^ r5; // 31 xor + r1 = r1 ^ ds[r0 & mask]; // 32 load + r2 = r2 - r3; // 33 sub + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 2); // 34 shfl + r0 = r0 - r6; // 35 sub + r4 = r4 * r7; // 36 mul + r5 = r5 ^ r6; // 37 xor + r0 = __umulhi(r0, r6); // 38 mulhi + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 39 shfl + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 40 load + r2 = r2 ^ ds[r4 & mask]; // 41 load + r5 = r5 ^ ds[r6 & mask]; // 42 load + r1 = rotr_var(r1, r4); // 43 rotr + r0 = r0 | r5; // 44 or + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r2, 2); // 45 shfl + r7 = r7 * r4; // 46 mul + r3 = r3 ^ r4; // 47 xor + r2 = __umulhi(r2, r6); // 48 mulhi + r5 = r5 ^ r4; // 49 xor + r1 = rotr_var(r1, r6); // 50 rotr + r7 = r7 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0x0fe79cecu : 0x68d3a5c0u); // 51 add + r5 = r5 + r4 + ((((sel >> 8u) & 1u) != 0u) ? 0x04309e6fu : 0xeb764d91u); // 52 add + r7 = r7 ^ r0; // 53 xor + r4 = r4 + r0 + ((((sel >> 6u) & 1u) != 0u) ? 0xba0b81c6u : 0x1390b188u); // 54 add + r0 = r0 + r5 + ((((sel >> 15u) & 1u) != 0u) ? 0x5ec96dd8u : 0x4c21a6cdu); // 55 add + r7 = r5 * r3 + r7; // 56 mad + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 57 load + r5 = r5 + r1 + ((((sel >> 1u) & 1u) != 0u) ? 0x6c0ac4ddu : 0x68297b06u); // 58 add + r6 = rotr_var(r6, r7); // 59 rotr + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 60 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 61 shfl + r4 = rotl_imm(r4, 6u); // 62 rotl + r1 = r2 * r1 + r1; // 63 mad + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixB-0/memhard.h b/proto-cuda/packs-readwidth/mixB-0/memhard.h new file mode 100644 index 000000000..0bcf555b8 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-0/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/0". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/mixB-0/memhard.metal b/proto-cuda/packs-readwidth/mixB-0/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-0/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/mixB-0/program.h b/proto-cuda/packs-readwidth/mixB-0/program.h new file mode 100644 index 000000000..ec99493d2 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-0/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/0". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-readwidth/B/0" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d7265616477696474682f422f30" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x5404700e8a2db8feull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=11 shfl=8 xor=8 mul=6 mad=4 rotr=4 mulhi=3 sub=2 or=1 rotl=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "mix25-50-25" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 25, 50, 25 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 5, 8, 3 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 2720 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x3673211cu, 0xaae550b4u, 0x5a0ce2e0u, 0x1d2471cfu, 0xba944366u, 0xbdd4d1dbu, 0xcb1984a9u, 0x081ce12au } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/mixB-0/program.json b/proto-cuda/packs-readwidth/mixB-0/program.json new file mode 100644 index 000000000..7847b95eb --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-0/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x5404700e8a2db8fe", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-readwidth/B/0", + "seed_bytes": "69676e65756d2d7265616477696474682f422f30", + "seed_words": ["0x3673211c", "0xaae550b4", "0x5a0ce2e0", "0x1d2471cf", "0xba944366", "0xbdd4d1db", "0xcb1984a9", "0x081ce12a"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "mix25-50-25", + "load_slots": 16, + "load_mix_percent_4_16_64": [25, 50, 25], + "load_width_counts_4_16_64": [5, 8, 3], + "bytes_per_hash": 2720, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "add": 11, "shfl": 8, "xor": 8, "mul": 6, "mad": 4, "rotr": 4, "mulhi": 3, "sub": 2, "or": 1, "rotl": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "shfl", "dst": 0, "src": 2, "src2": 1, "imm": "0x4e24f8dc", "imm2": "0x287cd532", "rot": 25, "bit": 23, "mask": 1, "width": 1}, + {"i": 1, "op": "mul", "dst": 2, "src": 0, "src2": 2, "imm": "0x6268115c", "imm2": "0xfe52413e", "rot": 30, "bit": 14, "mask": 4, "width": 1}, + {"i": 2, "op": "add", "dst": 0, "src": 6, "src2": 3, "imm": "0x7fbf4ae6", "imm2": "0x03fa29f3", "rot": 30, "bit": 20, "mask": 4, "width": 1}, + {"i": 3, "op": "mul", "dst": 2, "src": 7, "src2": 5, "imm": "0xcd8625e1", "imm2": "0xd6540f7e", "rot": 15, "bit": 15, "mask": 4, "width": 1}, + {"i": 4, "op": "load", "dst": 6, "src": 2, "src2": 5, "imm": "0xfeb8e6c3", "imm2": "0x3a797ec1", "rot": 31, "bit": 17, "mask": 2, "width": 1}, + {"i": 5, "op": "load", "dst": 7, "src": 0, "src2": 4, "imm": "0x8f4f518d", "imm2": "0xa6ed0a9e", "rot": 2, "bit": 14, "mask": 1, "width": 4}, + {"i": 6, "op": "shfl", "dst": 6, "src": 4, "src2": 1, "imm": "0xddd60fee", "imm2": "0x9529e2c3", "rot": 21, "bit": 9, "mask": 16, "width": 1}, + {"i": 7, "op": "xor", "dst": 3, "src": 5, "src2": 3, "imm": "0x015e2388", "imm2": "0x1a59ccd7", "rot": 2, "bit": 12, "mask": 2, "width": 1}, + {"i": 8, "op": "load", "dst": 6, "src": 7, "src2": 5, "imm": "0x90060b55", "imm2": "0x2315bf19", "rot": 20, "bit": 1, "mask": 16, "width": 1}, + {"i": 9, "op": "add", "dst": 0, "src": 7, "src2": 7, "imm": "0xd59b21b0", "imm2": "0x308c81bf", "rot": 5, "bit": 25, "mask": 16, "width": 1}, + {"i": 10, "op": "rotr", "dst": 5, "src": 7, "src2": 0, "imm": "0x4d9af117", "imm2": "0x49369d76", "rot": 31, "bit": 21, "mask": 4, "width": 1}, + {"i": 11, "op": "load", "dst": 3, "src": 6, "src2": 2, "imm": "0x2ddeb7f6", "imm2": "0x855ff344", "rot": 25, "bit": 26, "mask": 1, "width": 4}, + {"i": 12, "op": "mul", "dst": 2, "src": 6, "src2": 6, "imm": "0x8b16ab7b", "imm2": "0x970dc008", "rot": 27, "bit": 10, "mask": 16, "width": 1}, + {"i": 13, "op": "load", "dst": 7, "src": 0, "src2": 7, "imm": "0xcfd7f002", "imm2": "0x1e1d817b", "rot": 1, "bit": 0, "mask": 2, "width": 16}, + {"i": 14, "op": "mad", "dst": 2, "src": 5, "src2": 2, "imm": "0xea959ce8", "imm2": "0x156549ed", "rot": 11, "bit": 2, "mask": 8, "width": 1}, + {"i": 15, "op": "add", "dst": 4, "src": 0, "src2": 7, "imm": "0x902dc661", "imm2": "0x67081e7e", "rot": 12, "bit": 24, "mask": 16, "width": 1}, + {"i": 16, "op": "load", "dst": 3, "src": 4, "src2": 4, "imm": "0x86963c37", "imm2": "0xcdea85b2", "rot": 8, "bit": 19, "mask": 8, "width": 4}, + {"i": 17, "op": "mulhi", "dst": 3, "src": 4, "src2": 5, "imm": "0x5c5720f0", "imm2": "0x551bd9d5", "rot": 20, "bit": 12, "mask": 1, "width": 1}, + {"i": 18, "op": "load", "dst": 4, "src": 7, "src2": 7, "imm": "0xa9fcc38a", "imm2": "0xca2468f8", "rot": 25, "bit": 31, "mask": 1, "width": 4}, + {"i": 19, "op": "add", "dst": 2, "src": 1, "src2": 7, "imm": "0x3ee3182c", "imm2": "0xd1e74db0", "rot": 6, "bit": 9, "mask": 16, "width": 1}, + {"i": 20, "op": "xor", "dst": 7, "src": 6, "src2": 0, "imm": "0x7e5446be", "imm2": "0x9e29589c", "rot": 30, "bit": 14, "mask": 4, "width": 1}, + {"i": 21, "op": "xor", "dst": 5, "src": 3, "src2": 1, "imm": "0x161aa453", "imm2": "0xf277d1f5", "rot": 22, "bit": 18, "mask": 4, "width": 1}, + {"i": 22, "op": "shfl", "dst": 6, "src": 2, "src2": 0, "imm": "0xde7d5591", "imm2": "0x6dce8eb2", "rot": 11, "bit": 15, "mask": 16, "width": 1}, + {"i": 23, "op": "load", "dst": 0, "src": 2, "src2": 1, "imm": "0xf1ab876f", "imm2": "0x976deeca", "rot": 15, "bit": 26, "mask": 8, "width": 4}, + {"i": 24, "op": "add", "dst": 2, "src": 5, "src2": 5, "imm": "0xd2d451c6", "imm2": "0x62ac9e52", "rot": 19, "bit": 22, "mask": 1, "width": 1}, + {"i": 25, "op": "load", "dst": 3, "src": 7, "src2": 0, "imm": "0xfed1b2fd", "imm2": "0x56ac9cfa", "rot": 1, "bit": 16, "mask": 4, "width": 4}, + {"i": 26, "op": "shfl", "dst": 4, "src": 1, "src2": 5, "imm": "0x41708a21", "imm2": "0x9c760382", "rot": 25, "bit": 30, "mask": 16, "width": 1}, + {"i": 27, "op": "load", "dst": 1, "src": 4, "src2": 6, "imm": "0xe3c92c63", "imm2": "0x6d61f62c", "rot": 13, "bit": 19, "mask": 1, "width": 4}, + {"i": 28, "op": "add", "dst": 1, "src": 5, "src2": 0, "imm": "0x111ff813", "imm2": "0x072cfabe", "rot": 16, "bit": 2, "mask": 1, "width": 1}, + {"i": 29, "op": "mad", "dst": 4, "src": 1, "src2": 2, "imm": "0xeae5ec79", "imm2": "0x537fa17e", "rot": 20, "bit": 10, "mask": 16, "width": 1}, + {"i": 30, "op": "mul", "dst": 2, "src": 0, "src2": 2, "imm": "0xf468725d", "imm2": "0xeb01a28a", "rot": 30, "bit": 11, "mask": 2, "width": 1}, + {"i": 31, "op": "xor", "dst": 0, "src": 5, "src2": 1, "imm": "0x4b6a048b", "imm2": "0xd15a1f0a", "rot": 26, "bit": 4, "mask": 4, "width": 1}, + {"i": 32, "op": "load", "dst": 1, "src": 0, "src2": 4, "imm": "0x80ec21f1", "imm2": "0x694ca047", "rot": 19, "bit": 26, "mask": 8, "width": 1}, + {"i": 33, "op": "sub", "dst": 2, "src": 3, "src2": 6, "imm": "0x11d13dc6", "imm2": "0x6c8931e8", "rot": 10, "bit": 13, "mask": 2, "width": 1}, + {"i": 34, "op": "shfl", "dst": 2, "src": 0, "src2": 2, "imm": "0x056d2ef8", "imm2": "0xa833af40", "rot": 22, "bit": 26, "mask": 2, "width": 1}, + {"i": 35, "op": "sub", "dst": 0, "src": 6, "src2": 6, "imm": "0x9e95fe74", "imm2": "0xbf36d23c", "rot": 1, "bit": 12, "mask": 2, "width": 1}, + {"i": 36, "op": "mul", "dst": 4, "src": 7, "src2": 5, "imm": "0x3d50d394", "imm2": "0x637cc722", "rot": 12, "bit": 4, "mask": 16, "width": 1}, + {"i": 37, "op": "xor", "dst": 5, "src": 6, "src2": 7, "imm": "0x3272395d", "imm2": "0x36dd11fb", "rot": 31, "bit": 4, "mask": 16, "width": 1}, + {"i": 38, "op": "mulhi", "dst": 0, "src": 6, "src2": 6, "imm": "0x67324b08", "imm2": "0xd825e2fa", "rot": 11, "bit": 21, "mask": 1, "width": 1}, + {"i": 39, "op": "shfl", "dst": 5, "src": 7, "src2": 6, "imm": "0x6aff9133", "imm2": "0x22d6b873", "rot": 17, "bit": 4, "mask": 4, "width": 1}, + {"i": 40, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0xac1e076f", "imm2": "0x0e0ff9a2", "rot": 21, "bit": 8, "mask": 2, "width": 4}, + {"i": 41, "op": "load", "dst": 2, "src": 4, "src2": 0, "imm": "0xd94d7e29", "imm2": "0xa18f1af4", "rot": 14, "bit": 4, "mask": 4, "width": 1}, + {"i": 42, "op": "load", "dst": 5, "src": 6, "src2": 4, "imm": "0xe02caafc", "imm2": "0xe37b52cf", "rot": 28, "bit": 24, "mask": 8, "width": 1}, + {"i": 43, "op": "rotr", "dst": 1, "src": 4, "src2": 5, "imm": "0xde64b2e0", "imm2": "0xeb2327ba", "rot": 19, "bit": 5, "mask": 1, "width": 1}, + {"i": 44, "op": "or", "dst": 0, "src": 5, "src2": 1, "imm": "0x23797ba8", "imm2": "0xdaa53dbe", "rot": 17, "bit": 0, "mask": 8, "width": 1}, + {"i": 45, "op": "shfl", "dst": 0, "src": 2, "src2": 5, "imm": "0x9bb8bc40", "imm2": "0x9a1d5dc9", "rot": 22, "bit": 4, "mask": 2, "width": 1}, + {"i": 46, "op": "mul", "dst": 7, "src": 4, "src2": 5, "imm": "0x238172b8", "imm2": "0x1a3008f9", "rot": 25, "bit": 20, "mask": 16, "width": 1}, + {"i": 47, "op": "xor", "dst": 3, "src": 4, "src2": 5, "imm": "0x934047d2", "imm2": "0xe984b304", "rot": 27, "bit": 2, "mask": 1, "width": 1}, + {"i": 48, "op": "mulhi", "dst": 2, "src": 6, "src2": 5, "imm": "0xf1c06c0c", "imm2": "0x0f679cae", "rot": 5, "bit": 31, "mask": 1, "width": 1}, + {"i": 49, "op": "xor", "dst": 5, "src": 4, "src2": 3, "imm": "0x4b9454cd", "imm2": "0x816b6e2d", "rot": 5, "bit": 22, "mask": 4, "width": 1}, + {"i": 50, "op": "rotr", "dst": 1, "src": 6, "src2": 3, "imm": "0xe234dec2", "imm2": "0xeddc8839", "rot": 6, "bit": 7, "mask": 1, "width": 1}, + {"i": 51, "op": "add", "dst": 7, "src": 0, "src2": 0, "imm": "0x68d3a5c0", "imm2": "0x0fe79cec", "rot": 5, "bit": 22, "mask": 2, "width": 1}, + {"i": 52, "op": "add", "dst": 5, "src": 4, "src2": 1, "imm": "0xeb764d91", "imm2": "0x04309e6f", "rot": 23, "bit": 8, "mask": 1, "width": 1}, + {"i": 53, "op": "xor", "dst": 7, "src": 0, "src2": 3, "imm": "0x6d420b6b", "imm2": "0xa4eb51ff", "rot": 27, "bit": 9, "mask": 1, "width": 1}, + {"i": 54, "op": "add", "dst": 4, "src": 0, "src2": 0, "imm": "0x1390b188", "imm2": "0xba0b81c6", "rot": 4, "bit": 6, "mask": 16, "width": 1}, + {"i": 55, "op": "add", "dst": 0, "src": 5, "src2": 0, "imm": "0x4c21a6cd", "imm2": "0x5ec96dd8", "rot": 18, "bit": 15, "mask": 2, "width": 1}, + {"i": 56, "op": "mad", "dst": 7, "src": 5, "src2": 3, "imm": "0xd0655c74", "imm2": "0x8dc4bd13", "rot": 24, "bit": 11, "mask": 1, "width": 1}, + {"i": 57, "op": "load", "dst": 6, "src": 0, "src2": 7, "imm": "0xe738815c", "imm2": "0xc546dbe7", "rot": 24, "bit": 29, "mask": 2, "width": 16}, + {"i": 58, "op": "add", "dst": 5, "src": 1, "src2": 6, "imm": "0x68297b06", "imm2": "0x6c0ac4dd", "rot": 19, "bit": 1, "mask": 8, "width": 1}, + {"i": 59, "op": "rotr", "dst": 6, "src": 7, "src2": 3, "imm": "0x83f974a4", "imm2": "0x13fcd20f", "rot": 3, "bit": 18, "mask": 2, "width": 1}, + {"i": 60, "op": "load", "dst": 7, "src": 4, "src2": 5, "imm": "0xbca249c5", "imm2": "0x2094b70b", "rot": 15, "bit": 24, "mask": 2, "width": 16}, + {"i": 61, "op": "shfl", "dst": 6, "src": 3, "src2": 0, "imm": "0xcd1113c9", "imm2": "0x475dbb4b", "rot": 11, "bit": 11, "mask": 4, "width": 1}, + {"i": 62, "op": "rotl", "dst": 4, "src": 5, "src2": 7, "imm": "0xdbc5d842", "imm2": "0xf8b7004c", "rot": 6, "bit": 25, "mask": 8, "width": 1}, + {"i": 63, "op": "mad", "dst": 1, "src": 2, "src2": 1, "imm": "0xd2538bab", "imm2": "0x6b862569", "rot": 9, "bit": 5, "mask": 1, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/mixB-0/program.metal b/proto-cuda/packs-readwidth/mixB-0/program.metal new file mode 100644 index 000000000..84fb475f9 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-0/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x3673211cu, 0xaae550b4u, 0x5a0ce2e0u, 0x1d2471cfu, 0xba944366u, 0xbdd4d1dbu, 0xcb1984a9u, 0x081ce12au }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = r0 ^ simd_shuffle_xor(r2, (ushort)1); // 0 + r2 = r2 * r0; // 1 + r0 = r0 + r6 + select(0x7fbf4ae6u, 0x03fa29f3u, ((sel >> 20u) & 1u) != 0u); // 2 + r2 = r2 * r7; // 3 + r6 = r6 ^ dataset[r2 & MASK]; // 4 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 5 + r6 = r6 ^ simd_shuffle_xor(r4, (ushort)16); // 6 + r3 = r3 ^ r5; // 7 + r6 = r6 ^ dataset[r7 & MASK]; // 8 + r0 = r0 + r7 + select(0xd59b21b0u, 0x308c81bfu, ((sel >> 25u) & 1u) != 0u); // 9 + r5 = rotr_var(r5, r7); // 10 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 + r2 = r2 * r6; // 12 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 13 + r2 = r5 * r2 + r2; // 14 + r4 = r4 + r0 + select(0x902dc661u, 0x67081e7eu, ((sel >> 24u) & 1u) != 0u); // 15 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 + r3 = mulhi(r3, r4); // 17 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 18 + r2 = r2 + r1 + select(0x3ee3182cu, 0xd1e74db0u, ((sel >> 9u) & 1u) != 0u); // 19 + r7 = r7 ^ r6; // 20 + r5 = r5 ^ r3; // 21 + r6 = r6 ^ simd_shuffle_xor(r2, (ushort)16); // 22 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 23 + r2 = r2 + r5 + select(0xd2d451c6u, 0x62ac9e52u, ((sel >> 22u) & 1u) != 0u); // 24 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 25 + r4 = r4 ^ simd_shuffle_xor(r1, (ushort)16); // 26 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 27 + r1 = r1 + r5 + select(0x111ff813u, 0x072cfabeu, ((sel >> 2u) & 1u) != 0u); // 28 + r4 = r1 * r2 + r4; // 29 + r2 = r2 * r0; // 30 + r0 = r0 ^ r5; // 31 + r1 = r1 ^ dataset[r0 & MASK]; // 32 + r2 = r2 - r3; // 33 + r2 = r2 ^ simd_shuffle_xor(r0, (ushort)2); // 34 + r0 = r0 - r6; // 35 + r4 = r4 * r7; // 36 + r5 = r5 ^ r6; // 37 + r0 = mulhi(r0, r6); // 38 + r5 = r5 ^ simd_shuffle_xor(r7, (ushort)4); // 39 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 40 + r2 = r2 ^ dataset[r4 & MASK]; // 41 + r5 = r5 ^ dataset[r6 & MASK]; // 42 + r1 = rotr_var(r1, r4); // 43 + r0 = r0 | r5; // 44 + r0 = r0 ^ simd_shuffle_xor(r2, (ushort)2); // 45 + r7 = r7 * r4; // 46 + r3 = r3 ^ r4; // 47 + r2 = mulhi(r2, r6); // 48 + r5 = r5 ^ r4; // 49 + r1 = rotr_var(r1, r6); // 50 + r7 = r7 + r0 + select(0x68d3a5c0u, 0x0fe79cecu, ((sel >> 22u) & 1u) != 0u); // 51 + r5 = r5 + r4 + select(0xeb764d91u, 0x04309e6fu, ((sel >> 8u) & 1u) != 0u); // 52 + r7 = r7 ^ r0; // 53 + r4 = r4 + r0 + select(0x1390b188u, 0xba0b81c6u, ((sel >> 6u) & 1u) != 0u); // 54 + r0 = r0 + r5 + select(0x4c21a6cdu, 0x5ec96dd8u, ((sel >> 15u) & 1u) != 0u); // 55 + r7 = r5 * r3 + r7; // 56 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 57 + r5 = r5 + r1 + select(0x68297b06u, 0x6c0ac4ddu, ((sel >> 1u) & 1u) != 0u); // 58 + r6 = rotr_var(r6, r7); // 59 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 60 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 61 + r4 = rotl_imm(r4, 6u); // 62 + r1 = r2 * r1 + r1; // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-0/program_bound.metal b/proto-cuda/packs-readwidth/mixB-0/program_bound.metal new file mode 100644 index 000000000..91347d02d --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-0/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x3673211cu, 0xaae550b4u, 0x5a0ce2e0u, 0x1d2471cfu, 0xba944366u, 0xbdd4d1dbu, 0xcb1984a9u, 0x081ce12au }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = r0 ^ simd_shuffle_xor(r2, (ushort)1); // 0 + r2 = r2 * r0; // 1 + r0 = r0 + r6 + select(0x7fbf4ae6u, 0x03fa29f3u, ((sel >> 20u) & 1u) != 0u); // 2 + r2 = r2 * r7; // 3 + r6 = r6 ^ dataset[r2 & MASK]; // 4 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 5 + r6 = r6 ^ simd_shuffle_xor(r4, (ushort)16); // 6 + r3 = r3 ^ r5; // 7 + r6 = r6 ^ dataset[r7 & MASK]; // 8 + r0 = r0 + r7 + select(0xd59b21b0u, 0x308c81bfu, ((sel >> 25u) & 1u) != 0u); // 9 + r5 = rotr_var(r5, r7); // 10 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 + r2 = r2 * r6; // 12 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 13 + r2 = r5 * r2 + r2; // 14 + r4 = r4 + r0 + select(0x902dc661u, 0x67081e7eu, ((sel >> 24u) & 1u) != 0u); // 15 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 + r3 = mulhi(r3, r4); // 17 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 18 + r2 = r2 + r1 + select(0x3ee3182cu, 0xd1e74db0u, ((sel >> 9u) & 1u) != 0u); // 19 + r7 = r7 ^ r6; // 20 + r5 = r5 ^ r3; // 21 + r6 = r6 ^ simd_shuffle_xor(r2, (ushort)16); // 22 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 23 + r2 = r2 + r5 + select(0xd2d451c6u, 0x62ac9e52u, ((sel >> 22u) & 1u) != 0u); // 24 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 25 + r4 = r4 ^ simd_shuffle_xor(r1, (ushort)16); // 26 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 27 + r1 = r1 + r5 + select(0x111ff813u, 0x072cfabeu, ((sel >> 2u) & 1u) != 0u); // 28 + r4 = r1 * r2 + r4; // 29 + r2 = r2 * r0; // 30 + r0 = r0 ^ r5; // 31 + r1 = r1 ^ dataset[r0 & MASK]; // 32 + r2 = r2 - r3; // 33 + r2 = r2 ^ simd_shuffle_xor(r0, (ushort)2); // 34 + r0 = r0 - r6; // 35 + r4 = r4 * r7; // 36 + r5 = r5 ^ r6; // 37 + r0 = mulhi(r0, r6); // 38 + r5 = r5 ^ simd_shuffle_xor(r7, (ushort)4); // 39 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 40 + r2 = r2 ^ dataset[r4 & MASK]; // 41 + r5 = r5 ^ dataset[r6 & MASK]; // 42 + r1 = rotr_var(r1, r4); // 43 + r0 = r0 | r5; // 44 + r0 = r0 ^ simd_shuffle_xor(r2, (ushort)2); // 45 + r7 = r7 * r4; // 46 + r3 = r3 ^ r4; // 47 + r2 = mulhi(r2, r6); // 48 + r5 = r5 ^ r4; // 49 + r1 = rotr_var(r1, r6); // 50 + r7 = r7 + r0 + select(0x68d3a5c0u, 0x0fe79cecu, ((sel >> 22u) & 1u) != 0u); // 51 + r5 = r5 + r4 + select(0xeb764d91u, 0x04309e6fu, ((sel >> 8u) & 1u) != 0u); // 52 + r7 = r7 ^ r0; // 53 + r4 = r4 + r0 + select(0x1390b188u, 0xba0b81c6u, ((sel >> 6u) & 1u) != 0u); // 54 + r0 = r0 + r5 + select(0x4c21a6cdu, 0x5ec96dd8u, ((sel >> 15u) & 1u) != 0u); // 55 + r7 = r5 * r3 + r7; // 56 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 57 + r5 = r5 + r1 + select(0x68297b06u, 0x6c0ac4ddu, ((sel >> 1u) & 1u) != 0u); // 58 + r6 = rotr_var(r6, r7); // 59 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 60 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 61 + r4 = rotl_imm(r4, 6u); // 62 + r1 = r2 * r1 + r1; // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-0/vectors.h b/proto-cuda/packs-readwidth/mixB-0/vectors.h new file mode 100644 index 000000000..2bfccfb0b --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-0/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/0". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x2b6c8e6b238224dbull, 0x94dca5778f870901ull, 0xb0c977094b3c1a7aull, 0xdc0f62ab6e8db552ull, 0x5581af7ea67bfd20ull, 0xaa10426d0e2aa72full, 0x60fd1bf45ef55751ull, 0x9ab74415a3de8b66ull, + 0x2cb6ee8ac9c0efc8ull, 0x2ae892465dff9094ull, 0x99c9c94b9d2874c7ull, 0xb12be63d01872ea3ull, 0x84024fac8312e908ull, 0xbb67e4878f174f0cull, 0x3b027e93ca233521ull, 0x8ba6cfeb9bfbbc0full, + 0x44c59d77e3247d78ull, 0xd4a765b48c283d2aull, 0xe53083a844015f24ull, 0xb0f7b8dad47d924bull, 0x185b4a0f6150375eull, 0x537ed78b8976846cull, 0xc9d4264a0fc2da7dull, 0x6c0ef3c92dd3e063ull, + 0x396461e1a90913e9ull, 0xc2ea0411a102f780ull, 0x0d8b3f485c471e7eull, 0xd8a5c0b7995246c7ull, 0x4507e1e49dbd3c35ull, 0x80145ecfc13419c0ull, 0x9c2f1f09b3ef4ed1ull, 0xd776b774670d63f1ull + }, + { // base nonce 4096 + 0xe094ec92f00bd1fdull, 0x14d3ad7d06eba7bfull, 0xd4b4ee3a4e95e930ull, 0x5656fc12f488f952ull, 0x4db64fa3172356b8ull, 0x7669809a83c52addull, 0x240c4920d64dd3ebull, 0xfcc0c08f3cebe318ull, + 0xd48c95d54226a468ull, 0x383acc22e4367a07ull, 0x1b16e6976aa7160eull, 0xc1d0216840450b35ull, 0xcfc19bac8283d5e0ull, 0xaecef9011360b2c8ull, 0x6a1dbf38ce8f3476ull, 0x813ecd9804528e50ull, + 0x95273d12cbfda3e0ull, 0xb6abb5bc816f81d2ull, 0x3d734251c7eb5488ull, 0x162710fc525e1a35ull, 0xa8485c7fa11f1d69ull, 0xf81a8cdc4b928aa5ull, 0x1ad63b6197331e30ull, 0x04ce2b194f3fe8adull, + 0xc76ba96148e01225ull, 0x26e24635c3556e18ull, 0xeb8c92b751bc5c70ull, 0xb8886d57d441fe5dull, 0x3f8d68c1f1f1fe2dull, 0x7c441c21089fb44eull, 0x27e2950b292fec5aull, 0x581d48ac2f0dc82eull + }, + { // base nonce 1000000 + 0xea40156c5e27f6a3ull, 0x1db25b47fd85ab8eull, 0xdf6bd10d80c1b46dull, 0x2e1da26fcbea64b0ull, 0xdd5e41310fcabbfbull, 0xdf9780cb1f87d4afull, 0x6028258e6529007full, 0x886fecbfe753c3d9ull, + 0x3c20a9a6f5bcb77eull, 0x605194089256f91cull, 0xd7546f5b5c79f188ull, 0xefc8ae5b45bd346bull, 0xd067b744f6659df8ull, 0x9c8fbb301da7e947ull, 0x238f22c8b511e7cbull, 0xa67a31f99be23902ull, + 0x5d9224b25e12fcffull, 0x0166de07227aaebbull, 0xfb50fea4c86552a7ull, 0xf7535d8fe65fe0d0ull, 0x161ab76c522de840ull, 0xaa3d2ddcd8e80d1aull, 0x9bdb9180dd7d39faull, 0xa4111c74b1d438c9ull, + 0x75ec298c4f067224ull, 0x9c923c3a46403618ull, 0x61ec429ff3c0283eull, 0x155cf3df9f82ff51ull, 0xd6f0279de6eec131ull, 0xc3ea627086d07a96ull, 0x1cb80fef0710e142ull, 0xfd4772b1cd071be4ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/mixB-0/vectors.json b/proto-cuda/packs-readwidth/mixB-0/vectors.json new file mode 100644 index 000000000..e8e9487f7 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-0/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-readwidth/B/0", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x2b6c8e6b238224db", "0x94dca5778f870901", "0xb0c977094b3c1a7a", "0xdc0f62ab6e8db552", "0x5581af7ea67bfd20", "0xaa10426d0e2aa72f", "0x60fd1bf45ef55751", "0x9ab74415a3de8b66", + "0x2cb6ee8ac9c0efc8", "0x2ae892465dff9094", "0x99c9c94b9d2874c7", "0xb12be63d01872ea3", "0x84024fac8312e908", "0xbb67e4878f174f0c", "0x3b027e93ca233521", "0x8ba6cfeb9bfbbc0f", + "0x44c59d77e3247d78", "0xd4a765b48c283d2a", "0xe53083a844015f24", "0xb0f7b8dad47d924b", "0x185b4a0f6150375e", "0x537ed78b8976846c", "0xc9d4264a0fc2da7d", "0x6c0ef3c92dd3e063", + "0x396461e1a90913e9", "0xc2ea0411a102f780", "0x0d8b3f485c471e7e", "0xd8a5c0b7995246c7", "0x4507e1e49dbd3c35", "0x80145ecfc13419c0", "0x9c2f1f09b3ef4ed1", "0xd776b774670d63f1" + ]}, + {"base_nonce": 4096, "expected": [ + "0xe094ec92f00bd1fd", "0x14d3ad7d06eba7bf", "0xd4b4ee3a4e95e930", "0x5656fc12f488f952", "0x4db64fa3172356b8", "0x7669809a83c52add", "0x240c4920d64dd3eb", "0xfcc0c08f3cebe318", + "0xd48c95d54226a468", "0x383acc22e4367a07", "0x1b16e6976aa7160e", "0xc1d0216840450b35", "0xcfc19bac8283d5e0", "0xaecef9011360b2c8", "0x6a1dbf38ce8f3476", "0x813ecd9804528e50", + "0x95273d12cbfda3e0", "0xb6abb5bc816f81d2", "0x3d734251c7eb5488", "0x162710fc525e1a35", "0xa8485c7fa11f1d69", "0xf81a8cdc4b928aa5", "0x1ad63b6197331e30", "0x04ce2b194f3fe8ad", + "0xc76ba96148e01225", "0x26e24635c3556e18", "0xeb8c92b751bc5c70", "0xb8886d57d441fe5d", "0x3f8d68c1f1f1fe2d", "0x7c441c21089fb44e", "0x27e2950b292fec5a", "0x581d48ac2f0dc82e" + ]}, + {"base_nonce": 1000000, "expected": [ + "0xea40156c5e27f6a3", "0x1db25b47fd85ab8e", "0xdf6bd10d80c1b46d", "0x2e1da26fcbea64b0", "0xdd5e41310fcabbfb", "0xdf9780cb1f87d4af", "0x6028258e6529007f", "0x886fecbfe753c3d9", + "0x3c20a9a6f5bcb77e", "0x605194089256f91c", "0xd7546f5b5c79f188", "0xefc8ae5b45bd346b", "0xd067b744f6659df8", "0x9c8fbb301da7e947", "0x238f22c8b511e7cb", "0xa67a31f99be23902", + "0x5d9224b25e12fcff", "0x0166de07227aaebb", "0xfb50fea4c86552a7", "0xf7535d8fe65fe0d0", "0x161ab76c522de840", "0xaa3d2ddcd8e80d1a", "0x9bdb9180dd7d39fa", "0xa4111c74b1d438c9", + "0x75ec298c4f067224", "0x9c923c3a46403618", "0x61ec429ff3c0283e", "0x155cf3df9f82ff51", "0xd6f0279de6eec131", "0xc3ea627086d07a96", "0x1cb80fef0710e142", "0xfd4772b1cd071be4" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/mixB-1/kernel.cl b/proto-cuda/packs-readwidth/mixB-1/kernel.cl new file mode 100644 index 000000000..e43468b46 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-1/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/1". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x3e345412u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0xfc2b3a0eu; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0xfc2b3a0eu; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x7740353du; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x7740353du; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x2df159dbu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x2df159dbu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x50bc6eedu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x50bc6eedu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x2e61acb2u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x2e61acb2u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x3c800bffu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x3c800bffu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xbaecd8c1u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0xbaecd8c1u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x3e345412u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r5 = r4 * r0 + r5; // 0 mad + r1 = rotl_imm(r1, 23u); // 1 rotl + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 2 load + r5 = r4 * r2 + r5; // 3 mad + r1 = r1 * r6; // 4 mul + r1 = r1 + r0 + ((((sel >> 6u) & 1u) != 0u) ? 0x9dd9fb05u : 0x9db46598u); // 5 add + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 16u); r1 = r1 ^ t_; } // 6 shfl + r6 = r6 * r1; // 7 mul + r1 = mul_hi(r1, r7); // 8 mulhi + r2 = mul_hi(r2, r1); // 9 mulhi + r0 = r0 + r4 + ((((sel >> 13u) & 1u) != 0u) ? 0xc15822b6u : 0xe45fccebu); // 10 add + r5 = r1 * r2 + r5; // 11 mad + r4 = r4 * r7; // 12 mul + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 13 load + r3 = r3 ^ r4; // 14 xor + r1 = r1 ^ r6; // 15 xor + r5 = rotr_var(r5, r6); // 16 rotr + r3 = rotr_var(r3, r7); // 17 rotr + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r1 = r1 ^ t_; } // 18 shfl + r5 = r5 ^ ds[r6 & mask]; // 19 load + r6 = rotl_imm(r6, 30u); // 20 rotl + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 21 load + r3 = r3 ^ r5; // 22 xor + r5 = r5 + r7 + ((((sel >> 6u) & 1u) != 0u) ? 0x950603c6u : 0x1d4f8db7u); // 23 add + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 24 load + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 load + r2 = r2 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x0ae476a4u : 0x55db0d92u); // 26 add + r5 = r5 ^ ds[r6 & mask]; // 27 load + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 28 load + r1 = r1 ^ r6; // 29 xor + r7 = mul_hi(r7, r4); // 30 mulhi + r6 = rotr_var(r6, r3); // 31 rotr + r0 = r0 ^ ds[r3 & mask]; // 32 load + r5 = rotr_var(r5, r0); // 33 rotr + r5 = rotl_imm(r5, 20u); // 34 rotl + r0 = mul_hi(r0, r4); // 35 mulhi + r6 = r6 + r7 + ((((sel >> 18u) & 1u) != 0u) ? 0xfe7a0454u : 0x5d6e3dc5u); // 36 add + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 37 load + r2 = rotl_imm(r2, 19u); // 38 rotl + r3 = r3 + r6 + ((((sel >> 20u) & 1u) != 0u) ? 0xc1d4ae24u : 0xe7e241bau); // 39 add + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 40 load + r3 = r3 + r1 + ((((sel >> 21u) & 1u) != 0u) ? 0x8f9f8556u : 0xb45cdd60u); // 41 add + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 42 shfl + r0 = r0 * r1; // 43 mul + r0 = rotr_var(r0, r7); // 44 rotr + r5 = r5 - r3; // 45 sub + r2 = r7 * r7 + r2; // 46 mad + r6 = r3 * r1 + r6; // 47 mad + r0 = r0 * r7; // 48 mul + r0 = r0 + r4 + ((((sel >> 1u) & 1u) != 0u) ? 0xb9083b6cu : 0x3fa0c5cdu); // 49 add + r4 = r4 + r6 + ((((sel >> 13u) & 1u) != 0u) ? 0x9eecc778u : 0x2d6873e8u); // 50 add + r2 = r2 - r7; // 51 sub + r6 = rotl_imm(r6, 6u); // 52 rotl + r3 = rotr_var(r3, r5); // 53 rotr + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 54 load + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 55 load + r2 = r2 | r6; // 56 or + r4 = r4 ^ ds[r0 & mask]; // 57 load + r3 = r3 ^ r6; // 58 xor + r0 = r0 ^ ds[r6 & mask]; // 59 load + r2 = r2 + r4 + ((((sel >> 3u) & 1u) != 0u) ? 0x19c76fb3u : 0x5052a1c3u); // 60 add + r7 = r7 ^ ds[r4 & mask]; // 61 load + r6 = r6 | r1; // 62 or + r5 = r5 + r1 + ((((sel >> 4u) & 1u) != 0u) ? 0x535fe3fau : 0x08c757c0u); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/mixB-1/kernel.cu b/proto-cuda/packs-readwidth/mixB-1/kernel.cu new file mode 100644 index 000000000..2b53b35c1 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-1/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/1". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x3e345412u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0xfc2b3a0eu; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0xfc2b3a0eu; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x7740353du; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0x7740353du; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x2df159dbu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x2df159dbu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x50bc6eedu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0x50bc6eedu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x2e61acb2u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0x2e61acb2u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x3c800bffu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x3c800bffu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xbaecd8c1u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0xbaecd8c1u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x3e345412u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r5 = r4 * r0 + r5; // 0 mad + r1 = rotl_imm(r1, 23u); // 1 rotl + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 2 load + r5 = r4 * r2 + r5; // 3 mad + r1 = r1 * r6; // 4 mul + r1 = r1 + r0 + ((((sel >> 6u) & 1u) != 0u) ? 0x9dd9fb05u : 0x9db46598u); // 5 add + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r7, 16); // 6 shfl + r6 = r6 * r1; // 7 mul + r1 = __umulhi(r1, r7); // 8 mulhi + r2 = __umulhi(r2, r1); // 9 mulhi + r0 = r0 + r4 + ((((sel >> 13u) & 1u) != 0u) ? 0xc15822b6u : 0xe45fccebu); // 10 add + r5 = r1 * r2 + r5; // 11 mad + r4 = r4 * r7; // 12 mul + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 13 load + r3 = r3 ^ r4; // 14 xor + r1 = r1 ^ r6; // 15 xor + r5 = rotr_var(r5, r6); // 16 rotr + r3 = rotr_var(r3, r7); // 17 rotr + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 18 shfl + r5 = r5 ^ ds[r6 & mask]; // 19 load + r6 = rotl_imm(r6, 30u); // 20 rotl + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 21 load + r3 = r3 ^ r5; // 22 xor + r5 = r5 + r7 + ((((sel >> 6u) & 1u) != 0u) ? 0x950603c6u : 0x1d4f8db7u); // 23 add + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 24 load + { uint32_t b_ = (r1 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 load + r2 = r2 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x0ae476a4u : 0x55db0d92u); // 26 add + r5 = r5 ^ ds[r6 & mask]; // 27 load + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 28 load + r1 = r1 ^ r6; // 29 xor + r7 = __umulhi(r7, r4); // 30 mulhi + r6 = rotr_var(r6, r3); // 31 rotr + r0 = r0 ^ ds[r3 & mask]; // 32 load + r5 = rotr_var(r5, r0); // 33 rotr + r5 = rotl_imm(r5, 20u); // 34 rotl + r0 = __umulhi(r0, r4); // 35 mulhi + r6 = r6 + r7 + ((((sel >> 18u) & 1u) != 0u) ? 0xfe7a0454u : 0x5d6e3dc5u); // 36 add + { uint32_t b_ = (r1 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 37 load + r2 = rotl_imm(r2, 19u); // 38 rotl + r3 = r3 + r6 + ((((sel >> 20u) & 1u) != 0u) ? 0xc1d4ae24u : 0xe7e241bau); // 39 add + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 40 load + r3 = r3 + r1 + ((((sel >> 21u) & 1u) != 0u) ? 0x8f9f8556u : 0xb45cdd60u); // 41 add + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 42 shfl + r0 = r0 * r1; // 43 mul + r0 = rotr_var(r0, r7); // 44 rotr + r5 = r5 - r3; // 45 sub + r2 = r7 * r7 + r2; // 46 mad + r6 = r3 * r1 + r6; // 47 mad + r0 = r0 * r7; // 48 mul + r0 = r0 + r4 + ((((sel >> 1u) & 1u) != 0u) ? 0xb9083b6cu : 0x3fa0c5cdu); // 49 add + r4 = r4 + r6 + ((((sel >> 13u) & 1u) != 0u) ? 0x9eecc778u : 0x2d6873e8u); // 50 add + r2 = r2 - r7; // 51 sub + r6 = rotl_imm(r6, 6u); // 52 rotl + r3 = rotr_var(r3, r5); // 53 rotr + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 54 load + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 55 load + r2 = r2 | r6; // 56 or + r4 = r4 ^ ds[r0 & mask]; // 57 load + r3 = r3 ^ r6; // 58 xor + r0 = r0 ^ ds[r6 & mask]; // 59 load + r2 = r2 + r4 + ((((sel >> 3u) & 1u) != 0u) ? 0x19c76fb3u : 0x5052a1c3u); // 60 add + r7 = r7 ^ ds[r4 & mask]; // 61 load + r6 = r6 | r1; // 62 or + r5 = r5 + r1 + ((((sel >> 4u) & 1u) != 0u) ? 0x535fe3fau : 0x08c757c0u); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixB-1/kernel_bound.cl b/proto-cuda/packs-readwidth/mixB-1/kernel_bound.cl new file mode 100644 index 000000000..58ac94d5d --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-1/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/1". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x3e345412u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0xfc2b3a0eu; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0xfc2b3a0eu; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x7740353du; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x7740353du; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x2df159dbu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x2df159dbu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x50bc6eedu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x50bc6eedu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x2e61acb2u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x2e61acb2u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x3c800bffu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x3c800bffu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xbaecd8c1u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0xbaecd8c1u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x3e345412u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r5 = r4 * r0 + r5; // 0 mad + r1 = rotl_imm(r1, 23u); // 1 rotl + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 2 load + r5 = r4 * r2 + r5; // 3 mad + r1 = r1 * r6; // 4 mul + r1 = r1 + r0 + ((((sel >> 6u) & 1u) != 0u) ? 0x9dd9fb05u : 0x9db46598u); // 5 add + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 16u); r1 = r1 ^ t_; } // 6 shfl + r6 = r6 * r1; // 7 mul + r1 = mul_hi(r1, r7); // 8 mulhi + r2 = mul_hi(r2, r1); // 9 mulhi + r0 = r0 + r4 + ((((sel >> 13u) & 1u) != 0u) ? 0xc15822b6u : 0xe45fccebu); // 10 add + r5 = r1 * r2 + r5; // 11 mad + r4 = r4 * r7; // 12 mul + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 13 load + r3 = r3 ^ r4; // 14 xor + r1 = r1 ^ r6; // 15 xor + r5 = rotr_var(r5, r6); // 16 rotr + r3 = rotr_var(r3, r7); // 17 rotr + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r1 = r1 ^ t_; } // 18 shfl + r5 = r5 ^ ds[r6 & mask]; // 19 load + r6 = rotl_imm(r6, 30u); // 20 rotl + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 21 load + r3 = r3 ^ r5; // 22 xor + r5 = r5 + r7 + ((((sel >> 6u) & 1u) != 0u) ? 0x950603c6u : 0x1d4f8db7u); // 23 add + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 24 load + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 load + r2 = r2 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x0ae476a4u : 0x55db0d92u); // 26 add + r5 = r5 ^ ds[r6 & mask]; // 27 load + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 28 load + r1 = r1 ^ r6; // 29 xor + r7 = mul_hi(r7, r4); // 30 mulhi + r6 = rotr_var(r6, r3); // 31 rotr + r0 = r0 ^ ds[r3 & mask]; // 32 load + r5 = rotr_var(r5, r0); // 33 rotr + r5 = rotl_imm(r5, 20u); // 34 rotl + r0 = mul_hi(r0, r4); // 35 mulhi + r6 = r6 + r7 + ((((sel >> 18u) & 1u) != 0u) ? 0xfe7a0454u : 0x5d6e3dc5u); // 36 add + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 37 load + r2 = rotl_imm(r2, 19u); // 38 rotl + r3 = r3 + r6 + ((((sel >> 20u) & 1u) != 0u) ? 0xc1d4ae24u : 0xe7e241bau); // 39 add + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 40 load + r3 = r3 + r1 + ((((sel >> 21u) & 1u) != 0u) ? 0x8f9f8556u : 0xb45cdd60u); // 41 add + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 42 shfl + r0 = r0 * r1; // 43 mul + r0 = rotr_var(r0, r7); // 44 rotr + r5 = r5 - r3; // 45 sub + r2 = r7 * r7 + r2; // 46 mad + r6 = r3 * r1 + r6; // 47 mad + r0 = r0 * r7; // 48 mul + r0 = r0 + r4 + ((((sel >> 1u) & 1u) != 0u) ? 0xb9083b6cu : 0x3fa0c5cdu); // 49 add + r4 = r4 + r6 + ((((sel >> 13u) & 1u) != 0u) ? 0x9eecc778u : 0x2d6873e8u); // 50 add + r2 = r2 - r7; // 51 sub + r6 = rotl_imm(r6, 6u); // 52 rotl + r3 = rotr_var(r3, r5); // 53 rotr + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 54 load + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 55 load + r2 = r2 | r6; // 56 or + r4 = r4 ^ ds[r0 & mask]; // 57 load + r3 = r3 ^ r6; // 58 xor + r0 = r0 ^ ds[r6 & mask]; // 59 load + r2 = r2 + r4 + ((((sel >> 3u) & 1u) != 0u) ? 0x19c76fb3u : 0x5052a1c3u); // 60 add + r7 = r7 ^ ds[r4 & mask]; // 61 load + r6 = r6 | r1; // 62 or + r5 = r5 + r1 + ((((sel >> 4u) & 1u) != 0u) ? 0x535fe3fau : 0x08c757c0u); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r5 = r4 * r0 + r5; // 0 mad + r1 = rotl_imm(r1, 23u); // 1 rotl + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 2 load + r5 = r4 * r2 + r5; // 3 mad + r1 = r1 * r6; // 4 mul + r1 = r1 + r0 + ((((sel >> 6u) & 1u) != 0u) ? 0x9dd9fb05u : 0x9db46598u); // 5 add + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 16u); r1 = r1 ^ t_; } // 6 shfl + r6 = r6 * r1; // 7 mul + r1 = mul_hi(r1, r7); // 8 mulhi + r2 = mul_hi(r2, r1); // 9 mulhi + r0 = r0 + r4 + ((((sel >> 13u) & 1u) != 0u) ? 0xc15822b6u : 0xe45fccebu); // 10 add + r5 = r1 * r2 + r5; // 11 mad + r4 = r4 * r7; // 12 mul + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 13 load + r3 = r3 ^ r4; // 14 xor + r1 = r1 ^ r6; // 15 xor + r5 = rotr_var(r5, r6); // 16 rotr + r3 = rotr_var(r3, r7); // 17 rotr + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r1 = r1 ^ t_; } // 18 shfl + r5 = r5 ^ ds[r6 & mask]; // 19 load + r6 = rotl_imm(r6, 30u); // 20 rotl + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 21 load + r3 = r3 ^ r5; // 22 xor + r5 = r5 + r7 + ((((sel >> 6u) & 1u) != 0u) ? 0x950603c6u : 0x1d4f8db7u); // 23 add + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 24 load + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 load + r2 = r2 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x0ae476a4u : 0x55db0d92u); // 26 add + r5 = r5 ^ ds[r6 & mask]; // 27 load + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 28 load + r1 = r1 ^ r6; // 29 xor + r7 = mul_hi(r7, r4); // 30 mulhi + r6 = rotr_var(r6, r3); // 31 rotr + r0 = r0 ^ ds[r3 & mask]; // 32 load + r5 = rotr_var(r5, r0); // 33 rotr + r5 = rotl_imm(r5, 20u); // 34 rotl + r0 = mul_hi(r0, r4); // 35 mulhi + r6 = r6 + r7 + ((((sel >> 18u) & 1u) != 0u) ? 0xfe7a0454u : 0x5d6e3dc5u); // 36 add + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 37 load + r2 = rotl_imm(r2, 19u); // 38 rotl + r3 = r3 + r6 + ((((sel >> 20u) & 1u) != 0u) ? 0xc1d4ae24u : 0xe7e241bau); // 39 add + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 40 load + r3 = r3 + r1 + ((((sel >> 21u) & 1u) != 0u) ? 0x8f9f8556u : 0xb45cdd60u); // 41 add + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 42 shfl + r0 = r0 * r1; // 43 mul + r0 = rotr_var(r0, r7); // 44 rotr + r5 = r5 - r3; // 45 sub + r2 = r7 * r7 + r2; // 46 mad + r6 = r3 * r1 + r6; // 47 mad + r0 = r0 * r7; // 48 mul + r0 = r0 + r4 + ((((sel >> 1u) & 1u) != 0u) ? 0xb9083b6cu : 0x3fa0c5cdu); // 49 add + r4 = r4 + r6 + ((((sel >> 13u) & 1u) != 0u) ? 0x9eecc778u : 0x2d6873e8u); // 50 add + r2 = r2 - r7; // 51 sub + r6 = rotl_imm(r6, 6u); // 52 rotl + r3 = rotr_var(r3, r5); // 53 rotr + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 54 load + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 55 load + r2 = r2 | r6; // 56 or + r4 = r4 ^ ds[r0 & mask]; // 57 load + r3 = r3 ^ r6; // 58 xor + r0 = r0 ^ ds[r6 & mask]; // 59 load + r2 = r2 + r4 + ((((sel >> 3u) & 1u) != 0u) ? 0x19c76fb3u : 0x5052a1c3u); // 60 add + r7 = r7 ^ ds[r4 & mask]; // 61 load + r6 = r6 | r1; // 62 or + r5 = r5 + r1 + ((((sel >> 4u) & 1u) != 0u) ? 0x535fe3fau : 0x08c757c0u); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-1/kernel_bound.cu b/proto-cuda/packs-readwidth/mixB-1/kernel_bound.cu new file mode 100644 index 000000000..a3d1173f5 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-1/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/1". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r5 = r4 * r0 + r5; // 0 mad + r1 = rotl_imm(r1, 23u); // 1 rotl + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 2 load + r5 = r4 * r2 + r5; // 3 mad + r1 = r1 * r6; // 4 mul + r1 = r1 + r0 + ((((sel >> 6u) & 1u) != 0u) ? 0x9dd9fb05u : 0x9db46598u); // 5 add + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r7, 16); // 6 shfl + r6 = r6 * r1; // 7 mul + r1 = __umulhi(r1, r7); // 8 mulhi + r2 = __umulhi(r2, r1); // 9 mulhi + r0 = r0 + r4 + ((((sel >> 13u) & 1u) != 0u) ? 0xc15822b6u : 0xe45fccebu); // 10 add + r5 = r1 * r2 + r5; // 11 mad + r4 = r4 * r7; // 12 mul + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 13 load + r3 = r3 ^ r4; // 14 xor + r1 = r1 ^ r6; // 15 xor + r5 = rotr_var(r5, r6); // 16 rotr + r3 = rotr_var(r3, r7); // 17 rotr + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 18 shfl + r5 = r5 ^ ds[r6 & mask]; // 19 load + r6 = rotl_imm(r6, 30u); // 20 rotl + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 21 load + r3 = r3 ^ r5; // 22 xor + r5 = r5 + r7 + ((((sel >> 6u) & 1u) != 0u) ? 0x950603c6u : 0x1d4f8db7u); // 23 add + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 24 load + { uint32_t b_ = (r1 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 load + r2 = r2 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x0ae476a4u : 0x55db0d92u); // 26 add + r5 = r5 ^ ds[r6 & mask]; // 27 load + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 28 load + r1 = r1 ^ r6; // 29 xor + r7 = __umulhi(r7, r4); // 30 mulhi + r6 = rotr_var(r6, r3); // 31 rotr + r0 = r0 ^ ds[r3 & mask]; // 32 load + r5 = rotr_var(r5, r0); // 33 rotr + r5 = rotl_imm(r5, 20u); // 34 rotl + r0 = __umulhi(r0, r4); // 35 mulhi + r6 = r6 + r7 + ((((sel >> 18u) & 1u) != 0u) ? 0xfe7a0454u : 0x5d6e3dc5u); // 36 add + { uint32_t b_ = (r1 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 37 load + r2 = rotl_imm(r2, 19u); // 38 rotl + r3 = r3 + r6 + ((((sel >> 20u) & 1u) != 0u) ? 0xc1d4ae24u : 0xe7e241bau); // 39 add + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 40 load + r3 = r3 + r1 + ((((sel >> 21u) & 1u) != 0u) ? 0x8f9f8556u : 0xb45cdd60u); // 41 add + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 42 shfl + r0 = r0 * r1; // 43 mul + r0 = rotr_var(r0, r7); // 44 rotr + r5 = r5 - r3; // 45 sub + r2 = r7 * r7 + r2; // 46 mad + r6 = r3 * r1 + r6; // 47 mad + r0 = r0 * r7; // 48 mul + r0 = r0 + r4 + ((((sel >> 1u) & 1u) != 0u) ? 0xb9083b6cu : 0x3fa0c5cdu); // 49 add + r4 = r4 + r6 + ((((sel >> 13u) & 1u) != 0u) ? 0x9eecc778u : 0x2d6873e8u); // 50 add + r2 = r2 - r7; // 51 sub + r6 = rotl_imm(r6, 6u); // 52 rotl + r3 = rotr_var(r3, r5); // 53 rotr + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 54 load + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 55 load + r2 = r2 | r6; // 56 or + r4 = r4 ^ ds[r0 & mask]; // 57 load + r3 = r3 ^ r6; // 58 xor + r0 = r0 ^ ds[r6 & mask]; // 59 load + r2 = r2 + r4 + ((((sel >> 3u) & 1u) != 0u) ? 0x19c76fb3u : 0x5052a1c3u); // 60 add + r7 = r7 ^ ds[r4 & mask]; // 61 load + r6 = r6 | r1; // 62 or + r5 = r5 + r1 + ((((sel >> 4u) & 1u) != 0u) ? 0x535fe3fau : 0x08c757c0u); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixB-1/memhard.h b/proto-cuda/packs-readwidth/mixB-1/memhard.h new file mode 100644 index 000000000..d5634bf3d --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-1/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/1". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/mixB-1/memhard.metal b/proto-cuda/packs-readwidth/mixB-1/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-1/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/mixB-1/program.h b/proto-cuda/packs-readwidth/mixB-1/program.h new file mode 100644 index 000000000..666784a72 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-1/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/1". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-readwidth/B/1" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d7265616477696474682f422f31" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x695b119dcb76931bull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=11 rotr=6 mad=5 mul=5 rotl=5 xor=5 mulhi=4 shfl=3 or=2 sub=2" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "mix25-50-25" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 25, 50, 25 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 6, 8, 2 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 2240 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x3e345412u, 0xfc2b3a0eu, 0x7740353du, 0x2df159dbu, 0x50bc6eedu, 0x2e61acb2u, 0x3c800bffu, 0xbaecd8c1u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/mixB-1/program.json b/proto-cuda/packs-readwidth/mixB-1/program.json new file mode 100644 index 000000000..dce602122 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-1/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x695b119dcb76931b", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-readwidth/B/1", + "seed_bytes": "69676e65756d2d7265616477696474682f422f31", + "seed_words": ["0x3e345412", "0xfc2b3a0e", "0x7740353d", "0x2df159db", "0x50bc6eed", "0x2e61acb2", "0x3c800bff", "0xbaecd8c1"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "mix25-50-25", + "load_slots": 16, + "load_mix_percent_4_16_64": [25, 50, 25], + "load_width_counts_4_16_64": [6, 8, 2], + "bytes_per_hash": 2240, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "add": 11, "rotr": 6, "mad": 5, "mul": 5, "rotl": 5, "xor": 5, "mulhi": 4, "shfl": 3, "or": 2, "sub": 2}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mad", "dst": 5, "src": 4, "src2": 0, "imm": "0x392cc69d", "imm2": "0xdff38f87", "rot": 18, "bit": 19, "mask": 2, "width": 1}, + {"i": 1, "op": "rotl", "dst": 1, "src": 7, "src2": 0, "imm": "0xa731596a", "imm2": "0x4ce406f5", "rot": 23, "bit": 27, "mask": 4, "width": 1}, + {"i": 2, "op": "load", "dst": 0, "src": 5, "src2": 5, "imm": "0x4d78cee5", "imm2": "0x7a624dd7", "rot": 16, "bit": 24, "mask": 16, "width": 4}, + {"i": 3, "op": "mad", "dst": 5, "src": 4, "src2": 2, "imm": "0x80a8d418", "imm2": "0x3f405868", "rot": 21, "bit": 4, "mask": 8, "width": 1}, + {"i": 4, "op": "mul", "dst": 1, "src": 6, "src2": 4, "imm": "0xc4e8867c", "imm2": "0xd2ad8c44", "rot": 26, "bit": 21, "mask": 16, "width": 1}, + {"i": 5, "op": "add", "dst": 1, "src": 0, "src2": 6, "imm": "0x9db46598", "imm2": "0x9dd9fb05", "rot": 14, "bit": 6, "mask": 16, "width": 1}, + {"i": 6, "op": "shfl", "dst": 1, "src": 7, "src2": 5, "imm": "0xc9049c0f", "imm2": "0x007f6621", "rot": 24, "bit": 19, "mask": 16, "width": 1}, + {"i": 7, "op": "mul", "dst": 6, "src": 1, "src2": 7, "imm": "0x073fea86", "imm2": "0x1ed9dff1", "rot": 26, "bit": 24, "mask": 2, "width": 1}, + {"i": 8, "op": "mulhi", "dst": 1, "src": 7, "src2": 5, "imm": "0x3faa8f96", "imm2": "0x01445abc", "rot": 26, "bit": 3, "mask": 4, "width": 1}, + {"i": 9, "op": "mulhi", "dst": 2, "src": 1, "src2": 1, "imm": "0x97d622e9", "imm2": "0xfd732e57", "rot": 28, "bit": 29, "mask": 1, "width": 1}, + {"i": 10, "op": "add", "dst": 0, "src": 4, "src2": 4, "imm": "0xe45fcceb", "imm2": "0xc15822b6", "rot": 15, "bit": 13, "mask": 2, "width": 1}, + {"i": 11, "op": "mad", "dst": 5, "src": 1, "src2": 2, "imm": "0x6148c1cc", "imm2": "0xdaa7b1d3", "rot": 3, "bit": 25, "mask": 16, "width": 1}, + {"i": 12, "op": "mul", "dst": 4, "src": 7, "src2": 1, "imm": "0x7b0965e5", "imm2": "0x5086e5aa", "rot": 30, "bit": 28, "mask": 16, "width": 1}, + {"i": 13, "op": "load", "dst": 0, "src": 5, "src2": 4, "imm": "0x4c87a2ab", "imm2": "0x5db5ba1c", "rot": 22, "bit": 19, "mask": 16, "width": 16}, + {"i": 14, "op": "xor", "dst": 3, "src": 4, "src2": 6, "imm": "0xdc246d40", "imm2": "0xa5cfff4a", "rot": 6, "bit": 19, "mask": 16, "width": 1}, + {"i": 15, "op": "xor", "dst": 1, "src": 6, "src2": 6, "imm": "0x6d3e113d", "imm2": "0x69e25584", "rot": 12, "bit": 27, "mask": 2, "width": 1}, + {"i": 16, "op": "rotr", "dst": 5, "src": 6, "src2": 2, "imm": "0x45569439", "imm2": "0x14d0f916", "rot": 10, "bit": 5, "mask": 16, "width": 1}, + {"i": 17, "op": "rotr", "dst": 3, "src": 7, "src2": 1, "imm": "0x412cb31a", "imm2": "0xe0713d9c", "rot": 17, "bit": 0, "mask": 2, "width": 1}, + {"i": 18, "op": "shfl", "dst": 1, "src": 4, "src2": 3, "imm": "0x2419b99c", "imm2": "0x9ab3a71e", "rot": 17, "bit": 15, "mask": 16, "width": 1}, + {"i": 19, "op": "load", "dst": 5, "src": 6, "src2": 7, "imm": "0xa1f88336", "imm2": "0xca05a80c", "rot": 17, "bit": 22, "mask": 8, "width": 1}, + {"i": 20, "op": "rotl", "dst": 6, "src": 7, "src2": 4, "imm": "0x51de29c5", "imm2": "0xfe2d2af0", "rot": 30, "bit": 20, "mask": 1, "width": 1}, + {"i": 21, "op": "load", "dst": 5, "src": 2, "src2": 6, "imm": "0x5db413b6", "imm2": "0x3dad8864", "rot": 20, "bit": 23, "mask": 4, "width": 4}, + {"i": 22, "op": "xor", "dst": 3, "src": 5, "src2": 0, "imm": "0xdeadb225", "imm2": "0x7397a1f2", "rot": 5, "bit": 3, "mask": 2, "width": 1}, + {"i": 23, "op": "add", "dst": 5, "src": 7, "src2": 7, "imm": "0x1d4f8db7", "imm2": "0x950603c6", "rot": 19, "bit": 6, "mask": 1, "width": 1}, + {"i": 24, "op": "load", "dst": 3, "src": 4, "src2": 5, "imm": "0x148ee6e2", "imm2": "0xfc0cbfef", "rot": 1, "bit": 23, "mask": 16, "width": 4}, + {"i": 25, "op": "load", "dst": 0, "src": 1, "src2": 1, "imm": "0x114e77c3", "imm2": "0x8b1c13f1", "rot": 4, "bit": 0, "mask": 8, "width": 16}, + {"i": 26, "op": "add", "dst": 2, "src": 3, "src2": 3, "imm": "0x55db0d92", "imm2": "0x0ae476a4", "rot": 24, "bit": 31, "mask": 4, "width": 1}, + {"i": 27, "op": "load", "dst": 5, "src": 6, "src2": 5, "imm": "0xee300382", "imm2": "0x1609cf99", "rot": 8, "bit": 11, "mask": 16, "width": 1}, + {"i": 28, "op": "load", "dst": 2, "src": 5, "src2": 2, "imm": "0x96783a44", "imm2": "0x0f090418", "rot": 11, "bit": 24, "mask": 1, "width": 4}, + {"i": 29, "op": "xor", "dst": 1, "src": 6, "src2": 7, "imm": "0x6c8d15c2", "imm2": "0x2cde0f35", "rot": 18, "bit": 14, "mask": 4, "width": 1}, + {"i": 30, "op": "mulhi", "dst": 7, "src": 4, "src2": 1, "imm": "0xc68ea9de", "imm2": "0xc45d9dfa", "rot": 29, "bit": 14, "mask": 2, "width": 1}, + {"i": 31, "op": "rotr", "dst": 6, "src": 3, "src2": 5, "imm": "0x2b816dd1", "imm2": "0xdf13f697", "rot": 17, "bit": 12, "mask": 2, "width": 1}, + {"i": 32, "op": "load", "dst": 0, "src": 3, "src2": 3, "imm": "0x6711453d", "imm2": "0x78ae13e9", "rot": 2, "bit": 6, "mask": 4, "width": 1}, + {"i": 33, "op": "rotr", "dst": 5, "src": 0, "src2": 3, "imm": "0xe0068514", "imm2": "0xba437440", "rot": 14, "bit": 2, "mask": 2, "width": 1}, + {"i": 34, "op": "rotl", "dst": 5, "src": 3, "src2": 7, "imm": "0xdde8ab9a", "imm2": "0x9b3f4073", "rot": 20, "bit": 26, "mask": 1, "width": 1}, + {"i": 35, "op": "mulhi", "dst": 0, "src": 4, "src2": 2, "imm": "0x2b47ae8c", "imm2": "0x2617710c", "rot": 30, "bit": 29, "mask": 1, "width": 1}, + {"i": 36, "op": "add", "dst": 6, "src": 7, "src2": 0, "imm": "0x5d6e3dc5", "imm2": "0xfe7a0454", "rot": 22, "bit": 18, "mask": 8, "width": 1}, + {"i": 37, "op": "load", "dst": 2, "src": 1, "src2": 7, "imm": "0x21116475", "imm2": "0xbe53b7b8", "rot": 21, "bit": 2, "mask": 8, "width": 4}, + {"i": 38, "op": "rotl", "dst": 2, "src": 7, "src2": 2, "imm": "0xd76a18c4", "imm2": "0xb7ed0d92", "rot": 19, "bit": 10, "mask": 16, "width": 1}, + {"i": 39, "op": "add", "dst": 3, "src": 6, "src2": 3, "imm": "0xe7e241ba", "imm2": "0xc1d4ae24", "rot": 10, "bit": 20, "mask": 16, "width": 1}, + {"i": 40, "op": "load", "dst": 4, "src": 7, "src2": 0, "imm": "0xb1006316", "imm2": "0x617eb11a", "rot": 1, "bit": 27, "mask": 1, "width": 4}, + {"i": 41, "op": "add", "dst": 3, "src": 1, "src2": 3, "imm": "0xb45cdd60", "imm2": "0x8f9f8556", "rot": 5, "bit": 21, "mask": 16, "width": 1}, + {"i": 42, "op": "shfl", "dst": 5, "src": 2, "src2": 1, "imm": "0x8fc745c7", "imm2": "0x98256336", "rot": 16, "bit": 7, "mask": 4, "width": 1}, + {"i": 43, "op": "mul", "dst": 0, "src": 1, "src2": 3, "imm": "0xe8aea431", "imm2": "0xfeed34f6", "rot": 30, "bit": 2, "mask": 8, "width": 1}, + {"i": 44, "op": "rotr", "dst": 0, "src": 7, "src2": 4, "imm": "0x8f1da3ee", "imm2": "0xa9c20edf", "rot": 19, "bit": 17, "mask": 2, "width": 1}, + {"i": 45, "op": "sub", "dst": 5, "src": 3, "src2": 7, "imm": "0x9da4c5e7", "imm2": "0xbcd50b0c", "rot": 7, "bit": 18, "mask": 1, "width": 1}, + {"i": 46, "op": "mad", "dst": 2, "src": 7, "src2": 7, "imm": "0x837be930", "imm2": "0x9b62650b", "rot": 5, "bit": 23, "mask": 1, "width": 1}, + {"i": 47, "op": "mad", "dst": 6, "src": 3, "src2": 1, "imm": "0x825afaa4", "imm2": "0x0cf9f1f0", "rot": 17, "bit": 9, "mask": 2, "width": 1}, + {"i": 48, "op": "mul", "dst": 0, "src": 7, "src2": 2, "imm": "0x4424f910", "imm2": "0xbaeba671", "rot": 23, "bit": 6, "mask": 4, "width": 1}, + {"i": 49, "op": "add", "dst": 0, "src": 4, "src2": 7, "imm": "0x3fa0c5cd", "imm2": "0xb9083b6c", "rot": 28, "bit": 1, "mask": 2, "width": 1}, + {"i": 50, "op": "add", "dst": 4, "src": 6, "src2": 4, "imm": "0x2d6873e8", "imm2": "0x9eecc778", "rot": 29, "bit": 13, "mask": 16, "width": 1}, + {"i": 51, "op": "sub", "dst": 2, "src": 7, "src2": 2, "imm": "0xace9966d", "imm2": "0xe4213c89", "rot": 30, "bit": 18, "mask": 8, "width": 1}, + {"i": 52, "op": "rotl", "dst": 6, "src": 7, "src2": 0, "imm": "0x03bc6330", "imm2": "0x32d9aee6", "rot": 6, "bit": 19, "mask": 1, "width": 1}, + {"i": 53, "op": "rotr", "dst": 3, "src": 5, "src2": 1, "imm": "0xef7e5b3e", "imm2": "0xb2aea388", "rot": 27, "bit": 31, "mask": 2, "width": 1}, + {"i": 54, "op": "load", "dst": 3, "src": 2, "src2": 2, "imm": "0x711b4f5d", "imm2": "0xac8e6837", "rot": 1, "bit": 5, "mask": 1, "width": 4}, + {"i": 55, "op": "load", "dst": 7, "src": 5, "src2": 7, "imm": "0xb0f4faf4", "imm2": "0xb13fe736", "rot": 16, "bit": 2, "mask": 4, "width": 4}, + {"i": 56, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0x4f416150", "imm2": "0x2a9f59e3", "rot": 31, "bit": 6, "mask": 4, "width": 1}, + {"i": 57, "op": "load", "dst": 4, "src": 0, "src2": 0, "imm": "0x54ce0fba", "imm2": "0x48d98570", "rot": 11, "bit": 2, "mask": 4, "width": 1}, + {"i": 58, "op": "xor", "dst": 3, "src": 6, "src2": 6, "imm": "0xf34a7cca", "imm2": "0xf345c73d", "rot": 18, "bit": 18, "mask": 16, "width": 1}, + {"i": 59, "op": "load", "dst": 0, "src": 6, "src2": 6, "imm": "0x730474e9", "imm2": "0xb90ac73f", "rot": 25, "bit": 31, "mask": 4, "width": 1}, + {"i": 60, "op": "add", "dst": 2, "src": 4, "src2": 1, "imm": "0x5052a1c3", "imm2": "0x19c76fb3", "rot": 12, "bit": 3, "mask": 1, "width": 1}, + {"i": 61, "op": "load", "dst": 7, "src": 4, "src2": 0, "imm": "0xbecfc16b", "imm2": "0xd2837365", "rot": 21, "bit": 22, "mask": 8, "width": 1}, + {"i": 62, "op": "or", "dst": 6, "src": 1, "src2": 3, "imm": "0x7353a546", "imm2": "0xcc33abd9", "rot": 24, "bit": 5, "mask": 4, "width": 1}, + {"i": 63, "op": "add", "dst": 5, "src": 1, "src2": 1, "imm": "0x08c757c0", "imm2": "0x535fe3fa", "rot": 21, "bit": 4, "mask": 4, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/mixB-1/program.metal b/proto-cuda/packs-readwidth/mixB-1/program.metal new file mode 100644 index 000000000..ea7fe67f5 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-1/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x3e345412u, 0xfc2b3a0eu, 0x7740353du, 0x2df159dbu, 0x50bc6eedu, 0x2e61acb2u, 0x3c800bffu, 0xbaecd8c1u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r5 = r4 * r0 + r5; // 0 + r1 = rotl_imm(r1, 23u); // 1 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 2 + r5 = r4 * r2 + r5; // 3 + r1 = r1 * r6; // 4 + r1 = r1 + r0 + select(0x9db46598u, 0x9dd9fb05u, ((sel >> 6u) & 1u) != 0u); // 5 + r1 = r1 ^ simd_shuffle_xor(r7, (ushort)16); // 6 + r6 = r6 * r1; // 7 + r1 = mulhi(r1, r7); // 8 + r2 = mulhi(r2, r1); // 9 + r0 = r0 + r4 + select(0xe45fccebu, 0xc15822b6u, ((sel >> 13u) & 1u) != 0u); // 10 + r5 = r1 * r2 + r5; // 11 + r4 = r4 * r7; // 12 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 13 + r3 = r3 ^ r4; // 14 + r1 = r1 ^ r6; // 15 + r5 = rotr_var(r5, r6); // 16 + r3 = rotr_var(r3, r7); // 17 + r1 = r1 ^ simd_shuffle_xor(r4, (ushort)16); // 18 + r5 = r5 ^ dataset[r6 & MASK]; // 19 + r6 = rotl_imm(r6, 30u); // 20 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 21 + r3 = r3 ^ r5; // 22 + r5 = r5 + r7 + select(0x1d4f8db7u, 0x950603c6u, ((sel >> 6u) & 1u) != 0u); // 23 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 24 + { uint b_ = (r1 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 + r2 = r2 + r3 + select(0x55db0d92u, 0x0ae476a4u, ((sel >> 31u) & 1u) != 0u); // 26 + r5 = r5 ^ dataset[r6 & MASK]; // 27 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 28 + r1 = r1 ^ r6; // 29 + r7 = mulhi(r7, r4); // 30 + r6 = rotr_var(r6, r3); // 31 + r0 = r0 ^ dataset[r3 & MASK]; // 32 + r5 = rotr_var(r5, r0); // 33 + r5 = rotl_imm(r5, 20u); // 34 + r0 = mulhi(r0, r4); // 35 + r6 = r6 + r7 + select(0x5d6e3dc5u, 0xfe7a0454u, ((sel >> 18u) & 1u) != 0u); // 36 + { uint b_ = (r1 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 37 + r2 = rotl_imm(r2, 19u); // 38 + r3 = r3 + r6 + select(0xe7e241bau, 0xc1d4ae24u, ((sel >> 20u) & 1u) != 0u); // 39 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 40 + r3 = r3 + r1 + select(0xb45cdd60u, 0x8f9f8556u, ((sel >> 21u) & 1u) != 0u); // 41 + r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 42 + r0 = r0 * r1; // 43 + r0 = rotr_var(r0, r7); // 44 + r5 = r5 - r3; // 45 + r2 = r7 * r7 + r2; // 46 + r6 = r3 * r1 + r6; // 47 + r0 = r0 * r7; // 48 + r0 = r0 + r4 + select(0x3fa0c5cdu, 0xb9083b6cu, ((sel >> 1u) & 1u) != 0u); // 49 + r4 = r4 + r6 + select(0x2d6873e8u, 0x9eecc778u, ((sel >> 13u) & 1u) != 0u); // 50 + r2 = r2 - r7; // 51 + r6 = rotl_imm(r6, 6u); // 52 + r3 = rotr_var(r3, r5); // 53 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 54 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 55 + r2 = r2 | r6; // 56 + r4 = r4 ^ dataset[r0 & MASK]; // 57 + r3 = r3 ^ r6; // 58 + r0 = r0 ^ dataset[r6 & MASK]; // 59 + r2 = r2 + r4 + select(0x5052a1c3u, 0x19c76fb3u, ((sel >> 3u) & 1u) != 0u); // 60 + r7 = r7 ^ dataset[r4 & MASK]; // 61 + r6 = r6 | r1; // 62 + r5 = r5 + r1 + select(0x08c757c0u, 0x535fe3fau, ((sel >> 4u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-1/program_bound.metal b/proto-cuda/packs-readwidth/mixB-1/program_bound.metal new file mode 100644 index 000000000..8d5db3a05 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-1/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x3e345412u, 0xfc2b3a0eu, 0x7740353du, 0x2df159dbu, 0x50bc6eedu, 0x2e61acb2u, 0x3c800bffu, 0xbaecd8c1u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r5 = r4 * r0 + r5; // 0 + r1 = rotl_imm(r1, 23u); // 1 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 2 + r5 = r4 * r2 + r5; // 3 + r1 = r1 * r6; // 4 + r1 = r1 + r0 + select(0x9db46598u, 0x9dd9fb05u, ((sel >> 6u) & 1u) != 0u); // 5 + r1 = r1 ^ simd_shuffle_xor(r7, (ushort)16); // 6 + r6 = r6 * r1; // 7 + r1 = mulhi(r1, r7); // 8 + r2 = mulhi(r2, r1); // 9 + r0 = r0 + r4 + select(0xe45fccebu, 0xc15822b6u, ((sel >> 13u) & 1u) != 0u); // 10 + r5 = r1 * r2 + r5; // 11 + r4 = r4 * r7; // 12 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 13 + r3 = r3 ^ r4; // 14 + r1 = r1 ^ r6; // 15 + r5 = rotr_var(r5, r6); // 16 + r3 = rotr_var(r3, r7); // 17 + r1 = r1 ^ simd_shuffle_xor(r4, (ushort)16); // 18 + r5 = r5 ^ dataset[r6 & MASK]; // 19 + r6 = rotl_imm(r6, 30u); // 20 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 21 + r3 = r3 ^ r5; // 22 + r5 = r5 + r7 + select(0x1d4f8db7u, 0x950603c6u, ((sel >> 6u) & 1u) != 0u); // 23 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 24 + { uint b_ = (r1 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 25 + r2 = r2 + r3 + select(0x55db0d92u, 0x0ae476a4u, ((sel >> 31u) & 1u) != 0u); // 26 + r5 = r5 ^ dataset[r6 & MASK]; // 27 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 28 + r1 = r1 ^ r6; // 29 + r7 = mulhi(r7, r4); // 30 + r6 = rotr_var(r6, r3); // 31 + r0 = r0 ^ dataset[r3 & MASK]; // 32 + r5 = rotr_var(r5, r0); // 33 + r5 = rotl_imm(r5, 20u); // 34 + r0 = mulhi(r0, r4); // 35 + r6 = r6 + r7 + select(0x5d6e3dc5u, 0xfe7a0454u, ((sel >> 18u) & 1u) != 0u); // 36 + { uint b_ = (r1 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 37 + r2 = rotl_imm(r2, 19u); // 38 + r3 = r3 + r6 + select(0xe7e241bau, 0xc1d4ae24u, ((sel >> 20u) & 1u) != 0u); // 39 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 40 + r3 = r3 + r1 + select(0xb45cdd60u, 0x8f9f8556u, ((sel >> 21u) & 1u) != 0u); // 41 + r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 42 + r0 = r0 * r1; // 43 + r0 = rotr_var(r0, r7); // 44 + r5 = r5 - r3; // 45 + r2 = r7 * r7 + r2; // 46 + r6 = r3 * r1 + r6; // 47 + r0 = r0 * r7; // 48 + r0 = r0 + r4 + select(0x3fa0c5cdu, 0xb9083b6cu, ((sel >> 1u) & 1u) != 0u); // 49 + r4 = r4 + r6 + select(0x2d6873e8u, 0x9eecc778u, ((sel >> 13u) & 1u) != 0u); // 50 + r2 = r2 - r7; // 51 + r6 = rotl_imm(r6, 6u); // 52 + r3 = rotr_var(r3, r5); // 53 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 54 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 55 + r2 = r2 | r6; // 56 + r4 = r4 ^ dataset[r0 & MASK]; // 57 + r3 = r3 ^ r6; // 58 + r0 = r0 ^ dataset[r6 & MASK]; // 59 + r2 = r2 + r4 + select(0x5052a1c3u, 0x19c76fb3u, ((sel >> 3u) & 1u) != 0u); // 60 + r7 = r7 ^ dataset[r4 & MASK]; // 61 + r6 = r6 | r1; // 62 + r5 = r5 + r1 + select(0x08c757c0u, 0x535fe3fau, ((sel >> 4u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-1/vectors.h b/proto-cuda/packs-readwidth/mixB-1/vectors.h new file mode 100644 index 000000000..c9edd36c4 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-1/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/1". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x60775ffcf5642b8aull, 0x81153513adbc62d9ull, 0x5f08ab8decbfdc53ull, 0x42da80d8c852cf66ull, 0xc67335da31bdbc1dull, 0x0e5c929f9e96f35cull, 0x7e18bf8430170008ull, 0x5d1c508b8f356820ull, + 0xd83eaf5de2ad239bull, 0xccd00a56f293e023ull, 0xdc76d8d77b3c39eaull, 0x5baf712642b867aeull, 0x0539c20a48108b15ull, 0x0b1e50072f7f4458ull, 0x4d4169021ad5dd68ull, 0x0dfb7783d8e7e38aull, + 0x1b4b3088be627391ull, 0xf8642f256eb7c612ull, 0xcf32c47692c137daull, 0xd12c35d740e483a9ull, 0x29040a2ef59aa3c0ull, 0x239347afbc6f9b62ull, 0x8a7630e3151d29e4ull, 0x139ee6f308b29f40ull, + 0xce6d8a4ddc846086ull, 0xf6707441e743095aull, 0x56b1f8eb03c684a5ull, 0xd81db5b4de54b416ull, 0xfc47e52c27bc553dull, 0x754ff2162ac44ef3ull, 0xdb0f41a8ae0176f2ull, 0xf4cca925351ea46dull + }, + { // base nonce 4096 + 0x47946de7679105a6ull, 0x23fcd01f25f295c5ull, 0x7b5c44b644169a6eull, 0xb33d0c92b0d1a4a6ull, 0x7ec35998c8ff4f8eull, 0xab8e6812ec821db3ull, 0xb9ad96e53eebad3full, 0x441c94af57e6f816ull, + 0x1ec6cff841c73fa2ull, 0x45302cc1681ef509ull, 0x38c2150a193a9c2bull, 0x16c44d87bfeaf351ull, 0x179f0f5a86b985afull, 0x1ff0cd7b7c80b34aull, 0x9037de8b3e9790f3ull, 0x53644ba4c9e5978cull, + 0xfce1a03e7d8b97a0ull, 0xab1c36fe3eeb1bb0ull, 0x318c491c6ee204d9ull, 0x1f442efcf0b852c4ull, 0xd43383b2ba9ca08bull, 0x6841bfc3c0393f51ull, 0x3b103946efd2c16bull, 0xbfbb60c4b3528925ull, + 0x4188a3ece0e41779ull, 0xe8bb37faa383cf45ull, 0xbb31bda0fbaa8118ull, 0x38f155eff657623cull, 0xe6244ebbb877c9feull, 0x8e39d2e5284a3260ull, 0x6262c417a65ecd26ull, 0x3358616a2def28ffull + }, + { // base nonce 1000000 + 0x13c4b5e8ff8fc9bdull, 0xced78f207a2feae6ull, 0x32a518c8613d08a7ull, 0x555ba41f5021fea8ull, 0xaaf5baab577f37faull, 0xc0dff1420678e1cbull, 0x4cc930dd9e5a3fb3ull, 0xc9eff4c0986a0b3bull, + 0xa35410edfd25cddeull, 0x72ac270980fd3cebull, 0xbc5a0885d08d9c10ull, 0xbc3d70e3ef208949ull, 0x5acf0e43a6c574d3ull, 0x06286239ed47b903ull, 0xccc41c90c92d4ff2ull, 0x7dadcf9e7b791bb4ull, + 0x5aa277c92b18f158ull, 0x9bdcd7eb338ac712ull, 0xbe55aca71c17ea7dull, 0x65f688425cf79e58ull, 0xcf171a7123934ec4ull, 0x246d0a06019cd4e8ull, 0x68187456196da641ull, 0x6c07711c642b42b0ull, + 0x98a57208698c592aull, 0x46687cb1c03c61a3ull, 0x1cb58f1307036ceaull, 0x8c52c84ccad02e69ull, 0x517bebcfde66fd0dull, 0x015c92931bf627d1ull, 0xb564a819574c6d14ull, 0x1f69f57d613a0d39ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/mixB-1/vectors.json b/proto-cuda/packs-readwidth/mixB-1/vectors.json new file mode 100644 index 000000000..1e6b8a8c6 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-1/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-readwidth/B/1", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x60775ffcf5642b8a", "0x81153513adbc62d9", "0x5f08ab8decbfdc53", "0x42da80d8c852cf66", "0xc67335da31bdbc1d", "0x0e5c929f9e96f35c", "0x7e18bf8430170008", "0x5d1c508b8f356820", + "0xd83eaf5de2ad239b", "0xccd00a56f293e023", "0xdc76d8d77b3c39ea", "0x5baf712642b867ae", "0x0539c20a48108b15", "0x0b1e50072f7f4458", "0x4d4169021ad5dd68", "0x0dfb7783d8e7e38a", + "0x1b4b3088be627391", "0xf8642f256eb7c612", "0xcf32c47692c137da", "0xd12c35d740e483a9", "0x29040a2ef59aa3c0", "0x239347afbc6f9b62", "0x8a7630e3151d29e4", "0x139ee6f308b29f40", + "0xce6d8a4ddc846086", "0xf6707441e743095a", "0x56b1f8eb03c684a5", "0xd81db5b4de54b416", "0xfc47e52c27bc553d", "0x754ff2162ac44ef3", "0xdb0f41a8ae0176f2", "0xf4cca925351ea46d" + ]}, + {"base_nonce": 4096, "expected": [ + "0x47946de7679105a6", "0x23fcd01f25f295c5", "0x7b5c44b644169a6e", "0xb33d0c92b0d1a4a6", "0x7ec35998c8ff4f8e", "0xab8e6812ec821db3", "0xb9ad96e53eebad3f", "0x441c94af57e6f816", + "0x1ec6cff841c73fa2", "0x45302cc1681ef509", "0x38c2150a193a9c2b", "0x16c44d87bfeaf351", "0x179f0f5a86b985af", "0x1ff0cd7b7c80b34a", "0x9037de8b3e9790f3", "0x53644ba4c9e5978c", + "0xfce1a03e7d8b97a0", "0xab1c36fe3eeb1bb0", "0x318c491c6ee204d9", "0x1f442efcf0b852c4", "0xd43383b2ba9ca08b", "0x6841bfc3c0393f51", "0x3b103946efd2c16b", "0xbfbb60c4b3528925", + "0x4188a3ece0e41779", "0xe8bb37faa383cf45", "0xbb31bda0fbaa8118", "0x38f155eff657623c", "0xe6244ebbb877c9fe", "0x8e39d2e5284a3260", "0x6262c417a65ecd26", "0x3358616a2def28ff" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x13c4b5e8ff8fc9bd", "0xced78f207a2feae6", "0x32a518c8613d08a7", "0x555ba41f5021fea8", "0xaaf5baab577f37fa", "0xc0dff1420678e1cb", "0x4cc930dd9e5a3fb3", "0xc9eff4c0986a0b3b", + "0xa35410edfd25cdde", "0x72ac270980fd3ceb", "0xbc5a0885d08d9c10", "0xbc3d70e3ef208949", "0x5acf0e43a6c574d3", "0x06286239ed47b903", "0xccc41c90c92d4ff2", "0x7dadcf9e7b791bb4", + "0x5aa277c92b18f158", "0x9bdcd7eb338ac712", "0xbe55aca71c17ea7d", "0x65f688425cf79e58", "0xcf171a7123934ec4", "0x246d0a06019cd4e8", "0x68187456196da641", "0x6c07711c642b42b0", + "0x98a57208698c592a", "0x46687cb1c03c61a3", "0x1cb58f1307036cea", "0x8c52c84ccad02e69", "0x517bebcfde66fd0d", "0x015c92931bf627d1", "0xb564a819574c6d14", "0x1f69f57d613a0d39" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/mixB-2/kernel.cl b/proto-cuda/packs-readwidth/mixB-2/kernel.cl new file mode 100644 index 000000000..897f1eb25 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-2/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/2". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x8c11f8d8u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x945e59f8u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x945e59f8u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x4ab18392u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x4ab18392u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0xa6381ed7u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0xa6381ed7u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x797fadefu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x797fadefu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x188a03acu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x188a03acu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xea7032e0u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xea7032e0u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xcaf097f4u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0xcaf097f4u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x8c11f8d8u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = r0 ^ r6; // 0 xor + r0 = r1 * r4 + r0; // 1 mad + r2 = rotl_imm(r2, 21u); // 2 rotl + r3 = rotr_var(r3, r0); // 3 rotr + r0 = r0 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x474ae392u : 0x7a34b651u); // 4 add + r0 = mul_hi(r0, r3); // 5 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r3 = r3 ^ t_; } // 6 shfl + r7 = rotr_var(r7, r2); // 7 rotr + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 8 load + r4 = r4 | r5; // 9 or + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 10 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 load + r4 = r4 + r0 + ((((sel >> 16u) & 1u) != 0u) ? 0xf58f0389u : 0x99426786u); // 13 add + r0 = r0 + r6 + ((((sel >> 31u) & 1u) != 0u) ? 0xb5bbd552u : 0x6c662222u); // 14 add + r4 = rotr_var(r4, r6); // 15 rotr + r4 = r4 ^ r1; // 16 xor + r3 = rotl_imm(r3, 22u); // 17 rotl + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 18 load + r3 = rotr_var(r3, r0); // 19 rotr + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 20 load + r4 = rotl_imm(r4, 8u); // 21 rotl + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 22 load + r0 = rotr_var(r0, r3); // 23 rotr + r4 = r4 + r1 + ((((sel >> 13u) & 1u) != 0u) ? 0xd1ff184bu : 0x05f4dd31u); // 24 add + r4 = r4 ^ r2; // 25 xor + r3 = mul_hi(r3, r2); // 26 mulhi + r5 = r5 - r7; // 27 sub + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 28 load + r5 = r5 | r7; // 29 or + r4 = rotl_imm(r4, 28u); // 30 rotl + r3 = r3 ^ r5; // 31 xor + r2 = r2 * r1; // 32 mul + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 33 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r5 = r5 ^ t_; } // 34 shfl + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 35 load + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 36 load + r6 = r2 * r2 + r6; // 37 mad + r6 = r6 ^ r3; // 38 xor + r2 = r2 ^ ds[r0 & mask]; // 39 load + r1 = r1 * r2; // 40 mul + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 41 load + r4 = rotl_imm(r4, 31u); // 42 rotl + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 43 load + r4 = r4 + r2 + ((((sel >> 24u) & 1u) != 0u) ? 0xdd1539e6u : 0xb045a668u); // 44 add + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 45 load + r7 = r7 + r6 + ((((sel >> 9u) & 1u) != 0u) ? 0x8a264d7du : 0x85c0036eu); // 46 add + r0 = r1 * r3 + r0; // 47 mad + r7 = r7 | r0; // 48 or + r6 = r6 ^ r4; // 49 xor + r1 = r1 + r7 + ((((sel >> 11u) & 1u) != 0u) ? 0x9d56aa91u : 0x376c63cdu); // 50 add + r5 = r5 * r4; // 51 mul + r7 = rotr_var(r7, r0); // 52 rotr + r2 = r2 ^ r6; // 53 xor + r0 = rotl_imm(r0, 22u); // 54 rotl + r3 = rotr_var(r3, r5); // 55 rotr + r6 = r6 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0x7fbd3c70u : 0x4face1f0u); // 56 add + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 2u); r0 = r0 ^ t_; } // 57 shfl + r2 = r6 * r0 + r2; // 58 mad + r2 = mul_hi(r2, r5); // 59 mulhi + r5 = rotl_imm(r5, 2u); // 60 rotl + r6 = rotl_imm(r6, 28u); // 61 rotl + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 62 load + r7 = mul_hi(r7, r1); // 63 mulhi + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/mixB-2/kernel.cu b/proto-cuda/packs-readwidth/mixB-2/kernel.cu new file mode 100644 index 000000000..fc88a3d71 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-2/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/2". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x8c11f8d8u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x945e59f8u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x945e59f8u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x4ab18392u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0x4ab18392u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0xa6381ed7u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0xa6381ed7u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x797fadefu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0x797fadefu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x188a03acu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0x188a03acu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xea7032e0u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0xea7032e0u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xcaf097f4u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0xcaf097f4u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x8c11f8d8u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r0 = r0 ^ r6; // 0 xor + r0 = r1 * r4 + r0; // 1 mad + r2 = rotl_imm(r2, 21u); // 2 rotl + r3 = rotr_var(r3, r0); // 3 rotr + r0 = r0 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x474ae392u : 0x7a34b651u); // 4 add + r0 = __umulhi(r0, r3); // 5 mulhi + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 4); // 6 shfl + r7 = rotr_var(r7, r2); // 7 rotr + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 8 load + r4 = r4 | r5; // 9 or + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 10 load + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 load + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 load + r4 = r4 + r0 + ((((sel >> 16u) & 1u) != 0u) ? 0xf58f0389u : 0x99426786u); // 13 add + r0 = r0 + r6 + ((((sel >> 31u) & 1u) != 0u) ? 0xb5bbd552u : 0x6c662222u); // 14 add + r4 = rotr_var(r4, r6); // 15 rotr + r4 = r4 ^ r1; // 16 xor + r3 = rotl_imm(r3, 22u); // 17 rotl + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 18 load + r3 = rotr_var(r3, r0); // 19 rotr + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 20 load + r4 = rotl_imm(r4, 8u); // 21 rotl + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 22 load + r0 = rotr_var(r0, r3); // 23 rotr + r4 = r4 + r1 + ((((sel >> 13u) & 1u) != 0u) ? 0xd1ff184bu : 0x05f4dd31u); // 24 add + r4 = r4 ^ r2; // 25 xor + r3 = __umulhi(r3, r2); // 26 mulhi + r5 = r5 - r7; // 27 sub + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 28 load + r5 = r5 | r7; // 29 or + r4 = rotl_imm(r4, 28u); // 30 rotl + r3 = r3 ^ r5; // 31 xor + r2 = r2 * r1; // 32 mul + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 33 load + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 34 shfl + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 35 load + { uint32_t b_ = (r6 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 36 load + r6 = r2 * r2 + r6; // 37 mad + r6 = r6 ^ r3; // 38 xor + r2 = r2 ^ ds[r0 & mask]; // 39 load + r1 = r1 * r2; // 40 mul + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 41 load + r4 = rotl_imm(r4, 31u); // 42 rotl + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 43 load + r4 = r4 + r2 + ((((sel >> 24u) & 1u) != 0u) ? 0xdd1539e6u : 0xb045a668u); // 44 add + { uint32_t b_ = (r1 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 45 load + r7 = r7 + r6 + ((((sel >> 9u) & 1u) != 0u) ? 0x8a264d7du : 0x85c0036eu); // 46 add + r0 = r1 * r3 + r0; // 47 mad + r7 = r7 | r0; // 48 or + r6 = r6 ^ r4; // 49 xor + r1 = r1 + r7 + ((((sel >> 11u) & 1u) != 0u) ? 0x9d56aa91u : 0x376c63cdu); // 50 add + r5 = r5 * r4; // 51 mul + r7 = rotr_var(r7, r0); // 52 rotr + r2 = r2 ^ r6; // 53 xor + r0 = rotl_imm(r0, 22u); // 54 rotl + r3 = rotr_var(r3, r5); // 55 rotr + r6 = r6 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0x7fbd3c70u : 0x4face1f0u); // 56 add + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r1, 2); // 57 shfl + r2 = r6 * r0 + r2; // 58 mad + r2 = __umulhi(r2, r5); // 59 mulhi + r5 = rotl_imm(r5, 2u); // 60 rotl + r6 = rotl_imm(r6, 28u); // 61 rotl + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 62 load + r7 = __umulhi(r7, r1); // 63 mulhi + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixB-2/kernel_bound.cl b/proto-cuda/packs-readwidth/mixB-2/kernel_bound.cl new file mode 100644 index 000000000..896a73b85 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-2/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/2". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x8c11f8d8u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x945e59f8u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x945e59f8u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x4ab18392u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x4ab18392u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0xa6381ed7u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0xa6381ed7u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x797fadefu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x797fadefu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x188a03acu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x188a03acu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xea7032e0u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xea7032e0u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xcaf097f4u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0xcaf097f4u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x8c11f8d8u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = r0 ^ r6; // 0 xor + r0 = r1 * r4 + r0; // 1 mad + r2 = rotl_imm(r2, 21u); // 2 rotl + r3 = rotr_var(r3, r0); // 3 rotr + r0 = r0 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x474ae392u : 0x7a34b651u); // 4 add + r0 = mul_hi(r0, r3); // 5 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r3 = r3 ^ t_; } // 6 shfl + r7 = rotr_var(r7, r2); // 7 rotr + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 8 load + r4 = r4 | r5; // 9 or + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 10 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 load + r4 = r4 + r0 + ((((sel >> 16u) & 1u) != 0u) ? 0xf58f0389u : 0x99426786u); // 13 add + r0 = r0 + r6 + ((((sel >> 31u) & 1u) != 0u) ? 0xb5bbd552u : 0x6c662222u); // 14 add + r4 = rotr_var(r4, r6); // 15 rotr + r4 = r4 ^ r1; // 16 xor + r3 = rotl_imm(r3, 22u); // 17 rotl + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 18 load + r3 = rotr_var(r3, r0); // 19 rotr + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 20 load + r4 = rotl_imm(r4, 8u); // 21 rotl + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 22 load + r0 = rotr_var(r0, r3); // 23 rotr + r4 = r4 + r1 + ((((sel >> 13u) & 1u) != 0u) ? 0xd1ff184bu : 0x05f4dd31u); // 24 add + r4 = r4 ^ r2; // 25 xor + r3 = mul_hi(r3, r2); // 26 mulhi + r5 = r5 - r7; // 27 sub + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 28 load + r5 = r5 | r7; // 29 or + r4 = rotl_imm(r4, 28u); // 30 rotl + r3 = r3 ^ r5; // 31 xor + r2 = r2 * r1; // 32 mul + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 33 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r5 = r5 ^ t_; } // 34 shfl + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 35 load + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 36 load + r6 = r2 * r2 + r6; // 37 mad + r6 = r6 ^ r3; // 38 xor + r2 = r2 ^ ds[r0 & mask]; // 39 load + r1 = r1 * r2; // 40 mul + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 41 load + r4 = rotl_imm(r4, 31u); // 42 rotl + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 43 load + r4 = r4 + r2 + ((((sel >> 24u) & 1u) != 0u) ? 0xdd1539e6u : 0xb045a668u); // 44 add + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 45 load + r7 = r7 + r6 + ((((sel >> 9u) & 1u) != 0u) ? 0x8a264d7du : 0x85c0036eu); // 46 add + r0 = r1 * r3 + r0; // 47 mad + r7 = r7 | r0; // 48 or + r6 = r6 ^ r4; // 49 xor + r1 = r1 + r7 + ((((sel >> 11u) & 1u) != 0u) ? 0x9d56aa91u : 0x376c63cdu); // 50 add + r5 = r5 * r4; // 51 mul + r7 = rotr_var(r7, r0); // 52 rotr + r2 = r2 ^ r6; // 53 xor + r0 = rotl_imm(r0, 22u); // 54 rotl + r3 = rotr_var(r3, r5); // 55 rotr + r6 = r6 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0x7fbd3c70u : 0x4face1f0u); // 56 add + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 2u); r0 = r0 ^ t_; } // 57 shfl + r2 = r6 * r0 + r2; // 58 mad + r2 = mul_hi(r2, r5); // 59 mulhi + r5 = rotl_imm(r5, 2u); // 60 rotl + r6 = rotl_imm(r6, 28u); // 61 rotl + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 62 load + r7 = mul_hi(r7, r1); // 63 mulhi + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = r0 ^ r6; // 0 xor + r0 = r1 * r4 + r0; // 1 mad + r2 = rotl_imm(r2, 21u); // 2 rotl + r3 = rotr_var(r3, r0); // 3 rotr + r0 = r0 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x474ae392u : 0x7a34b651u); // 4 add + r0 = mul_hi(r0, r3); // 5 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r3 = r3 ^ t_; } // 6 shfl + r7 = rotr_var(r7, r2); // 7 rotr + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 8 load + r4 = r4 | r5; // 9 or + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 10 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 load + r4 = r4 + r0 + ((((sel >> 16u) & 1u) != 0u) ? 0xf58f0389u : 0x99426786u); // 13 add + r0 = r0 + r6 + ((((sel >> 31u) & 1u) != 0u) ? 0xb5bbd552u : 0x6c662222u); // 14 add + r4 = rotr_var(r4, r6); // 15 rotr + r4 = r4 ^ r1; // 16 xor + r3 = rotl_imm(r3, 22u); // 17 rotl + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 18 load + r3 = rotr_var(r3, r0); // 19 rotr + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 20 load + r4 = rotl_imm(r4, 8u); // 21 rotl + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 22 load + r0 = rotr_var(r0, r3); // 23 rotr + r4 = r4 + r1 + ((((sel >> 13u) & 1u) != 0u) ? 0xd1ff184bu : 0x05f4dd31u); // 24 add + r4 = r4 ^ r2; // 25 xor + r3 = mul_hi(r3, r2); // 26 mulhi + r5 = r5 - r7; // 27 sub + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 28 load + r5 = r5 | r7; // 29 or + r4 = rotl_imm(r4, 28u); // 30 rotl + r3 = r3 ^ r5; // 31 xor + r2 = r2 * r1; // 32 mul + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 33 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r5 = r5 ^ t_; } // 34 shfl + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 35 load + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 36 load + r6 = r2 * r2 + r6; // 37 mad + r6 = r6 ^ r3; // 38 xor + r2 = r2 ^ ds[r0 & mask]; // 39 load + r1 = r1 * r2; // 40 mul + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 41 load + r4 = rotl_imm(r4, 31u); // 42 rotl + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 43 load + r4 = r4 + r2 + ((((sel >> 24u) & 1u) != 0u) ? 0xdd1539e6u : 0xb045a668u); // 44 add + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 45 load + r7 = r7 + r6 + ((((sel >> 9u) & 1u) != 0u) ? 0x8a264d7du : 0x85c0036eu); // 46 add + r0 = r1 * r3 + r0; // 47 mad + r7 = r7 | r0; // 48 or + r6 = r6 ^ r4; // 49 xor + r1 = r1 + r7 + ((((sel >> 11u) & 1u) != 0u) ? 0x9d56aa91u : 0x376c63cdu); // 50 add + r5 = r5 * r4; // 51 mul + r7 = rotr_var(r7, r0); // 52 rotr + r2 = r2 ^ r6; // 53 xor + r0 = rotl_imm(r0, 22u); // 54 rotl + r3 = rotr_var(r3, r5); // 55 rotr + r6 = r6 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0x7fbd3c70u : 0x4face1f0u); // 56 add + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 2u); r0 = r0 ^ t_; } // 57 shfl + r2 = r6 * r0 + r2; // 58 mad + r2 = mul_hi(r2, r5); // 59 mulhi + r5 = rotl_imm(r5, 2u); // 60 rotl + r6 = rotl_imm(r6, 28u); // 61 rotl + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 62 load + r7 = mul_hi(r7, r1); // 63 mulhi + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-2/kernel_bound.cu b/proto-cuda/packs-readwidth/mixB-2/kernel_bound.cu new file mode 100644 index 000000000..cc282bfbe --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-2/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/2". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r0 = r0 ^ r6; // 0 xor + r0 = r1 * r4 + r0; // 1 mad + r2 = rotl_imm(r2, 21u); // 2 rotl + r3 = rotr_var(r3, r0); // 3 rotr + r0 = r0 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x474ae392u : 0x7a34b651u); // 4 add + r0 = __umulhi(r0, r3); // 5 mulhi + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 4); // 6 shfl + r7 = rotr_var(r7, r2); // 7 rotr + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 8 load + r4 = r4 | r5; // 9 or + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 10 load + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 load + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 load + r4 = r4 + r0 + ((((sel >> 16u) & 1u) != 0u) ? 0xf58f0389u : 0x99426786u); // 13 add + r0 = r0 + r6 + ((((sel >> 31u) & 1u) != 0u) ? 0xb5bbd552u : 0x6c662222u); // 14 add + r4 = rotr_var(r4, r6); // 15 rotr + r4 = r4 ^ r1; // 16 xor + r3 = rotl_imm(r3, 22u); // 17 rotl + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 18 load + r3 = rotr_var(r3, r0); // 19 rotr + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 20 load + r4 = rotl_imm(r4, 8u); // 21 rotl + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 22 load + r0 = rotr_var(r0, r3); // 23 rotr + r4 = r4 + r1 + ((((sel >> 13u) & 1u) != 0u) ? 0xd1ff184bu : 0x05f4dd31u); // 24 add + r4 = r4 ^ r2; // 25 xor + r3 = __umulhi(r3, r2); // 26 mulhi + r5 = r5 - r7; // 27 sub + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 28 load + r5 = r5 | r7; // 29 or + r4 = rotl_imm(r4, 28u); // 30 rotl + r3 = r3 ^ r5; // 31 xor + r2 = r2 * r1; // 32 mul + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 33 load + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 34 shfl + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 35 load + { uint32_t b_ = (r6 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 36 load + r6 = r2 * r2 + r6; // 37 mad + r6 = r6 ^ r3; // 38 xor + r2 = r2 ^ ds[r0 & mask]; // 39 load + r1 = r1 * r2; // 40 mul + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 41 load + r4 = rotl_imm(r4, 31u); // 42 rotl + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 43 load + r4 = r4 + r2 + ((((sel >> 24u) & 1u) != 0u) ? 0xdd1539e6u : 0xb045a668u); // 44 add + { uint32_t b_ = (r1 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 45 load + r7 = r7 + r6 + ((((sel >> 9u) & 1u) != 0u) ? 0x8a264d7du : 0x85c0036eu); // 46 add + r0 = r1 * r3 + r0; // 47 mad + r7 = r7 | r0; // 48 or + r6 = r6 ^ r4; // 49 xor + r1 = r1 + r7 + ((((sel >> 11u) & 1u) != 0u) ? 0x9d56aa91u : 0x376c63cdu); // 50 add + r5 = r5 * r4; // 51 mul + r7 = rotr_var(r7, r0); // 52 rotr + r2 = r2 ^ r6; // 53 xor + r0 = rotl_imm(r0, 22u); // 54 rotl + r3 = rotr_var(r3, r5); // 55 rotr + r6 = r6 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0x7fbd3c70u : 0x4face1f0u); // 56 add + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r1, 2); // 57 shfl + r2 = r6 * r0 + r2; // 58 mad + r2 = __umulhi(r2, r5); // 59 mulhi + r5 = rotl_imm(r5, 2u); // 60 rotl + r6 = rotl_imm(r6, 28u); // 61 rotl + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 62 load + r7 = __umulhi(r7, r1); // 63 mulhi + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixB-2/memhard.h b/proto-cuda/packs-readwidth/mixB-2/memhard.h new file mode 100644 index 000000000..8df3dd70f --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-2/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/2". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/mixB-2/memhard.metal b/proto-cuda/packs-readwidth/mixB-2/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-2/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/mixB-2/program.h b/proto-cuda/packs-readwidth/mixB-2/program.h new file mode 100644 index 000000000..e1ba9226f --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-2/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/2". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-readwidth/B/2" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d7265616477696474682f422f32" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0xc183ffe6ab211605ull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=8 rotl=8 rotr=7 xor=7 mad=4 mulhi=4 mul=3 or=3 shfl=3 sub=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "mix25-50-25" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 25, 50, 25 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 1, 7, 8 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 5024 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x8c11f8d8u, 0x945e59f8u, 0x4ab18392u, 0xa6381ed7u, 0x797fadefu, 0x188a03acu, 0xea7032e0u, 0xcaf097f4u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/mixB-2/program.json b/proto-cuda/packs-readwidth/mixB-2/program.json new file mode 100644 index 000000000..339120da0 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-2/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0xc183ffe6ab211605", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-readwidth/B/2", + "seed_bytes": "69676e65756d2d7265616477696474682f422f32", + "seed_words": ["0x8c11f8d8", "0x945e59f8", "0x4ab18392", "0xa6381ed7", "0x797fadef", "0x188a03ac", "0xea7032e0", "0xcaf097f4"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "mix25-50-25", + "load_slots": 16, + "load_mix_percent_4_16_64": [25, 50, 25], + "load_width_counts_4_16_64": [1, 7, 8], + "bytes_per_hash": 5024, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "add": 8, "rotl": 8, "rotr": 7, "xor": 7, "mad": 4, "mulhi": 4, "mul": 3, "or": 3, "shfl": 3, "sub": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "xor", "dst": 0, "src": 6, "src2": 6, "imm": "0xb5a6acaf", "imm2": "0x3a16d365", "rot": 22, "bit": 0, "mask": 1, "width": 1}, + {"i": 1, "op": "mad", "dst": 0, "src": 1, "src2": 4, "imm": "0x629ea39c", "imm2": "0x2bd64262", "rot": 21, "bit": 24, "mask": 1, "width": 1}, + {"i": 2, "op": "rotl", "dst": 2, "src": 5, "src2": 5, "imm": "0x4f98c110", "imm2": "0x0d82f00c", "rot": 21, "bit": 16, "mask": 4, "width": 1}, + {"i": 3, "op": "rotr", "dst": 3, "src": 0, "src2": 2, "imm": "0xa6685e68", "imm2": "0x70585b41", "rot": 19, "bit": 0, "mask": 8, "width": 1}, + {"i": 4, "op": "add", "dst": 0, "src": 2, "src2": 2, "imm": "0x7a34b651", "imm2": "0x474ae392", "rot": 3, "bit": 7, "mask": 16, "width": 1}, + {"i": 5, "op": "mulhi", "dst": 0, "src": 3, "src2": 2, "imm": "0x0e2e05ef", "imm2": "0x0950cfa8", "rot": 31, "bit": 8, "mask": 2, "width": 1}, + {"i": 6, "op": "shfl", "dst": 3, "src": 6, "src2": 7, "imm": "0x41d636b4", "imm2": "0xa1978647", "rot": 13, "bit": 19, "mask": 4, "width": 1}, + {"i": 7, "op": "rotr", "dst": 7, "src": 2, "src2": 6, "imm": "0x2a34c680", "imm2": "0x8b079edd", "rot": 27, "bit": 4, "mask": 1, "width": 1}, + {"i": 8, "op": "load", "dst": 2, "src": 0, "src2": 2, "imm": "0x6ca9f362", "imm2": "0x38680556", "rot": 21, "bit": 5, "mask": 8, "width": 16}, + {"i": 9, "op": "or", "dst": 4, "src": 5, "src2": 6, "imm": "0x0ede3e17", "imm2": "0xab200cc3", "rot": 27, "bit": 8, "mask": 16, "width": 1}, + {"i": 10, "op": "load", "dst": 0, "src": 4, "src2": 0, "imm": "0xb3a5ec67", "imm2": "0x0ce446f7", "rot": 1, "bit": 8, "mask": 16, "width": 16}, + {"i": 11, "op": "load", "dst": 5, "src": 0, "src2": 1, "imm": "0x08d3c266", "imm2": "0x2d844647", "rot": 12, "bit": 20, "mask": 16, "width": 4}, + {"i": 12, "op": "load", "dst": 4, "src": 7, "src2": 7, "imm": "0x0b891e32", "imm2": "0x8b22c632", "rot": 2, "bit": 10, "mask": 1, "width": 4}, + {"i": 13, "op": "add", "dst": 4, "src": 0, "src2": 1, "imm": "0x99426786", "imm2": "0xf58f0389", "rot": 10, "bit": 16, "mask": 1, "width": 1}, + {"i": 14, "op": "add", "dst": 0, "src": 6, "src2": 7, "imm": "0x6c662222", "imm2": "0xb5bbd552", "rot": 29, "bit": 31, "mask": 8, "width": 1}, + {"i": 15, "op": "rotr", "dst": 4, "src": 6, "src2": 3, "imm": "0x0a5bcff8", "imm2": "0xb73018d3", "rot": 10, "bit": 8, "mask": 4, "width": 1}, + {"i": 16, "op": "xor", "dst": 4, "src": 1, "src2": 4, "imm": "0xf986d47e", "imm2": "0x9212f429", "rot": 28, "bit": 21, "mask": 16, "width": 1}, + {"i": 17, "op": "rotl", "dst": 3, "src": 7, "src2": 4, "imm": "0x8de91805", "imm2": "0xc7b33ed8", "rot": 22, "bit": 27, "mask": 4, "width": 1}, + {"i": 18, "op": "load", "dst": 3, "src": 4, "src2": 2, "imm": "0x084e3922", "imm2": "0x2cdb1d35", "rot": 29, "bit": 1, "mask": 8, "width": 4}, + {"i": 19, "op": "rotr", "dst": 3, "src": 0, "src2": 5, "imm": "0x5e1362d5", "imm2": "0x3a804b96", "rot": 4, "bit": 11, "mask": 8, "width": 1}, + {"i": 20, "op": "load", "dst": 1, "src": 3, "src2": 6, "imm": "0x5269dadb", "imm2": "0x9c710481", "rot": 31, "bit": 16, "mask": 4, "width": 4}, + {"i": 21, "op": "rotl", "dst": 4, "src": 2, "src2": 6, "imm": "0x527da520", "imm2": "0x380edfcd", "rot": 8, "bit": 25, "mask": 16, "width": 1}, + {"i": 22, "op": "load", "dst": 1, "src": 5, "src2": 7, "imm": "0xd9616dc0", "imm2": "0xcc1e9bb1", "rot": 25, "bit": 12, "mask": 8, "width": 16}, + {"i": 23, "op": "rotr", "dst": 0, "src": 3, "src2": 7, "imm": "0x7e7e2717", "imm2": "0x4b659383", "rot": 4, "bit": 3, "mask": 1, "width": 1}, + {"i": 24, "op": "add", "dst": 4, "src": 1, "src2": 6, "imm": "0x05f4dd31", "imm2": "0xd1ff184b", "rot": 27, "bit": 13, "mask": 8, "width": 1}, + {"i": 25, "op": "xor", "dst": 4, "src": 2, "src2": 5, "imm": "0xd73346b3", "imm2": "0x2e9e82b4", "rot": 27, "bit": 28, "mask": 4, "width": 1}, + {"i": 26, "op": "mulhi", "dst": 3, "src": 2, "src2": 4, "imm": "0xd9afc8a7", "imm2": "0xc38e4e80", "rot": 23, "bit": 24, "mask": 4, "width": 1}, + {"i": 27, "op": "sub", "dst": 5, "src": 7, "src2": 1, "imm": "0x1b5d4625", "imm2": "0xc275a42b", "rot": 30, "bit": 11, "mask": 2, "width": 1}, + {"i": 28, "op": "load", "dst": 3, "src": 5, "src2": 7, "imm": "0x329571c7", "imm2": "0x000b807e", "rot": 25, "bit": 26, "mask": 8, "width": 16}, + {"i": 29, "op": "or", "dst": 5, "src": 7, "src2": 3, "imm": "0xc5afee75", "imm2": "0xc73ec965", "rot": 15, "bit": 17, "mask": 8, "width": 1}, + {"i": 30, "op": "rotl", "dst": 4, "src": 3, "src2": 6, "imm": "0x0e74c55a", "imm2": "0xad6ac031", "rot": 28, "bit": 3, "mask": 4, "width": 1}, + {"i": 31, "op": "xor", "dst": 3, "src": 5, "src2": 4, "imm": "0x79d4fac3", "imm2": "0xf48a2d2e", "rot": 24, "bit": 12, "mask": 1, "width": 1}, + {"i": 32, "op": "mul", "dst": 2, "src": 1, "src2": 4, "imm": "0x1dba16cc", "imm2": "0xd0cf73d1", "rot": 26, "bit": 3, "mask": 1, "width": 1}, + {"i": 33, "op": "load", "dst": 3, "src": 4, "src2": 1, "imm": "0xd4328498", "imm2": "0x0140c7e0", "rot": 28, "bit": 24, "mask": 8, "width": 4}, + {"i": 34, "op": "shfl", "dst": 5, "src": 4, "src2": 7, "imm": "0x7f9d6623", "imm2": "0x063550be", "rot": 14, "bit": 2, "mask": 16, "width": 1}, + {"i": 35, "op": "load", "dst": 6, "src": 5, "src2": 4, "imm": "0xec4cd413", "imm2": "0xe62d9347", "rot": 12, "bit": 13, "mask": 2, "width": 4}, + {"i": 36, "op": "load", "dst": 4, "src": 6, "src2": 0, "imm": "0x4fa040f1", "imm2": "0xc537e766", "rot": 12, "bit": 23, "mask": 4, "width": 16}, + {"i": 37, "op": "mad", "dst": 6, "src": 2, "src2": 2, "imm": "0x2fa75f28", "imm2": "0xf2a66942", "rot": 8, "bit": 3, "mask": 4, "width": 1}, + {"i": 38, "op": "xor", "dst": 6, "src": 3, "src2": 4, "imm": "0xb201f75e", "imm2": "0xa8afa3f4", "rot": 22, "bit": 26, "mask": 1, "width": 1}, + {"i": 39, "op": "load", "dst": 2, "src": 0, "src2": 1, "imm": "0x524d8b9f", "imm2": "0x2197ac78", "rot": 4, "bit": 2, "mask": 4, "width": 1}, + {"i": 40, "op": "mul", "dst": 1, "src": 2, "src2": 7, "imm": "0x3972b861", "imm2": "0x7e4f91f1", "rot": 9, "bit": 23, "mask": 4, "width": 1}, + {"i": 41, "op": "load", "dst": 4, "src": 2, "src2": 2, "imm": "0xcdab329b", "imm2": "0xc5cb8830", "rot": 21, "bit": 10, "mask": 16, "width": 16}, + {"i": 42, "op": "rotl", "dst": 4, "src": 5, "src2": 7, "imm": "0xc1aa2f50", "imm2": "0x8864642e", "rot": 31, "bit": 5, "mask": 1, "width": 1}, + {"i": 43, "op": "load", "dst": 6, "src": 4, "src2": 6, "imm": "0x3805326c", "imm2": "0x75a9eaf4", "rot": 14, "bit": 16, "mask": 8, "width": 4}, + {"i": 44, "op": "add", "dst": 4, "src": 2, "src2": 0, "imm": "0xb045a668", "imm2": "0xdd1539e6", "rot": 16, "bit": 24, "mask": 1, "width": 1}, + {"i": 45, "op": "load", "dst": 5, "src": 1, "src2": 0, "imm": "0xde27aabb", "imm2": "0x256b267e", "rot": 24, "bit": 14, "mask": 2, "width": 16}, + {"i": 46, "op": "add", "dst": 7, "src": 6, "src2": 3, "imm": "0x85c0036e", "imm2": "0x8a264d7d", "rot": 31, "bit": 9, "mask": 4, "width": 1}, + {"i": 47, "op": "mad", "dst": 0, "src": 1, "src2": 3, "imm": "0x844552ff", "imm2": "0x8ee51674", "rot": 23, "bit": 23, "mask": 4, "width": 1}, + {"i": 48, "op": "or", "dst": 7, "src": 0, "src2": 6, "imm": "0x20f7430b", "imm2": "0xf89bf854", "rot": 1, "bit": 10, "mask": 8, "width": 1}, + {"i": 49, "op": "xor", "dst": 6, "src": 4, "src2": 1, "imm": "0x7919ad9e", "imm2": "0xd4e43dd3", "rot": 23, "bit": 22, "mask": 4, "width": 1}, + {"i": 50, "op": "add", "dst": 1, "src": 7, "src2": 3, "imm": "0x376c63cd", "imm2": "0x9d56aa91", "rot": 21, "bit": 11, "mask": 2, "width": 1}, + {"i": 51, "op": "mul", "dst": 5, "src": 4, "src2": 5, "imm": "0xc747bed9", "imm2": "0x204c0638", "rot": 21, "bit": 3, "mask": 2, "width": 1}, + {"i": 52, "op": "rotr", "dst": 7, "src": 0, "src2": 5, "imm": "0x3ef36afd", "imm2": "0x0e3ea570", "rot": 28, "bit": 14, "mask": 1, "width": 1}, + {"i": 53, "op": "xor", "dst": 2, "src": 6, "src2": 7, "imm": "0x9e99cc31", "imm2": "0x78899080", "rot": 6, "bit": 17, "mask": 16, "width": 1}, + {"i": 54, "op": "rotl", "dst": 0, "src": 7, "src2": 6, "imm": "0x25f152bf", "imm2": "0x6ccbd353", "rot": 22, "bit": 23, "mask": 4, "width": 1}, + {"i": 55, "op": "rotr", "dst": 3, "src": 5, "src2": 5, "imm": "0x8187728d", "imm2": "0x6f4f287a", "rot": 2, "bit": 18, "mask": 2, "width": 1}, + {"i": 56, "op": "add", "dst": 6, "src": 2, "src2": 3, "imm": "0x4face1f0", "imm2": "0x7fbd3c70", "rot": 2, "bit": 21, "mask": 2, "width": 1}, + {"i": 57, "op": "shfl", "dst": 0, "src": 1, "src2": 0, "imm": "0x17c29c31", "imm2": "0x5438a0c8", "rot": 29, "bit": 16, "mask": 2, "width": 1}, + {"i": 58, "op": "mad", "dst": 2, "src": 6, "src2": 0, "imm": "0x305dd740", "imm2": "0x5e196392", "rot": 4, "bit": 29, "mask": 2, "width": 1}, + {"i": 59, "op": "mulhi", "dst": 2, "src": 5, "src2": 3, "imm": "0x364a67f1", "imm2": "0x717ceeaa", "rot": 5, "bit": 7, "mask": 4, "width": 1}, + {"i": 60, "op": "rotl", "dst": 5, "src": 6, "src2": 3, "imm": "0x0f0b8fb5", "imm2": "0x257c95f1", "rot": 2, "bit": 13, "mask": 1, "width": 1}, + {"i": 61, "op": "rotl", "dst": 6, "src": 3, "src2": 1, "imm": "0xa3590558", "imm2": "0x372dea9c", "rot": 28, "bit": 19, "mask": 16, "width": 1}, + {"i": 62, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0x2c2192e7", "imm2": "0xc8096ee6", "rot": 16, "bit": 7, "mask": 1, "width": 16}, + {"i": 63, "op": "mulhi", "dst": 7, "src": 1, "src2": 5, "imm": "0x52f561e6", "imm2": "0x62d77ee8", "rot": 13, "bit": 22, "mask": 8, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/mixB-2/program.metal b/proto-cuda/packs-readwidth/mixB-2/program.metal new file mode 100644 index 000000000..50315bb9b --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-2/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x8c11f8d8u, 0x945e59f8u, 0x4ab18392u, 0xa6381ed7u, 0x797fadefu, 0x188a03acu, 0xea7032e0u, 0xcaf097f4u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = r0 ^ r6; // 0 + r0 = r1 * r4 + r0; // 1 + r2 = rotl_imm(r2, 21u); // 2 + r3 = rotr_var(r3, r0); // 3 + r0 = r0 + r2 + select(0x7a34b651u, 0x474ae392u, ((sel >> 7u) & 1u) != 0u); // 4 + r0 = mulhi(r0, r3); // 5 + r3 = r3 ^ simd_shuffle_xor(r6, (ushort)4); // 6 + r7 = rotr_var(r7, r2); // 7 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 8 + r4 = r4 | r5; // 9 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 10 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 + r4 = r4 + r0 + select(0x99426786u, 0xf58f0389u, ((sel >> 16u) & 1u) != 0u); // 13 + r0 = r0 + r6 + select(0x6c662222u, 0xb5bbd552u, ((sel >> 31u) & 1u) != 0u); // 14 + r4 = rotr_var(r4, r6); // 15 + r4 = r4 ^ r1; // 16 + r3 = rotl_imm(r3, 22u); // 17 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 18 + r3 = rotr_var(r3, r0); // 19 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 20 + r4 = rotl_imm(r4, 8u); // 21 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 22 + r0 = rotr_var(r0, r3); // 23 + r4 = r4 + r1 + select(0x05f4dd31u, 0xd1ff184bu, ((sel >> 13u) & 1u) != 0u); // 24 + r4 = r4 ^ r2; // 25 + r3 = mulhi(r3, r2); // 26 + r5 = r5 - r7; // 27 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 28 + r5 = r5 | r7; // 29 + r4 = rotl_imm(r4, 28u); // 30 + r3 = r3 ^ r5; // 31 + r2 = r2 * r1; // 32 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 33 + r5 = r5 ^ simd_shuffle_xor(r4, (ushort)16); // 34 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 35 + { uint b_ = (r6 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 36 + r6 = r2 * r2 + r6; // 37 + r6 = r6 ^ r3; // 38 + r2 = r2 ^ dataset[r0 & MASK]; // 39 + r1 = r1 * r2; // 40 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 41 + r4 = rotl_imm(r4, 31u); // 42 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 43 + r4 = r4 + r2 + select(0xb045a668u, 0xdd1539e6u, ((sel >> 24u) & 1u) != 0u); // 44 + { uint b_ = (r1 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 45 + r7 = r7 + r6 + select(0x85c0036eu, 0x8a264d7du, ((sel >> 9u) & 1u) != 0u); // 46 + r0 = r1 * r3 + r0; // 47 + r7 = r7 | r0; // 48 + r6 = r6 ^ r4; // 49 + r1 = r1 + r7 + select(0x376c63cdu, 0x9d56aa91u, ((sel >> 11u) & 1u) != 0u); // 50 + r5 = r5 * r4; // 51 + r7 = rotr_var(r7, r0); // 52 + r2 = r2 ^ r6; // 53 + r0 = rotl_imm(r0, 22u); // 54 + r3 = rotr_var(r3, r5); // 55 + r6 = r6 + r2 + select(0x4face1f0u, 0x7fbd3c70u, ((sel >> 21u) & 1u) != 0u); // 56 + r0 = r0 ^ simd_shuffle_xor(r1, (ushort)2); // 57 + r2 = r6 * r0 + r2; // 58 + r2 = mulhi(r2, r5); // 59 + r5 = rotl_imm(r5, 2u); // 60 + r6 = rotl_imm(r6, 28u); // 61 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 62 + r7 = mulhi(r7, r1); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-2/program_bound.metal b/proto-cuda/packs-readwidth/mixB-2/program_bound.metal new file mode 100644 index 000000000..0440394f2 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-2/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x8c11f8d8u, 0x945e59f8u, 0x4ab18392u, 0xa6381ed7u, 0x797fadefu, 0x188a03acu, 0xea7032e0u, 0xcaf097f4u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = r0 ^ r6; // 0 + r0 = r1 * r4 + r0; // 1 + r2 = rotl_imm(r2, 21u); // 2 + r3 = rotr_var(r3, r0); // 3 + r0 = r0 + r2 + select(0x7a34b651u, 0x474ae392u, ((sel >> 7u) & 1u) != 0u); // 4 + r0 = mulhi(r0, r3); // 5 + r3 = r3 ^ simd_shuffle_xor(r6, (ushort)4); // 6 + r7 = rotr_var(r7, r2); // 7 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 8 + r4 = r4 | r5; // 9 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 10 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 + r4 = r4 + r0 + select(0x99426786u, 0xf58f0389u, ((sel >> 16u) & 1u) != 0u); // 13 + r0 = r0 + r6 + select(0x6c662222u, 0xb5bbd552u, ((sel >> 31u) & 1u) != 0u); // 14 + r4 = rotr_var(r4, r6); // 15 + r4 = r4 ^ r1; // 16 + r3 = rotl_imm(r3, 22u); // 17 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 18 + r3 = rotr_var(r3, r0); // 19 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 20 + r4 = rotl_imm(r4, 8u); // 21 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 22 + r0 = rotr_var(r0, r3); // 23 + r4 = r4 + r1 + select(0x05f4dd31u, 0xd1ff184bu, ((sel >> 13u) & 1u) != 0u); // 24 + r4 = r4 ^ r2; // 25 + r3 = mulhi(r3, r2); // 26 + r5 = r5 - r7; // 27 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 28 + r5 = r5 | r7; // 29 + r4 = rotl_imm(r4, 28u); // 30 + r3 = r3 ^ r5; // 31 + r2 = r2 * r1; // 32 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 33 + r5 = r5 ^ simd_shuffle_xor(r4, (ushort)16); // 34 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 35 + { uint b_ = (r6 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 36 + r6 = r2 * r2 + r6; // 37 + r6 = r6 ^ r3; // 38 + r2 = r2 ^ dataset[r0 & MASK]; // 39 + r1 = r1 * r2; // 40 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 41 + r4 = rotl_imm(r4, 31u); // 42 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 43 + r4 = r4 + r2 + select(0xb045a668u, 0xdd1539e6u, ((sel >> 24u) & 1u) != 0u); // 44 + { uint b_ = (r1 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 45 + r7 = r7 + r6 + select(0x85c0036eu, 0x8a264d7du, ((sel >> 9u) & 1u) != 0u); // 46 + r0 = r1 * r3 + r0; // 47 + r7 = r7 | r0; // 48 + r6 = r6 ^ r4; // 49 + r1 = r1 + r7 + select(0x376c63cdu, 0x9d56aa91u, ((sel >> 11u) & 1u) != 0u); // 50 + r5 = r5 * r4; // 51 + r7 = rotr_var(r7, r0); // 52 + r2 = r2 ^ r6; // 53 + r0 = rotl_imm(r0, 22u); // 54 + r3 = rotr_var(r3, r5); // 55 + r6 = r6 + r2 + select(0x4face1f0u, 0x7fbd3c70u, ((sel >> 21u) & 1u) != 0u); // 56 + r0 = r0 ^ simd_shuffle_xor(r1, (ushort)2); // 57 + r2 = r6 * r0 + r2; // 58 + r2 = mulhi(r2, r5); // 59 + r5 = rotl_imm(r5, 2u); // 60 + r6 = rotl_imm(r6, 28u); // 61 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 62 + r7 = mulhi(r7, r1); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-2/vectors.h b/proto-cuda/packs-readwidth/mixB-2/vectors.h new file mode 100644 index 000000000..4523b5f7f --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-2/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/2". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x8e0770915c34f8cbull, 0x7eeadbde999c8490ull, 0xdbf3c010876b4910ull, 0x9fc729efb92b0a53ull, 0x50bd1e0032bd3610ull, 0xf91e1b4208faadfaull, 0x13450a43f1d11c11ull, 0x187080bf93323917ull, + 0x73291f56ff3a9634ull, 0x59cf29b759edad45ull, 0xe0d234bb0f48a66bull, 0x63bd1d6a09c913a5ull, 0x4815316cd4fcc14bull, 0xf3b12bfcfd809133ull, 0xb2d98152d7b01eb4ull, 0x6279fde5cf9e0d8cull, + 0x68da28afa62210dcull, 0x9e2e655e9d98b522ull, 0xc4eea5e470ec187full, 0x5d84305bd94955ecull, 0x3c5c32e071f66493ull, 0x2ba8f595f2229b91ull, 0x20edca1ec4c7921full, 0x44b0a796cbbca675ull, + 0x1de1c69461951d8aull, 0xac14088b8372eabaull, 0xa9454e87941712b7ull, 0x66c22d2c3b112f4cull, 0x9550ed5301345735ull, 0x294058f2ec7e054eull, 0xa92a7d3750e73e5bull, 0xd70d21d337d5dafbull + }, + { // base nonce 4096 + 0x6d7d508b316ac720ull, 0x675452634595d29cull, 0xf120abc94c7b9aeaull, 0x811ee0f2c2a3cde2ull, 0x4480d96bc6e69471ull, 0x9358f1cb0563e439ull, 0xa1830bff69bd8825ull, 0x67e609de775cb3adull, + 0x8ae717337f04693dull, 0x232a79466e3f836cull, 0xea0437198a50160dull, 0xc058d269fb4d5359ull, 0xbb1d542bcc392436ull, 0x0953f45e56e01a14ull, 0x2be8a38bb9af773eull, 0x6e166f5e49b09d7full, + 0x80d90cb523c0d0ddull, 0x65f829c21be28018ull, 0xa78c082db4d53fd3ull, 0x44a83bc5839de423ull, 0x2b36d552820a12baull, 0xb59d330743ad6185ull, 0xab56118c9612bb4cull, 0x37047607898c3e24ull, + 0x16ca2838e4546f5bull, 0xdf5da44ace7ac75eull, 0xbd339b416b338c3cull, 0xfddcec4e27c95307ull, 0xd182f21c7b19eb01ull, 0xc90553bbeb148431ull, 0xb3733d3e73213f67ull, 0xbdc8baa0f7f3cba5ull + }, + { // base nonce 1000000 + 0x16afbb50e1037248ull, 0x99d1b552e8ca9a75ull, 0x77b62683fe9fce6eull, 0xd42e7993d6b006c7ull, 0xabfd95ae21d3be44ull, 0x8745867d1ab2d0f4ull, 0x5b13b4fe24d422e5ull, 0x9bf02944f69a727full, + 0xab4cb6fc57054fb7ull, 0xed930abaabf1ed87ull, 0xa19a42ff991d8caeull, 0x31864706ab6251e8ull, 0x31026daed2943ab7ull, 0x680f91b927699537ull, 0xb440c27741077eeeull, 0x3bf9208f1d8556e4ull, + 0x8d44ca689dea20adull, 0x56e6b9afc31a1c4dull, 0xa16676237b406812ull, 0x12117a7974319c20ull, 0x8cd93c1c722cfef5ull, 0x4c3497d3bc073cb2ull, 0xf7f0dfd9f741a6d6ull, 0xf8302f51da875d34ull, + 0x4dc6d3def7e8edbaull, 0x74bf07551eaca6f4ull, 0x9e0569e0e4ff420dull, 0xce2734adf4a32b89ull, 0xc88c70346bd05e2bull, 0xa006afd72b05c173ull, 0x397698e62a815243ull, 0x5b8bfd6849164a3bull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/mixB-2/vectors.json b/proto-cuda/packs-readwidth/mixB-2/vectors.json new file mode 100644 index 000000000..11da036c4 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-2/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-readwidth/B/2", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x8e0770915c34f8cb", "0x7eeadbde999c8490", "0xdbf3c010876b4910", "0x9fc729efb92b0a53", "0x50bd1e0032bd3610", "0xf91e1b4208faadfa", "0x13450a43f1d11c11", "0x187080bf93323917", + "0x73291f56ff3a9634", "0x59cf29b759edad45", "0xe0d234bb0f48a66b", "0x63bd1d6a09c913a5", "0x4815316cd4fcc14b", "0xf3b12bfcfd809133", "0xb2d98152d7b01eb4", "0x6279fde5cf9e0d8c", + "0x68da28afa62210dc", "0x9e2e655e9d98b522", "0xc4eea5e470ec187f", "0x5d84305bd94955ec", "0x3c5c32e071f66493", "0x2ba8f595f2229b91", "0x20edca1ec4c7921f", "0x44b0a796cbbca675", + "0x1de1c69461951d8a", "0xac14088b8372eaba", "0xa9454e87941712b7", "0x66c22d2c3b112f4c", "0x9550ed5301345735", "0x294058f2ec7e054e", "0xa92a7d3750e73e5b", "0xd70d21d337d5dafb" + ]}, + {"base_nonce": 4096, "expected": [ + "0x6d7d508b316ac720", "0x675452634595d29c", "0xf120abc94c7b9aea", "0x811ee0f2c2a3cde2", "0x4480d96bc6e69471", "0x9358f1cb0563e439", "0xa1830bff69bd8825", "0x67e609de775cb3ad", + "0x8ae717337f04693d", "0x232a79466e3f836c", "0xea0437198a50160d", "0xc058d269fb4d5359", "0xbb1d542bcc392436", "0x0953f45e56e01a14", "0x2be8a38bb9af773e", "0x6e166f5e49b09d7f", + "0x80d90cb523c0d0dd", "0x65f829c21be28018", "0xa78c082db4d53fd3", "0x44a83bc5839de423", "0x2b36d552820a12ba", "0xb59d330743ad6185", "0xab56118c9612bb4c", "0x37047607898c3e24", + "0x16ca2838e4546f5b", "0xdf5da44ace7ac75e", "0xbd339b416b338c3c", "0xfddcec4e27c95307", "0xd182f21c7b19eb01", "0xc90553bbeb148431", "0xb3733d3e73213f67", "0xbdc8baa0f7f3cba5" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x16afbb50e1037248", "0x99d1b552e8ca9a75", "0x77b62683fe9fce6e", "0xd42e7993d6b006c7", "0xabfd95ae21d3be44", "0x8745867d1ab2d0f4", "0x5b13b4fe24d422e5", "0x9bf02944f69a727f", + "0xab4cb6fc57054fb7", "0xed930abaabf1ed87", "0xa19a42ff991d8cae", "0x31864706ab6251e8", "0x31026daed2943ab7", "0x680f91b927699537", "0xb440c27741077eee", "0x3bf9208f1d8556e4", + "0x8d44ca689dea20ad", "0x56e6b9afc31a1c4d", "0xa16676237b406812", "0x12117a7974319c20", "0x8cd93c1c722cfef5", "0x4c3497d3bc073cb2", "0xf7f0dfd9f741a6d6", "0xf8302f51da875d34", + "0x4dc6d3def7e8edba", "0x74bf07551eaca6f4", "0x9e0569e0e4ff420d", "0xce2734adf4a32b89", "0xc88c70346bd05e2b", "0xa006afd72b05c173", "0x397698e62a815243", "0x5b8bfd6849164a3b" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/mixB-3/kernel.cl b/proto-cuda/packs-readwidth/mixB-3/kernel.cl new file mode 100644 index 000000000..a5096a8b2 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-3/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/3". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x641145b6u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0xd54149dau; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0xd54149dau; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xb9deccd5u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xb9deccd5u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0xd7a898afu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0xd7a898afu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xb3350422u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xb3350422u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xe03937b0u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xe03937b0u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xd4c29221u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xd4c29221u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x3188f3b9u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x3188f3b9u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x641145b6u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 8u); r4 = r4 ^ t_; } // 0 shfl + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 1 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 8u); r5 = r5 ^ t_; } // 2 shfl + r2 = r2 ^ r4; // 3 xor + r2 = mul_hi(r2, r3); // 4 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r1 = r1 ^ t_; } // 5 shfl + r2 = r2 + r0 + ((((sel >> 23u) & 1u) != 0u) ? 0xd05f6c1eu : 0x846ac225u); // 6 add + r6 = r6 | r4; // 7 or + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 2u); r7 = r7 ^ t_; } // 8 shfl + r0 = r0 ^ ds[r5 & mask]; // 9 load + r2 = r2 + r3 + ((((sel >> 0u) & 1u) != 0u) ? 0xd712e0e2u : 0x610f3311u); // 10 add + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 load + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 12 load + r2 = r7 * r2 + r2; // 13 mad + r4 = rotl_imm(r4, 16u); // 14 rotl + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 15 load + r2 = r2 * r0; // 16 mul + r3 = r3 + r2 + ((((sel >> 25u) & 1u) != 0u) ? 0xbd25a0e1u : 0x4b2e8512u); // 17 add + r7 = r7 - r3; // 18 sub + r0 = rotr_var(r0, r1); // 19 rotr + r7 = r3 * r2 + r7; // 20 mad + r5 = r5 | r6; // 21 or + r4 = r4 | r7; // 22 or + r1 = r1 + r5 + ((((sel >> 24u) & 1u) != 0u) ? 0x81c807e5u : 0xc4e71168u); // 23 add + r6 = rotr_var(r6, r4); // 24 rotr + r7 = rotl_imm(r7, 8u); // 25 rotl + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 26 load + r3 = r3 + r2 + ((((sel >> 24u) & 1u) != 0u) ? 0xce63be51u : 0x31062899u); // 27 add + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 28 load + r4 = r4 ^ r6; // 29 xor + r6 = r6 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x20c63d72u : 0x0b8b0fbcu); // 30 add + r1 = r1 + r4 + ((((sel >> 9u) & 1u) != 0u) ? 0x8c39cfe4u : 0xb98107c2u); // 31 add + r4 = r4 - r7; // 32 sub + r5 = rotl_imm(r5, 1u); // 33 rotl + r5 = r5 * r1; // 34 mul + r3 = r3 + r0 + ((((sel >> 9u) & 1u) != 0u) ? 0xe2ce0632u : 0x1917b710u); // 35 add + r4 = r4 | r0; // 36 or + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 37 load + r0 = mul_hi(r0, r3); // 38 mulhi + r0 = r0 + r1 + ((((sel >> 8u) & 1u) != 0u) ? 0xd525e51du : 0xb6c2e64eu); // 39 add + r7 = rotl_imm(r7, 14u); // 40 rotl + r5 = r7 * r5 + r5; // 41 mad + r0 = r0 ^ ds[r3 & mask]; // 42 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 43 load + r6 = mul_hi(r6, r7); // 44 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 45 shfl + r6 = r6 + r7 + ((((sel >> 24u) & 1u) != 0u) ? 0xa60938beu : 0xa80ad691u); // 46 add + r7 = rotr_var(r7, r4); // 47 rotr + r0 = r7 * r3 + r0; // 48 mad + r7 = r7 * r4; // 49 mul + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 50 load + r7 = rotl_imm(r7, 29u); // 51 rotl + r1 = r1 * r6; // 52 mul + r6 = r6 ^ r3; // 53 xor + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 54 load + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 55 load + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 56 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 2u); r2 = r2 ^ t_; } // 57 shfl + r4 = r4 + r0 + ((((sel >> 13u) & 1u) != 0u) ? 0x62cbd03eu : 0x684cdfc9u); // 58 add + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 59 load + r2 = mul_hi(r2, r6); // 60 mulhi + r5 = r5 - r3; // 61 sub + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 62 load + r7 = r7 + r3 + ((((sel >> 13u) & 1u) != 0u) ? 0xdf1bb20eu : 0xd13cc0b2u); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/mixB-3/kernel.cu b/proto-cuda/packs-readwidth/mixB-3/kernel.cu new file mode 100644 index 000000000..10e0d1352 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-3/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/3". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x641145b6u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0xd54149dau; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0xd54149dau; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xb9deccd5u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xb9deccd5u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0xd7a898afu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0xd7a898afu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xb3350422u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xb3350422u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xe03937b0u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xe03937b0u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xd4c29221u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0xd4c29221u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x3188f3b9u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x3188f3b9u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x641145b6u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r4 = r4 ^ __shfl_xor_sync(0xffffffffu, r2, 8); // 0 shfl + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 1 load + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 8); // 2 shfl + r2 = r2 ^ r4; // 3 xor + r2 = __umulhi(r2, r3); // 4 mulhi + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 5 shfl + r2 = r2 + r0 + ((((sel >> 23u) & 1u) != 0u) ? 0xd05f6c1eu : 0x846ac225u); // 6 add + r6 = r6 | r4; // 7 or + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r1, 2); // 8 shfl + r0 = r0 ^ ds[r5 & mask]; // 9 load + r2 = r2 + r3 + ((((sel >> 0u) & 1u) != 0u) ? 0xd712e0e2u : 0x610f3311u); // 10 add + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 load + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 12 load + r2 = r7 * r2 + r2; // 13 mad + r4 = rotl_imm(r4, 16u); // 14 rotl + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 15 load + r2 = r2 * r0; // 16 mul + r3 = r3 + r2 + ((((sel >> 25u) & 1u) != 0u) ? 0xbd25a0e1u : 0x4b2e8512u); // 17 add + r7 = r7 - r3; // 18 sub + r0 = rotr_var(r0, r1); // 19 rotr + r7 = r3 * r2 + r7; // 20 mad + r5 = r5 | r6; // 21 or + r4 = r4 | r7; // 22 or + r1 = r1 + r5 + ((((sel >> 24u) & 1u) != 0u) ? 0x81c807e5u : 0xc4e71168u); // 23 add + r6 = rotr_var(r6, r4); // 24 rotr + r7 = rotl_imm(r7, 8u); // 25 rotl + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 26 load + r3 = r3 + r2 + ((((sel >> 24u) & 1u) != 0u) ? 0xce63be51u : 0x31062899u); // 27 add + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 28 load + r4 = r4 ^ r6; // 29 xor + r6 = r6 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x20c63d72u : 0x0b8b0fbcu); // 30 add + r1 = r1 + r4 + ((((sel >> 9u) & 1u) != 0u) ? 0x8c39cfe4u : 0xb98107c2u); // 31 add + r4 = r4 - r7; // 32 sub + r5 = rotl_imm(r5, 1u); // 33 rotl + r5 = r5 * r1; // 34 mul + r3 = r3 + r0 + ((((sel >> 9u) & 1u) != 0u) ? 0xe2ce0632u : 0x1917b710u); // 35 add + r4 = r4 | r0; // 36 or + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 37 load + r0 = __umulhi(r0, r3); // 38 mulhi + r0 = r0 + r1 + ((((sel >> 8u) & 1u) != 0u) ? 0xd525e51du : 0xb6c2e64eu); // 39 add + r7 = rotl_imm(r7, 14u); // 40 rotl + r5 = r7 * r5 + r5; // 41 mad + r0 = r0 ^ ds[r3 & mask]; // 42 load + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 43 load + r6 = __umulhi(r6, r7); // 44 mulhi + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 45 shfl + r6 = r6 + r7 + ((((sel >> 24u) & 1u) != 0u) ? 0xa60938beu : 0xa80ad691u); // 46 add + r7 = rotr_var(r7, r4); // 47 rotr + r0 = r7 * r3 + r0; // 48 mad + r7 = r7 * r4; // 49 mul + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 50 load + r7 = rotl_imm(r7, 29u); // 51 rotl + r1 = r1 * r6; // 52 mul + r6 = r6 ^ r3; // 53 xor + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 54 load + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 55 load + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 56 load + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r3, 2); // 57 shfl + r4 = r4 + r0 + ((((sel >> 13u) & 1u) != 0u) ? 0x62cbd03eu : 0x684cdfc9u); // 58 add + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 59 load + r2 = __umulhi(r2, r6); // 60 mulhi + r5 = r5 - r3; // 61 sub + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 62 load + r7 = r7 + r3 + ((((sel >> 13u) & 1u) != 0u) ? 0xdf1bb20eu : 0xd13cc0b2u); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixB-3/kernel_bound.cl b/proto-cuda/packs-readwidth/mixB-3/kernel_bound.cl new file mode 100644 index 000000000..1d7402896 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-3/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/3". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x641145b6u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0xd54149dau; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0xd54149dau; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xb9deccd5u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xb9deccd5u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0xd7a898afu; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0xd7a898afu; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xb3350422u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xb3350422u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xe03937b0u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xe03937b0u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xd4c29221u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xd4c29221u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x3188f3b9u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x3188f3b9u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x641145b6u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 8u); r4 = r4 ^ t_; } // 0 shfl + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 1 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 8u); r5 = r5 ^ t_; } // 2 shfl + r2 = r2 ^ r4; // 3 xor + r2 = mul_hi(r2, r3); // 4 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r1 = r1 ^ t_; } // 5 shfl + r2 = r2 + r0 + ((((sel >> 23u) & 1u) != 0u) ? 0xd05f6c1eu : 0x846ac225u); // 6 add + r6 = r6 | r4; // 7 or + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 2u); r7 = r7 ^ t_; } // 8 shfl + r0 = r0 ^ ds[r5 & mask]; // 9 load + r2 = r2 + r3 + ((((sel >> 0u) & 1u) != 0u) ? 0xd712e0e2u : 0x610f3311u); // 10 add + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 load + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 12 load + r2 = r7 * r2 + r2; // 13 mad + r4 = rotl_imm(r4, 16u); // 14 rotl + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 15 load + r2 = r2 * r0; // 16 mul + r3 = r3 + r2 + ((((sel >> 25u) & 1u) != 0u) ? 0xbd25a0e1u : 0x4b2e8512u); // 17 add + r7 = r7 - r3; // 18 sub + r0 = rotr_var(r0, r1); // 19 rotr + r7 = r3 * r2 + r7; // 20 mad + r5 = r5 | r6; // 21 or + r4 = r4 | r7; // 22 or + r1 = r1 + r5 + ((((sel >> 24u) & 1u) != 0u) ? 0x81c807e5u : 0xc4e71168u); // 23 add + r6 = rotr_var(r6, r4); // 24 rotr + r7 = rotl_imm(r7, 8u); // 25 rotl + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 26 load + r3 = r3 + r2 + ((((sel >> 24u) & 1u) != 0u) ? 0xce63be51u : 0x31062899u); // 27 add + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 28 load + r4 = r4 ^ r6; // 29 xor + r6 = r6 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x20c63d72u : 0x0b8b0fbcu); // 30 add + r1 = r1 + r4 + ((((sel >> 9u) & 1u) != 0u) ? 0x8c39cfe4u : 0xb98107c2u); // 31 add + r4 = r4 - r7; // 32 sub + r5 = rotl_imm(r5, 1u); // 33 rotl + r5 = r5 * r1; // 34 mul + r3 = r3 + r0 + ((((sel >> 9u) & 1u) != 0u) ? 0xe2ce0632u : 0x1917b710u); // 35 add + r4 = r4 | r0; // 36 or + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 37 load + r0 = mul_hi(r0, r3); // 38 mulhi + r0 = r0 + r1 + ((((sel >> 8u) & 1u) != 0u) ? 0xd525e51du : 0xb6c2e64eu); // 39 add + r7 = rotl_imm(r7, 14u); // 40 rotl + r5 = r7 * r5 + r5; // 41 mad + r0 = r0 ^ ds[r3 & mask]; // 42 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 43 load + r6 = mul_hi(r6, r7); // 44 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 45 shfl + r6 = r6 + r7 + ((((sel >> 24u) & 1u) != 0u) ? 0xa60938beu : 0xa80ad691u); // 46 add + r7 = rotr_var(r7, r4); // 47 rotr + r0 = r7 * r3 + r0; // 48 mad + r7 = r7 * r4; // 49 mul + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 50 load + r7 = rotl_imm(r7, 29u); // 51 rotl + r1 = r1 * r6; // 52 mul + r6 = r6 ^ r3; // 53 xor + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 54 load + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 55 load + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 56 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 2u); r2 = r2 ^ t_; } // 57 shfl + r4 = r4 + r0 + ((((sel >> 13u) & 1u) != 0u) ? 0x62cbd03eu : 0x684cdfc9u); // 58 add + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 59 load + r2 = mul_hi(r2, r6); // 60 mulhi + r5 = r5 - r3; // 61 sub + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 62 load + r7 = r7 + r3 + ((((sel >> 13u) & 1u) != 0u) ? 0xdf1bb20eu : 0xd13cc0b2u); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 8u); r4 = r4 ^ t_; } // 0 shfl + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 1 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 8u); r5 = r5 ^ t_; } // 2 shfl + r2 = r2 ^ r4; // 3 xor + r2 = mul_hi(r2, r3); // 4 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r1 = r1 ^ t_; } // 5 shfl + r2 = r2 + r0 + ((((sel >> 23u) & 1u) != 0u) ? 0xd05f6c1eu : 0x846ac225u); // 6 add + r6 = r6 | r4; // 7 or + { uint t_; IGNEUM_SHFL_XOR(t_, r1, 2u); r7 = r7 ^ t_; } // 8 shfl + r0 = r0 ^ ds[r5 & mask]; // 9 load + r2 = r2 + r3 + ((((sel >> 0u) & 1u) != 0u) ? 0xd712e0e2u : 0x610f3311u); // 10 add + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 load + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 12 load + r2 = r7 * r2 + r2; // 13 mad + r4 = rotl_imm(r4, 16u); // 14 rotl + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 15 load + r2 = r2 * r0; // 16 mul + r3 = r3 + r2 + ((((sel >> 25u) & 1u) != 0u) ? 0xbd25a0e1u : 0x4b2e8512u); // 17 add + r7 = r7 - r3; // 18 sub + r0 = rotr_var(r0, r1); // 19 rotr + r7 = r3 * r2 + r7; // 20 mad + r5 = r5 | r6; // 21 or + r4 = r4 | r7; // 22 or + r1 = r1 + r5 + ((((sel >> 24u) & 1u) != 0u) ? 0x81c807e5u : 0xc4e71168u); // 23 add + r6 = rotr_var(r6, r4); // 24 rotr + r7 = rotl_imm(r7, 8u); // 25 rotl + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 26 load + r3 = r3 + r2 + ((((sel >> 24u) & 1u) != 0u) ? 0xce63be51u : 0x31062899u); // 27 add + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 28 load + r4 = r4 ^ r6; // 29 xor + r6 = r6 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x20c63d72u : 0x0b8b0fbcu); // 30 add + r1 = r1 + r4 + ((((sel >> 9u) & 1u) != 0u) ? 0x8c39cfe4u : 0xb98107c2u); // 31 add + r4 = r4 - r7; // 32 sub + r5 = rotl_imm(r5, 1u); // 33 rotl + r5 = r5 * r1; // 34 mul + r3 = r3 + r0 + ((((sel >> 9u) & 1u) != 0u) ? 0xe2ce0632u : 0x1917b710u); // 35 add + r4 = r4 | r0; // 36 or + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 37 load + r0 = mul_hi(r0, r3); // 38 mulhi + r0 = r0 + r1 + ((((sel >> 8u) & 1u) != 0u) ? 0xd525e51du : 0xb6c2e64eu); // 39 add + r7 = rotl_imm(r7, 14u); // 40 rotl + r5 = r7 * r5 + r5; // 41 mad + r0 = r0 ^ ds[r3 & mask]; // 42 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 43 load + r6 = mul_hi(r6, r7); // 44 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 45 shfl + r6 = r6 + r7 + ((((sel >> 24u) & 1u) != 0u) ? 0xa60938beu : 0xa80ad691u); // 46 add + r7 = rotr_var(r7, r4); // 47 rotr + r0 = r7 * r3 + r0; // 48 mad + r7 = r7 * r4; // 49 mul + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 50 load + r7 = rotl_imm(r7, 29u); // 51 rotl + r1 = r1 * r6; // 52 mul + r6 = r6 ^ r3; // 53 xor + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 54 load + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 55 load + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 56 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 2u); r2 = r2 ^ t_; } // 57 shfl + r4 = r4 + r0 + ((((sel >> 13u) & 1u) != 0u) ? 0x62cbd03eu : 0x684cdfc9u); // 58 add + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 59 load + r2 = mul_hi(r2, r6); // 60 mulhi + r5 = r5 - r3; // 61 sub + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 62 load + r7 = r7 + r3 + ((((sel >> 13u) & 1u) != 0u) ? 0xdf1bb20eu : 0xd13cc0b2u); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-3/kernel_bound.cu b/proto-cuda/packs-readwidth/mixB-3/kernel_bound.cu new file mode 100644 index 000000000..bb20548a0 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-3/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/3". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r4 = r4 ^ __shfl_xor_sync(0xffffffffu, r2, 8); // 0 shfl + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 1 load + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 8); // 2 shfl + r2 = r2 ^ r4; // 3 xor + r2 = __umulhi(r2, r3); // 4 mulhi + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 5 shfl + r2 = r2 + r0 + ((((sel >> 23u) & 1u) != 0u) ? 0xd05f6c1eu : 0x846ac225u); // 6 add + r6 = r6 | r4; // 7 or + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r1, 2); // 8 shfl + r0 = r0 ^ ds[r5 & mask]; // 9 load + r2 = r2 + r3 + ((((sel >> 0u) & 1u) != 0u) ? 0xd712e0e2u : 0x610f3311u); // 10 add + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 load + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 12 load + r2 = r7 * r2 + r2; // 13 mad + r4 = rotl_imm(r4, 16u); // 14 rotl + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 15 load + r2 = r2 * r0; // 16 mul + r3 = r3 + r2 + ((((sel >> 25u) & 1u) != 0u) ? 0xbd25a0e1u : 0x4b2e8512u); // 17 add + r7 = r7 - r3; // 18 sub + r0 = rotr_var(r0, r1); // 19 rotr + r7 = r3 * r2 + r7; // 20 mad + r5 = r5 | r6; // 21 or + r4 = r4 | r7; // 22 or + r1 = r1 + r5 + ((((sel >> 24u) & 1u) != 0u) ? 0x81c807e5u : 0xc4e71168u); // 23 add + r6 = rotr_var(r6, r4); // 24 rotr + r7 = rotl_imm(r7, 8u); // 25 rotl + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 26 load + r3 = r3 + r2 + ((((sel >> 24u) & 1u) != 0u) ? 0xce63be51u : 0x31062899u); // 27 add + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 28 load + r4 = r4 ^ r6; // 29 xor + r6 = r6 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x20c63d72u : 0x0b8b0fbcu); // 30 add + r1 = r1 + r4 + ((((sel >> 9u) & 1u) != 0u) ? 0x8c39cfe4u : 0xb98107c2u); // 31 add + r4 = r4 - r7; // 32 sub + r5 = rotl_imm(r5, 1u); // 33 rotl + r5 = r5 * r1; // 34 mul + r3 = r3 + r0 + ((((sel >> 9u) & 1u) != 0u) ? 0xe2ce0632u : 0x1917b710u); // 35 add + r4 = r4 | r0; // 36 or + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 37 load + r0 = __umulhi(r0, r3); // 38 mulhi + r0 = r0 + r1 + ((((sel >> 8u) & 1u) != 0u) ? 0xd525e51du : 0xb6c2e64eu); // 39 add + r7 = rotl_imm(r7, 14u); // 40 rotl + r5 = r7 * r5 + r5; // 41 mad + r0 = r0 ^ ds[r3 & mask]; // 42 load + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 43 load + r6 = __umulhi(r6, r7); // 44 mulhi + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 45 shfl + r6 = r6 + r7 + ((((sel >> 24u) & 1u) != 0u) ? 0xa60938beu : 0xa80ad691u); // 46 add + r7 = rotr_var(r7, r4); // 47 rotr + r0 = r7 * r3 + r0; // 48 mad + r7 = r7 * r4; // 49 mul + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 50 load + r7 = rotl_imm(r7, 29u); // 51 rotl + r1 = r1 * r6; // 52 mul + r6 = r6 ^ r3; // 53 xor + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 54 load + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 55 load + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 56 load + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r3, 2); // 57 shfl + r4 = r4 + r0 + ((((sel >> 13u) & 1u) != 0u) ? 0x62cbd03eu : 0x684cdfc9u); // 58 add + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 59 load + r2 = __umulhi(r2, r6); // 60 mulhi + r5 = r5 - r3; // 61 sub + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 62 load + r7 = r7 + r3 + ((((sel >> 13u) & 1u) != 0u) ? 0xdf1bb20eu : 0xd13cc0b2u); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixB-3/memhard.h b/proto-cuda/packs-readwidth/mixB-3/memhard.h new file mode 100644 index 000000000..a4be69e3e --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-3/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/3". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/mixB-3/memhard.metal b/proto-cuda/packs-readwidth/mixB-3/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-3/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/mixB-3/program.h b/proto-cuda/packs-readwidth/mixB-3/program.h new file mode 100644 index 000000000..7b029af5d --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-3/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/3". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-readwidth/B/3" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d7265616477696474682f422f33" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0xdbdcc6fcf26f6a3full +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=12 shfl=6 rotl=5 mad=4 mul=4 mulhi=4 or=4 rotr=3 sub=3 xor=3" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "mix25-50-25" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 25, 50, 25 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 2, 8, 6 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 4160 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x641145b6u, 0xd54149dau, 0xb9deccd5u, 0xd7a898afu, 0xb3350422u, 0xe03937b0u, 0xd4c29221u, 0x3188f3b9u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/mixB-3/program.json b/proto-cuda/packs-readwidth/mixB-3/program.json new file mode 100644 index 000000000..bf21630ff --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-3/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0xdbdcc6fcf26f6a3f", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-readwidth/B/3", + "seed_bytes": "69676e65756d2d7265616477696474682f422f33", + "seed_words": ["0x641145b6", "0xd54149da", "0xb9deccd5", "0xd7a898af", "0xb3350422", "0xe03937b0", "0xd4c29221", "0x3188f3b9"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "mix25-50-25", + "load_slots": 16, + "load_mix_percent_4_16_64": [25, 50, 25], + "load_width_counts_4_16_64": [2, 8, 6], + "bytes_per_hash": 4160, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "add": 12, "shfl": 6, "rotl": 5, "mad": 4, "mul": 4, "mulhi": 4, "or": 4, "rotr": 3, "sub": 3, "xor": 3}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "shfl", "dst": 4, "src": 2, "src2": 5, "imm": "0xe2a27ec9", "imm2": "0x52997d6e", "rot": 9, "bit": 19, "mask": 8, "width": 1}, + {"i": 1, "op": "load", "dst": 3, "src": 4, "src2": 2, "imm": "0x018bff2b", "imm2": "0x938f0b29", "rot": 3, "bit": 8, "mask": 1, "width": 4}, + {"i": 2, "op": "shfl", "dst": 5, "src": 2, "src2": 1, "imm": "0x661b4328", "imm2": "0xab2aa600", "rot": 27, "bit": 31, "mask": 8, "width": 1}, + {"i": 3, "op": "xor", "dst": 2, "src": 4, "src2": 5, "imm": "0xd4843df7", "imm2": "0xd97d6521", "rot": 20, "bit": 11, "mask": 4, "width": 1}, + {"i": 4, "op": "mulhi", "dst": 2, "src": 3, "src2": 7, "imm": "0xd07bc032", "imm2": "0x3e9064c2", "rot": 9, "bit": 14, "mask": 4, "width": 1}, + {"i": 5, "op": "shfl", "dst": 1, "src": 2, "src2": 7, "imm": "0xa4121cbc", "imm2": "0x79f0bc9c", "rot": 20, "bit": 28, "mask": 4, "width": 1}, + {"i": 6, "op": "add", "dst": 2, "src": 0, "src2": 3, "imm": "0x846ac225", "imm2": "0xd05f6c1e", "rot": 25, "bit": 23, "mask": 16, "width": 1}, + {"i": 7, "op": "or", "dst": 6, "src": 4, "src2": 2, "imm": "0x71a059df", "imm2": "0x44ea7a0e", "rot": 31, "bit": 17, "mask": 2, "width": 1}, + {"i": 8, "op": "shfl", "dst": 7, "src": 1, "src2": 6, "imm": "0x1d2ba406", "imm2": "0xc168b9d6", "rot": 24, "bit": 28, "mask": 2, "width": 1}, + {"i": 9, "op": "load", "dst": 0, "src": 5, "src2": 1, "imm": "0x1d17335c", "imm2": "0xb600a287", "rot": 8, "bit": 18, "mask": 1, "width": 1}, + {"i": 10, "op": "add", "dst": 2, "src": 3, "src2": 6, "imm": "0x610f3311", "imm2": "0xd712e0e2", "rot": 1, "bit": 0, "mask": 1, "width": 1}, + {"i": 11, "op": "load", "dst": 5, "src": 0, "src2": 6, "imm": "0x37535f59", "imm2": "0xae703e33", "rot": 7, "bit": 0, "mask": 16, "width": 4}, + {"i": 12, "op": "load", "dst": 0, "src": 7, "src2": 0, "imm": "0x092eda3f", "imm2": "0xde068ac2", "rot": 11, "bit": 7, "mask": 4, "width": 16}, + {"i": 13, "op": "mad", "dst": 2, "src": 7, "src2": 2, "imm": "0x455e127e", "imm2": "0x4866fdfe", "rot": 5, "bit": 23, "mask": 8, "width": 1}, + {"i": 14, "op": "rotl", "dst": 4, "src": 7, "src2": 6, "imm": "0x4ede57ee", "imm2": "0x3af2491c", "rot": 16, "bit": 5, "mask": 1, "width": 1}, + {"i": 15, "op": "load", "dst": 5, "src": 6, "src2": 3, "imm": "0xa67fa362", "imm2": "0x21df0701", "rot": 12, "bit": 1, "mask": 2, "width": 4}, + {"i": 16, "op": "mul", "dst": 2, "src": 0, "src2": 0, "imm": "0x5c8127f4", "imm2": "0xafc8d3ef", "rot": 8, "bit": 7, "mask": 4, "width": 1}, + {"i": 17, "op": "add", "dst": 3, "src": 2, "src2": 7, "imm": "0x4b2e8512", "imm2": "0xbd25a0e1", "rot": 7, "bit": 25, "mask": 8, "width": 1}, + {"i": 18, "op": "sub", "dst": 7, "src": 3, "src2": 2, "imm": "0xe79e2662", "imm2": "0xcb84ce6e", "rot": 27, "bit": 25, "mask": 8, "width": 1}, + {"i": 19, "op": "rotr", "dst": 0, "src": 1, "src2": 2, "imm": "0x7a4b785f", "imm2": "0x434d64a6", "rot": 31, "bit": 22, "mask": 8, "width": 1}, + {"i": 20, "op": "mad", "dst": 7, "src": 3, "src2": 2, "imm": "0x34a1c288", "imm2": "0x4cf8fb63", "rot": 17, "bit": 0, "mask": 8, "width": 1}, + {"i": 21, "op": "or", "dst": 5, "src": 6, "src2": 4, "imm": "0x9bd1b2f1", "imm2": "0x2e6abc9b", "rot": 22, "bit": 21, "mask": 16, "width": 1}, + {"i": 22, "op": "or", "dst": 4, "src": 7, "src2": 4, "imm": "0x841bc7f8", "imm2": "0x42650812", "rot": 31, "bit": 15, "mask": 16, "width": 1}, + {"i": 23, "op": "add", "dst": 1, "src": 5, "src2": 0, "imm": "0xc4e71168", "imm2": "0x81c807e5", "rot": 27, "bit": 24, "mask": 8, "width": 1}, + {"i": 24, "op": "rotr", "dst": 6, "src": 4, "src2": 2, "imm": "0x38b95599", "imm2": "0x644166d5", "rot": 9, "bit": 27, "mask": 2, "width": 1}, + {"i": 25, "op": "rotl", "dst": 7, "src": 1, "src2": 1, "imm": "0x76dd04ae", "imm2": "0x41afc946", "rot": 8, "bit": 9, "mask": 4, "width": 1}, + {"i": 26, "op": "load", "dst": 5, "src": 0, "src2": 4, "imm": "0x12c82e95", "imm2": "0x7a8ba490", "rot": 2, "bit": 25, "mask": 8, "width": 4}, + {"i": 27, "op": "add", "dst": 3, "src": 2, "src2": 6, "imm": "0x31062899", "imm2": "0xce63be51", "rot": 24, "bit": 24, "mask": 4, "width": 1}, + {"i": 28, "op": "load", "dst": 5, "src": 3, "src2": 1, "imm": "0x643a93a2", "imm2": "0x797b06a0", "rot": 27, "bit": 8, "mask": 4, "width": 16}, + {"i": 29, "op": "xor", "dst": 4, "src": 6, "src2": 7, "imm": "0xb6b6f83b", "imm2": "0x8f6dcf41", "rot": 22, "bit": 11, "mask": 2, "width": 1}, + {"i": 30, "op": "add", "dst": 6, "src": 3, "src2": 7, "imm": "0x0b8b0fbc", "imm2": "0x20c63d72", "rot": 4, "bit": 4, "mask": 16, "width": 1}, + {"i": 31, "op": "add", "dst": 1, "src": 4, "src2": 1, "imm": "0xb98107c2", "imm2": "0x8c39cfe4", "rot": 30, "bit": 9, "mask": 4, "width": 1}, + {"i": 32, "op": "sub", "dst": 4, "src": 7, "src2": 4, "imm": "0x77045c96", "imm2": "0xe325885b", "rot": 20, "bit": 10, "mask": 16, "width": 1}, + {"i": 33, "op": "rotl", "dst": 5, "src": 6, "src2": 1, "imm": "0xc0d8c504", "imm2": "0xa5a87427", "rot": 1, "bit": 4, "mask": 1, "width": 1}, + {"i": 34, "op": "mul", "dst": 5, "src": 1, "src2": 2, "imm": "0xb8601550", "imm2": "0x58384e6f", "rot": 5, "bit": 20, "mask": 1, "width": 1}, + {"i": 35, "op": "add", "dst": 3, "src": 0, "src2": 7, "imm": "0x1917b710", "imm2": "0xe2ce0632", "rot": 18, "bit": 9, "mask": 2, "width": 1}, + {"i": 36, "op": "or", "dst": 4, "src": 0, "src2": 6, "imm": "0x83e99926", "imm2": "0x25af16ba", "rot": 6, "bit": 18, "mask": 8, "width": 1}, + {"i": 37, "op": "load", "dst": 0, "src": 2, "src2": 4, "imm": "0x6e767327", "imm2": "0xdbbe352d", "rot": 29, "bit": 17, "mask": 2, "width": 16}, + {"i": 38, "op": "mulhi", "dst": 0, "src": 3, "src2": 6, "imm": "0xcfc2910c", "imm2": "0x98c9a2c8", "rot": 13, "bit": 10, "mask": 2, "width": 1}, + {"i": 39, "op": "add", "dst": 0, "src": 1, "src2": 3, "imm": "0xb6c2e64e", "imm2": "0xd525e51d", "rot": 17, "bit": 8, "mask": 2, "width": 1}, + {"i": 40, "op": "rotl", "dst": 7, "src": 2, "src2": 1, "imm": "0x3f2b4296", "imm2": "0x5dee0e9e", "rot": 14, "bit": 11, "mask": 2, "width": 1}, + {"i": 41, "op": "mad", "dst": 5, "src": 7, "src2": 5, "imm": "0xf8cb9bb4", "imm2": "0xbeb81676", "rot": 22, "bit": 14, "mask": 1, "width": 1}, + {"i": 42, "op": "load", "dst": 0, "src": 3, "src2": 6, "imm": "0xac55c404", "imm2": "0x0ac01b55", "rot": 27, "bit": 30, "mask": 2, "width": 1}, + {"i": 43, "op": "load", "dst": 0, "src": 7, "src2": 6, "imm": "0xfb99e0c0", "imm2": "0x20fb35b2", "rot": 13, "bit": 8, "mask": 4, "width": 4}, + {"i": 44, "op": "mulhi", "dst": 6, "src": 7, "src2": 4, "imm": "0xbfada51f", "imm2": "0xc7388304", "rot": 17, "bit": 30, "mask": 16, "width": 1}, + {"i": 45, "op": "shfl", "dst": 3, "src": 5, "src2": 0, "imm": "0x6323d6a6", "imm2": "0x2bfea98b", "rot": 24, "bit": 27, "mask": 1, "width": 1}, + {"i": 46, "op": "add", "dst": 6, "src": 7, "src2": 1, "imm": "0xa80ad691", "imm2": "0xa60938be", "rot": 22, "bit": 24, "mask": 1, "width": 1}, + {"i": 47, "op": "rotr", "dst": 7, "src": 4, "src2": 2, "imm": "0x0c2eb5f8", "imm2": "0xaa639df2", "rot": 15, "bit": 2, "mask": 4, "width": 1}, + {"i": 48, "op": "mad", "dst": 0, "src": 7, "src2": 3, "imm": "0x231aa6be", "imm2": "0x5b6ee205", "rot": 13, "bit": 23, "mask": 2, "width": 1}, + {"i": 49, "op": "mul", "dst": 7, "src": 4, "src2": 7, "imm": "0x1a40720a", "imm2": "0x4c6033ff", "rot": 5, "bit": 8, "mask": 4, "width": 1}, + {"i": 50, "op": "load", "dst": 2, "src": 7, "src2": 5, "imm": "0x027bb9e7", "imm2": "0xda3cb47c", "rot": 22, "bit": 8, "mask": 4, "width": 4}, + {"i": 51, "op": "rotl", "dst": 7, "src": 0, "src2": 7, "imm": "0xd7338678", "imm2": "0xaf250992", "rot": 29, "bit": 18, "mask": 4, "width": 1}, + {"i": 52, "op": "mul", "dst": 1, "src": 6, "src2": 7, "imm": "0xe1c24369", "imm2": "0x8f12532b", "rot": 11, "bit": 13, "mask": 1, "width": 1}, + {"i": 53, "op": "xor", "dst": 6, "src": 3, "src2": 0, "imm": "0x1ca8117f", "imm2": "0x4e0c544e", "rot": 16, "bit": 27, "mask": 8, "width": 1}, + {"i": 54, "op": "load", "dst": 3, "src": 7, "src2": 4, "imm": "0x97c7b342", "imm2": "0xcfcc6b7f", "rot": 7, "bit": 12, "mask": 4, "width": 16}, + {"i": 55, "op": "load", "dst": 3, "src": 2, "src2": 0, "imm": "0xbc3595d8", "imm2": "0x9ece5818", "rot": 13, "bit": 29, "mask": 8, "width": 4}, + {"i": 56, "op": "load", "dst": 4, "src": 0, "src2": 2, "imm": "0xb24abdf1", "imm2": "0xc45a2ce8", "rot": 27, "bit": 17, "mask": 2, "width": 16}, + {"i": 57, "op": "shfl", "dst": 2, "src": 3, "src2": 2, "imm": "0xbfdde7e8", "imm2": "0x51be0c90", "rot": 3, "bit": 17, "mask": 2, "width": 1}, + {"i": 58, "op": "add", "dst": 4, "src": 0, "src2": 3, "imm": "0x684cdfc9", "imm2": "0x62cbd03e", "rot": 1, "bit": 13, "mask": 16, "width": 1}, + {"i": 59, "op": "load", "dst": 3, "src": 2, "src2": 6, "imm": "0xed925765", "imm2": "0x4fbd58fd", "rot": 27, "bit": 17, "mask": 16, "width": 16}, + {"i": 60, "op": "mulhi", "dst": 2, "src": 6, "src2": 1, "imm": "0x960864aa", "imm2": "0xa93da8c7", "rot": 2, "bit": 28, "mask": 8, "width": 1}, + {"i": 61, "op": "sub", "dst": 5, "src": 3, "src2": 0, "imm": "0xf9aafd21", "imm2": "0x4e69b9ec", "rot": 24, "bit": 28, "mask": 8, "width": 1}, + {"i": 62, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xc9b9f7c4", "imm2": "0x70b87e1b", "rot": 27, "bit": 27, "mask": 2, "width": 4}, + {"i": 63, "op": "add", "dst": 7, "src": 3, "src2": 2, "imm": "0xd13cc0b2", "imm2": "0xdf1bb20e", "rot": 12, "bit": 13, "mask": 8, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/mixB-3/program.metal b/proto-cuda/packs-readwidth/mixB-3/program.metal new file mode 100644 index 000000000..c1240a5db --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-3/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x641145b6u, 0xd54149dau, 0xb9deccd5u, 0xd7a898afu, 0xb3350422u, 0xe03937b0u, 0xd4c29221u, 0x3188f3b9u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r4 = r4 ^ simd_shuffle_xor(r2, (ushort)8); // 0 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 1 + r5 = r5 ^ simd_shuffle_xor(r2, (ushort)8); // 2 + r2 = r2 ^ r4; // 3 + r2 = mulhi(r2, r3); // 4 + r1 = r1 ^ simd_shuffle_xor(r2, (ushort)4); // 5 + r2 = r2 + r0 + select(0x846ac225u, 0xd05f6c1eu, ((sel >> 23u) & 1u) != 0u); // 6 + r6 = r6 | r4; // 7 + r7 = r7 ^ simd_shuffle_xor(r1, (ushort)2); // 8 + r0 = r0 ^ dataset[r5 & MASK]; // 9 + r2 = r2 + r3 + select(0x610f3311u, 0xd712e0e2u, ((sel >> 0u) & 1u) != 0u); // 10 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 12 + r2 = r7 * r2 + r2; // 13 + r4 = rotl_imm(r4, 16u); // 14 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 15 + r2 = r2 * r0; // 16 + r3 = r3 + r2 + select(0x4b2e8512u, 0xbd25a0e1u, ((sel >> 25u) & 1u) != 0u); // 17 + r7 = r7 - r3; // 18 + r0 = rotr_var(r0, r1); // 19 + r7 = r3 * r2 + r7; // 20 + r5 = r5 | r6; // 21 + r4 = r4 | r7; // 22 + r1 = r1 + r5 + select(0xc4e71168u, 0x81c807e5u, ((sel >> 24u) & 1u) != 0u); // 23 + r6 = rotr_var(r6, r4); // 24 + r7 = rotl_imm(r7, 8u); // 25 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 26 + r3 = r3 + r2 + select(0x31062899u, 0xce63be51u, ((sel >> 24u) & 1u) != 0u); // 27 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 28 + r4 = r4 ^ r6; // 29 + r6 = r6 + r3 + select(0x0b8b0fbcu, 0x20c63d72u, ((sel >> 4u) & 1u) != 0u); // 30 + r1 = r1 + r4 + select(0xb98107c2u, 0x8c39cfe4u, ((sel >> 9u) & 1u) != 0u); // 31 + r4 = r4 - r7; // 32 + r5 = rotl_imm(r5, 1u); // 33 + r5 = r5 * r1; // 34 + r3 = r3 + r0 + select(0x1917b710u, 0xe2ce0632u, ((sel >> 9u) & 1u) != 0u); // 35 + r4 = r4 | r0; // 36 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 37 + r0 = mulhi(r0, r3); // 38 + r0 = r0 + r1 + select(0xb6c2e64eu, 0xd525e51du, ((sel >> 8u) & 1u) != 0u); // 39 + r7 = rotl_imm(r7, 14u); // 40 + r5 = r7 * r5 + r5; // 41 + r0 = r0 ^ dataset[r3 & MASK]; // 42 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 43 + r6 = mulhi(r6, r7); // 44 + r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 45 + r6 = r6 + r7 + select(0xa80ad691u, 0xa60938beu, ((sel >> 24u) & 1u) != 0u); // 46 + r7 = rotr_var(r7, r4); // 47 + r0 = r7 * r3 + r0; // 48 + r7 = r7 * r4; // 49 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 50 + r7 = rotl_imm(r7, 29u); // 51 + r1 = r1 * r6; // 52 + r6 = r6 ^ r3; // 53 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 54 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 55 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 56 + r2 = r2 ^ simd_shuffle_xor(r3, (ushort)2); // 57 + r4 = r4 + r0 + select(0x684cdfc9u, 0x62cbd03eu, ((sel >> 13u) & 1u) != 0u); // 58 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 59 + r2 = mulhi(r2, r6); // 60 + r5 = r5 - r3; // 61 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 62 + r7 = r7 + r3 + select(0xd13cc0b2u, 0xdf1bb20eu, ((sel >> 13u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-3/program_bound.metal b/proto-cuda/packs-readwidth/mixB-3/program_bound.metal new file mode 100644 index 000000000..9c4b97235 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-3/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x641145b6u, 0xd54149dau, 0xb9deccd5u, 0xd7a898afu, 0xb3350422u, 0xe03937b0u, 0xd4c29221u, 0x3188f3b9u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r4 = r4 ^ simd_shuffle_xor(r2, (ushort)8); // 0 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 1 + r5 = r5 ^ simd_shuffle_xor(r2, (ushort)8); // 2 + r2 = r2 ^ r4; // 3 + r2 = mulhi(r2, r3); // 4 + r1 = r1 ^ simd_shuffle_xor(r2, (ushort)4); // 5 + r2 = r2 + r0 + select(0x846ac225u, 0xd05f6c1eu, ((sel >> 23u) & 1u) != 0u); // 6 + r6 = r6 | r4; // 7 + r7 = r7 ^ simd_shuffle_xor(r1, (ushort)2); // 8 + r0 = r0 ^ dataset[r5 & MASK]; // 9 + r2 = r2 + r3 + select(0x610f3311u, 0xd712e0e2u, ((sel >> 0u) & 1u) != 0u); // 10 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 11 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 12 + r2 = r7 * r2 + r2; // 13 + r4 = rotl_imm(r4, 16u); // 14 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 15 + r2 = r2 * r0; // 16 + r3 = r3 + r2 + select(0x4b2e8512u, 0xbd25a0e1u, ((sel >> 25u) & 1u) != 0u); // 17 + r7 = r7 - r3; // 18 + r0 = rotr_var(r0, r1); // 19 + r7 = r3 * r2 + r7; // 20 + r5 = r5 | r6; // 21 + r4 = r4 | r7; // 22 + r1 = r1 + r5 + select(0xc4e71168u, 0x81c807e5u, ((sel >> 24u) & 1u) != 0u); // 23 + r6 = rotr_var(r6, r4); // 24 + r7 = rotl_imm(r7, 8u); // 25 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 26 + r3 = r3 + r2 + select(0x31062899u, 0xce63be51u, ((sel >> 24u) & 1u) != 0u); // 27 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 28 + r4 = r4 ^ r6; // 29 + r6 = r6 + r3 + select(0x0b8b0fbcu, 0x20c63d72u, ((sel >> 4u) & 1u) != 0u); // 30 + r1 = r1 + r4 + select(0xb98107c2u, 0x8c39cfe4u, ((sel >> 9u) & 1u) != 0u); // 31 + r4 = r4 - r7; // 32 + r5 = rotl_imm(r5, 1u); // 33 + r5 = r5 * r1; // 34 + r3 = r3 + r0 + select(0x1917b710u, 0xe2ce0632u, ((sel >> 9u) & 1u) != 0u); // 35 + r4 = r4 | r0; // 36 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 37 + r0 = mulhi(r0, r3); // 38 + r0 = r0 + r1 + select(0xb6c2e64eu, 0xd525e51du, ((sel >> 8u) & 1u) != 0u); // 39 + r7 = rotl_imm(r7, 14u); // 40 + r5 = r7 * r5 + r5; // 41 + r0 = r0 ^ dataset[r3 & MASK]; // 42 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 43 + r6 = mulhi(r6, r7); // 44 + r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 45 + r6 = r6 + r7 + select(0xa80ad691u, 0xa60938beu, ((sel >> 24u) & 1u) != 0u); // 46 + r7 = rotr_var(r7, r4); // 47 + r0 = r7 * r3 + r0; // 48 + r7 = r7 * r4; // 49 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 50 + r7 = rotl_imm(r7, 29u); // 51 + r1 = r1 * r6; // 52 + r6 = r6 ^ r3; // 53 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 54 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 55 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 56 + r2 = r2 ^ simd_shuffle_xor(r3, (ushort)2); // 57 + r4 = r4 + r0 + select(0x684cdfc9u, 0x62cbd03eu, ((sel >> 13u) & 1u) != 0u); // 58 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 59 + r2 = mulhi(r2, r6); // 60 + r5 = r5 - r3; // 61 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 62 + r7 = r7 + r3 + select(0xd13cc0b2u, 0xdf1bb20eu, ((sel >> 13u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-3/vectors.h b/proto-cuda/packs-readwidth/mixB-3/vectors.h new file mode 100644 index 000000000..dc87a3f7f --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-3/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/3". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x4e8737306b3cfc38ull, 0xc231cdf825c74168ull, 0xa43dcbb2bb1506f6ull, 0xf9f34341f7058336ull, 0xcc74f41f8a648441ull, 0xc7c830fbacc55514ull, 0xb1be6fbab47c2f93ull, 0x85bb299656d21468ull, + 0x7f907ad7cdb54241ull, 0x036ad520a4e9d12aull, 0x5cd0cb5fafedc21cull, 0x323f8717958fed0full, 0x02a33f13dd97c08eull, 0x40caf1dcff2d7640ull, 0xbcf4a95b85a7e21cull, 0xf9d85cedb80cc647ull, + 0x4a7a32fc4cd069b6ull, 0xfdf98e4f8b31875eull, 0x6ed4d84bd867e3a4ull, 0x9d31839b5dfc2b07ull, 0xc723794ff44db212ull, 0x8b6798c5bbbae7aeull, 0x59240a005100b8a9ull, 0x4a3a517b44c3e520ull, + 0x8d78475b924fb0bdull, 0x0f134e2add939186ull, 0xff55688e0948be5dull, 0x29d2e6f622ba1ed1ull, 0x36544678d523021dull, 0xee17c23f966e5b16ull, 0xf4498af9abb2aae3ull, 0x02e167756380ce09ull + }, + { // base nonce 4096 + 0x1db019fd49b7d4cfull, 0x63272832e0367923ull, 0x56cbcc1213658525ull, 0x19935ebf3c265d21ull, 0xa0aef9843a03a870ull, 0xfbd0416aff2a0babull, 0xbf553f70a7a1be59ull, 0xf6a97dbeb7a7c05dull, + 0xa6e39ec1b1ff5ff7ull, 0xb1256208dee8b890ull, 0x5e1971e1d19a7ad7ull, 0xe2f6bb06bdddb1eaull, 0x5afc47580b0331b5ull, 0x08d01d967a92f449ull, 0x1aecebed8066afbbull, 0xdda960d0b52fbda1ull, + 0x363a2fa528261dc5ull, 0xaf6ab1fc1c709944ull, 0xea6ad7ee53099757ull, 0x2fe2f3d49114a71full, 0x3dcee95f3c3ea504ull, 0x3721e06985b8e75dull, 0x030b8cfda8806eb5ull, 0xe6b0b272a04876d0ull, + 0x020f94c4253a82e0ull, 0xbfec18cba32bc55aull, 0x228743dbbf4401fdull, 0xfac6cdd28436685aull, 0x1ec551010a39ff0eull, 0x3de13c246acd4aafull, 0x93d678f6aa9be46bull, 0x5dc5df14b47957a9ull + }, + { // base nonce 1000000 + 0xd2b37babfd478e77ull, 0x78a6be49ce7f374dull, 0xd6719a3b1eec99adull, 0xc1466ec0fc84be57ull, 0x521f1a72b6f4e1cfull, 0xfabcd5b90f10af94ull, 0xffd895c86ee4e4d1ull, 0x09afafeba0b04339ull, + 0xf378ad96a86d56a6ull, 0x0df21d4de5b4bb66ull, 0x2ff00005a8b6c642ull, 0x0191e72d236cf845ull, 0xb9d679a98ad736f2ull, 0x1083d968d5f47081ull, 0x9b3163f22dc6f426ull, 0xbc6cbebc6fe536c0ull, + 0xbc9601ef6fdc8521ull, 0xa6d7041fbee05d19ull, 0x8e63547b22e00360ull, 0x58f46fb454037d72ull, 0x48f1c90efe9e8df3ull, 0x783c16d05a88ca7dull, 0x92975e0d3930ee02ull, 0x9f3d28afbd51b033ull, + 0x3f3db070327cf46cull, 0x9ebefffc0e234b2dull, 0xc737687909a104deull, 0x19dabbf178b2a350ull, 0x27266ba188cc08c2ull, 0xd485cd7f14c3e1adull, 0xc6021d069d226c52ull, 0x43af641dc4dd9016ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/mixB-3/vectors.json b/proto-cuda/packs-readwidth/mixB-3/vectors.json new file mode 100644 index 000000000..dfee6616a --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-3/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-readwidth/B/3", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x4e8737306b3cfc38", "0xc231cdf825c74168", "0xa43dcbb2bb1506f6", "0xf9f34341f7058336", "0xcc74f41f8a648441", "0xc7c830fbacc55514", "0xb1be6fbab47c2f93", "0x85bb299656d21468", + "0x7f907ad7cdb54241", "0x036ad520a4e9d12a", "0x5cd0cb5fafedc21c", "0x323f8717958fed0f", "0x02a33f13dd97c08e", "0x40caf1dcff2d7640", "0xbcf4a95b85a7e21c", "0xf9d85cedb80cc647", + "0x4a7a32fc4cd069b6", "0xfdf98e4f8b31875e", "0x6ed4d84bd867e3a4", "0x9d31839b5dfc2b07", "0xc723794ff44db212", "0x8b6798c5bbbae7ae", "0x59240a005100b8a9", "0x4a3a517b44c3e520", + "0x8d78475b924fb0bd", "0x0f134e2add939186", "0xff55688e0948be5d", "0x29d2e6f622ba1ed1", "0x36544678d523021d", "0xee17c23f966e5b16", "0xf4498af9abb2aae3", "0x02e167756380ce09" + ]}, + {"base_nonce": 4096, "expected": [ + "0x1db019fd49b7d4cf", "0x63272832e0367923", "0x56cbcc1213658525", "0x19935ebf3c265d21", "0xa0aef9843a03a870", "0xfbd0416aff2a0bab", "0xbf553f70a7a1be59", "0xf6a97dbeb7a7c05d", + "0xa6e39ec1b1ff5ff7", "0xb1256208dee8b890", "0x5e1971e1d19a7ad7", "0xe2f6bb06bdddb1ea", "0x5afc47580b0331b5", "0x08d01d967a92f449", "0x1aecebed8066afbb", "0xdda960d0b52fbda1", + "0x363a2fa528261dc5", "0xaf6ab1fc1c709944", "0xea6ad7ee53099757", "0x2fe2f3d49114a71f", "0x3dcee95f3c3ea504", "0x3721e06985b8e75d", "0x030b8cfda8806eb5", "0xe6b0b272a04876d0", + "0x020f94c4253a82e0", "0xbfec18cba32bc55a", "0x228743dbbf4401fd", "0xfac6cdd28436685a", "0x1ec551010a39ff0e", "0x3de13c246acd4aaf", "0x93d678f6aa9be46b", "0x5dc5df14b47957a9" + ]}, + {"base_nonce": 1000000, "expected": [ + "0xd2b37babfd478e77", "0x78a6be49ce7f374d", "0xd6719a3b1eec99ad", "0xc1466ec0fc84be57", "0x521f1a72b6f4e1cf", "0xfabcd5b90f10af94", "0xffd895c86ee4e4d1", "0x09afafeba0b04339", + "0xf378ad96a86d56a6", "0x0df21d4de5b4bb66", "0x2ff00005a8b6c642", "0x0191e72d236cf845", "0xb9d679a98ad736f2", "0x1083d968d5f47081", "0x9b3163f22dc6f426", "0xbc6cbebc6fe536c0", + "0xbc9601ef6fdc8521", "0xa6d7041fbee05d19", "0x8e63547b22e00360", "0x58f46fb454037d72", "0x48f1c90efe9e8df3", "0x783c16d05a88ca7d", "0x92975e0d3930ee02", "0x9f3d28afbd51b033", + "0x3f3db070327cf46c", "0x9ebefffc0e234b2d", "0xc737687909a104de", "0x19dabbf178b2a350", "0x27266ba188cc08c2", "0xd485cd7f14c3e1ad", "0xc6021d069d226c52", "0x43af641dc4dd9016" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/mixB-4/kernel.cl b/proto-cuda/packs-readwidth/mixB-4/kernel.cl new file mode 100644 index 000000000..f23fc88c7 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-4/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/4". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x9f289c8cu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x0b002fccu; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x0b002fccu; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x7b25b8ffu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x7b25b8ffu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x1b37f538u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x1b37f538u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x4f9ab04fu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x4f9ab04fu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x1120ee60u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x1120ee60u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xfa16504cu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xfa16504cu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xcf94f6e9u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0xcf94f6e9u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x9f289c8cu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r4 = mul_hi(r4, r6); // 0 mulhi + r1 = r1 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x69326459u : 0xf82e323du); // 1 add + r0 = r0 * r7; // 2 mul + r7 = r7 ^ ds[r4 & mask]; // 3 load + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 4 load + r7 = rotl_imm(r7, 13u); // 5 rotl + r4 = r4 | r1; // 6 or + r1 = r1 ^ r2; // 7 xor + r5 = r5 * r1; // 8 mul + r5 = mul_hi(r5, r0); // 9 mulhi + r0 = mul_hi(r0, r4); // 10 mulhi + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 load + r5 = mul_hi(r5, r3); // 12 mulhi + r0 = r0 ^ ds[r5 & mask]; // 13 load + r1 = mul_hi(r1, r5); // 14 mulhi + r0 = mul_hi(r0, r6); // 15 mulhi + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 16 load + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 17 load + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 18 load + r1 = r1 | r7; // 19 or + r6 = rotr_var(r6, r7); // 20 rotr + r1 = r1 + r7 + ((((sel >> 10u) & 1u) != 0u) ? 0xa65a329cu : 0x8e775c1eu); // 21 add + r7 = mul_hi(r7, r3); // 22 mulhi + r2 = mul_hi(r2, r0); // 23 mulhi + r7 = r7 ^ ds[r1 & mask]; // 24 load + r3 = r3 ^ r6; // 25 xor + r0 = rotl_imm(r0, 14u); // 26 rotl + r6 = r6 ^ ds[r5 & mask]; // 27 load + r3 = r3 ^ r6; // 28 xor + r2 = r2 * r4; // 29 mul + r7 = r7 + r6 + ((((sel >> 30u) & 1u) != 0u) ? 0x1027531du : 0xfcc52962u); // 30 add + r0 = r0 - r3; // 31 sub + r5 = r5 * r6; // 32 mul + r0 = r0 ^ ds[r7 & mask]; // 33 load + r4 = r4 + r0 + ((((sel >> 29u) & 1u) != 0u) ? 0x42903198u : 0xb7efd271u); // 34 add + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 16u); r2 = r2 ^ t_; } // 35 shfl + r7 = r1 * r2 + r7; // 36 mad + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 load + r6 = rotr_var(r6, r7); // 38 rotr + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 39 load + r4 = r2 * r5 + r4; // 40 mad + r3 = r3 ^ r0; // 41 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r5 = r5 ^ t_; } // 42 shfl + r7 = r7 + r4 + ((((sel >> 31u) & 1u) != 0u) ? 0xbfb91116u : 0x1a6db0eau); // 43 add + r1 = r3 * r1 + r1; // 44 mad + r1 = r1 ^ r7; // 45 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 16u); r1 = r1 ^ t_; } // 46 shfl + r7 = r7 + r3 + ((((sel >> 28u) & 1u) != 0u) ? 0x3d0b1fbau : 0xe9abdbdau); // 47 add + r6 = r6 * r4; // 48 mul + r5 = r6 * r1 + r5; // 49 mad + r4 = r4 + r1 + ((((sel >> 31u) & 1u) != 0u) ? 0x16f36013u : 0x50fe456fu); // 50 add + r3 = r3 | r4; // 51 or + r3 = r3 ^ r7; // 52 xor + r5 = rotr_var(r5, r3); // 53 rotr + r5 = r5 * r2; // 54 mul + r1 = r1 + r6 + ((((sel >> 7u) & 1u) != 0u) ? 0x829ed387u : 0xd7d21551u); // 55 add + r3 = r3 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x44ee9a98u : 0x52871989u); // 56 add + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 57 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 58 load + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 2u); r2 = r2 ^ t_; } // 59 shfl + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 load + r1 = r1 ^ ds[r4 & mask]; // 61 load + r0 = r0 * r4; // 62 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 1u); r1 = r1 ^ t_; } // 63 shfl + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/mixB-4/kernel.cu b/proto-cuda/packs-readwidth/mixB-4/kernel.cu new file mode 100644 index 000000000..7a62bc982 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-4/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/4". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x9f289c8cu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x0b002fccu; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x0b002fccu; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x7b25b8ffu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0x7b25b8ffu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x1b37f538u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x1b37f538u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x4f9ab04fu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0x4f9ab04fu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x1120ee60u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0x1120ee60u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xfa16504cu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0xfa16504cu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xcf94f6e9u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0xcf94f6e9u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x9f289c8cu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r4 = __umulhi(r4, r6); // 0 mulhi + r1 = r1 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x69326459u : 0xf82e323du); // 1 add + r0 = r0 * r7; // 2 mul + r7 = r7 ^ ds[r4 & mask]; // 3 load + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 4 load + r7 = rotl_imm(r7, 13u); // 5 rotl + r4 = r4 | r1; // 6 or + r1 = r1 ^ r2; // 7 xor + r5 = r5 * r1; // 8 mul + r5 = __umulhi(r5, r0); // 9 mulhi + r0 = __umulhi(r0, r4); // 10 mulhi + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 load + r5 = __umulhi(r5, r3); // 12 mulhi + r0 = r0 ^ ds[r5 & mask]; // 13 load + r1 = __umulhi(r1, r5); // 14 mulhi + r0 = __umulhi(r0, r6); // 15 mulhi + { uint32_t b_ = (r1 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 16 load + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 17 load + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 18 load + r1 = r1 | r7; // 19 or + r6 = rotr_var(r6, r7); // 20 rotr + r1 = r1 + r7 + ((((sel >> 10u) & 1u) != 0u) ? 0xa65a329cu : 0x8e775c1eu); // 21 add + r7 = __umulhi(r7, r3); // 22 mulhi + r2 = __umulhi(r2, r0); // 23 mulhi + r7 = r7 ^ ds[r1 & mask]; // 24 load + r3 = r3 ^ r6; // 25 xor + r0 = rotl_imm(r0, 14u); // 26 rotl + r6 = r6 ^ ds[r5 & mask]; // 27 load + r3 = r3 ^ r6; // 28 xor + r2 = r2 * r4; // 29 mul + r7 = r7 + r6 + ((((sel >> 30u) & 1u) != 0u) ? 0x1027531du : 0xfcc52962u); // 30 add + r0 = r0 - r3; // 31 sub + r5 = r5 * r6; // 32 mul + r0 = r0 ^ ds[r7 & mask]; // 33 load + r4 = r4 + r0 + ((((sel >> 29u) & 1u) != 0u) ? 0x42903198u : 0xb7efd271u); // 34 add + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 16); // 35 shfl + r7 = r1 * r2 + r7; // 36 mad + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 load + r6 = rotr_var(r6, r7); // 38 rotr + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 39 load + r4 = r2 * r5 + r4; // 40 mad + r3 = r3 ^ r0; // 41 xor + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 42 shfl + r7 = r7 + r4 + ((((sel >> 31u) & 1u) != 0u) ? 0xbfb91116u : 0x1a6db0eau); // 43 add + r1 = r3 * r1 + r1; // 44 mad + r1 = r1 ^ r7; // 45 xor + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r3, 16); // 46 shfl + r7 = r7 + r3 + ((((sel >> 28u) & 1u) != 0u) ? 0x3d0b1fbau : 0xe9abdbdau); // 47 add + r6 = r6 * r4; // 48 mul + r5 = r6 * r1 + r5; // 49 mad + r4 = r4 + r1 + ((((sel >> 31u) & 1u) != 0u) ? 0x16f36013u : 0x50fe456fu); // 50 add + r3 = r3 | r4; // 51 or + r3 = r3 ^ r7; // 52 xor + r5 = rotr_var(r5, r3); // 53 rotr + r5 = r5 * r2; // 54 mul + r1 = r1 + r6 + ((((sel >> 7u) & 1u) != 0u) ? 0x829ed387u : 0xd7d21551u); // 55 add + r3 = r3 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x44ee9a98u : 0x52871989u); // 56 add + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 57 load + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 58 load + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 2); // 59 shfl + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 load + r1 = r1 ^ ds[r4 & mask]; // 61 load + r0 = r0 * r4; // 62 mul + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r0, 1); // 63 shfl + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixB-4/kernel_bound.cl b/proto-cuda/packs-readwidth/mixB-4/kernel_bound.cl new file mode 100644 index 000000000..06a748975 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-4/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/4". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x9f289c8cu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x0b002fccu; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x0b002fccu; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x7b25b8ffu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x7b25b8ffu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x1b37f538u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x1b37f538u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x4f9ab04fu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0x4f9ab04fu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x1120ee60u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x1120ee60u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xfa16504cu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0xfa16504cu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xcf94f6e9u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0xcf94f6e9u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x9f289c8cu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r4 = mul_hi(r4, r6); // 0 mulhi + r1 = r1 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x69326459u : 0xf82e323du); // 1 add + r0 = r0 * r7; // 2 mul + r7 = r7 ^ ds[r4 & mask]; // 3 load + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 4 load + r7 = rotl_imm(r7, 13u); // 5 rotl + r4 = r4 | r1; // 6 or + r1 = r1 ^ r2; // 7 xor + r5 = r5 * r1; // 8 mul + r5 = mul_hi(r5, r0); // 9 mulhi + r0 = mul_hi(r0, r4); // 10 mulhi + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 load + r5 = mul_hi(r5, r3); // 12 mulhi + r0 = r0 ^ ds[r5 & mask]; // 13 load + r1 = mul_hi(r1, r5); // 14 mulhi + r0 = mul_hi(r0, r6); // 15 mulhi + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 16 load + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 17 load + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 18 load + r1 = r1 | r7; // 19 or + r6 = rotr_var(r6, r7); // 20 rotr + r1 = r1 + r7 + ((((sel >> 10u) & 1u) != 0u) ? 0xa65a329cu : 0x8e775c1eu); // 21 add + r7 = mul_hi(r7, r3); // 22 mulhi + r2 = mul_hi(r2, r0); // 23 mulhi + r7 = r7 ^ ds[r1 & mask]; // 24 load + r3 = r3 ^ r6; // 25 xor + r0 = rotl_imm(r0, 14u); // 26 rotl + r6 = r6 ^ ds[r5 & mask]; // 27 load + r3 = r3 ^ r6; // 28 xor + r2 = r2 * r4; // 29 mul + r7 = r7 + r6 + ((((sel >> 30u) & 1u) != 0u) ? 0x1027531du : 0xfcc52962u); // 30 add + r0 = r0 - r3; // 31 sub + r5 = r5 * r6; // 32 mul + r0 = r0 ^ ds[r7 & mask]; // 33 load + r4 = r4 + r0 + ((((sel >> 29u) & 1u) != 0u) ? 0x42903198u : 0xb7efd271u); // 34 add + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 16u); r2 = r2 ^ t_; } // 35 shfl + r7 = r1 * r2 + r7; // 36 mad + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 load + r6 = rotr_var(r6, r7); // 38 rotr + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 39 load + r4 = r2 * r5 + r4; // 40 mad + r3 = r3 ^ r0; // 41 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r5 = r5 ^ t_; } // 42 shfl + r7 = r7 + r4 + ((((sel >> 31u) & 1u) != 0u) ? 0xbfb91116u : 0x1a6db0eau); // 43 add + r1 = r3 * r1 + r1; // 44 mad + r1 = r1 ^ r7; // 45 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 16u); r1 = r1 ^ t_; } // 46 shfl + r7 = r7 + r3 + ((((sel >> 28u) & 1u) != 0u) ? 0x3d0b1fbau : 0xe9abdbdau); // 47 add + r6 = r6 * r4; // 48 mul + r5 = r6 * r1 + r5; // 49 mad + r4 = r4 + r1 + ((((sel >> 31u) & 1u) != 0u) ? 0x16f36013u : 0x50fe456fu); // 50 add + r3 = r3 | r4; // 51 or + r3 = r3 ^ r7; // 52 xor + r5 = rotr_var(r5, r3); // 53 rotr + r5 = r5 * r2; // 54 mul + r1 = r1 + r6 + ((((sel >> 7u) & 1u) != 0u) ? 0x829ed387u : 0xd7d21551u); // 55 add + r3 = r3 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x44ee9a98u : 0x52871989u); // 56 add + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 57 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 58 load + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 2u); r2 = r2 ^ t_; } // 59 shfl + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 load + r1 = r1 ^ ds[r4 & mask]; // 61 load + r0 = r0 * r4; // 62 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 1u); r1 = r1 ^ t_; } // 63 shfl + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r4 = mul_hi(r4, r6); // 0 mulhi + r1 = r1 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x69326459u : 0xf82e323du); // 1 add + r0 = r0 * r7; // 2 mul + r7 = r7 ^ ds[r4 & mask]; // 3 load + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 4 load + r7 = rotl_imm(r7, 13u); // 5 rotl + r4 = r4 | r1; // 6 or + r1 = r1 ^ r2; // 7 xor + r5 = r5 * r1; // 8 mul + r5 = mul_hi(r5, r0); // 9 mulhi + r0 = mul_hi(r0, r4); // 10 mulhi + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 load + r5 = mul_hi(r5, r3); // 12 mulhi + r0 = r0 ^ ds[r5 & mask]; // 13 load + r1 = mul_hi(r1, r5); // 14 mulhi + r0 = mul_hi(r0, r6); // 15 mulhi + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 16 load + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 17 load + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 18 load + r1 = r1 | r7; // 19 or + r6 = rotr_var(r6, r7); // 20 rotr + r1 = r1 + r7 + ((((sel >> 10u) & 1u) != 0u) ? 0xa65a329cu : 0x8e775c1eu); // 21 add + r7 = mul_hi(r7, r3); // 22 mulhi + r2 = mul_hi(r2, r0); // 23 mulhi + r7 = r7 ^ ds[r1 & mask]; // 24 load + r3 = r3 ^ r6; // 25 xor + r0 = rotl_imm(r0, 14u); // 26 rotl + r6 = r6 ^ ds[r5 & mask]; // 27 load + r3 = r3 ^ r6; // 28 xor + r2 = r2 * r4; // 29 mul + r7 = r7 + r6 + ((((sel >> 30u) & 1u) != 0u) ? 0x1027531du : 0xfcc52962u); // 30 add + r0 = r0 - r3; // 31 sub + r5 = r5 * r6; // 32 mul + r0 = r0 ^ ds[r7 & mask]; // 33 load + r4 = r4 + r0 + ((((sel >> 29u) & 1u) != 0u) ? 0x42903198u : 0xb7efd271u); // 34 add + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 16u); r2 = r2 ^ t_; } // 35 shfl + r7 = r1 * r2 + r7; // 36 mad + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 load + r6 = rotr_var(r6, r7); // 38 rotr + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 39 load + r4 = r2 * r5 + r4; // 40 mad + r3 = r3 ^ r0; // 41 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r5 = r5 ^ t_; } // 42 shfl + r7 = r7 + r4 + ((((sel >> 31u) & 1u) != 0u) ? 0xbfb91116u : 0x1a6db0eau); // 43 add + r1 = r3 * r1 + r1; // 44 mad + r1 = r1 ^ r7; // 45 xor + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 16u); r1 = r1 ^ t_; } // 46 shfl + r7 = r7 + r3 + ((((sel >> 28u) & 1u) != 0u) ? 0x3d0b1fbau : 0xe9abdbdau); // 47 add + r6 = r6 * r4; // 48 mul + r5 = r6 * r1 + r5; // 49 mad + r4 = r4 + r1 + ((((sel >> 31u) & 1u) != 0u) ? 0x16f36013u : 0x50fe456fu); // 50 add + r3 = r3 | r4; // 51 or + r3 = r3 ^ r7; // 52 xor + r5 = rotr_var(r5, r3); // 53 rotr + r5 = r5 * r2; // 54 mul + r1 = r1 + r6 + ((((sel >> 7u) & 1u) != 0u) ? 0x829ed387u : 0xd7d21551u); // 55 add + r3 = r3 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x44ee9a98u : 0x52871989u); // 56 add + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 57 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 58 load + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 2u); r2 = r2 ^ t_; } // 59 shfl + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 load + r1 = r1 ^ ds[r4 & mask]; // 61 load + r0 = r0 * r4; // 62 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 1u); r1 = r1 ^ t_; } // 63 shfl + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-4/kernel_bound.cu b/proto-cuda/packs-readwidth/mixB-4/kernel_bound.cu new file mode 100644 index 000000000..3245e2619 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-4/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/4". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r4 = __umulhi(r4, r6); // 0 mulhi + r1 = r1 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x69326459u : 0xf82e323du); // 1 add + r0 = r0 * r7; // 2 mul + r7 = r7 ^ ds[r4 & mask]; // 3 load + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 4 load + r7 = rotl_imm(r7, 13u); // 5 rotl + r4 = r4 | r1; // 6 or + r1 = r1 ^ r2; // 7 xor + r5 = r5 * r1; // 8 mul + r5 = __umulhi(r5, r0); // 9 mulhi + r0 = __umulhi(r0, r4); // 10 mulhi + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 load + r5 = __umulhi(r5, r3); // 12 mulhi + r0 = r0 ^ ds[r5 & mask]; // 13 load + r1 = __umulhi(r1, r5); // 14 mulhi + r0 = __umulhi(r0, r6); // 15 mulhi + { uint32_t b_ = (r1 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 16 load + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 17 load + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 18 load + r1 = r1 | r7; // 19 or + r6 = rotr_var(r6, r7); // 20 rotr + r1 = r1 + r7 + ((((sel >> 10u) & 1u) != 0u) ? 0xa65a329cu : 0x8e775c1eu); // 21 add + r7 = __umulhi(r7, r3); // 22 mulhi + r2 = __umulhi(r2, r0); // 23 mulhi + r7 = r7 ^ ds[r1 & mask]; // 24 load + r3 = r3 ^ r6; // 25 xor + r0 = rotl_imm(r0, 14u); // 26 rotl + r6 = r6 ^ ds[r5 & mask]; // 27 load + r3 = r3 ^ r6; // 28 xor + r2 = r2 * r4; // 29 mul + r7 = r7 + r6 + ((((sel >> 30u) & 1u) != 0u) ? 0x1027531du : 0xfcc52962u); // 30 add + r0 = r0 - r3; // 31 sub + r5 = r5 * r6; // 32 mul + r0 = r0 ^ ds[r7 & mask]; // 33 load + r4 = r4 + r0 + ((((sel >> 29u) & 1u) != 0u) ? 0x42903198u : 0xb7efd271u); // 34 add + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 16); // 35 shfl + r7 = r1 * r2 + r7; // 36 mad + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 load + r6 = rotr_var(r6, r7); // 38 rotr + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 39 load + r4 = r2 * r5 + r4; // 40 mad + r3 = r3 ^ r0; // 41 xor + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 42 shfl + r7 = r7 + r4 + ((((sel >> 31u) & 1u) != 0u) ? 0xbfb91116u : 0x1a6db0eau); // 43 add + r1 = r3 * r1 + r1; // 44 mad + r1 = r1 ^ r7; // 45 xor + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r3, 16); // 46 shfl + r7 = r7 + r3 + ((((sel >> 28u) & 1u) != 0u) ? 0x3d0b1fbau : 0xe9abdbdau); // 47 add + r6 = r6 * r4; // 48 mul + r5 = r6 * r1 + r5; // 49 mad + r4 = r4 + r1 + ((((sel >> 31u) & 1u) != 0u) ? 0x16f36013u : 0x50fe456fu); // 50 add + r3 = r3 | r4; // 51 or + r3 = r3 ^ r7; // 52 xor + r5 = rotr_var(r5, r3); // 53 rotr + r5 = r5 * r2; // 54 mul + r1 = r1 + r6 + ((((sel >> 7u) & 1u) != 0u) ? 0x829ed387u : 0xd7d21551u); // 55 add + r3 = r3 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x44ee9a98u : 0x52871989u); // 56 add + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 57 load + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 58 load + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 2); // 59 shfl + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 load + r1 = r1 ^ ds[r4 & mask]; // 61 load + r0 = r0 * r4; // 62 mul + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r0, 1); // 63 shfl + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixB-4/memhard.h b/proto-cuda/packs-readwidth/mixB-4/memhard.h new file mode 100644 index 000000000..499a479a4 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-4/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/4". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/mixB-4/memhard.metal b/proto-cuda/packs-readwidth/mixB-4/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-4/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/mixB-4/program.h b/proto-cuda/packs-readwidth/mixB-4/program.h new file mode 100644 index 000000000..e5f36e034 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-4/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/4". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-readwidth/B/4" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d7265616477696474682f422f34" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x0e24fd2065c7407eull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=9 mulhi=8 mul=7 xor=6 shfl=5 mad=4 or=3 rotr=3 rotl=2 sub=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "mix25-50-25" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 25, 50, 25 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 6, 5, 5 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 3392 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x9f289c8cu, 0x0b002fccu, 0x7b25b8ffu, 0x1b37f538u, 0x4f9ab04fu, 0x1120ee60u, 0xfa16504cu, 0xcf94f6e9u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/mixB-4/program.json b/proto-cuda/packs-readwidth/mixB-4/program.json new file mode 100644 index 000000000..f93a6a9f9 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-4/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x0e24fd2065c7407e", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-readwidth/B/4", + "seed_bytes": "69676e65756d2d7265616477696474682f422f34", + "seed_words": ["0x9f289c8c", "0x0b002fcc", "0x7b25b8ff", "0x1b37f538", "0x4f9ab04f", "0x1120ee60", "0xfa16504c", "0xcf94f6e9"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "mix25-50-25", + "load_slots": 16, + "load_mix_percent_4_16_64": [25, 50, 25], + "load_width_counts_4_16_64": [6, 5, 5], + "bytes_per_hash": 3392, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "add": 9, "mulhi": 8, "mul": 7, "xor": 6, "shfl": 5, "mad": 4, "or": 3, "rotr": 3, "rotl": 2, "sub": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mulhi", "dst": 4, "src": 6, "src2": 2, "imm": "0xc8d13b7c", "imm2": "0x67f31d11", "rot": 28, "bit": 4, "mask": 1, "width": 1}, + {"i": 1, "op": "add", "dst": 1, "src": 2, "src2": 4, "imm": "0xf82e323d", "imm2": "0x69326459", "rot": 14, "bit": 7, "mask": 4, "width": 1}, + {"i": 2, "op": "mul", "dst": 0, "src": 7, "src2": 1, "imm": "0x64bd3191", "imm2": "0x8ee2d349", "rot": 22, "bit": 19, "mask": 2, "width": 1}, + {"i": 3, "op": "load", "dst": 7, "src": 4, "src2": 4, "imm": "0x286232e4", "imm2": "0x7ae714ac", "rot": 16, "bit": 31, "mask": 8, "width": 1}, + {"i": 4, "op": "load", "dst": 4, "src": 7, "src2": 0, "imm": "0x418a79c7", "imm2": "0x12f7dcdf", "rot": 3, "bit": 13, "mask": 2, "width": 16}, + {"i": 5, "op": "rotl", "dst": 7, "src": 3, "src2": 2, "imm": "0x56cae3ac", "imm2": "0x773171bb", "rot": 13, "bit": 29, "mask": 1, "width": 1}, + {"i": 6, "op": "or", "dst": 4, "src": 1, "src2": 7, "imm": "0x2bec3323", "imm2": "0x5f7e3dfd", "rot": 28, "bit": 29, "mask": 1, "width": 1}, + {"i": 7, "op": "xor", "dst": 1, "src": 2, "src2": 1, "imm": "0x65342190", "imm2": "0xf10b5286", "rot": 19, "bit": 16, "mask": 1, "width": 1}, + {"i": 8, "op": "mul", "dst": 5, "src": 1, "src2": 2, "imm": "0x0162462a", "imm2": "0xfc8e9470", "rot": 27, "bit": 10, "mask": 8, "width": 1}, + {"i": 9, "op": "mulhi", "dst": 5, "src": 0, "src2": 5, "imm": "0x8635054b", "imm2": "0x0ecc646f", "rot": 12, "bit": 3, "mask": 16, "width": 1}, + {"i": 10, "op": "mulhi", "dst": 0, "src": 4, "src2": 3, "imm": "0x49e93064", "imm2": "0x8f584303", "rot": 3, "bit": 14, "mask": 1, "width": 1}, + {"i": 11, "op": "load", "dst": 3, "src": 5, "src2": 1, "imm": "0x0ca1992a", "imm2": "0x190c01d4", "rot": 10, "bit": 12, "mask": 16, "width": 4}, + {"i": 12, "op": "mulhi", "dst": 5, "src": 3, "src2": 3, "imm": "0x9c684edd", "imm2": "0x329e5370", "rot": 2, "bit": 22, "mask": 8, "width": 1}, + {"i": 13, "op": "load", "dst": 0, "src": 5, "src2": 4, "imm": "0x406e4e11", "imm2": "0x8bec4797", "rot": 18, "bit": 12, "mask": 1, "width": 1}, + {"i": 14, "op": "mulhi", "dst": 1, "src": 5, "src2": 5, "imm": "0x841636d5", "imm2": "0xdb69e272", "rot": 6, "bit": 24, "mask": 2, "width": 1}, + {"i": 15, "op": "mulhi", "dst": 0, "src": 6, "src2": 6, "imm": "0xf606eb08", "imm2": "0x07d1cc36", "rot": 23, "bit": 10, "mask": 1, "width": 1}, + {"i": 16, "op": "load", "dst": 5, "src": 1, "src2": 7, "imm": "0x90f8f480", "imm2": "0x06bf7942", "rot": 14, "bit": 1, "mask": 8, "width": 4}, + {"i": 17, "op": "load", "dst": 5, "src": 0, "src2": 4, "imm": "0x2799861a", "imm2": "0x78b2bbab", "rot": 9, "bit": 19, "mask": 8, "width": 16}, + {"i": 18, "op": "load", "dst": 7, "src": 3, "src2": 5, "imm": "0xc5abfe75", "imm2": "0xd0261b62", "rot": 28, "bit": 31, "mask": 16, "width": 16}, + {"i": 19, "op": "or", "dst": 1, "src": 7, "src2": 0, "imm": "0x86b04a7e", "imm2": "0xe0cc1eff", "rot": 20, "bit": 29, "mask": 2, "width": 1}, + {"i": 20, "op": "rotr", "dst": 6, "src": 7, "src2": 6, "imm": "0x4ce4fb67", "imm2": "0x2a149c0e", "rot": 11, "bit": 5, "mask": 4, "width": 1}, + {"i": 21, "op": "add", "dst": 1, "src": 7, "src2": 2, "imm": "0x8e775c1e", "imm2": "0xa65a329c", "rot": 12, "bit": 10, "mask": 1, "width": 1}, + {"i": 22, "op": "mulhi", "dst": 7, "src": 3, "src2": 3, "imm": "0x4eab7705", "imm2": "0xb367431a", "rot": 12, "bit": 28, "mask": 8, "width": 1}, + {"i": 23, "op": "mulhi", "dst": 2, "src": 0, "src2": 7, "imm": "0x37cb29dd", "imm2": "0x52bb3023", "rot": 24, "bit": 4, "mask": 1, "width": 1}, + {"i": 24, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x7e09f54d", "imm2": "0xb1379d43", "rot": 10, "bit": 18, "mask": 1, "width": 1}, + {"i": 25, "op": "xor", "dst": 3, "src": 6, "src2": 2, "imm": "0xbd9ff25d", "imm2": "0xc7cdff19", "rot": 22, "bit": 21, "mask": 16, "width": 1}, + {"i": 26, "op": "rotl", "dst": 0, "src": 5, "src2": 5, "imm": "0xbafba43b", "imm2": "0x4cb28b03", "rot": 14, "bit": 24, "mask": 2, "width": 1}, + {"i": 27, "op": "load", "dst": 6, "src": 5, "src2": 6, "imm": "0xee205237", "imm2": "0x7c0ad7be", "rot": 6, "bit": 4, "mask": 4, "width": 1}, + {"i": 28, "op": "xor", "dst": 3, "src": 6, "src2": 6, "imm": "0xce371837", "imm2": "0x2df2fcc2", "rot": 19, "bit": 20, "mask": 16, "width": 1}, + {"i": 29, "op": "mul", "dst": 2, "src": 4, "src2": 7, "imm": "0x21540a5a", "imm2": "0xd48ad1a7", "rot": 3, "bit": 5, "mask": 2, "width": 1}, + {"i": 30, "op": "add", "dst": 7, "src": 6, "src2": 0, "imm": "0xfcc52962", "imm2": "0x1027531d", "rot": 25, "bit": 30, "mask": 1, "width": 1}, + {"i": 31, "op": "sub", "dst": 0, "src": 3, "src2": 1, "imm": "0xe84b5faa", "imm2": "0x57b97bb0", "rot": 21, "bit": 1, "mask": 16, "width": 1}, + {"i": 32, "op": "mul", "dst": 5, "src": 6, "src2": 4, "imm": "0x6a127e78", "imm2": "0x96a0438d", "rot": 30, "bit": 30, "mask": 2, "width": 1}, + {"i": 33, "op": "load", "dst": 0, "src": 7, "src2": 7, "imm": "0xe4d0ff0d", "imm2": "0x8060f0bc", "rot": 4, "bit": 29, "mask": 16, "width": 1}, + {"i": 34, "op": "add", "dst": 4, "src": 0, "src2": 7, "imm": "0xb7efd271", "imm2": "0x42903198", "rot": 31, "bit": 29, "mask": 4, "width": 1}, + {"i": 35, "op": "shfl", "dst": 2, "src": 0, "src2": 6, "imm": "0x935b7779", "imm2": "0x923f6275", "rot": 3, "bit": 10, "mask": 16, "width": 1}, + {"i": 36, "op": "mad", "dst": 7, "src": 1, "src2": 2, "imm": "0x4f147f69", "imm2": "0xa3b64762", "rot": 15, "bit": 12, "mask": 16, "width": 1}, + {"i": 37, "op": "load", "dst": 4, "src": 5, "src2": 3, "imm": "0xaecb1a67", "imm2": "0x6fd6e94d", "rot": 18, "bit": 8, "mask": 1, "width": 16}, + {"i": 38, "op": "rotr", "dst": 6, "src": 7, "src2": 2, "imm": "0x78f67e50", "imm2": "0x94d71142", "rot": 18, "bit": 9, "mask": 4, "width": 1}, + {"i": 39, "op": "load", "dst": 0, "src": 7, "src2": 2, "imm": "0x881803d2", "imm2": "0x3c58d2ef", "rot": 26, "bit": 25, "mask": 16, "width": 4}, + {"i": 40, "op": "mad", "dst": 4, "src": 2, "src2": 5, "imm": "0x6ec96693", "imm2": "0x734eb929", "rot": 27, "bit": 9, "mask": 8, "width": 1}, + {"i": 41, "op": "xor", "dst": 3, "src": 0, "src2": 0, "imm": "0xfc0a4ea0", "imm2": "0xb693b2f9", "rot": 1, "bit": 23, "mask": 1, "width": 1}, + {"i": 42, "op": "shfl", "dst": 5, "src": 3, "src2": 1, "imm": "0x14a3151f", "imm2": "0x698db5f9", "rot": 16, "bit": 7, "mask": 4, "width": 1}, + {"i": 43, "op": "add", "dst": 7, "src": 4, "src2": 5, "imm": "0x1a6db0ea", "imm2": "0xbfb91116", "rot": 3, "bit": 31, "mask": 4, "width": 1}, + {"i": 44, "op": "mad", "dst": 1, "src": 3, "src2": 1, "imm": "0x1fe40044", "imm2": "0xb2e1fb95", "rot": 12, "bit": 27, "mask": 4, "width": 1}, + {"i": 45, "op": "xor", "dst": 1, "src": 7, "src2": 6, "imm": "0x8416c247", "imm2": "0x78cb20c7", "rot": 29, "bit": 14, "mask": 2, "width": 1}, + {"i": 46, "op": "shfl", "dst": 1, "src": 3, "src2": 2, "imm": "0x1f3d42cb", "imm2": "0x1a52df69", "rot": 15, "bit": 29, "mask": 16, "width": 1}, + {"i": 47, "op": "add", "dst": 7, "src": 3, "src2": 7, "imm": "0xe9abdbda", "imm2": "0x3d0b1fba", "rot": 24, "bit": 28, "mask": 8, "width": 1}, + {"i": 48, "op": "mul", "dst": 6, "src": 4, "src2": 6, "imm": "0x0d5cd4f5", "imm2": "0xedc72d13", "rot": 3, "bit": 18, "mask": 8, "width": 1}, + {"i": 49, "op": "mad", "dst": 5, "src": 6, "src2": 1, "imm": "0x649f46b1", "imm2": "0x7d8bbb8e", "rot": 5, "bit": 17, "mask": 16, "width": 1}, + {"i": 50, "op": "add", "dst": 4, "src": 1, "src2": 7, "imm": "0x50fe456f", "imm2": "0x16f36013", "rot": 26, "bit": 31, "mask": 2, "width": 1}, + {"i": 51, "op": "or", "dst": 3, "src": 4, "src2": 6, "imm": "0xdaf27536", "imm2": "0xec4501d0", "rot": 31, "bit": 2, "mask": 4, "width": 1}, + {"i": 52, "op": "xor", "dst": 3, "src": 7, "src2": 5, "imm": "0xb867c9ba", "imm2": "0x714db26a", "rot": 23, "bit": 11, "mask": 4, "width": 1}, + {"i": 53, "op": "rotr", "dst": 5, "src": 3, "src2": 4, "imm": "0xd72c153c", "imm2": "0x278a713a", "rot": 19, "bit": 25, "mask": 1, "width": 1}, + {"i": 54, "op": "mul", "dst": 5, "src": 2, "src2": 4, "imm": "0xc4d09a76", "imm2": "0xfb7f9671", "rot": 21, "bit": 17, "mask": 8, "width": 1}, + {"i": 55, "op": "add", "dst": 1, "src": 6, "src2": 6, "imm": "0xd7d21551", "imm2": "0x829ed387", "rot": 28, "bit": 7, "mask": 16, "width": 1}, + {"i": 56, "op": "add", "dst": 3, "src": 2, "src2": 2, "imm": "0x52871989", "imm2": "0x44ee9a98", "rot": 31, "bit": 7, "mask": 4, "width": 1}, + {"i": 57, "op": "load", "dst": 4, "src": 5, "src2": 6, "imm": "0x04312083", "imm2": "0x2afd5fd2", "rot": 4, "bit": 26, "mask": 2, "width": 4}, + {"i": 58, "op": "load", "dst": 0, "src": 7, "src2": 7, "imm": "0x40dd5442", "imm2": "0x43d19eaf", "rot": 20, "bit": 10, "mask": 2, "width": 4}, + {"i": 59, "op": "shfl", "dst": 2, "src": 0, "src2": 4, "imm": "0x657f86e0", "imm2": "0x832515b8", "rot": 31, "bit": 23, "mask": 2, "width": 1}, + {"i": 60, "op": "load", "dst": 6, "src": 3, "src2": 2, "imm": "0x25beb616", "imm2": "0x29aa75a1", "rot": 14, "bit": 6, "mask": 4, "width": 16}, + {"i": 61, "op": "load", "dst": 1, "src": 4, "src2": 7, "imm": "0x78340ea4", "imm2": "0x7e6a6323", "rot": 4, "bit": 16, "mask": 1, "width": 1}, + {"i": 62, "op": "mul", "dst": 0, "src": 4, "src2": 4, "imm": "0xce786bf8", "imm2": "0x28f94fc6", "rot": 1, "bit": 20, "mask": 16, "width": 1}, + {"i": 63, "op": "shfl", "dst": 1, "src": 0, "src2": 7, "imm": "0x338e1b8f", "imm2": "0x3020bfc3", "rot": 8, "bit": 17, "mask": 1, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/mixB-4/program.metal b/proto-cuda/packs-readwidth/mixB-4/program.metal new file mode 100644 index 000000000..a5f8c8e8a --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-4/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x9f289c8cu, 0x0b002fccu, 0x7b25b8ffu, 0x1b37f538u, 0x4f9ab04fu, 0x1120ee60u, 0xfa16504cu, 0xcf94f6e9u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r4 = mulhi(r4, r6); // 0 + r1 = r1 + r2 + select(0xf82e323du, 0x69326459u, ((sel >> 7u) & 1u) != 0u); // 1 + r0 = r0 * r7; // 2 + r7 = r7 ^ dataset[r4 & MASK]; // 3 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 4 + r7 = rotl_imm(r7, 13u); // 5 + r4 = r4 | r1; // 6 + r1 = r1 ^ r2; // 7 + r5 = r5 * r1; // 8 + r5 = mulhi(r5, r0); // 9 + r0 = mulhi(r0, r4); // 10 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 + r5 = mulhi(r5, r3); // 12 + r0 = r0 ^ dataset[r5 & MASK]; // 13 + r1 = mulhi(r1, r5); // 14 + r0 = mulhi(r0, r6); // 15 + { uint b_ = (r1 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 16 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 17 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 18 + r1 = r1 | r7; // 19 + r6 = rotr_var(r6, r7); // 20 + r1 = r1 + r7 + select(0x8e775c1eu, 0xa65a329cu, ((sel >> 10u) & 1u) != 0u); // 21 + r7 = mulhi(r7, r3); // 22 + r2 = mulhi(r2, r0); // 23 + r7 = r7 ^ dataset[r1 & MASK]; // 24 + r3 = r3 ^ r6; // 25 + r0 = rotl_imm(r0, 14u); // 26 + r6 = r6 ^ dataset[r5 & MASK]; // 27 + r3 = r3 ^ r6; // 28 + r2 = r2 * r4; // 29 + r7 = r7 + r6 + select(0xfcc52962u, 0x1027531du, ((sel >> 30u) & 1u) != 0u); // 30 + r0 = r0 - r3; // 31 + r5 = r5 * r6; // 32 + r0 = r0 ^ dataset[r7 & MASK]; // 33 + r4 = r4 + r0 + select(0xb7efd271u, 0x42903198u, ((sel >> 29u) & 1u) != 0u); // 34 + r2 = r2 ^ simd_shuffle_xor(r0, (ushort)16); // 35 + r7 = r1 * r2 + r7; // 36 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 + r6 = rotr_var(r6, r7); // 38 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 39 + r4 = r2 * r5 + r4; // 40 + r3 = r3 ^ r0; // 41 + r5 = r5 ^ simd_shuffle_xor(r3, (ushort)4); // 42 + r7 = r7 + r4 + select(0x1a6db0eau, 0xbfb91116u, ((sel >> 31u) & 1u) != 0u); // 43 + r1 = r3 * r1 + r1; // 44 + r1 = r1 ^ r7; // 45 + r1 = r1 ^ simd_shuffle_xor(r3, (ushort)16); // 46 + r7 = r7 + r3 + select(0xe9abdbdau, 0x3d0b1fbau, ((sel >> 28u) & 1u) != 0u); // 47 + r6 = r6 * r4; // 48 + r5 = r6 * r1 + r5; // 49 + r4 = r4 + r1 + select(0x50fe456fu, 0x16f36013u, ((sel >> 31u) & 1u) != 0u); // 50 + r3 = r3 | r4; // 51 + r3 = r3 ^ r7; // 52 + r5 = rotr_var(r5, r3); // 53 + r5 = r5 * r2; // 54 + r1 = r1 + r6 + select(0xd7d21551u, 0x829ed387u, ((sel >> 7u) & 1u) != 0u); // 55 + r3 = r3 + r2 + select(0x52871989u, 0x44ee9a98u, ((sel >> 7u) & 1u) != 0u); // 56 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 57 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 58 + r2 = r2 ^ simd_shuffle_xor(r0, (ushort)2); // 59 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 + r1 = r1 ^ dataset[r4 & MASK]; // 61 + r0 = r0 * r4; // 62 + r1 = r1 ^ simd_shuffle_xor(r0, (ushort)1); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-4/program_bound.metal b/proto-cuda/packs-readwidth/mixB-4/program_bound.metal new file mode 100644 index 000000000..9513c41b5 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-4/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x9f289c8cu, 0x0b002fccu, 0x7b25b8ffu, 0x1b37f538u, 0x4f9ab04fu, 0x1120ee60u, 0xfa16504cu, 0xcf94f6e9u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r4 = mulhi(r4, r6); // 0 + r1 = r1 + r2 + select(0xf82e323du, 0x69326459u, ((sel >> 7u) & 1u) != 0u); // 1 + r0 = r0 * r7; // 2 + r7 = r7 ^ dataset[r4 & MASK]; // 3 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 4 + r7 = rotl_imm(r7, 13u); // 5 + r4 = r4 | r1; // 6 + r1 = r1 ^ r2; // 7 + r5 = r5 * r1; // 8 + r5 = mulhi(r5, r0); // 9 + r0 = mulhi(r0, r4); // 10 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 11 + r5 = mulhi(r5, r3); // 12 + r0 = r0 ^ dataset[r5 & MASK]; // 13 + r1 = mulhi(r1, r5); // 14 + r0 = mulhi(r0, r6); // 15 + { uint b_ = (r1 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r5 = x_; } // 16 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 17 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 18 + r1 = r1 | r7; // 19 + r6 = rotr_var(r6, r7); // 20 + r1 = r1 + r7 + select(0x8e775c1eu, 0xa65a329cu, ((sel >> 10u) & 1u) != 0u); // 21 + r7 = mulhi(r7, r3); // 22 + r2 = mulhi(r2, r0); // 23 + r7 = r7 ^ dataset[r1 & MASK]; // 24 + r3 = r3 ^ r6; // 25 + r0 = rotl_imm(r0, 14u); // 26 + r6 = r6 ^ dataset[r5 & MASK]; // 27 + r3 = r3 ^ r6; // 28 + r2 = r2 * r4; // 29 + r7 = r7 + r6 + select(0xfcc52962u, 0x1027531du, ((sel >> 30u) & 1u) != 0u); // 30 + r0 = r0 - r3; // 31 + r5 = r5 * r6; // 32 + r0 = r0 ^ dataset[r7 & MASK]; // 33 + r4 = r4 + r0 + select(0xb7efd271u, 0x42903198u, ((sel >> 29u) & 1u) != 0u); // 34 + r2 = r2 ^ simd_shuffle_xor(r0, (ushort)16); // 35 + r7 = r1 * r2 + r7; // 36 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 + r6 = rotr_var(r6, r7); // 38 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 39 + r4 = r2 * r5 + r4; // 40 + r3 = r3 ^ r0; // 41 + r5 = r5 ^ simd_shuffle_xor(r3, (ushort)4); // 42 + r7 = r7 + r4 + select(0x1a6db0eau, 0xbfb91116u, ((sel >> 31u) & 1u) != 0u); // 43 + r1 = r3 * r1 + r1; // 44 + r1 = r1 ^ r7; // 45 + r1 = r1 ^ simd_shuffle_xor(r3, (ushort)16); // 46 + r7 = r7 + r3 + select(0xe9abdbdau, 0x3d0b1fbau, ((sel >> 28u) & 1u) != 0u); // 47 + r6 = r6 * r4; // 48 + r5 = r6 * r1 + r5; // 49 + r4 = r4 + r1 + select(0x50fe456fu, 0x16f36013u, ((sel >> 31u) & 1u) != 0u); // 50 + r3 = r3 | r4; // 51 + r3 = r3 ^ r7; // 52 + r5 = rotr_var(r5, r3); // 53 + r5 = r5 * r2; // 54 + r1 = r1 + r6 + select(0xd7d21551u, 0x829ed387u, ((sel >> 7u) & 1u) != 0u); // 55 + r3 = r3 + r2 + select(0x52871989u, 0x44ee9a98u, ((sel >> 7u) & 1u) != 0u); // 56 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 57 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 58 + r2 = r2 ^ simd_shuffle_xor(r0, (ushort)2); // 59 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 60 + r1 = r1 ^ dataset[r4 & MASK]; // 61 + r0 = r0 * r4; // 62 + r1 = r1 ^ simd_shuffle_xor(r0, (ushort)1); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-4/vectors.h b/proto-cuda/packs-readwidth/mixB-4/vectors.h new file mode 100644 index 000000000..a46c9ff81 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-4/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/4". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0xccbf13523eb98e1cull, 0xde00c7742d2be1bdull, 0x1d8e9bec867f27cfull, 0x87c642bfc4555964ull, 0xebd3b684bd25e326ull, 0x26e8d44084dd46c4ull, 0xebc22fc70706e7cdull, 0x710511a2650bf010ull, + 0x3af8b4cc45f29169ull, 0x80894f97c7cd4df6ull, 0xa4af33c8527e7764ull, 0x5f1fbfb7fd9564cbull, 0x2c773cb706a1a42bull, 0x98162b9127313ca2ull, 0x4dfe411a5e479e8bull, 0xae8f817240359de1ull, + 0x886278be11111cd1ull, 0x7c863987232717e8ull, 0xb03012760076a479ull, 0x1e3e89fe818ecdd8ull, 0x3dea83b2d3110fb9ull, 0xa795058b55327e1cull, 0xa2f14aa0db5ac196ull, 0xd71d13d14e002d4cull, + 0x38839cf729cb4cdfull, 0x0dc77b3b788e8b2bull, 0xca52aec855a52bc7ull, 0x806f41b443370c67ull, 0x319bdd6da3b77edbull, 0x491a6b0cb0596827ull, 0x6af898d470e784faull, 0xb69a05cefd5c2354ull + }, + { // base nonce 4096 + 0xbe6b33d9002e7a42ull, 0xc8ade69410f7235eull, 0x1636381d403e413eull, 0xea26edc627bf86c0ull, 0x1c1c8babbcaf751aull, 0xc02f878509a1b636ull, 0xacef8546587013ecull, 0xf64e85d17acce074ull, + 0x18476ed6e485f0baull, 0xf4007c7c9b551c11ull, 0x183620eb05cd7e88ull, 0xe0e667a7da7537a9ull, 0x3b4a261784591e9full, 0x30592938cc61c375ull, 0xba67cff643ea4de0ull, 0x3000c4a026b1826bull, + 0x645e4ec8aeaac7acull, 0x7780ab7eb4c6bf32ull, 0xcf0b591e5ac0856dull, 0x84f09b5464fe996cull, 0x936a3d3f7108b502ull, 0x174c405a72705f38ull, 0x7763e1873df980e3ull, 0x4e20703bc3332e49ull, + 0x9a15e8e23198c6faull, 0x966a4e0d516aa4ecull, 0xbd2c22a47b37241cull, 0x0ef2de2005c7f6c4ull, 0x21f02f3ecc40ce28ull, 0x79d1edb3b550499aull, 0xa2a7385ddb3aa83dull, 0xfaf71ea406207d8bull + }, + { // base nonce 1000000 + 0x9436ad8852cc7f1full, 0x620b41289d11d943ull, 0x0e73d532ee712f8full, 0xb1d049a282b91035ull, 0x5beab7fa33aa30b9ull, 0x13a47d879b415947ull, 0x6c90273b32b21bf3ull, 0x5fa9091d559222f1ull, + 0x0e29995ccce2565bull, 0x79901ecb9830ea23ull, 0x7d054d5bfa4a4527ull, 0xda08a3d847db493dull, 0x1e4b25a9d5d97088ull, 0xa807e3bdb22db815ull, 0x885236dcd4d3ed2eull, 0x7465fa54dac84bfcull, + 0x0fc4ed75b3f5247eull, 0xae4a6167dec77d59ull, 0xed6fbd7a487e3d54ull, 0x8ba573ae293a8110ull, 0x45227aed5a5bf69eull, 0x8b972b670a4b4d2full, 0xc3a8b1f10c03de08ull, 0x4a848502e46f9af4ull, + 0x469751222c8bef10ull, 0x496a8d3d03001f77ull, 0xab94b15d15daef78ull, 0x140bd97c8bf30d21ull, 0x014c4403cf19290dull, 0xea63daba67f689f0ull, 0xda55cf1de361660eull, 0xd61daa2c678c9830ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/mixB-4/vectors.json b/proto-cuda/packs-readwidth/mixB-4/vectors.json new file mode 100644 index 000000000..69a432bbb --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-4/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-readwidth/B/4", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0xccbf13523eb98e1c", "0xde00c7742d2be1bd", "0x1d8e9bec867f27cf", "0x87c642bfc4555964", "0xebd3b684bd25e326", "0x26e8d44084dd46c4", "0xebc22fc70706e7cd", "0x710511a2650bf010", + "0x3af8b4cc45f29169", "0x80894f97c7cd4df6", "0xa4af33c8527e7764", "0x5f1fbfb7fd9564cb", "0x2c773cb706a1a42b", "0x98162b9127313ca2", "0x4dfe411a5e479e8b", "0xae8f817240359de1", + "0x886278be11111cd1", "0x7c863987232717e8", "0xb03012760076a479", "0x1e3e89fe818ecdd8", "0x3dea83b2d3110fb9", "0xa795058b55327e1c", "0xa2f14aa0db5ac196", "0xd71d13d14e002d4c", + "0x38839cf729cb4cdf", "0x0dc77b3b788e8b2b", "0xca52aec855a52bc7", "0x806f41b443370c67", "0x319bdd6da3b77edb", "0x491a6b0cb0596827", "0x6af898d470e784fa", "0xb69a05cefd5c2354" + ]}, + {"base_nonce": 4096, "expected": [ + "0xbe6b33d9002e7a42", "0xc8ade69410f7235e", "0x1636381d403e413e", "0xea26edc627bf86c0", "0x1c1c8babbcaf751a", "0xc02f878509a1b636", "0xacef8546587013ec", "0xf64e85d17acce074", + "0x18476ed6e485f0ba", "0xf4007c7c9b551c11", "0x183620eb05cd7e88", "0xe0e667a7da7537a9", "0x3b4a261784591e9f", "0x30592938cc61c375", "0xba67cff643ea4de0", "0x3000c4a026b1826b", + "0x645e4ec8aeaac7ac", "0x7780ab7eb4c6bf32", "0xcf0b591e5ac0856d", "0x84f09b5464fe996c", "0x936a3d3f7108b502", "0x174c405a72705f38", "0x7763e1873df980e3", "0x4e20703bc3332e49", + "0x9a15e8e23198c6fa", "0x966a4e0d516aa4ec", "0xbd2c22a47b37241c", "0x0ef2de2005c7f6c4", "0x21f02f3ecc40ce28", "0x79d1edb3b550499a", "0xa2a7385ddb3aa83d", "0xfaf71ea406207d8b" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x9436ad8852cc7f1f", "0x620b41289d11d943", "0x0e73d532ee712f8f", "0xb1d049a282b91035", "0x5beab7fa33aa30b9", "0x13a47d879b415947", "0x6c90273b32b21bf3", "0x5fa9091d559222f1", + "0x0e29995ccce2565b", "0x79901ecb9830ea23", "0x7d054d5bfa4a4527", "0xda08a3d847db493d", "0x1e4b25a9d5d97088", "0xa807e3bdb22db815", "0x885236dcd4d3ed2e", "0x7465fa54dac84bfc", + "0x0fc4ed75b3f5247e", "0xae4a6167dec77d59", "0xed6fbd7a487e3d54", "0x8ba573ae293a8110", "0x45227aed5a5bf69e", "0x8b972b670a4b4d2f", "0xc3a8b1f10c03de08", "0x4a848502e46f9af4", + "0x469751222c8bef10", "0x496a8d3d03001f77", "0xab94b15d15daef78", "0x140bd97c8bf30d21", "0x014c4403cf19290d", "0xea63daba67f689f0", "0xda55cf1de361660e", "0xd61daa2c678c9830" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/mixB-5/kernel.cl b/proto-cuda/packs-readwidth/mixB-5/kernel.cl new file mode 100644 index 000000000..98119d453 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-5/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/5". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0xd572f9e9u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x30b20f21u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x30b20f21u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x91d2d3fcu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x91d2d3fcu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x9395a5a1u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x9395a5a1u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef3dcd1fu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xef3dcd1fu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x8c9e416eu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x8c9e416eu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x47c6146eu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x47c6146eu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x360af18au; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x360af18au; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0xd572f9e9u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = r3 * r1; // 0 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 4u); r2 = r2 ^ t_; } // 1 shfl + r0 = r0 ^ r6; // 2 xor + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 3 load + r6 = r6 ^ ds[r0 & mask]; // 4 load + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 2u); r2 = r2 ^ t_; } // 5 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r7 = r7 ^ t_; } // 6 shfl + r5 = rotl_imm(r5, 8u); // 7 rotl + r1 = r5 * r7 + r1; // 8 mad + r0 = r0 * r3; // 9 mul + r0 = r0 * r4; // 10 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 16u); r3 = r3 ^ t_; } // 11 shfl + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 16u); r0 = r0 ^ t_; } // 13 shfl + r4 = r4 - r6; // 14 sub + r7 = r7 | r5; // 15 or + r0 = rotr_var(r0, r6); // 16 rotr + r1 = r1 - r7; // 17 sub + r3 = r3 - r6; // 18 sub + r5 = r5 - r4; // 19 sub + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 20 shfl + r1 = r1 | r2; // 21 or + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r1 = r1 ^ t_; } // 22 shfl + r5 = rotr_var(r5, r3); // 23 rotr + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 24 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 25 load + r7 = r7 | r1; // 26 or + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 load + r3 = r3 ^ r4; // 28 xor + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 29 load + r5 = r5 + r4 + ((((sel >> 15u) & 1u) != 0u) ? 0x90aaff78u : 0xd0651e01u); // 30 add + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 31 load + r0 = r0 * r4; // 32 mul + r0 = r0 + r3 + ((((sel >> 2u) & 1u) != 0u) ? 0x3d45d6cfu : 0x8ac5ddc2u); // 33 add + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 34 load + r6 = r6 ^ r2; // 35 xor + r1 = r1 ^ r6; // 36 xor + r0 = r4 * r3 + r0; // 37 mad + r7 = rotl_imm(r7, 17u); // 38 rotl + r5 = r5 + r7 + ((((sel >> 0u) & 1u) != 0u) ? 0x4b7eb232u : 0x651ff080u); // 39 add + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 40 load + r6 = mul_hi(r6, r1); // 41 mulhi + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 42 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 43 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 44 load + r3 = r3 * r2; // 45 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 2u); r6 = r6 ^ t_; } // 46 shfl + r4 = r2 * r3 + r4; // 47 mad + r4 = r4 + r2 + ((((sel >> 20u) & 1u) != 0u) ? 0x97f34047u : 0x4f0330b4u); // 48 add + r6 = r6 ^ r3; // 49 xor + r0 = r0 + r2 + ((((sel >> 17u) & 1u) != 0u) ? 0xb24ddcb5u : 0x6e73a5d8u); // 50 add + r6 = rotl_imm(r6, 22u); // 51 rotl + r4 = r4 * r0; // 52 mul + r6 = rotr_var(r6, r3); // 53 rotr + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 54 load + r3 = r3 - r1; // 55 sub + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 1u); r3 = r3 ^ t_; } // 56 shfl + r3 = r3 | r1; // 57 or + r4 = r7 * r4 + r4; // 58 mad + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 load + r4 = r6 * r3 + r4; // 60 mad + r2 = rotl_imm(r2, 30u); // 61 rotl + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 62 load + r7 = r7 + r4 + ((((sel >> 3u) & 1u) != 0u) ? 0x4bcca5cau : 0x12a17f52u); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/mixB-5/kernel.cu b/proto-cuda/packs-readwidth/mixB-5/kernel.cu new file mode 100644 index 000000000..05dd88ab5 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-5/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/5". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0xd572f9e9u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x30b20f21u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x30b20f21u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x91d2d3fcu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0x91d2d3fcu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x9395a5a1u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x9395a5a1u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef3dcd1fu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xef3dcd1fu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x8c9e416eu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0x8c9e416eu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x47c6146eu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x47c6146eu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x360af18au; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x360af18au; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0xd572f9e9u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r3 = r3 * r1; // 0 mul + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 4); // 1 shfl + r0 = r0 ^ r6; // 2 xor + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 3 load + r6 = r6 ^ ds[r0 & mask]; // 4 load + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 2); // 5 shfl + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 4); // 6 shfl + r5 = rotl_imm(r5, 8u); // 7 rotl + r1 = r5 * r7 + r1; // 8 mad + r0 = r0 * r3; // 9 mul + r0 = r0 * r4; // 10 mul + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r0, 16); // 11 shfl + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 load + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r2, 16); // 13 shfl + r4 = r4 - r6; // 14 sub + r7 = r7 | r5; // 15 or + r0 = rotr_var(r0, r6); // 16 rotr + r1 = r1 - r7; // 17 sub + r3 = r3 - r6; // 18 sub + r5 = r5 - r4; // 19 sub + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 20 shfl + r1 = r1 | r2; // 21 or + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 22 shfl + r5 = rotr_var(r5, r3); // 23 rotr + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 24 load + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 25 load + r7 = r7 | r1; // 26 or + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 load + r3 = r3 ^ r4; // 28 xor + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 29 load + r5 = r5 + r4 + ((((sel >> 15u) & 1u) != 0u) ? 0x90aaff78u : 0xd0651e01u); // 30 add + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 31 load + r0 = r0 * r4; // 32 mul + r0 = r0 + r3 + ((((sel >> 2u) & 1u) != 0u) ? 0x3d45d6cfu : 0x8ac5ddc2u); // 33 add + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 34 load + r6 = r6 ^ r2; // 35 xor + r1 = r1 ^ r6; // 36 xor + r0 = r4 * r3 + r0; // 37 mad + r7 = rotl_imm(r7, 17u); // 38 rotl + r5 = r5 + r7 + ((((sel >> 0u) & 1u) != 0u) ? 0x4b7eb232u : 0x651ff080u); // 39 add + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 40 load + r6 = __umulhi(r6, r1); // 41 mulhi + { uint32_t b_ = (r6 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 42 load + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 43 load + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 44 load + r3 = r3 * r2; // 45 mul + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 2); // 46 shfl + r4 = r2 * r3 + r4; // 47 mad + r4 = r4 + r2 + ((((sel >> 20u) & 1u) != 0u) ? 0x97f34047u : 0x4f0330b4u); // 48 add + r6 = r6 ^ r3; // 49 xor + r0 = r0 + r2 + ((((sel >> 17u) & 1u) != 0u) ? 0xb24ddcb5u : 0x6e73a5d8u); // 50 add + r6 = rotl_imm(r6, 22u); // 51 rotl + r4 = r4 * r0; // 52 mul + r6 = rotr_var(r6, r3); // 53 rotr + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 54 load + r3 = r3 - r1; // 55 sub + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 1); // 56 shfl + r3 = r3 | r1; // 57 or + r4 = r7 * r4 + r4; // 58 mad + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 load + r4 = r6 * r3 + r4; // 60 mad + r2 = rotl_imm(r2, 30u); // 61 rotl + { uint32_t b_ = (r1 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 62 load + r7 = r7 + r4 + ((((sel >> 3u) & 1u) != 0u) ? 0x4bcca5cau : 0x12a17f52u); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixB-5/kernel_bound.cl b/proto-cuda/packs-readwidth/mixB-5/kernel_bound.cl new file mode 100644 index 000000000..c435d1213 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-5/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/5". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0xd572f9e9u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x30b20f21u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x30b20f21u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x91d2d3fcu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0x91d2d3fcu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x9395a5a1u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x9395a5a1u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef3dcd1fu; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xef3dcd1fu; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x8c9e416eu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0x8c9e416eu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x47c6146eu; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x47c6146eu; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x360af18au; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x360af18au; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0xd572f9e9u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = r3 * r1; // 0 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 4u); r2 = r2 ^ t_; } // 1 shfl + r0 = r0 ^ r6; // 2 xor + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 3 load + r6 = r6 ^ ds[r0 & mask]; // 4 load + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 2u); r2 = r2 ^ t_; } // 5 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r7 = r7 ^ t_; } // 6 shfl + r5 = rotl_imm(r5, 8u); // 7 rotl + r1 = r5 * r7 + r1; // 8 mad + r0 = r0 * r3; // 9 mul + r0 = r0 * r4; // 10 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 16u); r3 = r3 ^ t_; } // 11 shfl + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 16u); r0 = r0 ^ t_; } // 13 shfl + r4 = r4 - r6; // 14 sub + r7 = r7 | r5; // 15 or + r0 = rotr_var(r0, r6); // 16 rotr + r1 = r1 - r7; // 17 sub + r3 = r3 - r6; // 18 sub + r5 = r5 - r4; // 19 sub + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 20 shfl + r1 = r1 | r2; // 21 or + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r1 = r1 ^ t_; } // 22 shfl + r5 = rotr_var(r5, r3); // 23 rotr + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 24 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 25 load + r7 = r7 | r1; // 26 or + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 load + r3 = r3 ^ r4; // 28 xor + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 29 load + r5 = r5 + r4 + ((((sel >> 15u) & 1u) != 0u) ? 0x90aaff78u : 0xd0651e01u); // 30 add + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 31 load + r0 = r0 * r4; // 32 mul + r0 = r0 + r3 + ((((sel >> 2u) & 1u) != 0u) ? 0x3d45d6cfu : 0x8ac5ddc2u); // 33 add + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 34 load + r6 = r6 ^ r2; // 35 xor + r1 = r1 ^ r6; // 36 xor + r0 = r4 * r3 + r0; // 37 mad + r7 = rotl_imm(r7, 17u); // 38 rotl + r5 = r5 + r7 + ((((sel >> 0u) & 1u) != 0u) ? 0x4b7eb232u : 0x651ff080u); // 39 add + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 40 load + r6 = mul_hi(r6, r1); // 41 mulhi + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 42 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 43 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 44 load + r3 = r3 * r2; // 45 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 2u); r6 = r6 ^ t_; } // 46 shfl + r4 = r2 * r3 + r4; // 47 mad + r4 = r4 + r2 + ((((sel >> 20u) & 1u) != 0u) ? 0x97f34047u : 0x4f0330b4u); // 48 add + r6 = r6 ^ r3; // 49 xor + r0 = r0 + r2 + ((((sel >> 17u) & 1u) != 0u) ? 0xb24ddcb5u : 0x6e73a5d8u); // 50 add + r6 = rotl_imm(r6, 22u); // 51 rotl + r4 = r4 * r0; // 52 mul + r6 = rotr_var(r6, r3); // 53 rotr + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 54 load + r3 = r3 - r1; // 55 sub + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 1u); r3 = r3 ^ t_; } // 56 shfl + r3 = r3 | r1; // 57 or + r4 = r7 * r4 + r4; // 58 mad + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 load + r4 = r6 * r3 + r4; // 60 mad + r2 = rotl_imm(r2, 30u); // 61 rotl + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 62 load + r7 = r7 + r4 + ((((sel >> 3u) & 1u) != 0u) ? 0x4bcca5cau : 0x12a17f52u); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = r3 * r1; // 0 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 4u); r2 = r2 ^ t_; } // 1 shfl + r0 = r0 ^ r6; // 2 xor + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 3 load + r6 = r6 ^ ds[r0 & mask]; // 4 load + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 2u); r2 = r2 ^ t_; } // 5 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r7 = r7 ^ t_; } // 6 shfl + r5 = rotl_imm(r5, 8u); // 7 rotl + r1 = r5 * r7 + r1; // 8 mad + r0 = r0 * r3; // 9 mul + r0 = r0 * r4; // 10 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 16u); r3 = r3 ^ t_; } // 11 shfl + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 load + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 16u); r0 = r0 ^ t_; } // 13 shfl + r4 = r4 - r6; // 14 sub + r7 = r7 | r5; // 15 or + r0 = rotr_var(r0, r6); // 16 rotr + r1 = r1 - r7; // 17 sub + r3 = r3 - r6; // 18 sub + r5 = r5 - r4; // 19 sub + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 20 shfl + r1 = r1 | r2; // 21 or + { uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r1 = r1 ^ t_; } // 22 shfl + r5 = rotr_var(r5, r3); // 23 rotr + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 24 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 25 load + r7 = r7 | r1; // 26 or + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 load + r3 = r3 ^ r4; // 28 xor + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 29 load + r5 = r5 + r4 + ((((sel >> 15u) & 1u) != 0u) ? 0x90aaff78u : 0xd0651e01u); // 30 add + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 31 load + r0 = r0 * r4; // 32 mul + r0 = r0 + r3 + ((((sel >> 2u) & 1u) != 0u) ? 0x3d45d6cfu : 0x8ac5ddc2u); // 33 add + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 34 load + r6 = r6 ^ r2; // 35 xor + r1 = r1 ^ r6; // 36 xor + r0 = r4 * r3 + r0; // 37 mad + r7 = rotl_imm(r7, 17u); // 38 rotl + r5 = r5 + r7 + ((((sel >> 0u) & 1u) != 0u) ? 0x4b7eb232u : 0x651ff080u); // 39 add + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 40 load + r6 = mul_hi(r6, r1); // 41 mulhi + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 42 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 43 load + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 44 load + r3 = r3 * r2; // 45 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 2u); r6 = r6 ^ t_; } // 46 shfl + r4 = r2 * r3 + r4; // 47 mad + r4 = r4 + r2 + ((((sel >> 20u) & 1u) != 0u) ? 0x97f34047u : 0x4f0330b4u); // 48 add + r6 = r6 ^ r3; // 49 xor + r0 = r0 + r2 + ((((sel >> 17u) & 1u) != 0u) ? 0xb24ddcb5u : 0x6e73a5d8u); // 50 add + r6 = rotl_imm(r6, 22u); // 51 rotl + r4 = r4 * r0; // 52 mul + r6 = rotr_var(r6, r3); // 53 rotr + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 54 load + r3 = r3 - r1; // 55 sub + { uint t_; IGNEUM_SHFL_XOR(t_, r7, 1u); r3 = r3 ^ t_; } // 56 shfl + r3 = r3 | r1; // 57 or + r4 = r7 * r4 + r4; // 58 mad + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 load + r4 = r6 * r3 + r4; // 60 mad + r2 = rotl_imm(r2, 30u); // 61 rotl + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 62 load + r7 = r7 + r4 + ((((sel >> 3u) & 1u) != 0u) ? 0x4bcca5cau : 0x12a17f52u); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-5/kernel_bound.cu b/proto-cuda/packs-readwidth/mixB-5/kernel_bound.cu new file mode 100644 index 000000000..1c413a248 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-5/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/5". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r3 = r3 * r1; // 0 mul + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 4); // 1 shfl + r0 = r0 ^ r6; // 2 xor + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 3 load + r6 = r6 ^ ds[r0 & mask]; // 4 load + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 2); // 5 shfl + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 4); // 6 shfl + r5 = rotl_imm(r5, 8u); // 7 rotl + r1 = r5 * r7 + r1; // 8 mad + r0 = r0 * r3; // 9 mul + r0 = r0 * r4; // 10 mul + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r0, 16); // 11 shfl + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 load + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r2, 16); // 13 shfl + r4 = r4 - r6; // 14 sub + r7 = r7 | r5; // 15 or + r0 = rotr_var(r0, r6); // 16 rotr + r1 = r1 - r7; // 17 sub + r3 = r3 - r6; // 18 sub + r5 = r5 - r4; // 19 sub + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 20 shfl + r1 = r1 | r2; // 21 or + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 22 shfl + r5 = rotr_var(r5, r3); // 23 rotr + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 24 load + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 25 load + r7 = r7 | r1; // 26 or + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 load + r3 = r3 ^ r4; // 28 xor + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 29 load + r5 = r5 + r4 + ((((sel >> 15u) & 1u) != 0u) ? 0x90aaff78u : 0xd0651e01u); // 30 add + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 31 load + r0 = r0 * r4; // 32 mul + r0 = r0 + r3 + ((((sel >> 2u) & 1u) != 0u) ? 0x3d45d6cfu : 0x8ac5ddc2u); // 33 add + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 34 load + r6 = r6 ^ r2; // 35 xor + r1 = r1 ^ r6; // 36 xor + r0 = r4 * r3 + r0; // 37 mad + r7 = rotl_imm(r7, 17u); // 38 rotl + r5 = r5 + r7 + ((((sel >> 0u) & 1u) != 0u) ? 0x4b7eb232u : 0x651ff080u); // 39 add + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 40 load + r6 = __umulhi(r6, r1); // 41 mulhi + { uint32_t b_ = (r6 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 42 load + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 43 load + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 44 load + r3 = r3 * r2; // 45 mul + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 2); // 46 shfl + r4 = r2 * r3 + r4; // 47 mad + r4 = r4 + r2 + ((((sel >> 20u) & 1u) != 0u) ? 0x97f34047u : 0x4f0330b4u); // 48 add + r6 = r6 ^ r3; // 49 xor + r0 = r0 + r2 + ((((sel >> 17u) & 1u) != 0u) ? 0xb24ddcb5u : 0x6e73a5d8u); // 50 add + r6 = rotl_imm(r6, 22u); // 51 rotl + r4 = r4 * r0; // 52 mul + r6 = rotr_var(r6, r3); // 53 rotr + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 54 load + r3 = r3 - r1; // 55 sub + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 1); // 56 shfl + r3 = r3 | r1; // 57 or + r4 = r7 * r4 + r4; // 58 mad + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 load + r4 = r6 * r3 + r4; // 60 mad + r2 = rotl_imm(r2, 30u); // 61 rotl + { uint32_t b_ = (r1 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 62 load + r7 = r7 + r4 + ((((sel >> 3u) & 1u) != 0u) ? 0x4bcca5cau : 0x12a17f52u); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/mixB-5/memhard.h b/proto-cuda/packs-readwidth/mixB-5/memhard.h new file mode 100644 index 000000000..5541ff7aa --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-5/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/5". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/mixB-5/memhard.metal b/proto-cuda/packs-readwidth/mixB-5/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-5/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/mixB-5/program.h b/proto-cuda/packs-readwidth/mixB-5/program.h new file mode 100644 index 000000000..2b588fc66 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-5/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/5". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-readwidth/B/5" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d7265616477696474682f422f35" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x6a607f9221afe5daull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 shfl=9 add=6 mul=6 mad=5 sub=5 xor=5 or=4 rotl=4 rotr=3 mulhi=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "mix25-50-25" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 25, 50, 25 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 1, 11, 4 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 3488 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0xd572f9e9u, 0x30b20f21u, 0x91d2d3fcu, 0x9395a5a1u, 0xef3dcd1fu, 0x8c9e416eu, 0x47c6146eu, 0x360af18au } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/mixB-5/program.json b/proto-cuda/packs-readwidth/mixB-5/program.json new file mode 100644 index 000000000..faa289072 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-5/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x6a607f9221afe5da", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-readwidth/B/5", + "seed_bytes": "69676e65756d2d7265616477696474682f422f35", + "seed_words": ["0xd572f9e9", "0x30b20f21", "0x91d2d3fc", "0x9395a5a1", "0xef3dcd1f", "0x8c9e416e", "0x47c6146e", "0x360af18a"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "mix25-50-25", + "load_slots": 16, + "load_mix_percent_4_16_64": [25, 50, 25], + "load_width_counts_4_16_64": [1, 11, 4], + "bytes_per_hash": 3488, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "shfl": 9, "add": 6, "mul": 6, "mad": 5, "sub": 5, "xor": 5, "or": 4, "rotl": 4, "rotr": 3, "mulhi": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mul", "dst": 3, "src": 1, "src2": 5, "imm": "0x8d99879f", "imm2": "0xa0b4a6c1", "rot": 7, "bit": 29, "mask": 16, "width": 1}, + {"i": 1, "op": "shfl", "dst": 2, "src": 4, "src2": 0, "imm": "0x6eb233b5", "imm2": "0xa60c7aee", "rot": 2, "bit": 18, "mask": 4, "width": 1}, + {"i": 2, "op": "xor", "dst": 0, "src": 6, "src2": 6, "imm": "0x08bb9aa0", "imm2": "0xc3c30d7b", "rot": 12, "bit": 13, "mask": 16, "width": 1}, + {"i": 3, "op": "load", "dst": 4, "src": 3, "src2": 2, "imm": "0x65cf244b", "imm2": "0xc8e345dd", "rot": 13, "bit": 14, "mask": 1, "width": 16}, + {"i": 4, "op": "load", "dst": 6, "src": 0, "src2": 6, "imm": "0xbd2e02d9", "imm2": "0x12c4bffd", "rot": 8, "bit": 20, "mask": 8, "width": 1}, + {"i": 5, "op": "shfl", "dst": 2, "src": 0, "src2": 3, "imm": "0x45017107", "imm2": "0xe7972d4b", "rot": 11, "bit": 26, "mask": 2, "width": 1}, + {"i": 6, "op": "shfl", "dst": 7, "src": 6, "src2": 3, "imm": "0x6fee330f", "imm2": "0x63572bf7", "rot": 30, "bit": 28, "mask": 4, "width": 1}, + {"i": 7, "op": "rotl", "dst": 5, "src": 2, "src2": 3, "imm": "0x655cf6dd", "imm2": "0x9f4cbce8", "rot": 8, "bit": 30, "mask": 1, "width": 1}, + {"i": 8, "op": "mad", "dst": 1, "src": 5, "src2": 7, "imm": "0xc2d75c8a", "imm2": "0x5dda61cd", "rot": 6, "bit": 14, "mask": 16, "width": 1}, + {"i": 9, "op": "mul", "dst": 0, "src": 3, "src2": 5, "imm": "0xd185bb8e", "imm2": "0x19fef814", "rot": 4, "bit": 7, "mask": 2, "width": 1}, + {"i": 10, "op": "mul", "dst": 0, "src": 4, "src2": 3, "imm": "0xc448790d", "imm2": "0x5eccbec1", "rot": 17, "bit": 5, "mask": 16, "width": 1}, + {"i": 11, "op": "shfl", "dst": 3, "src": 0, "src2": 5, "imm": "0x2124d54a", "imm2": "0xb61e62fc", "rot": 16, "bit": 17, "mask": 16, "width": 1}, + {"i": 12, "op": "load", "dst": 4, "src": 5, "src2": 4, "imm": "0x1210ee04", "imm2": "0x6c0239db", "rot": 16, "bit": 18, "mask": 8, "width": 4}, + {"i": 13, "op": "shfl", "dst": 0, "src": 2, "src2": 2, "imm": "0x8a4ec78b", "imm2": "0x9fff9bdd", "rot": 19, "bit": 0, "mask": 16, "width": 1}, + {"i": 14, "op": "sub", "dst": 4, "src": 6, "src2": 0, "imm": "0x1efbdf5f", "imm2": "0x2b3fbc72", "rot": 24, "bit": 31, "mask": 16, "width": 1}, + {"i": 15, "op": "or", "dst": 7, "src": 5, "src2": 5, "imm": "0x7c272d45", "imm2": "0xcb7860aa", "rot": 20, "bit": 13, "mask": 2, "width": 1}, + {"i": 16, "op": "rotr", "dst": 0, "src": 6, "src2": 3, "imm": "0x133400bd", "imm2": "0x5b4307e2", "rot": 15, "bit": 30, "mask": 2, "width": 1}, + {"i": 17, "op": "sub", "dst": 1, "src": 7, "src2": 6, "imm": "0x60b98288", "imm2": "0x7266fa37", "rot": 5, "bit": 8, "mask": 2, "width": 1}, + {"i": 18, "op": "sub", "dst": 3, "src": 6, "src2": 2, "imm": "0x81238503", "imm2": "0xada4d549", "rot": 8, "bit": 10, "mask": 8, "width": 1}, + {"i": 19, "op": "sub", "dst": 5, "src": 4, "src2": 3, "imm": "0x97d5351f", "imm2": "0x87da996e", "rot": 16, "bit": 0, "mask": 16, "width": 1}, + {"i": 20, "op": "shfl", "dst": 3, "src": 5, "src2": 4, "imm": "0x61a9fb98", "imm2": "0x5d6413a9", "rot": 27, "bit": 2, "mask": 1, "width": 1}, + {"i": 21, "op": "or", "dst": 1, "src": 2, "src2": 0, "imm": "0x76659fd4", "imm2": "0x1b52a4fd", "rot": 3, "bit": 7, "mask": 16, "width": 1}, + {"i": 22, "op": "shfl", "dst": 1, "src": 0, "src2": 0, "imm": "0x7072d886", "imm2": "0x31c1c1e4", "rot": 5, "bit": 17, "mask": 4, "width": 1}, + {"i": 23, "op": "rotr", "dst": 5, "src": 3, "src2": 2, "imm": "0x4e913f3c", "imm2": "0x650bcb89", "rot": 4, "bit": 7, "mask": 4, "width": 1}, + {"i": 24, "op": "load", "dst": 7, "src": 2, "src2": 2, "imm": "0xd693898d", "imm2": "0x4eafd5cf", "rot": 26, "bit": 15, "mask": 8, "width": 4}, + {"i": 25, "op": "load", "dst": 1, "src": 0, "src2": 3, "imm": "0x78fc3704", "imm2": "0xf64c3a71", "rot": 12, "bit": 8, "mask": 8, "width": 4}, + {"i": 26, "op": "or", "dst": 7, "src": 1, "src2": 3, "imm": "0xbea69d0d", "imm2": "0x6827d6f3", "rot": 13, "bit": 9, "mask": 2, "width": 1}, + {"i": 27, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0xd0d39c63", "imm2": "0x4f108768", "rot": 10, "bit": 14, "mask": 4, "width": 4}, + {"i": 28, "op": "xor", "dst": 3, "src": 4, "src2": 3, "imm": "0xe4e057c0", "imm2": "0x16f1be43", "rot": 24, "bit": 23, "mask": 1, "width": 1}, + {"i": 29, "op": "load", "dst": 6, "src": 5, "src2": 5, "imm": "0xb89861c5", "imm2": "0x8e5ee7bd", "rot": 13, "bit": 6, "mask": 16, "width": 16}, + {"i": 30, "op": "add", "dst": 5, "src": 4, "src2": 0, "imm": "0xd0651e01", "imm2": "0x90aaff78", "rot": 10, "bit": 15, "mask": 16, "width": 1}, + {"i": 31, "op": "load", "dst": 1, "src": 3, "src2": 1, "imm": "0x966081ac", "imm2": "0x12902378", "rot": 18, "bit": 22, "mask": 16, "width": 16}, + {"i": 32, "op": "mul", "dst": 0, "src": 4, "src2": 1, "imm": "0x8191bcaf", "imm2": "0x6e9c9179", "rot": 21, "bit": 18, "mask": 8, "width": 1}, + {"i": 33, "op": "add", "dst": 0, "src": 3, "src2": 2, "imm": "0x8ac5ddc2", "imm2": "0x3d45d6cf", "rot": 15, "bit": 2, "mask": 1, "width": 1}, + {"i": 34, "op": "load", "dst": 6, "src": 4, "src2": 5, "imm": "0xfe5e8e98", "imm2": "0x42f50042", "rot": 10, "bit": 8, "mask": 4, "width": 4}, + {"i": 35, "op": "xor", "dst": 6, "src": 2, "src2": 0, "imm": "0x91c0ab45", "imm2": "0xd04ac64a", "rot": 17, "bit": 19, "mask": 1, "width": 1}, + {"i": 36, "op": "xor", "dst": 1, "src": 6, "src2": 4, "imm": "0x33ee7c8a", "imm2": "0x95a5252f", "rot": 25, "bit": 4, "mask": 2, "width": 1}, + {"i": 37, "op": "mad", "dst": 0, "src": 4, "src2": 3, "imm": "0xecf08fbc", "imm2": "0x410bb43e", "rot": 23, "bit": 4, "mask": 8, "width": 1}, + {"i": 38, "op": "rotl", "dst": 7, "src": 1, "src2": 6, "imm": "0xce8d1fbd", "imm2": "0xa4eb0972", "rot": 17, "bit": 22, "mask": 16, "width": 1}, + {"i": 39, "op": "add", "dst": 5, "src": 7, "src2": 6, "imm": "0x651ff080", "imm2": "0x4b7eb232", "rot": 13, "bit": 0, "mask": 2, "width": 1}, + {"i": 40, "op": "load", "dst": 1, "src": 2, "src2": 0, "imm": "0x9b6de222", "imm2": "0xa65a3196", "rot": 19, "bit": 3, "mask": 8, "width": 4}, + {"i": 41, "op": "mulhi", "dst": 6, "src": 1, "src2": 5, "imm": "0x2eb1b89c", "imm2": "0xcd51c7e9", "rot": 22, "bit": 28, "mask": 16, "width": 1}, + {"i": 42, "op": "load", "dst": 1, "src": 6, "src2": 0, "imm": "0x34c83029", "imm2": "0xc7caa918", "rot": 23, "bit": 23, "mask": 2, "width": 16}, + {"i": 43, "op": "load", "dst": 3, "src": 0, "src2": 7, "imm": "0x5fdcdab2", "imm2": "0x091f2301", "rot": 10, "bit": 18, "mask": 8, "width": 4}, + {"i": 44, "op": "load", "dst": 6, "src": 7, "src2": 7, "imm": "0x12f2c079", "imm2": "0x9f9f8573", "rot": 11, "bit": 8, "mask": 1, "width": 4}, + {"i": 45, "op": "mul", "dst": 3, "src": 2, "src2": 2, "imm": "0xca1cb045", "imm2": "0xbefca4ec", "rot": 5, "bit": 23, "mask": 8, "width": 1}, + {"i": 46, "op": "shfl", "dst": 6, "src": 5, "src2": 6, "imm": "0x511f46fb", "imm2": "0x11028be8", "rot": 4, "bit": 25, "mask": 2, "width": 1}, + {"i": 47, "op": "mad", "dst": 4, "src": 2, "src2": 3, "imm": "0xe1547c3a", "imm2": "0x6513ccb5", "rot": 5, "bit": 29, "mask": 8, "width": 1}, + {"i": 48, "op": "add", "dst": 4, "src": 2, "src2": 5, "imm": "0x4f0330b4", "imm2": "0x97f34047", "rot": 26, "bit": 20, "mask": 16, "width": 1}, + {"i": 49, "op": "xor", "dst": 6, "src": 3, "src2": 0, "imm": "0x8fa8489b", "imm2": "0xf72c5554", "rot": 6, "bit": 5, "mask": 16, "width": 1}, + {"i": 50, "op": "add", "dst": 0, "src": 2, "src2": 3, "imm": "0x6e73a5d8", "imm2": "0xb24ddcb5", "rot": 19, "bit": 17, "mask": 4, "width": 1}, + {"i": 51, "op": "rotl", "dst": 6, "src": 3, "src2": 6, "imm": "0x5a2a99f0", "imm2": "0xff22310a", "rot": 22, "bit": 19, "mask": 1, "width": 1}, + {"i": 52, "op": "mul", "dst": 4, "src": 0, "src2": 4, "imm": "0x2f1f4fd7", "imm2": "0xb39caa61", "rot": 21, "bit": 8, "mask": 4, "width": 1}, + {"i": 53, "op": "rotr", "dst": 6, "src": 3, "src2": 0, "imm": "0x11dec3cc", "imm2": "0xf2c4a139", "rot": 3, "bit": 10, "mask": 2, "width": 1}, + {"i": 54, "op": "load", "dst": 0, "src": 6, "src2": 5, "imm": "0x31111698", "imm2": "0x3ccec9cd", "rot": 22, "bit": 9, "mask": 2, "width": 4}, + {"i": 55, "op": "sub", "dst": 3, "src": 1, "src2": 0, "imm": "0x579deaba", "imm2": "0x71c63728", "rot": 25, "bit": 12, "mask": 16, "width": 1}, + {"i": 56, "op": "shfl", "dst": 3, "src": 7, "src2": 6, "imm": "0xa537855c", "imm2": "0x93db1532", "rot": 2, "bit": 5, "mask": 1, "width": 1}, + {"i": 57, "op": "or", "dst": 3, "src": 1, "src2": 4, "imm": "0x8ba9bc0c", "imm2": "0xfdda17c1", "rot": 29, "bit": 18, "mask": 8, "width": 1}, + {"i": 58, "op": "mad", "dst": 4, "src": 7, "src2": 4, "imm": "0x872fe127", "imm2": "0x47a6166c", "rot": 10, "bit": 10, "mask": 1, "width": 1}, + {"i": 59, "op": "load", "dst": 1, "src": 0, "src2": 7, "imm": "0x0089e437", "imm2": "0xcf6c172e", "rot": 14, "bit": 17, "mask": 1, "width": 4}, + {"i": 60, "op": "mad", "dst": 4, "src": 6, "src2": 3, "imm": "0x9a9b124b", "imm2": "0x25a02ef4", "rot": 7, "bit": 19, "mask": 4, "width": 1}, + {"i": 61, "op": "rotl", "dst": 2, "src": 6, "src2": 2, "imm": "0x0489abc1", "imm2": "0x13809620", "rot": 30, "bit": 4, "mask": 8, "width": 1}, + {"i": 62, "op": "load", "dst": 2, "src": 1, "src2": 5, "imm": "0xb1f6105a", "imm2": "0xefbc6585", "rot": 8, "bit": 4, "mask": 4, "width": 4}, + {"i": 63, "op": "add", "dst": 7, "src": 4, "src2": 5, "imm": "0x12a17f52", "imm2": "0x4bcca5ca", "rot": 23, "bit": 3, "mask": 1, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/mixB-5/program.metal b/proto-cuda/packs-readwidth/mixB-5/program.metal new file mode 100644 index 000000000..9d731e8e7 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-5/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0xd572f9e9u, 0x30b20f21u, 0x91d2d3fcu, 0x9395a5a1u, 0xef3dcd1fu, 0x8c9e416eu, 0x47c6146eu, 0x360af18au }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = r3 * r1; // 0 + r2 = r2 ^ simd_shuffle_xor(r4, (ushort)4); // 1 + r0 = r0 ^ r6; // 2 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 3 + r6 = r6 ^ dataset[r0 & MASK]; // 4 + r2 = r2 ^ simd_shuffle_xor(r0, (ushort)2); // 5 + r7 = r7 ^ simd_shuffle_xor(r6, (ushort)4); // 6 + r5 = rotl_imm(r5, 8u); // 7 + r1 = r5 * r7 + r1; // 8 + r0 = r0 * r3; // 9 + r0 = r0 * r4; // 10 + r3 = r3 ^ simd_shuffle_xor(r0, (ushort)16); // 11 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 + r0 = r0 ^ simd_shuffle_xor(r2, (ushort)16); // 13 + r4 = r4 - r6; // 14 + r7 = r7 | r5; // 15 + r0 = rotr_var(r0, r6); // 16 + r1 = r1 - r7; // 17 + r3 = r3 - r6; // 18 + r5 = r5 - r4; // 19 + r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 20 + r1 = r1 | r2; // 21 + r1 = r1 ^ simd_shuffle_xor(r0, (ushort)4); // 22 + r5 = rotr_var(r5, r3); // 23 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 24 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 25 + r7 = r7 | r1; // 26 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 + r3 = r3 ^ r4; // 28 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 29 + r5 = r5 + r4 + select(0xd0651e01u, 0x90aaff78u, ((sel >> 15u) & 1u) != 0u); // 30 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 31 + r0 = r0 * r4; // 32 + r0 = r0 + r3 + select(0x8ac5ddc2u, 0x3d45d6cfu, ((sel >> 2u) & 1u) != 0u); // 33 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 34 + r6 = r6 ^ r2; // 35 + r1 = r1 ^ r6; // 36 + r0 = r4 * r3 + r0; // 37 + r7 = rotl_imm(r7, 17u); // 38 + r5 = r5 + r7 + select(0x651ff080u, 0x4b7eb232u, ((sel >> 0u) & 1u) != 0u); // 39 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 40 + r6 = mulhi(r6, r1); // 41 + { uint b_ = (r6 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 42 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 43 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 44 + r3 = r3 * r2; // 45 + r6 = r6 ^ simd_shuffle_xor(r5, (ushort)2); // 46 + r4 = r2 * r3 + r4; // 47 + r4 = r4 + r2 + select(0x4f0330b4u, 0x97f34047u, ((sel >> 20u) & 1u) != 0u); // 48 + r6 = r6 ^ r3; // 49 + r0 = r0 + r2 + select(0x6e73a5d8u, 0xb24ddcb5u, ((sel >> 17u) & 1u) != 0u); // 50 + r6 = rotl_imm(r6, 22u); // 51 + r4 = r4 * r0; // 52 + r6 = rotr_var(r6, r3); // 53 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 54 + r3 = r3 - r1; // 55 + r3 = r3 ^ simd_shuffle_xor(r7, (ushort)1); // 56 + r3 = r3 | r1; // 57 + r4 = r7 * r4 + r4; // 58 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 + r4 = r6 * r3 + r4; // 60 + r2 = rotl_imm(r2, 30u); // 61 + { uint b_ = (r1 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 62 + r7 = r7 + r4 + select(0x12a17f52u, 0x4bcca5cau, ((sel >> 3u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-5/program_bound.metal b/proto-cuda/packs-readwidth/mixB-5/program_bound.metal new file mode 100644 index 000000000..6f7e29e2e --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-5/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0xd572f9e9u, 0x30b20f21u, 0x91d2d3fcu, 0x9395a5a1u, 0xef3dcd1fu, 0x8c9e416eu, 0x47c6146eu, 0x360af18au }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r3 = r3 * r1; // 0 + r2 = r2 ^ simd_shuffle_xor(r4, (ushort)4); // 1 + r0 = r0 ^ r6; // 2 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 3 + r6 = r6 ^ dataset[r0 & MASK]; // 4 + r2 = r2 ^ simd_shuffle_xor(r0, (ushort)2); // 5 + r7 = r7 ^ simd_shuffle_xor(r6, (ushort)4); // 6 + r5 = rotl_imm(r5, 8u); // 7 + r1 = r5 * r7 + r1; // 8 + r0 = r0 * r3; // 9 + r0 = r0 * r4; // 10 + r3 = r3 ^ simd_shuffle_xor(r0, (ushort)16); // 11 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 12 + r0 = r0 ^ simd_shuffle_xor(r2, (ushort)16); // 13 + r4 = r4 - r6; // 14 + r7 = r7 | r5; // 15 + r0 = rotr_var(r0, r6); // 16 + r1 = r1 - r7; // 17 + r3 = r3 - r6; // 18 + r5 = r5 - r4; // 19 + r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 20 + r1 = r1 | r2; // 21 + r1 = r1 ^ simd_shuffle_xor(r0, (ushort)4); // 22 + r5 = rotr_var(r5, r3); // 23 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 24 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 25 + r7 = r7 | r1; // 26 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 27 + r3 = r3 ^ r4; // 28 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 29 + r5 = r5 + r4 + select(0xd0651e01u, 0x90aaff78u, ((sel >> 15u) & 1u) != 0u); // 30 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 31 + r0 = r0 * r4; // 32 + r0 = r0 + r3 + select(0x8ac5ddc2u, 0x3d45d6cfu, ((sel >> 2u) & 1u) != 0u); // 33 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 34 + r6 = r6 ^ r2; // 35 + r1 = r1 ^ r6; // 36 + r0 = r4 * r3 + r0; // 37 + r7 = rotl_imm(r7, 17u); // 38 + r5 = r5 + r7 + select(0x651ff080u, 0x4b7eb232u, ((sel >> 0u) & 1u) != 0u); // 39 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 40 + r6 = mulhi(r6, r1); // 41 + { uint b_ = (r6 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 42 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 43 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r6 = x_; } // 44 + r3 = r3 * r2; // 45 + r6 = r6 ^ simd_shuffle_xor(r5, (ushort)2); // 46 + r4 = r2 * r3 + r4; // 47 + r4 = r4 + r2 + select(0x4f0330b4u, 0x97f34047u, ((sel >> 20u) & 1u) != 0u); // 48 + r6 = r6 ^ r3; // 49 + r0 = r0 + r2 + select(0x6e73a5d8u, 0xb24ddcb5u, ((sel >> 17u) & 1u) != 0u); // 50 + r6 = rotl_imm(r6, 22u); // 51 + r4 = r4 * r0; // 52 + r6 = rotr_var(r6, r3); // 53 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 54 + r3 = r3 - r1; // 55 + r3 = r3 ^ simd_shuffle_xor(r7, (ushort)1); // 56 + r3 = r3 | r1; // 57 + r4 = r7 * r4 + r4; // 58 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 + r4 = r6 * r3 + r4; // 60 + r2 = rotl_imm(r2, 30u); // 61 + { uint b_ = (r1 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 62 + r7 = r7 + r4 + select(0x12a17f52u, 0x4bcca5cau, ((sel >> 3u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/mixB-5/vectors.h b/proto-cuda/packs-readwidth/mixB-5/vectors.h new file mode 100644 index 000000000..294507154 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-5/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-readwidth/B/5". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0xb08d5d7189d56e94ull, 0xab5b75dd880b7295ull, 0x26bb92e29f3e71e3ull, 0xa666e3257334a8abull, 0xcd243526b9659fd9ull, 0x5f8280786af5626cull, 0xd2ff443e6729e841ull, 0xb27ab6b717855802ull, + 0x2da4f7cb38c0ff8cull, 0xf6209f0089cd3bb2ull, 0xae1558a9ff2cfe7bull, 0x1e6d6fdc3a83c2b2ull, 0x1c6218028959af92ull, 0x26897968e6dc646full, 0x3f76b5a3a3039a94ull, 0xefe6034ccc554eb3ull, + 0x1bf0f378688bfda1ull, 0xec29e8055259a75cull, 0xc61a1da068f19f99ull, 0xeea05bf7ec651606ull, 0x270fd486c72c09bdull, 0x3aa2a4a7c961df6full, 0x66322a0646d9c393ull, 0x478056eb7217a331ull, + 0x754ef3ad248cf343ull, 0xca39f0508d80787bull, 0x63ed748a40628aa8ull, 0x26705cf47a417e28ull, 0x983306333881de0full, 0x04f275174fbac096ull, 0xfa645124238644dcull, 0xab45dc7c708e978aull + }, + { // base nonce 4096 + 0xe7c7d262a04473ddull, 0x4e9f01a5c1d2a4efull, 0x2d3348e999898c29ull, 0x29a4e27ddca4638aull, 0xe781b9fce52b39deull, 0x27dc37e476b22244ull, 0x45ebd4e62d962ef5ull, 0x61179d19d487824aull, + 0x0bbd6ab97315881cull, 0x67bc43c05b267aebull, 0x98cbdad1d881d68eull, 0x94c5936a6a4bc847ull, 0x65e30de6119340f1ull, 0xe53c3d01bbcbf563ull, 0x9c714bbb0f201cf2ull, 0x83368bd0bed14ebbull, + 0x5ecbb99386943683ull, 0x789eb6f9df88fd10ull, 0xdb863e9e8b176f85ull, 0x93f1c16de8ef25cbull, 0x7967c6caee510090ull, 0x1111b6c026de48e1ull, 0x04c65ab39b766ba1ull, 0x9f9baf77e03d1997ull, + 0x1fee295ff5d15b5aull, 0xa6e4c15e1bf0ebacull, 0xa7428cd8cec2eff2ull, 0x2f1ca539ccb1f2e0ull, 0x8dc1d3623c8eeb5cull, 0x89bf0d1f967f8508ull, 0x7c7f5b462c97d4f7ull, 0x82c4b25aa1d4abfeull + }, + { // base nonce 1000000 + 0x37841a68b29fb1b7ull, 0x8ee48fcf088cd0adull, 0x1414c8f175fb4064ull, 0xaeb75ecf4fd94977ull, 0x6a2df0aef7920cb0ull, 0xd8fcee5c9a453d44ull, 0xd1ec4bdf30de6c8aull, 0x10f9d4896113f0c9ull, + 0xf6a189f4230e3351ull, 0x84ba6f8aff2659bfull, 0x65afd2b39c0b9c74ull, 0x2621b2d24af1026aull, 0x2c82f513d44616a4ull, 0xc0fd4e133df90c89ull, 0xd09cc88053c5d0cbull, 0x8e741e232dcf13a8ull, + 0x0bcecb0607611669ull, 0x8e47ae615099cd64ull, 0x5fc5ce2fe2575680ull, 0xf93a4d95b195e930ull, 0x208ced09b346163cull, 0x05092ef2e765bae5ull, 0x7d06298ad52cad4dull, 0x6a6075eca24b7210ull, + 0xd22405915ee4dc77ull, 0x2bdd6431ab621d07ull, 0xb8100d4a8c50208full, 0xe05f85f37b897bc2ull, 0xd03aed7e1a66910eull, 0xe6a9acd67726eec1ull, 0xb0509594e9a7b011ull, 0x0ea38ba4b390b72cull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/mixB-5/vectors.json b/proto-cuda/packs-readwidth/mixB-5/vectors.json new file mode 100644 index 000000000..f1a7d87f9 --- /dev/null +++ b/proto-cuda/packs-readwidth/mixB-5/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-readwidth/B/5", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0xb08d5d7189d56e94", "0xab5b75dd880b7295", "0x26bb92e29f3e71e3", "0xa666e3257334a8ab", "0xcd243526b9659fd9", "0x5f8280786af5626c", "0xd2ff443e6729e841", "0xb27ab6b717855802", + "0x2da4f7cb38c0ff8c", "0xf6209f0089cd3bb2", "0xae1558a9ff2cfe7b", "0x1e6d6fdc3a83c2b2", "0x1c6218028959af92", "0x26897968e6dc646f", "0x3f76b5a3a3039a94", "0xefe6034ccc554eb3", + "0x1bf0f378688bfda1", "0xec29e8055259a75c", "0xc61a1da068f19f99", "0xeea05bf7ec651606", "0x270fd486c72c09bd", "0x3aa2a4a7c961df6f", "0x66322a0646d9c393", "0x478056eb7217a331", + "0x754ef3ad248cf343", "0xca39f0508d80787b", "0x63ed748a40628aa8", "0x26705cf47a417e28", "0x983306333881de0f", "0x04f275174fbac096", "0xfa645124238644dc", "0xab45dc7c708e978a" + ]}, + {"base_nonce": 4096, "expected": [ + "0xe7c7d262a04473dd", "0x4e9f01a5c1d2a4ef", "0x2d3348e999898c29", "0x29a4e27ddca4638a", "0xe781b9fce52b39de", "0x27dc37e476b22244", "0x45ebd4e62d962ef5", "0x61179d19d487824a", + "0x0bbd6ab97315881c", "0x67bc43c05b267aeb", "0x98cbdad1d881d68e", "0x94c5936a6a4bc847", "0x65e30de6119340f1", "0xe53c3d01bbcbf563", "0x9c714bbb0f201cf2", "0x83368bd0bed14ebb", + "0x5ecbb99386943683", "0x789eb6f9df88fd10", "0xdb863e9e8b176f85", "0x93f1c16de8ef25cb", "0x7967c6caee510090", "0x1111b6c026de48e1", "0x04c65ab39b766ba1", "0x9f9baf77e03d1997", + "0x1fee295ff5d15b5a", "0xa6e4c15e1bf0ebac", "0xa7428cd8cec2eff2", "0x2f1ca539ccb1f2e0", "0x8dc1d3623c8eeb5c", "0x89bf0d1f967f8508", "0x7c7f5b462c97d4f7", "0x82c4b25aa1d4abfe" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x37841a68b29fb1b7", "0x8ee48fcf088cd0ad", "0x1414c8f175fb4064", "0xaeb75ecf4fd94977", "0x6a2df0aef7920cb0", "0xd8fcee5c9a453d44", "0xd1ec4bdf30de6c8a", "0x10f9d4896113f0c9", + "0xf6a189f4230e3351", "0x84ba6f8aff2659bf", "0x65afd2b39c0b9c74", "0x2621b2d24af1026a", "0x2c82f513d44616a4", "0xc0fd4e133df90c89", "0xd09cc88053c5d0cb", "0x8e741e232dcf13a8", + "0x0bcecb0607611669", "0x8e47ae615099cd64", "0x5fc5ce2fe2575680", "0xf93a4d95b195e930", "0x208ced09b346163c", "0x05092ef2e765bae5", "0x7d06298ad52cad4d", "0x6a6075eca24b7210", + "0xd22405915ee4dc77", "0x2bdd6431ab621d07", "0xb8100d4a8c50208f", "0xe05f85f37b897bc2", "0xd03aed7e1a66910e", "0xe6a9acd67726eec1", "0xb0509594e9a7b011", "0x0ea38ba4b390b72c" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/scr0/kernel.cl b/proto-cuda/packs-readwidth/scr0/kernel.cl new file mode 100644 index 000000000..5ab647860 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr0/kernel.cl @@ -0,0 +1,291 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + r3 = r3 ^ ds[r6 & mask]; // 16 load + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + r0 = r0 ^ ds[r2 & mask]; // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + r1 = r1 ^ ds[r4 & mask]; // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/scr0/kernel.cu b/proto-cuda/packs-readwidth/scr0/kernel.cu new file mode 100644 index 000000000..f7ac2c32d --- /dev/null +++ b/proto-cuda/packs-readwidth/scr0/kernel.cu @@ -0,0 +1,177 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + r3 = r3 ^ ds[r6 & mask]; // 16 load + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + r0 = r0 ^ ds[r2 & mask]; // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + r1 = r1 ^ ds[r4 & mask]; // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr0/kernel_bound.cl b/proto-cuda/packs-readwidth/scr0/kernel_bound.cl new file mode 100644 index 000000000..eaba1546c --- /dev/null +++ b/proto-cuda/packs-readwidth/scr0/kernel_bound.cl @@ -0,0 +1,393 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + r3 = r3 ^ ds[r6 & mask]; // 16 load + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + r0 = r0 ^ ds[r2 & mask]; // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + r1 = r1 ^ ds[r4 & mask]; // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + r3 = r3 ^ ds[r6 & mask]; // 16 load + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + r0 = r0 ^ ds[r2 & mask]; // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + r1 = r1 ^ ds[r4 & mask]; // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr0/kernel_bound.cu b/proto-cuda/packs-readwidth/scr0/kernel_bound.cu new file mode 100644 index 000000000..5f1b9aa92 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr0/kernel_bound.cu @@ -0,0 +1,136 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + r3 = r3 ^ ds[r6 & mask]; // 16 load + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + r0 = r0 ^ ds[r2 & mask]; // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + r1 = r1 ^ ds[r4 & mask]; // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr0/memhard.h b/proto-cuda/packs-readwidth/scr0/memhard.h new file mode 100644 index 000000000..4803d8e40 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr0/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/scr0/memhard.metal b/proto-cuda/packs-readwidth/scr0/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr0/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/scr0/program.h b/proto-cuda/packs-readwidth/scr0/program.h new file mode 100644 index 000000000..2409c10b0 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr0/program.h @@ -0,0 +1,67 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-genesis" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x2f098ae568f38029ull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=9 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 shfl=4 rotr=2 rotl=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "scr0" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 100, 0, 0 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 512 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// Variant 5: persistent warps, a 1 MiB scratch per launched warp (the host launches N warps and passes scratch, +// groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). +#define IGNEUM_PERSISTENT_WARPS 1 +#define IGNEUM_SCRATCH_OPS 0 // scratch read-modify-writes per program (0 per hash) +#define IGNEUM_SCRATCH_SLOTS 2048u +#define IGNEUM_SCRATCH_WORDS_PER_LANE 8192u +#define IGNEUM_SCRATCH_BYTES_PER_WARP 1048576u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/scr0/program.json b/proto-cuda/packs-readwidth/scr0/program.json new file mode 100644 index 000000000..ad5bdb8c0 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr0/program.json @@ -0,0 +1,129 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x2f098ae568f38029", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-genesis", + "seed_bytes": "69676e65756d2d67656e65736973", + "seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "scr0", + "load_slots": 16, + "load_mix_percent_4_16_64": [100, 0, 0], + "load_width_counts_4_16_64": [16, 0, 0], + "bytes_per_hash": 512, + "scratch_ops_per_hash": 0, + "scratch": "variant 5 (measurement only): persistent warps; a 1 MiB scratch per warp of 2048 16-byte slots per lane (lane-major); slot = src & 0x7ff; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "add": 9, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "shfl": 4, "rotr": 2, "rotl": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2, "width": 1}, + {"i": 1, "op": "add", "dst": 1, "src": 7, "src2": 2, "imm": "0x42da7657", "imm2": "0xc3bd2355", "rot": 25, "bit": 4, "mask": 16, "width": 1}, + {"i": 2, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x61f0b51c", "imm2": "0x2735a174", "rot": 4, "bit": 26, "mask": 2, "width": 1}, + {"i": 3, "op": "mad", "dst": 4, "src": 0, "src2": 6, "imm": "0x679648a8", "imm2": "0x3044ba32", "rot": 31, "bit": 31, "mask": 4, "width": 1}, + {"i": 4, "op": "load", "dst": 7, "src": 2, "src2": 6, "imm": "0x5d1ca2a2", "imm2": "0xe2481807", "rot": 24, "bit": 3, "mask": 1, "width": 1}, + {"i": 5, "op": "load", "dst": 4, "src": 1, "src2": 2, "imm": "0x987c017a", "imm2": "0xf4d60559", "rot": 2, "bit": 0, "mask": 4, "width": 1}, + {"i": 6, "op": "shfl", "dst": 6, "src": 3, "src2": 7, "imm": "0x6ea7b2df", "imm2": "0x9fce5071", "rot": 7, "bit": 15, "mask": 4, "width": 1}, + {"i": 7, "op": "shfl", "dst": 1, "src": 5, "src2": 1, "imm": "0x26a2ecde", "imm2": "0xfec6ad22", "rot": 15, "bit": 11, "mask": 8, "width": 1}, + {"i": 8, "op": "xor", "dst": 7, "src": 5, "src2": 2, "imm": "0xbe4b445c", "imm2": "0x17a5a9c7", "rot": 8, "bit": 8, "mask": 1, "width": 1}, + {"i": 9, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4, "width": 1}, + {"i": 10, "op": "or", "dst": 1, "src": 2, "src2": 3, "imm": "0x4e7dc10d", "imm2": "0x196d165c", "rot": 14, "bit": 27, "mask": 16, "width": 1}, + {"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 1, "imm": "0xc5c3b55d", "imm2": "0xec061424", "rot": 26, "bit": 27, "mask": 8, "width": 1}, + {"i": 12, "op": "or", "dst": 6, "src": 2, "src2": 3, "imm": "0x306542fe", "imm2": "0x1bb1b429", "rot": 31, "bit": 0, "mask": 2, "width": 1}, + {"i": 13, "op": "mul", "dst": 2, "src": 5, "src2": 6, "imm": "0xa672cdd3", "imm2": "0x59a4829c", "rot": 22, "bit": 13, "mask": 16, "width": 1}, + {"i": 14, "op": "load", "dst": 1, "src": 2, "src2": 5, "imm": "0x028b4d37", "imm2": "0x7bbd78ea", "rot": 15, "bit": 2, "mask": 8, "width": 1}, + {"i": 15, "op": "rotl", "dst": 7, "src": 6, "src2": 6, "imm": "0x5c88a1a7", "imm2": "0x5c628769", "rot": 1, "bit": 3, "mask": 8, "width": 1}, + {"i": 16, "op": "load", "dst": 3, "src": 6, "src2": 7, "imm": "0xbac2ae81", "imm2": "0xcbbc7bdb", "rot": 18, "bit": 8, "mask": 8, "width": 1}, + {"i": 17, "op": "load", "dst": 7, "src": 4, "src2": 2, "imm": "0xe8ab93e9", "imm2": "0xa00de107", "rot": 2, "bit": 1, "mask": 16, "width": 1}, + {"i": 18, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2, "width": 1}, + {"i": 19, "op": "mad", "dst": 4, "src": 0, "src2": 2, "imm": "0x5fba7bc2", "imm2": "0xdf099cfb", "rot": 4, "bit": 15, "mask": 16, "width": 1}, + {"i": 20, "op": "shfl", "dst": 0, "src": 6, "src2": 3, "imm": "0x0a3056de", "imm2": "0x7f0c25c3", "rot": 27, "bit": 13, "mask": 8, "width": 1}, + {"i": 21, "op": "xor", "dst": 5, "src": 7, "src2": 4, "imm": "0xbd066e1d", "imm2": "0x6d3ddc5a", "rot": 2, "bit": 29, "mask": 1, "width": 1}, + {"i": 22, "op": "mulhi", "dst": 2, "src": 5, "src2": 0, "imm": "0xc7e9887a", "imm2": "0x19ec898f", "rot": 14, "bit": 9, "mask": 1, "width": 1}, + {"i": 23, "op": "load", "dst": 3, "src": 7, "src2": 2, "imm": "0xc7fcfc8f", "imm2": "0x8528b94f", "rot": 17, "bit": 13, "mask": 4, "width": 1}, + {"i": 24, "op": "mulhi", "dst": 7, "src": 3, "src2": 5, "imm": "0xd91641e8", "imm2": "0xaf77faf2", "rot": 22, "bit": 21, "mask": 1, "width": 1}, + {"i": 25, "op": "or", "dst": 5, "src": 4, "src2": 0, "imm": "0x84c03868", "imm2": "0xf6c691b7", "rot": 29, "bit": 14, "mask": 8, "width": 1}, + {"i": 26, "op": "mad", "dst": 4, "src": 5, "src2": 2, "imm": "0x3bb2b6ba", "imm2": "0x49d95fd5", "rot": 1, "bit": 5, "mask": 8, "width": 1}, + {"i": 27, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4, "width": 1}, + {"i": 28, "op": "mulhi", "dst": 6, "src": 7, "src2": 6, "imm": "0xd69c4715", "imm2": "0xe0ebc4ce", "rot": 29, "bit": 2, "mask": 8, "width": 1}, + {"i": 29, "op": "add", "dst": 6, "src": 1, "src2": 2, "imm": "0x3b2d2124", "imm2": "0x187a9128", "rot": 1, "bit": 9, "mask": 16, "width": 1}, + {"i": 30, "op": "rotr", "dst": 6, "src": 7, "src2": 0, "imm": "0x5c64a589", "imm2": "0x61c9a38d", "rot": 17, "bit": 21, "mask": 16, "width": 1}, + {"i": 31, "op": "load", "dst": 3, "src": 1, "src2": 7, "imm": "0xc37723fa", "imm2": "0xf3b024da", "rot": 16, "bit": 27, "mask": 16, "width": 1}, + {"i": 32, "op": "load", "dst": 1, "src": 0, "src2": 7, "imm": "0xcc7972c4", "imm2": "0xad098d15", "rot": 30, "bit": 21, "mask": 8, "width": 1}, + {"i": 33, "op": "add", "dst": 0, "src": 4, "src2": 4, "imm": "0x2c35699f", "imm2": "0x351dde38", "rot": 21, "bit": 18, "mask": 4, "width": 1}, + {"i": 34, "op": "load", "dst": 0, "src": 2, "src2": 3, "imm": "0xfae8902b", "imm2": "0x5cd8306f", "rot": 5, "bit": 28, "mask": 16, "width": 1}, + {"i": 35, "op": "mul", "dst": 0, "src": 3, "src2": 1, "imm": "0x4fa3f3db", "imm2": "0xdbf37e75", "rot": 7, "bit": 18, "mask": 4, "width": 1}, + {"i": 36, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2, "width": 1}, + {"i": 37, "op": "load", "dst": 4, "src": 0, "src2": 0, "imm": "0x04cc1d55", "imm2": "0x35c52d04", "rot": 11, "bit": 14, "mask": 2, "width": 1}, + {"i": 38, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0x3958f280", "imm2": "0x8713c7e1", "rot": 5, "bit": 23, "mask": 16, "width": 1}, + {"i": 39, "op": "add", "dst": 0, "src": 3, "src2": 3, "imm": "0xa907b90b", "imm2": "0x1b053acf", "rot": 30, "bit": 25, "mask": 16, "width": 1}, + {"i": 40, "op": "rotr", "dst": 2, "src": 5, "src2": 4, "imm": "0xf8662282", "imm2": "0x10bb9e30", "rot": 8, "bit": 6, "mask": 2, "width": 1}, + {"i": 41, "op": "mul", "dst": 3, "src": 2, "src2": 4, "imm": "0x49087d74", "imm2": "0x6348b489", "rot": 17, "bit": 9, "mask": 16, "width": 1}, + {"i": 42, "op": "add", "dst": 1, "src": 5, "src2": 1, "imm": "0xa32e000c", "imm2": "0x6058c2e3", "rot": 25, "bit": 20, "mask": 8, "width": 1}, + {"i": 43, "op": "xor", "dst": 3, "src": 4, "src2": 2, "imm": "0x3dad0eb6", "imm2": "0xb97578cb", "rot": 3, "bit": 27, "mask": 1, "width": 1}, + {"i": 44, "op": "load", "dst": 3, "src": 5, "src2": 7, "imm": "0x374aec92", "imm2": "0x626f11df", "rot": 20, "bit": 18, "mask": 8, "width": 1}, + {"i": 45, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4, "width": 1}, + {"i": 46, "op": "xor", "dst": 7, "src": 1, "src2": 0, "imm": "0xef6ac348", "imm2": "0x963bb7e6", "rot": 26, "bit": 3, "mask": 8, "width": 1}, + {"i": 47, "op": "add", "dst": 0, "src": 3, "src2": 0, "imm": "0x838b5065", "imm2": "0x36360066", "rot": 3, "bit": 31, "mask": 4, "width": 1}, + {"i": 48, "op": "mulhi", "dst": 7, "src": 5, "src2": 0, "imm": "0x8458f7ac", "imm2": "0xc1c15026", "rot": 27, "bit": 15, "mask": 8, "width": 1}, + {"i": 49, "op": "load", "dst": 0, "src": 2, "src2": 4, "imm": "0x636a9dc4", "imm2": "0xac023d9b", "rot": 22, "bit": 29, "mask": 1, "width": 1}, + {"i": 50, "op": "sub", "dst": 2, "src": 6, "src2": 0, "imm": "0x2baec8c9", "imm2": "0x4390f156", "rot": 3, "bit": 12, "mask": 8, "width": 1}, + {"i": 51, "op": "sub", "dst": 7, "src": 5, "src2": 7, "imm": "0x19234061", "imm2": "0xe84dfade", "rot": 4, "bit": 19, "mask": 1, "width": 1}, + {"i": 52, "op": "xor", "dst": 2, "src": 3, "src2": 5, "imm": "0xdc2cd71e", "imm2": "0x1b5d334b", "rot": 9, "bit": 8, "mask": 8, "width": 1}, + {"i": 53, "op": "sub", "dst": 7, "src": 0, "src2": 4, "imm": "0x605c31ec", "imm2": "0x9923ff88", "rot": 28, "bit": 25, "mask": 4, "width": 1}, + {"i": 54, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4, "width": 1}, + {"i": 55, "op": "xor", "dst": 7, "src": 5, "src2": 5, "imm": "0xad7493e7", "imm2": "0x3e400372", "rot": 13, "bit": 8, "mask": 1, "width": 1}, + {"i": 56, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0x87e933c9", "imm2": "0x8c854c1b", "rot": 17, "bit": 3, "mask": 8, "width": 1}, + {"i": 57, "op": "sub", "dst": 5, "src": 6, "src2": 5, "imm": "0x11be3bc9", "imm2": "0xbbaa8e24", "rot": 6, "bit": 5, "mask": 16, "width": 1}, + {"i": 58, "op": "load", "dst": 1, "src": 3, "src2": 2, "imm": "0xa732351a", "imm2": "0xc01349cd", "rot": 14, "bit": 17, "mask": 16, "width": 1}, + {"i": 59, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0xb20547b2", "imm2": "0xc94655de", "rot": 27, "bit": 30, "mask": 1, "width": 1}, + {"i": 60, "op": "sub", "dst": 4, "src": 6, "src2": 7, "imm": "0x67cf904c", "imm2": "0x6873b216", "rot": 27, "bit": 7, "mask": 16, "width": 1}, + {"i": 61, "op": "mul", "dst": 1, "src": 2, "src2": 7, "imm": "0x93ab0bf4", "imm2": "0x96158375", "rot": 14, "bit": 0, "mask": 16, "width": 1}, + {"i": 62, "op": "mad", "dst": 3, "src": 6, "src2": 0, "imm": "0x41a443a3", "imm2": "0xe69d7919", "rot": 9, "bit": 0, "mask": 16, "width": 1}, + {"i": 63, "op": "add", "dst": 0, "src": 1, "src2": 3, "imm": "0x2fe0e98b", "imm2": "0xc88e2942", "rot": 5, "bit": 16, "mask": 16, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/scr0/program.metal b/proto-cuda/packs-readwidth/scr0/program.metal new file mode 100644 index 000000000..9835a937b --- /dev/null +++ b/proto-cuda/packs-readwidth/scr0/program.metal @@ -0,0 +1,126 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + device uint* scratch [[buffer(3)]], + constant uint& groups [[buffer(4)]], + constant uint& salt [[buffer(5)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + r7 = r7 ^ dataset[r2 & MASK]; // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + r3 = r3 ^ dataset[r6 & MASK]; // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + r3 = r3 ^ dataset[r7 & MASK]; // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + r1 = r1 ^ dataset[r0 & MASK]; // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + r0 = r0 ^ dataset[r2 & MASK]; // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + r4 = r4 ^ dataset[r0 & MASK]; // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + r3 = r3 ^ dataset[r5 & MASK]; // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + r1 = r1 ^ dataset[r4 & MASK]; // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr0/program_bound.metal b/proto-cuda/packs-readwidth/scr0/program_bound.metal new file mode 100644 index 000000000..fbfa7581b --- /dev/null +++ b/proto-cuda/packs-readwidth/scr0/program_bound.metal @@ -0,0 +1,128 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + device uint* scratch [[buffer(4)]], + constant uint& groups [[buffer(5)]], + constant uint& salt [[buffer(6)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + r7 = r7 ^ dataset[r2 & MASK]; // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + r3 = r3 ^ dataset[r6 & MASK]; // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + r3 = r3 ^ dataset[r7 & MASK]; // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + r1 = r1 ^ dataset[r0 & MASK]; // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + r0 = r0 ^ dataset[r2 & MASK]; // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + r4 = r4 ^ dataset[r0 & MASK]; // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + r3 = r3 ^ dataset[r5 & MASK]; // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + r1 = r1 ^ dataset[r4 & MASK]; // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr0/vectors.h b/proto-cuda/packs-readwidth/scr0/vectors.h new file mode 100644 index 000000000..629a6ecf1 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr0/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0xebc45fe35d0c5a21ull, 0x3916ee6a80fa5ba3ull, 0x7b57ffdfddba6d6bull, 0x8e200453850ee32cull, 0x7e4f2c67c580c008ull, 0xecf909597ffcf8f0ull, 0x806504370b4b6f01ull, 0xebecae9eaf55ce28ull, + 0xb6a3f9ad8b562282ull, 0xded5cef894fcf4a6ull, 0xa3438b87979ccf0bull, 0x2113b24404657113ull, 0xaf8a742c58164ca9ull, 0x65e773520b077004ull, 0xfd88561d07b3bd47ull, 0x35c14ad5e3ebc66eull, + 0xf907b0620fb66d47ull, 0xdae2e4b9f15b8b58ull, 0x8067e2b229867642ull, 0x9c69c587f463a385ull, 0x5c68af10704faeecull, 0x234d5f34dc11b782ull, 0x09dc00c7503bde86ull, 0x91401126f3da1072ull, + 0xf7a53d7e9cf11856ull, 0xc70d5b84f0fc4036ull, 0x62e509b6c9c9716eull, 0x4dc509d6888fbddaull, 0xac89daf888a6107eull, 0x0006d56934f5dbd6ull, 0xb9bee3c866a9df21ull, 0x21a61101a1b5a09full + }, + { // base nonce 4096 + 0xa53fa2af90595891ull, 0x79fec7bbf23956d9ull, 0x44b8b0df7d734784ull, 0xa61f407dec8561b6ull, 0xb6c833abe86b46afull, 0x1593b3125e926771ull, 0xa89a078f614261b9ull, 0x9fea8ad57bf9ec0eull, + 0xa461b8b079972418ull, 0xcb440cecdea5dacaull, 0x87a45183a598b312ull, 0x74cfca514c3c967cull, 0xcc894f0d7b7278b7ull, 0x7cf1ade93c5abf9eull, 0xddba95279ec571eaull, 0x202ea62a834f6a98ull, + 0x2bb8befca6688eb3ull, 0x449d6a53cba73a5full, 0x5de1012931ccf785ull, 0x613e75ca80e3c1bdull, 0x326759818c934daeull, 0x5c79544c5e4d51f2ull, 0x4b1db22556527e4bull, 0xa5a07e12093a1ca4ull, + 0xce534c73c24279e8ull, 0xab4ca540546295ecull, 0x640569a3111fddf3ull, 0x782d96b811744f48ull, 0xcd473bed4cb7c7baull, 0xa1446a3f1f0dda81ull, 0xbfdecb9439fbe0f7ull, 0x6509222e4fc011ccull + }, + { // base nonce 1000000 + 0x16bd0d2e158a9f05ull, 0x2124042da743d42aull, 0xd8f06d223cdace97ull, 0xd62d36756a072f2dull, 0x2c3a9202512802aeull, 0x7af8f5508a349b0bull, 0x8167e84acf74e305ull, 0xbbdba57b508376cbull, + 0xd557e50340318fe8ull, 0x9b03054615d2b5b7ull, 0x0ce96b5c522cf6c1ull, 0xc7bafbcd73334cfdull, 0xc304b2f24b8afce8ull, 0x9ad58fa19c6fbf9eull, 0x5ae991b005be9f67ull, 0xf6a6dd02b0cb96e8ull, + 0xc183b6e2a21911daull, 0x2ae64b0e9fea5270ull, 0xaaadb378d5500fccull, 0x663cb59d6d5d1a56ull, 0xf9a433e3539e1345ull, 0xa8636b3e29e507b4ull, 0x06ad3fdd4f035a29ull, 0x959d7b0f3c345b88ull, + 0x67a959a794c4f5cdull, 0xa6d0ccd0fc27a3c7ull, 0x38cbaf6b6e4627cdull, 0xf0e552767a0c79a6ull, 0x64052dfffb7d6863ull, 0x761ef6c8d379da60ull, 0x2314d09da3c98751ull, 0x8972ceca84eff258ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/scr0/vectors.json b/proto-cuda/packs-readwidth/scr0/vectors.json new file mode 100644 index 000000000..d0e0a8039 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr0/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-genesis", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0xebc45fe35d0c5a21", "0x3916ee6a80fa5ba3", "0x7b57ffdfddba6d6b", "0x8e200453850ee32c", "0x7e4f2c67c580c008", "0xecf909597ffcf8f0", "0x806504370b4b6f01", "0xebecae9eaf55ce28", + "0xb6a3f9ad8b562282", "0xded5cef894fcf4a6", "0xa3438b87979ccf0b", "0x2113b24404657113", "0xaf8a742c58164ca9", "0x65e773520b077004", "0xfd88561d07b3bd47", "0x35c14ad5e3ebc66e", + "0xf907b0620fb66d47", "0xdae2e4b9f15b8b58", "0x8067e2b229867642", "0x9c69c587f463a385", "0x5c68af10704faeec", "0x234d5f34dc11b782", "0x09dc00c7503bde86", "0x91401126f3da1072", + "0xf7a53d7e9cf11856", "0xc70d5b84f0fc4036", "0x62e509b6c9c9716e", "0x4dc509d6888fbdda", "0xac89daf888a6107e", "0x0006d56934f5dbd6", "0xb9bee3c866a9df21", "0x21a61101a1b5a09f" + ]}, + {"base_nonce": 4096, "expected": [ + "0xa53fa2af90595891", "0x79fec7bbf23956d9", "0x44b8b0df7d734784", "0xa61f407dec8561b6", "0xb6c833abe86b46af", "0x1593b3125e926771", "0xa89a078f614261b9", "0x9fea8ad57bf9ec0e", + "0xa461b8b079972418", "0xcb440cecdea5daca", "0x87a45183a598b312", "0x74cfca514c3c967c", "0xcc894f0d7b7278b7", "0x7cf1ade93c5abf9e", "0xddba95279ec571ea", "0x202ea62a834f6a98", + "0x2bb8befca6688eb3", "0x449d6a53cba73a5f", "0x5de1012931ccf785", "0x613e75ca80e3c1bd", "0x326759818c934dae", "0x5c79544c5e4d51f2", "0x4b1db22556527e4b", "0xa5a07e12093a1ca4", + "0xce534c73c24279e8", "0xab4ca540546295ec", "0x640569a3111fddf3", "0x782d96b811744f48", "0xcd473bed4cb7c7ba", "0xa1446a3f1f0dda81", "0xbfdecb9439fbe0f7", "0x6509222e4fc011cc" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x16bd0d2e158a9f05", "0x2124042da743d42a", "0xd8f06d223cdace97", "0xd62d36756a072f2d", "0x2c3a9202512802ae", "0x7af8f5508a349b0b", "0x8167e84acf74e305", "0xbbdba57b508376cb", + "0xd557e50340318fe8", "0x9b03054615d2b5b7", "0x0ce96b5c522cf6c1", "0xc7bafbcd73334cfd", "0xc304b2f24b8afce8", "0x9ad58fa19c6fbf9e", "0x5ae991b005be9f67", "0xf6a6dd02b0cb96e8", + "0xc183b6e2a21911da", "0x2ae64b0e9fea5270", "0xaaadb378d5500fcc", "0x663cb59d6d5d1a56", "0xf9a433e3539e1345", "0xa8636b3e29e507b4", "0x06ad3fdd4f035a29", "0x959d7b0f3c345b88", + "0x67a959a794c4f5cd", "0xa6d0ccd0fc27a3c7", "0x38cbaf6b6e4627cd", "0xf0e552767a0c79a6", "0x64052dfffb7d6863", "0x761ef6c8d379da60", "0x2314d09da3c98751", "0x8972ceca84eff258" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/scr2/kernel.cl b/proto-cuda/packs-readwidth/scr2/kernel.cl new file mode 100644 index 000000000..b7d17259d --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2/kernel.cl @@ -0,0 +1,291 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + r1 = r1 ^ ds[r4 & mask]; // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/scr2/kernel.cu b/proto-cuda/packs-readwidth/scr2/kernel.cu new file mode 100644 index 000000000..b1b566dea --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2/kernel.cu @@ -0,0 +1,177 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t s_ = r6 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + r1 = r1 ^ ds[r4 & mask]; // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr2/kernel_bound.cl b/proto-cuda/packs-readwidth/scr2/kernel_bound.cl new file mode 100644 index 000000000..9889162fc --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2/kernel_bound.cl @@ -0,0 +1,393 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + r1 = r1 ^ ds[r4 & mask]; // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + r1 = r1 ^ ds[r4 & mask]; // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr2/kernel_bound.cu b/proto-cuda/packs-readwidth/scr2/kernel_bound.cu new file mode 100644 index 000000000..b012ad62e --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2/kernel_bound.cu @@ -0,0 +1,136 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t s_ = r6 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + r1 = r1 ^ ds[r4 & mask]; // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr2/memhard.h b/proto-cuda/packs-readwidth/scr2/memhard.h new file mode 100644 index 000000000..4803d8e40 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/scr2/memhard.metal b/proto-cuda/packs-readwidth/scr2/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/scr2/program.h b/proto-cuda/packs-readwidth/scr2/program.h new file mode 100644 index 000000000..4892b7988 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2/program.h @@ -0,0 +1,67 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-genesis" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x2f0988e568f37cc3ull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=14 add=9 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 shfl=4 rotr=2 scratch=2 rotl=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "scr2" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 100, 0, 0 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 14, 0, 0 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 448 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// Variant 5: persistent warps, a 1 MiB scratch per launched warp (the host launches N warps and passes scratch, +// groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). +#define IGNEUM_PERSISTENT_WARPS 1 +#define IGNEUM_SCRATCH_OPS 2 // scratch read-modify-writes per program (16 per hash) +#define IGNEUM_SCRATCH_SLOTS 2048u +#define IGNEUM_SCRATCH_WORDS_PER_LANE 8192u +#define IGNEUM_SCRATCH_BYTES_PER_WARP 1048576u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/scr2/program.json b/proto-cuda/packs-readwidth/scr2/program.json new file mode 100644 index 000000000..b9b6b935c --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2/program.json @@ -0,0 +1,129 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x2f0988e568f37cc3", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-genesis", + "seed_bytes": "69676e65756d2d67656e65736973", + "seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "scr2", + "load_slots": 16, + "load_mix_percent_4_16_64": [100, 0, 0], + "load_width_counts_4_16_64": [14, 0, 0], + "bytes_per_hash": 448, + "scratch_ops_per_hash": 16, + "scratch": "variant 5 (measurement only): persistent warps; a 1 MiB scratch per warp of 2048 16-byte slots per lane (lane-major); slot = src & 0x7ff; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 14, "add": 9, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "shfl": 4, "rotr": 2, "scratch": 2, "rotl": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2, "width": 1}, + {"i": 1, "op": "add", "dst": 1, "src": 7, "src2": 2, "imm": "0x42da7657", "imm2": "0xc3bd2355", "rot": 25, "bit": 4, "mask": 16, "width": 1}, + {"i": 2, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x61f0b51c", "imm2": "0x2735a174", "rot": 4, "bit": 26, "mask": 2, "width": 1}, + {"i": 3, "op": "mad", "dst": 4, "src": 0, "src2": 6, "imm": "0x679648a8", "imm2": "0x3044ba32", "rot": 31, "bit": 31, "mask": 4, "width": 1}, + {"i": 4, "op": "load", "dst": 7, "src": 2, "src2": 6, "imm": "0x5d1ca2a2", "imm2": "0xe2481807", "rot": 24, "bit": 3, "mask": 1, "width": 1}, + {"i": 5, "op": "load", "dst": 4, "src": 1, "src2": 2, "imm": "0x987c017a", "imm2": "0xf4d60559", "rot": 2, "bit": 0, "mask": 4, "width": 1}, + {"i": 6, "op": "shfl", "dst": 6, "src": 3, "src2": 7, "imm": "0x6ea7b2df", "imm2": "0x9fce5071", "rot": 7, "bit": 15, "mask": 4, "width": 1}, + {"i": 7, "op": "shfl", "dst": 1, "src": 5, "src2": 1, "imm": "0x26a2ecde", "imm2": "0xfec6ad22", "rot": 15, "bit": 11, "mask": 8, "width": 1}, + {"i": 8, "op": "xor", "dst": 7, "src": 5, "src2": 2, "imm": "0xbe4b445c", "imm2": "0x17a5a9c7", "rot": 8, "bit": 8, "mask": 1, "width": 1}, + {"i": 9, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4, "width": 1}, + {"i": 10, "op": "or", "dst": 1, "src": 2, "src2": 3, "imm": "0x4e7dc10d", "imm2": "0x196d165c", "rot": 14, "bit": 27, "mask": 16, "width": 1}, + {"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 1, "imm": "0xc5c3b55d", "imm2": "0xec061424", "rot": 26, "bit": 27, "mask": 8, "width": 1}, + {"i": 12, "op": "or", "dst": 6, "src": 2, "src2": 3, "imm": "0x306542fe", "imm2": "0x1bb1b429", "rot": 31, "bit": 0, "mask": 2, "width": 1}, + {"i": 13, "op": "mul", "dst": 2, "src": 5, "src2": 6, "imm": "0xa672cdd3", "imm2": "0x59a4829c", "rot": 22, "bit": 13, "mask": 16, "width": 1}, + {"i": 14, "op": "load", "dst": 1, "src": 2, "src2": 5, "imm": "0x028b4d37", "imm2": "0x7bbd78ea", "rot": 15, "bit": 2, "mask": 8, "width": 1}, + {"i": 15, "op": "rotl", "dst": 7, "src": 6, "src2": 6, "imm": "0x5c88a1a7", "imm2": "0x5c628769", "rot": 1, "bit": 3, "mask": 8, "width": 1}, + {"i": 16, "op": "scratch", "dst": 3, "src": 6, "src2": 7, "imm": "0xbac2ae81", "imm2": "0xcbbc7bdb", "rot": 18, "bit": 8, "mask": 8, "width": 1}, + {"i": 17, "op": "load", "dst": 7, "src": 4, "src2": 2, "imm": "0xe8ab93e9", "imm2": "0xa00de107", "rot": 2, "bit": 1, "mask": 16, "width": 1}, + {"i": 18, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2, "width": 1}, + {"i": 19, "op": "mad", "dst": 4, "src": 0, "src2": 2, "imm": "0x5fba7bc2", "imm2": "0xdf099cfb", "rot": 4, "bit": 15, "mask": 16, "width": 1}, + {"i": 20, "op": "shfl", "dst": 0, "src": 6, "src2": 3, "imm": "0x0a3056de", "imm2": "0x7f0c25c3", "rot": 27, "bit": 13, "mask": 8, "width": 1}, + {"i": 21, "op": "xor", "dst": 5, "src": 7, "src2": 4, "imm": "0xbd066e1d", "imm2": "0x6d3ddc5a", "rot": 2, "bit": 29, "mask": 1, "width": 1}, + {"i": 22, "op": "mulhi", "dst": 2, "src": 5, "src2": 0, "imm": "0xc7e9887a", "imm2": "0x19ec898f", "rot": 14, "bit": 9, "mask": 1, "width": 1}, + {"i": 23, "op": "load", "dst": 3, "src": 7, "src2": 2, "imm": "0xc7fcfc8f", "imm2": "0x8528b94f", "rot": 17, "bit": 13, "mask": 4, "width": 1}, + {"i": 24, "op": "mulhi", "dst": 7, "src": 3, "src2": 5, "imm": "0xd91641e8", "imm2": "0xaf77faf2", "rot": 22, "bit": 21, "mask": 1, "width": 1}, + {"i": 25, "op": "or", "dst": 5, "src": 4, "src2": 0, "imm": "0x84c03868", "imm2": "0xf6c691b7", "rot": 29, "bit": 14, "mask": 8, "width": 1}, + {"i": 26, "op": "mad", "dst": 4, "src": 5, "src2": 2, "imm": "0x3bb2b6ba", "imm2": "0x49d95fd5", "rot": 1, "bit": 5, "mask": 8, "width": 1}, + {"i": 27, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4, "width": 1}, + {"i": 28, "op": "mulhi", "dst": 6, "src": 7, "src2": 6, "imm": "0xd69c4715", "imm2": "0xe0ebc4ce", "rot": 29, "bit": 2, "mask": 8, "width": 1}, + {"i": 29, "op": "add", "dst": 6, "src": 1, "src2": 2, "imm": "0x3b2d2124", "imm2": "0x187a9128", "rot": 1, "bit": 9, "mask": 16, "width": 1}, + {"i": 30, "op": "rotr", "dst": 6, "src": 7, "src2": 0, "imm": "0x5c64a589", "imm2": "0x61c9a38d", "rot": 17, "bit": 21, "mask": 16, "width": 1}, + {"i": 31, "op": "load", "dst": 3, "src": 1, "src2": 7, "imm": "0xc37723fa", "imm2": "0xf3b024da", "rot": 16, "bit": 27, "mask": 16, "width": 1}, + {"i": 32, "op": "load", "dst": 1, "src": 0, "src2": 7, "imm": "0xcc7972c4", "imm2": "0xad098d15", "rot": 30, "bit": 21, "mask": 8, "width": 1}, + {"i": 33, "op": "add", "dst": 0, "src": 4, "src2": 4, "imm": "0x2c35699f", "imm2": "0x351dde38", "rot": 21, "bit": 18, "mask": 4, "width": 1}, + {"i": 34, "op": "scratch", "dst": 0, "src": 2, "src2": 3, "imm": "0xfae8902b", "imm2": "0x5cd8306f", "rot": 5, "bit": 28, "mask": 16, "width": 1}, + {"i": 35, "op": "mul", "dst": 0, "src": 3, "src2": 1, "imm": "0x4fa3f3db", "imm2": "0xdbf37e75", "rot": 7, "bit": 18, "mask": 4, "width": 1}, + {"i": 36, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2, "width": 1}, + {"i": 37, "op": "load", "dst": 4, "src": 0, "src2": 0, "imm": "0x04cc1d55", "imm2": "0x35c52d04", "rot": 11, "bit": 14, "mask": 2, "width": 1}, + {"i": 38, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0x3958f280", "imm2": "0x8713c7e1", "rot": 5, "bit": 23, "mask": 16, "width": 1}, + {"i": 39, "op": "add", "dst": 0, "src": 3, "src2": 3, "imm": "0xa907b90b", "imm2": "0x1b053acf", "rot": 30, "bit": 25, "mask": 16, "width": 1}, + {"i": 40, "op": "rotr", "dst": 2, "src": 5, "src2": 4, "imm": "0xf8662282", "imm2": "0x10bb9e30", "rot": 8, "bit": 6, "mask": 2, "width": 1}, + {"i": 41, "op": "mul", "dst": 3, "src": 2, "src2": 4, "imm": "0x49087d74", "imm2": "0x6348b489", "rot": 17, "bit": 9, "mask": 16, "width": 1}, + {"i": 42, "op": "add", "dst": 1, "src": 5, "src2": 1, "imm": "0xa32e000c", "imm2": "0x6058c2e3", "rot": 25, "bit": 20, "mask": 8, "width": 1}, + {"i": 43, "op": "xor", "dst": 3, "src": 4, "src2": 2, "imm": "0x3dad0eb6", "imm2": "0xb97578cb", "rot": 3, "bit": 27, "mask": 1, "width": 1}, + {"i": 44, "op": "load", "dst": 3, "src": 5, "src2": 7, "imm": "0x374aec92", "imm2": "0x626f11df", "rot": 20, "bit": 18, "mask": 8, "width": 1}, + {"i": 45, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4, "width": 1}, + {"i": 46, "op": "xor", "dst": 7, "src": 1, "src2": 0, "imm": "0xef6ac348", "imm2": "0x963bb7e6", "rot": 26, "bit": 3, "mask": 8, "width": 1}, + {"i": 47, "op": "add", "dst": 0, "src": 3, "src2": 0, "imm": "0x838b5065", "imm2": "0x36360066", "rot": 3, "bit": 31, "mask": 4, "width": 1}, + {"i": 48, "op": "mulhi", "dst": 7, "src": 5, "src2": 0, "imm": "0x8458f7ac", "imm2": "0xc1c15026", "rot": 27, "bit": 15, "mask": 8, "width": 1}, + {"i": 49, "op": "load", "dst": 0, "src": 2, "src2": 4, "imm": "0x636a9dc4", "imm2": "0xac023d9b", "rot": 22, "bit": 29, "mask": 1, "width": 1}, + {"i": 50, "op": "sub", "dst": 2, "src": 6, "src2": 0, "imm": "0x2baec8c9", "imm2": "0x4390f156", "rot": 3, "bit": 12, "mask": 8, "width": 1}, + {"i": 51, "op": "sub", "dst": 7, "src": 5, "src2": 7, "imm": "0x19234061", "imm2": "0xe84dfade", "rot": 4, "bit": 19, "mask": 1, "width": 1}, + {"i": 52, "op": "xor", "dst": 2, "src": 3, "src2": 5, "imm": "0xdc2cd71e", "imm2": "0x1b5d334b", "rot": 9, "bit": 8, "mask": 8, "width": 1}, + {"i": 53, "op": "sub", "dst": 7, "src": 0, "src2": 4, "imm": "0x605c31ec", "imm2": "0x9923ff88", "rot": 28, "bit": 25, "mask": 4, "width": 1}, + {"i": 54, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4, "width": 1}, + {"i": 55, "op": "xor", "dst": 7, "src": 5, "src2": 5, "imm": "0xad7493e7", "imm2": "0x3e400372", "rot": 13, "bit": 8, "mask": 1, "width": 1}, + {"i": 56, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0x87e933c9", "imm2": "0x8c854c1b", "rot": 17, "bit": 3, "mask": 8, "width": 1}, + {"i": 57, "op": "sub", "dst": 5, "src": 6, "src2": 5, "imm": "0x11be3bc9", "imm2": "0xbbaa8e24", "rot": 6, "bit": 5, "mask": 16, "width": 1}, + {"i": 58, "op": "load", "dst": 1, "src": 3, "src2": 2, "imm": "0xa732351a", "imm2": "0xc01349cd", "rot": 14, "bit": 17, "mask": 16, "width": 1}, + {"i": 59, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0xb20547b2", "imm2": "0xc94655de", "rot": 27, "bit": 30, "mask": 1, "width": 1}, + {"i": 60, "op": "sub", "dst": 4, "src": 6, "src2": 7, "imm": "0x67cf904c", "imm2": "0x6873b216", "rot": 27, "bit": 7, "mask": 16, "width": 1}, + {"i": 61, "op": "mul", "dst": 1, "src": 2, "src2": 7, "imm": "0x93ab0bf4", "imm2": "0x96158375", "rot": 14, "bit": 0, "mask": 16, "width": 1}, + {"i": 62, "op": "mad", "dst": 3, "src": 6, "src2": 0, "imm": "0x41a443a3", "imm2": "0xe69d7919", "rot": 9, "bit": 0, "mask": 16, "width": 1}, + {"i": 63, "op": "add", "dst": 0, "src": 1, "src2": 3, "imm": "0x2fe0e98b", "imm2": "0xc88e2942", "rot": 5, "bit": 16, "mask": 16, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/scr2/program.metal b/proto-cuda/packs-readwidth/scr2/program.metal new file mode 100644 index 000000000..5c8835ec2 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2/program.metal @@ -0,0 +1,126 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + device uint* scratch [[buffer(3)]], + constant uint& groups [[buffer(4)]], + constant uint& salt [[buffer(5)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + r7 = r7 ^ dataset[r2 & MASK]; // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint s_ = r6 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + r3 = r3 ^ dataset[r7 & MASK]; // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + r1 = r1 ^ dataset[r0 & MASK]; // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + r4 = r4 ^ dataset[r0 & MASK]; // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + r3 = r3 ^ dataset[r5 & MASK]; // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + r1 = r1 ^ dataset[r4 & MASK]; // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr2/program_bound.metal b/proto-cuda/packs-readwidth/scr2/program_bound.metal new file mode 100644 index 000000000..b521473ef --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2/program_bound.metal @@ -0,0 +1,128 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + device uint* scratch [[buffer(4)]], + constant uint& groups [[buffer(5)]], + constant uint& salt [[buffer(6)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + r7 = r7 ^ dataset[r2 & MASK]; // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint s_ = r6 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + r3 = r3 ^ dataset[r7 & MASK]; // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + r1 = r1 ^ dataset[r0 & MASK]; // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + r4 = r4 ^ dataset[r0 & MASK]; // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + r3 = r3 ^ dataset[r5 & MASK]; // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + r1 = r1 ^ dataset[r4 & MASK]; // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr2/vectors.h b/proto-cuda/packs-readwidth/scr2/vectors.h new file mode 100644 index 000000000..820c17e88 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x2273e2732203e32aull, 0xa513d354bd107990ull, 0xe005b7515054c85full, 0x18a61b37b30cd1fbull, 0xa21d5b98e8d07e9cull, 0x5c24171a391d5ed0ull, 0x0f29e7583e1794b8ull, 0x7eca8a374d1a4f70ull, + 0x260ca011cd9ea10cull, 0xce1050798fce3d43ull, 0x9560d939dca19041ull, 0x8480a8440b80ecc3ull, 0xfaa99aac459b739eull, 0x7f083e72458e08abull, 0x78d876842f68672bull, 0x3b9bcf6275d3575cull, + 0x0256af61bdbf11b3ull, 0xefc6771cae646cbdull, 0xbc44f1c9f9f54d87ull, 0x6caedd783487eb7dull, 0x001b31fcb4fe0d4dull, 0x947a7ba1057e25b6ull, 0xb9e5a0204d68c22aull, 0x50489bed25d42661ull, + 0xb1019bff6d1057cdull, 0xd1442990562ce940ull, 0xcd986a47f98801dbull, 0x9c8796b6df23300full, 0xbd53ad05d2c877a9ull, 0xc95e863774a15b0aull, 0x132d8a91fb2fa67aull, 0x53dd38e8eadc8a24ull + }, + { // base nonce 4096 + 0x78c93312a03fb0eeull, 0x184fea638ec9b5fbull, 0x5687e8dcc4301dbfull, 0xed02c94f23681dfcull, 0x326d70162241ff6dull, 0x452017eb4ed2dfcfull, 0xc10b0e016f1e28c9ull, 0x691ce0cecf2a99baull, + 0x6c9506f34e0e63ceull, 0x447a98c2b7fdfa40ull, 0x07486b0e4055b2c9ull, 0x41781460bd47fd5cull, 0x01db316e35198291ull, 0xccd7e727f139a880ull, 0xdd7bd9efd16bf21cull, 0x8285d37966656366ull, + 0x383deade15fe0ecbull, 0x5fd64f5873c8e324ull, 0xad584cb6839c5e1dull, 0xbb842707fb5e9460ull, 0x4e8bc8f87978fcbdull, 0x18eb56f4a1fae881ull, 0x4c3b731a6b0c47a1ull, 0xda52cf9d69b252ebull, + 0xb5ff19b2b3eeb13eull, 0xe2595cea2afe42ddull, 0x3ff108424c9e6e38ull, 0x3a8a9e1995f359caull, 0x6a6b1da662cf2126ull, 0x54e684c127bb181full, 0x2018caa81f1a7d50ull, 0x33e94d2c92d9d148ull + }, + { // base nonce 1000000 + 0x04a41389bf3dfd3dull, 0xc509164def9207dfull, 0x4a8ffdbdf46e429dull, 0xff13bf0dc1b39aebull, 0xb852acc8e24133d7ull, 0x4bdd991ae56252acull, 0xa7739e74b3a054e9ull, 0xb4e36218d4b45fdcull, + 0x8bbd323155f5edc5ull, 0xb7b56a90659e7fd2ull, 0xdff7c495b7027480ull, 0xffa8adb5c0302b06ull, 0xe97d7967d89a5672ull, 0x0d0c2d4e6493926eull, 0xe9a5cda333cf2043ull, 0xdc95256d0986e5d8ull, + 0xd0dc211b811d6843ull, 0x68dfa3d0fb9a569bull, 0xa9e0028dfd9178c0ull, 0x4a36ca1fc40b20a9ull, 0xe7c765c5a735294bull, 0xf08954b015cb2628ull, 0xc69ee66ecf2740c5ull, 0xe3d01e899e46b089ull, + 0xc3558c74159c8603ull, 0x4c7aeb196bd01b04ull, 0x13c17119385f1910ull, 0xda7fca98e0989b8aull, 0x6f95baf340817945ull, 0x1af52756fd3afcabull, 0xb8eefc370bbe7e4bull, 0x94a55055b48bd4dbull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/scr2/vectors.json b/proto-cuda/packs-readwidth/scr2/vectors.json new file mode 100644 index 000000000..d30ae7d93 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-genesis", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x2273e2732203e32a", "0xa513d354bd107990", "0xe005b7515054c85f", "0x18a61b37b30cd1fb", "0xa21d5b98e8d07e9c", "0x5c24171a391d5ed0", "0x0f29e7583e1794b8", "0x7eca8a374d1a4f70", + "0x260ca011cd9ea10c", "0xce1050798fce3d43", "0x9560d939dca19041", "0x8480a8440b80ecc3", "0xfaa99aac459b739e", "0x7f083e72458e08ab", "0x78d876842f68672b", "0x3b9bcf6275d3575c", + "0x0256af61bdbf11b3", "0xefc6771cae646cbd", "0xbc44f1c9f9f54d87", "0x6caedd783487eb7d", "0x001b31fcb4fe0d4d", "0x947a7ba1057e25b6", "0xb9e5a0204d68c22a", "0x50489bed25d42661", + "0xb1019bff6d1057cd", "0xd1442990562ce940", "0xcd986a47f98801db", "0x9c8796b6df23300f", "0xbd53ad05d2c877a9", "0xc95e863774a15b0a", "0x132d8a91fb2fa67a", "0x53dd38e8eadc8a24" + ]}, + {"base_nonce": 4096, "expected": [ + "0x78c93312a03fb0ee", "0x184fea638ec9b5fb", "0x5687e8dcc4301dbf", "0xed02c94f23681dfc", "0x326d70162241ff6d", "0x452017eb4ed2dfcf", "0xc10b0e016f1e28c9", "0x691ce0cecf2a99ba", + "0x6c9506f34e0e63ce", "0x447a98c2b7fdfa40", "0x07486b0e4055b2c9", "0x41781460bd47fd5c", "0x01db316e35198291", "0xccd7e727f139a880", "0xdd7bd9efd16bf21c", "0x8285d37966656366", + "0x383deade15fe0ecb", "0x5fd64f5873c8e324", "0xad584cb6839c5e1d", "0xbb842707fb5e9460", "0x4e8bc8f87978fcbd", "0x18eb56f4a1fae881", "0x4c3b731a6b0c47a1", "0xda52cf9d69b252eb", + "0xb5ff19b2b3eeb13e", "0xe2595cea2afe42dd", "0x3ff108424c9e6e38", "0x3a8a9e1995f359ca", "0x6a6b1da662cf2126", "0x54e684c127bb181f", "0x2018caa81f1a7d50", "0x33e94d2c92d9d148" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x04a41389bf3dfd3d", "0xc509164def9207df", "0x4a8ffdbdf46e429d", "0xff13bf0dc1b39aeb", "0xb852acc8e24133d7", "0x4bdd991ae56252ac", "0xa7739e74b3a054e9", "0xb4e36218d4b45fdc", + "0x8bbd323155f5edc5", "0xb7b56a90659e7fd2", "0xdff7c495b7027480", "0xffa8adb5c0302b06", "0xe97d7967d89a5672", "0x0d0c2d4e6493926e", "0xe9a5cda333cf2043", "0xdc95256d0986e5d8", + "0xd0dc211b811d6843", "0x68dfa3d0fb9a569b", "0xa9e0028dfd9178c0", "0x4a36ca1fc40b20a9", "0xe7c765c5a735294b", "0xf08954b015cb2628", "0xc69ee66ecf2740c5", "0xe3d01e899e46b089", + "0xc3558c74159c8603", "0x4c7aeb196bd01b04", "0x13c17119385f1910", "0xda7fca98e0989b8a", "0x6f95baf340817945", "0x1af52756fd3afcab", "0xb8eefc370bbe7e4b", "0x94a55055b48bd4db" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/scr4/kernel.cl b/proto-cuda/packs-readwidth/scr4/kernel.cl new file mode 100644 index 000000000..a75793f6f --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4/kernel.cl @@ -0,0 +1,291 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/scr4/kernel.cu b/proto-cuda/packs-readwidth/scr4/kernel.cu new file mode 100644 index 000000000..fd9584208 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4/kernel.cu @@ -0,0 +1,177 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t s_ = r6 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint32_t s_ = r0 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint32_t s_ = r4 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr4/kernel_bound.cl b/proto-cuda/packs-readwidth/scr4/kernel_bound.cl new file mode 100644 index 000000000..e33afd4bc --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4/kernel_bound.cl @@ -0,0 +1,393 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr4/kernel_bound.cu b/proto-cuda/packs-readwidth/scr4/kernel_bound.cu new file mode 100644 index 000000000..a4d184a0a --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4/kernel_bound.cu @@ -0,0 +1,136 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t s_ = r6 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint32_t s_ = r0 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint32_t s_ = r4 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr4/memhard.h b/proto-cuda/packs-readwidth/scr4/memhard.h new file mode 100644 index 000000000..4803d8e40 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/scr4/memhard.metal b/proto-cuda/packs-readwidth/scr4/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/scr4/program.h b/proto-cuda/packs-readwidth/scr4/program.h new file mode 100644 index 000000000..f8f7a6823 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4/program.h @@ -0,0 +1,67 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-genesis" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x2f098ee568f386f5ull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=12 add=9 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 scratch=4 shfl=4 rotr=2 rotl=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "scr4" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 100, 0, 0 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 12, 0, 0 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 384 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// Variant 5: persistent warps, a 1 MiB scratch per launched warp (the host launches N warps and passes scratch, +// groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). +#define IGNEUM_PERSISTENT_WARPS 1 +#define IGNEUM_SCRATCH_OPS 4 // scratch read-modify-writes per program (32 per hash) +#define IGNEUM_SCRATCH_SLOTS 2048u +#define IGNEUM_SCRATCH_WORDS_PER_LANE 8192u +#define IGNEUM_SCRATCH_BYTES_PER_WARP 1048576u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/scr4/program.json b/proto-cuda/packs-readwidth/scr4/program.json new file mode 100644 index 000000000..61fe86581 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4/program.json @@ -0,0 +1,129 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x2f098ee568f386f5", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-genesis", + "seed_bytes": "69676e65756d2d67656e65736973", + "seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "scr4", + "load_slots": 16, + "load_mix_percent_4_16_64": [100, 0, 0], + "load_width_counts_4_16_64": [12, 0, 0], + "bytes_per_hash": 384, + "scratch_ops_per_hash": 32, + "scratch": "variant 5 (measurement only): persistent warps; a 1 MiB scratch per warp of 2048 16-byte slots per lane (lane-major); slot = src & 0x7ff; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 12, "add": 9, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "scratch": 4, "shfl": 4, "rotr": 2, "rotl": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2, "width": 1}, + {"i": 1, "op": "add", "dst": 1, "src": 7, "src2": 2, "imm": "0x42da7657", "imm2": "0xc3bd2355", "rot": 25, "bit": 4, "mask": 16, "width": 1}, + {"i": 2, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x61f0b51c", "imm2": "0x2735a174", "rot": 4, "bit": 26, "mask": 2, "width": 1}, + {"i": 3, "op": "mad", "dst": 4, "src": 0, "src2": 6, "imm": "0x679648a8", "imm2": "0x3044ba32", "rot": 31, "bit": 31, "mask": 4, "width": 1}, + {"i": 4, "op": "load", "dst": 7, "src": 2, "src2": 6, "imm": "0x5d1ca2a2", "imm2": "0xe2481807", "rot": 24, "bit": 3, "mask": 1, "width": 1}, + {"i": 5, "op": "load", "dst": 4, "src": 1, "src2": 2, "imm": "0x987c017a", "imm2": "0xf4d60559", "rot": 2, "bit": 0, "mask": 4, "width": 1}, + {"i": 6, "op": "shfl", "dst": 6, "src": 3, "src2": 7, "imm": "0x6ea7b2df", "imm2": "0x9fce5071", "rot": 7, "bit": 15, "mask": 4, "width": 1}, + {"i": 7, "op": "shfl", "dst": 1, "src": 5, "src2": 1, "imm": "0x26a2ecde", "imm2": "0xfec6ad22", "rot": 15, "bit": 11, "mask": 8, "width": 1}, + {"i": 8, "op": "xor", "dst": 7, "src": 5, "src2": 2, "imm": "0xbe4b445c", "imm2": "0x17a5a9c7", "rot": 8, "bit": 8, "mask": 1, "width": 1}, + {"i": 9, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4, "width": 1}, + {"i": 10, "op": "or", "dst": 1, "src": 2, "src2": 3, "imm": "0x4e7dc10d", "imm2": "0x196d165c", "rot": 14, "bit": 27, "mask": 16, "width": 1}, + {"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 1, "imm": "0xc5c3b55d", "imm2": "0xec061424", "rot": 26, "bit": 27, "mask": 8, "width": 1}, + {"i": 12, "op": "or", "dst": 6, "src": 2, "src2": 3, "imm": "0x306542fe", "imm2": "0x1bb1b429", "rot": 31, "bit": 0, "mask": 2, "width": 1}, + {"i": 13, "op": "mul", "dst": 2, "src": 5, "src2": 6, "imm": "0xa672cdd3", "imm2": "0x59a4829c", "rot": 22, "bit": 13, "mask": 16, "width": 1}, + {"i": 14, "op": "load", "dst": 1, "src": 2, "src2": 5, "imm": "0x028b4d37", "imm2": "0x7bbd78ea", "rot": 15, "bit": 2, "mask": 8, "width": 1}, + {"i": 15, "op": "rotl", "dst": 7, "src": 6, "src2": 6, "imm": "0x5c88a1a7", "imm2": "0x5c628769", "rot": 1, "bit": 3, "mask": 8, "width": 1}, + {"i": 16, "op": "scratch", "dst": 3, "src": 6, "src2": 7, "imm": "0xbac2ae81", "imm2": "0xcbbc7bdb", "rot": 18, "bit": 8, "mask": 8, "width": 1}, + {"i": 17, "op": "load", "dst": 7, "src": 4, "src2": 2, "imm": "0xe8ab93e9", "imm2": "0xa00de107", "rot": 2, "bit": 1, "mask": 16, "width": 1}, + {"i": 18, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2, "width": 1}, + {"i": 19, "op": "mad", "dst": 4, "src": 0, "src2": 2, "imm": "0x5fba7bc2", "imm2": "0xdf099cfb", "rot": 4, "bit": 15, "mask": 16, "width": 1}, + {"i": 20, "op": "shfl", "dst": 0, "src": 6, "src2": 3, "imm": "0x0a3056de", "imm2": "0x7f0c25c3", "rot": 27, "bit": 13, "mask": 8, "width": 1}, + {"i": 21, "op": "xor", "dst": 5, "src": 7, "src2": 4, "imm": "0xbd066e1d", "imm2": "0x6d3ddc5a", "rot": 2, "bit": 29, "mask": 1, "width": 1}, + {"i": 22, "op": "mulhi", "dst": 2, "src": 5, "src2": 0, "imm": "0xc7e9887a", "imm2": "0x19ec898f", "rot": 14, "bit": 9, "mask": 1, "width": 1}, + {"i": 23, "op": "load", "dst": 3, "src": 7, "src2": 2, "imm": "0xc7fcfc8f", "imm2": "0x8528b94f", "rot": 17, "bit": 13, "mask": 4, "width": 1}, + {"i": 24, "op": "mulhi", "dst": 7, "src": 3, "src2": 5, "imm": "0xd91641e8", "imm2": "0xaf77faf2", "rot": 22, "bit": 21, "mask": 1, "width": 1}, + {"i": 25, "op": "or", "dst": 5, "src": 4, "src2": 0, "imm": "0x84c03868", "imm2": "0xf6c691b7", "rot": 29, "bit": 14, "mask": 8, "width": 1}, + {"i": 26, "op": "mad", "dst": 4, "src": 5, "src2": 2, "imm": "0x3bb2b6ba", "imm2": "0x49d95fd5", "rot": 1, "bit": 5, "mask": 8, "width": 1}, + {"i": 27, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4, "width": 1}, + {"i": 28, "op": "mulhi", "dst": 6, "src": 7, "src2": 6, "imm": "0xd69c4715", "imm2": "0xe0ebc4ce", "rot": 29, "bit": 2, "mask": 8, "width": 1}, + {"i": 29, "op": "add", "dst": 6, "src": 1, "src2": 2, "imm": "0x3b2d2124", "imm2": "0x187a9128", "rot": 1, "bit": 9, "mask": 16, "width": 1}, + {"i": 30, "op": "rotr", "dst": 6, "src": 7, "src2": 0, "imm": "0x5c64a589", "imm2": "0x61c9a38d", "rot": 17, "bit": 21, "mask": 16, "width": 1}, + {"i": 31, "op": "load", "dst": 3, "src": 1, "src2": 7, "imm": "0xc37723fa", "imm2": "0xf3b024da", "rot": 16, "bit": 27, "mask": 16, "width": 1}, + {"i": 32, "op": "scratch", "dst": 1, "src": 0, "src2": 7, "imm": "0xcc7972c4", "imm2": "0xad098d15", "rot": 30, "bit": 21, "mask": 8, "width": 1}, + {"i": 33, "op": "add", "dst": 0, "src": 4, "src2": 4, "imm": "0x2c35699f", "imm2": "0x351dde38", "rot": 21, "bit": 18, "mask": 4, "width": 1}, + {"i": 34, "op": "scratch", "dst": 0, "src": 2, "src2": 3, "imm": "0xfae8902b", "imm2": "0x5cd8306f", "rot": 5, "bit": 28, "mask": 16, "width": 1}, + {"i": 35, "op": "mul", "dst": 0, "src": 3, "src2": 1, "imm": "0x4fa3f3db", "imm2": "0xdbf37e75", "rot": 7, "bit": 18, "mask": 4, "width": 1}, + {"i": 36, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2, "width": 1}, + {"i": 37, "op": "load", "dst": 4, "src": 0, "src2": 0, "imm": "0x04cc1d55", "imm2": "0x35c52d04", "rot": 11, "bit": 14, "mask": 2, "width": 1}, + {"i": 38, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0x3958f280", "imm2": "0x8713c7e1", "rot": 5, "bit": 23, "mask": 16, "width": 1}, + {"i": 39, "op": "add", "dst": 0, "src": 3, "src2": 3, "imm": "0xa907b90b", "imm2": "0x1b053acf", "rot": 30, "bit": 25, "mask": 16, "width": 1}, + {"i": 40, "op": "rotr", "dst": 2, "src": 5, "src2": 4, "imm": "0xf8662282", "imm2": "0x10bb9e30", "rot": 8, "bit": 6, "mask": 2, "width": 1}, + {"i": 41, "op": "mul", "dst": 3, "src": 2, "src2": 4, "imm": "0x49087d74", "imm2": "0x6348b489", "rot": 17, "bit": 9, "mask": 16, "width": 1}, + {"i": 42, "op": "add", "dst": 1, "src": 5, "src2": 1, "imm": "0xa32e000c", "imm2": "0x6058c2e3", "rot": 25, "bit": 20, "mask": 8, "width": 1}, + {"i": 43, "op": "xor", "dst": 3, "src": 4, "src2": 2, "imm": "0x3dad0eb6", "imm2": "0xb97578cb", "rot": 3, "bit": 27, "mask": 1, "width": 1}, + {"i": 44, "op": "load", "dst": 3, "src": 5, "src2": 7, "imm": "0x374aec92", "imm2": "0x626f11df", "rot": 20, "bit": 18, "mask": 8, "width": 1}, + {"i": 45, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4, "width": 1}, + {"i": 46, "op": "xor", "dst": 7, "src": 1, "src2": 0, "imm": "0xef6ac348", "imm2": "0x963bb7e6", "rot": 26, "bit": 3, "mask": 8, "width": 1}, + {"i": 47, "op": "add", "dst": 0, "src": 3, "src2": 0, "imm": "0x838b5065", "imm2": "0x36360066", "rot": 3, "bit": 31, "mask": 4, "width": 1}, + {"i": 48, "op": "mulhi", "dst": 7, "src": 5, "src2": 0, "imm": "0x8458f7ac", "imm2": "0xc1c15026", "rot": 27, "bit": 15, "mask": 8, "width": 1}, + {"i": 49, "op": "load", "dst": 0, "src": 2, "src2": 4, "imm": "0x636a9dc4", "imm2": "0xac023d9b", "rot": 22, "bit": 29, "mask": 1, "width": 1}, + {"i": 50, "op": "sub", "dst": 2, "src": 6, "src2": 0, "imm": "0x2baec8c9", "imm2": "0x4390f156", "rot": 3, "bit": 12, "mask": 8, "width": 1}, + {"i": 51, "op": "sub", "dst": 7, "src": 5, "src2": 7, "imm": "0x19234061", "imm2": "0xe84dfade", "rot": 4, "bit": 19, "mask": 1, "width": 1}, + {"i": 52, "op": "xor", "dst": 2, "src": 3, "src2": 5, "imm": "0xdc2cd71e", "imm2": "0x1b5d334b", "rot": 9, "bit": 8, "mask": 8, "width": 1}, + {"i": 53, "op": "sub", "dst": 7, "src": 0, "src2": 4, "imm": "0x605c31ec", "imm2": "0x9923ff88", "rot": 28, "bit": 25, "mask": 4, "width": 1}, + {"i": 54, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4, "width": 1}, + {"i": 55, "op": "xor", "dst": 7, "src": 5, "src2": 5, "imm": "0xad7493e7", "imm2": "0x3e400372", "rot": 13, "bit": 8, "mask": 1, "width": 1}, + {"i": 56, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0x87e933c9", "imm2": "0x8c854c1b", "rot": 17, "bit": 3, "mask": 8, "width": 1}, + {"i": 57, "op": "sub", "dst": 5, "src": 6, "src2": 5, "imm": "0x11be3bc9", "imm2": "0xbbaa8e24", "rot": 6, "bit": 5, "mask": 16, "width": 1}, + {"i": 58, "op": "load", "dst": 1, "src": 3, "src2": 2, "imm": "0xa732351a", "imm2": "0xc01349cd", "rot": 14, "bit": 17, "mask": 16, "width": 1}, + {"i": 59, "op": "scratch", "dst": 1, "src": 4, "src2": 0, "imm": "0xb20547b2", "imm2": "0xc94655de", "rot": 27, "bit": 30, "mask": 1, "width": 1}, + {"i": 60, "op": "sub", "dst": 4, "src": 6, "src2": 7, "imm": "0x67cf904c", "imm2": "0x6873b216", "rot": 27, "bit": 7, "mask": 16, "width": 1}, + {"i": 61, "op": "mul", "dst": 1, "src": 2, "src2": 7, "imm": "0x93ab0bf4", "imm2": "0x96158375", "rot": 14, "bit": 0, "mask": 16, "width": 1}, + {"i": 62, "op": "mad", "dst": 3, "src": 6, "src2": 0, "imm": "0x41a443a3", "imm2": "0xe69d7919", "rot": 9, "bit": 0, "mask": 16, "width": 1}, + {"i": 63, "op": "add", "dst": 0, "src": 1, "src2": 3, "imm": "0x2fe0e98b", "imm2": "0xc88e2942", "rot": 5, "bit": 16, "mask": 16, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/scr4/program.metal b/proto-cuda/packs-readwidth/scr4/program.metal new file mode 100644 index 000000000..4a8922643 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4/program.metal @@ -0,0 +1,126 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + device uint* scratch [[buffer(3)]], + constant uint& groups [[buffer(4)]], + constant uint& salt [[buffer(5)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + r7 = r7 ^ dataset[r2 & MASK]; // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint s_ = r6 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + r3 = r3 ^ dataset[r7 & MASK]; // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + { uint s_ = r0 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + r4 = r4 ^ dataset[r0 & MASK]; // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + r3 = r3 ^ dataset[r5 & MASK]; // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + { uint s_ = r4 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr4/program_bound.metal b/proto-cuda/packs-readwidth/scr4/program_bound.metal new file mode 100644 index 000000000..6235bce84 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4/program_bound.metal @@ -0,0 +1,128 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + device uint* scratch [[buffer(4)]], + constant uint& groups [[buffer(5)]], + constant uint& salt [[buffer(6)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + r7 = r7 ^ dataset[r2 & MASK]; // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint s_ = r6 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + r3 = r3 ^ dataset[r7 & MASK]; // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + { uint s_ = r0 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + r4 = r4 ^ dataset[r0 & MASK]; // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + r3 = r3 ^ dataset[r5 & MASK]; // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + { uint s_ = r4 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr4/vectors.h b/proto-cuda/packs-readwidth/scr4/vectors.h new file mode 100644 index 000000000..88076f2e3 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x66cffcc97c46e625ull, 0xbc9019f8df50fbfdull, 0x65629c90dde6016eull, 0xd647a41effa03d3bull, 0x86da3b6bbd751b99ull, 0x6ccf4240a0fb2d19ull, 0xeb39a1e06f17378cull, 0x2ea6b349b289fb10ull, + 0x6067211e6c220500ull, 0x6e6095dfedd1360full, 0xbd1190d8b50e1b48ull, 0x216dc72a0c08d5b5ull, 0x5be1f8c080836b0cull, 0x2a32932a5953ed73ull, 0xcc2a3d68be83c802ull, 0xe46daca15338278full, + 0xb2de43b96761e459ull, 0x9004acd06588cbeaull, 0x6a9a3543cf93004full, 0xff956d859cb6e408ull, 0x4397ec6e3c5fb045ull, 0x521dea569cd481d5ull, 0x89832b34108759f0ull, 0xf66e393836ffe4eaull, + 0xb4e39af6c40ea2f4ull, 0x3adc22085dd8d648ull, 0x27efe270958bbfbbull, 0x6c80be0e8dca60d8ull, 0xa0afbc6a60260d59ull, 0x5d9a257fb9189537ull, 0xeb837aeef55dc3edull, 0xd381174dc14f8951ull + }, + { // base nonce 4096 + 0xfb1f61aaeaeaef94ull, 0x184c2160963d8b57ull, 0x42a053c628625778ull, 0xeaad0e41c4770812ull, 0x1d6d389ceb462ce1ull, 0x4a639827672bdbd4ull, 0x857e42aa5a42dd6full, 0xb4ef399e5339979cull, + 0x497a29225b099233ull, 0x71d8b42862d81954ull, 0x0af995663313bf04ull, 0xf436fd126619d7a1ull, 0x199e4e3333cff269ull, 0x64077952f3775768ull, 0x51af1d126c5e8388ull, 0xffbaf44fe6b15cfdull, + 0xfc8fed86ecae34a7ull, 0x4cb548616f7a7d6bull, 0xc21d938c8b5bef35ull, 0x34789cbdd7088f71ull, 0xacb099a2c207d891ull, 0xfe1902d162374413ull, 0x26f7831c28f4020bull, 0xdf5192952b4af6b0ull, + 0xecab61fe88dbaff4ull, 0x941c491f7fdb86e5ull, 0x2b1900c53f746e77ull, 0x8c40507b1caffeb2ull, 0x7532a1ec2b9169efull, 0x1cf399b0c8bfb520ull, 0xdf003d2bb8a2cc0cull, 0x4da853307fc977a9ull + }, + { // base nonce 1000000 + 0x3d094bd04694b96full, 0xaecbd76cecd1a20aull, 0xbcb86febe56b17feull, 0x98082b557ba97517ull, 0xbb5f94108888564bull, 0xea3284877a30fc87ull, 0xc608fa4d5a8bb2adull, 0x946c721c511e0729ull, + 0x46c6eba292083aedull, 0x936cb97231eb6795ull, 0xb1413c434c712cbbull, 0xedfd554d3948c1bdull, 0xa8a20cbef2faccd5ull, 0x5fe39d756cadbcadull, 0x208b2627380791feull, 0xf52f9374ce480218ull, + 0xb9db7cd8814eb29eull, 0xf32ed2192b5a8719ull, 0x4f1b06a054940aefull, 0x406df498e4365eb5ull, 0x1982075caad345efull, 0x590f725623dbbbbdull, 0xa26d9192dedfefa5ull, 0x36219ec00da18980ull, + 0x5361d0dcb0f8b1a3ull, 0x35bdefa2fbb5ffc3ull, 0xba4c2a4e473a9c80ull, 0x107d9d3030f8b9d3ull, 0xa8bb094266d6b987ull, 0x86164fdfbb1426e8ull, 0xa6e8cb895021cbbdull, 0xfe12809e9d99a243ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/scr4/vectors.json b/proto-cuda/packs-readwidth/scr4/vectors.json new file mode 100644 index 000000000..a0f951221 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-genesis", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x66cffcc97c46e625", "0xbc9019f8df50fbfd", "0x65629c90dde6016e", "0xd647a41effa03d3b", "0x86da3b6bbd751b99", "0x6ccf4240a0fb2d19", "0xeb39a1e06f17378c", "0x2ea6b349b289fb10", + "0x6067211e6c220500", "0x6e6095dfedd1360f", "0xbd1190d8b50e1b48", "0x216dc72a0c08d5b5", "0x5be1f8c080836b0c", "0x2a32932a5953ed73", "0xcc2a3d68be83c802", "0xe46daca15338278f", + "0xb2de43b96761e459", "0x9004acd06588cbea", "0x6a9a3543cf93004f", "0xff956d859cb6e408", "0x4397ec6e3c5fb045", "0x521dea569cd481d5", "0x89832b34108759f0", "0xf66e393836ffe4ea", + "0xb4e39af6c40ea2f4", "0x3adc22085dd8d648", "0x27efe270958bbfbb", "0x6c80be0e8dca60d8", "0xa0afbc6a60260d59", "0x5d9a257fb9189537", "0xeb837aeef55dc3ed", "0xd381174dc14f8951" + ]}, + {"base_nonce": 4096, "expected": [ + "0xfb1f61aaeaeaef94", "0x184c2160963d8b57", "0x42a053c628625778", "0xeaad0e41c4770812", "0x1d6d389ceb462ce1", "0x4a639827672bdbd4", "0x857e42aa5a42dd6f", "0xb4ef399e5339979c", + "0x497a29225b099233", "0x71d8b42862d81954", "0x0af995663313bf04", "0xf436fd126619d7a1", "0x199e4e3333cff269", "0x64077952f3775768", "0x51af1d126c5e8388", "0xffbaf44fe6b15cfd", + "0xfc8fed86ecae34a7", "0x4cb548616f7a7d6b", "0xc21d938c8b5bef35", "0x34789cbdd7088f71", "0xacb099a2c207d891", "0xfe1902d162374413", "0x26f7831c28f4020b", "0xdf5192952b4af6b0", + "0xecab61fe88dbaff4", "0x941c491f7fdb86e5", "0x2b1900c53f746e77", "0x8c40507b1caffeb2", "0x7532a1ec2b9169ef", "0x1cf399b0c8bfb520", "0xdf003d2bb8a2cc0c", "0x4da853307fc977a9" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x3d094bd04694b96f", "0xaecbd76cecd1a20a", "0xbcb86febe56b17fe", "0x98082b557ba97517", "0xbb5f94108888564b", "0xea3284877a30fc87", "0xc608fa4d5a8bb2ad", "0x946c721c511e0729", + "0x46c6eba292083aed", "0x936cb97231eb6795", "0xb1413c434c712cbb", "0xedfd554d3948c1bd", "0xa8a20cbef2faccd5", "0x5fe39d756cadbcad", "0x208b2627380791fe", "0xf52f9374ce480218", + "0xb9db7cd8814eb29e", "0xf32ed2192b5a8719", "0x4f1b06a054940aef", "0x406df498e4365eb5", "0x1982075caad345ef", "0x590f725623dbbbbd", "0xa26d9192dedfefa5", "0x36219ec00da18980", + "0x5361d0dcb0f8b1a3", "0x35bdefa2fbb5ffc3", "0xba4c2a4e473a9c80", "0x107d9d3030f8b9d3", "0xa8bb094266d6b987", "0x86164fdfbb1426e8", "0xa6e8cb895021cbbd", "0xfe12809e9d99a243" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/scr8/kernel.cl b/proto-cuda/packs-readwidth/scr8/kernel.cl new file mode 100644 index 000000000..50015ca9b --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8/kernel.cl @@ -0,0 +1,291 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint s_ = r7 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 23 scratch + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint s_ = r5 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/scr8/kernel.cu b/proto-cuda/packs-readwidth/scr8/kernel.cu new file mode 100644 index 000000000..85212204b --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8/kernel.cu @@ -0,0 +1,177 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t s_ = r6 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + { uint32_t s_ = r7 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 scratch + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint32_t s_ = r0 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint32_t s_ = r0 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint32_t s_ = r5 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint32_t s_ = r4 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr8/kernel_bound.cl b/proto-cuda/packs-readwidth/scr8/kernel_bound.cl new file mode 100644 index 000000000..73b2b9c9f --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8/kernel_bound.cl @@ -0,0 +1,393 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint s_ = r7 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 23 scratch + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint s_ = r5 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint s_ = r7 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 23 scratch + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint s_ = r5 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr8/kernel_bound.cu b/proto-cuda/packs-readwidth/scr8/kernel_bound.cu new file mode 100644 index 000000000..13d40a9fc --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8/kernel_bound.cu @@ -0,0 +1,136 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t s_ = r6 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + { uint32_t s_ = r7 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 scratch + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint32_t s_ = r0 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint32_t s_ = r0 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint32_t s_ = r5 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint32_t s_ = r4 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr8/memhard.h b/proto-cuda/packs-readwidth/scr8/memhard.h new file mode 100644 index 000000000..4803d8e40 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/scr8/memhard.metal b/proto-cuda/packs-readwidth/scr8/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/scr8/program.h b/proto-cuda/packs-readwidth/scr8/program.h new file mode 100644 index 000000000..769c814b6 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8/program.h @@ -0,0 +1,67 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-genesis" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x2f0992e568f38dc1ull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "add=9 load=8 scratch=8 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 shfl=4 rotr=2 rotl=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "scr8" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 100, 0, 0 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 8, 0, 0 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 256 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// Variant 5: persistent warps, a 1 MiB scratch per launched warp (the host launches N warps and passes scratch, +// groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). +#define IGNEUM_PERSISTENT_WARPS 1 +#define IGNEUM_SCRATCH_OPS 8 // scratch read-modify-writes per program (64 per hash) +#define IGNEUM_SCRATCH_SLOTS 2048u +#define IGNEUM_SCRATCH_WORDS_PER_LANE 8192u +#define IGNEUM_SCRATCH_BYTES_PER_WARP 1048576u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/scr8/program.json b/proto-cuda/packs-readwidth/scr8/program.json new file mode 100644 index 000000000..25348b09f --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8/program.json @@ -0,0 +1,129 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x2f0992e568f38dc1", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-genesis", + "seed_bytes": "69676e65756d2d67656e65736973", + "seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "scr8", + "load_slots": 16, + "load_mix_percent_4_16_64": [100, 0, 0], + "load_width_counts_4_16_64": [8, 0, 0], + "bytes_per_hash": 256, + "scratch_ops_per_hash": 64, + "scratch": "variant 5 (measurement only): persistent warps; a 1 MiB scratch per warp of 2048 16-byte slots per lane (lane-major); slot = src & 0x7ff; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"add": 9, "load": 8, "scratch": 8, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "shfl": 4, "rotr": 2, "rotl": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2, "width": 1}, + {"i": 1, "op": "add", "dst": 1, "src": 7, "src2": 2, "imm": "0x42da7657", "imm2": "0xc3bd2355", "rot": 25, "bit": 4, "mask": 16, "width": 1}, + {"i": 2, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x61f0b51c", "imm2": "0x2735a174", "rot": 4, "bit": 26, "mask": 2, "width": 1}, + {"i": 3, "op": "mad", "dst": 4, "src": 0, "src2": 6, "imm": "0x679648a8", "imm2": "0x3044ba32", "rot": 31, "bit": 31, "mask": 4, "width": 1}, + {"i": 4, "op": "scratch", "dst": 7, "src": 2, "src2": 6, "imm": "0x5d1ca2a2", "imm2": "0xe2481807", "rot": 24, "bit": 3, "mask": 1, "width": 1}, + {"i": 5, "op": "load", "dst": 4, "src": 1, "src2": 2, "imm": "0x987c017a", "imm2": "0xf4d60559", "rot": 2, "bit": 0, "mask": 4, "width": 1}, + {"i": 6, "op": "shfl", "dst": 6, "src": 3, "src2": 7, "imm": "0x6ea7b2df", "imm2": "0x9fce5071", "rot": 7, "bit": 15, "mask": 4, "width": 1}, + {"i": 7, "op": "shfl", "dst": 1, "src": 5, "src2": 1, "imm": "0x26a2ecde", "imm2": "0xfec6ad22", "rot": 15, "bit": 11, "mask": 8, "width": 1}, + {"i": 8, "op": "xor", "dst": 7, "src": 5, "src2": 2, "imm": "0xbe4b445c", "imm2": "0x17a5a9c7", "rot": 8, "bit": 8, "mask": 1, "width": 1}, + {"i": 9, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4, "width": 1}, + {"i": 10, "op": "or", "dst": 1, "src": 2, "src2": 3, "imm": "0x4e7dc10d", "imm2": "0x196d165c", "rot": 14, "bit": 27, "mask": 16, "width": 1}, + {"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 1, "imm": "0xc5c3b55d", "imm2": "0xec061424", "rot": 26, "bit": 27, "mask": 8, "width": 1}, + {"i": 12, "op": "or", "dst": 6, "src": 2, "src2": 3, "imm": "0x306542fe", "imm2": "0x1bb1b429", "rot": 31, "bit": 0, "mask": 2, "width": 1}, + {"i": 13, "op": "mul", "dst": 2, "src": 5, "src2": 6, "imm": "0xa672cdd3", "imm2": "0x59a4829c", "rot": 22, "bit": 13, "mask": 16, "width": 1}, + {"i": 14, "op": "load", "dst": 1, "src": 2, "src2": 5, "imm": "0x028b4d37", "imm2": "0x7bbd78ea", "rot": 15, "bit": 2, "mask": 8, "width": 1}, + {"i": 15, "op": "rotl", "dst": 7, "src": 6, "src2": 6, "imm": "0x5c88a1a7", "imm2": "0x5c628769", "rot": 1, "bit": 3, "mask": 8, "width": 1}, + {"i": 16, "op": "scratch", "dst": 3, "src": 6, "src2": 7, "imm": "0xbac2ae81", "imm2": "0xcbbc7bdb", "rot": 18, "bit": 8, "mask": 8, "width": 1}, + {"i": 17, "op": "load", "dst": 7, "src": 4, "src2": 2, "imm": "0xe8ab93e9", "imm2": "0xa00de107", "rot": 2, "bit": 1, "mask": 16, "width": 1}, + {"i": 18, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2, "width": 1}, + {"i": 19, "op": "mad", "dst": 4, "src": 0, "src2": 2, "imm": "0x5fba7bc2", "imm2": "0xdf099cfb", "rot": 4, "bit": 15, "mask": 16, "width": 1}, + {"i": 20, "op": "shfl", "dst": 0, "src": 6, "src2": 3, "imm": "0x0a3056de", "imm2": "0x7f0c25c3", "rot": 27, "bit": 13, "mask": 8, "width": 1}, + {"i": 21, "op": "xor", "dst": 5, "src": 7, "src2": 4, "imm": "0xbd066e1d", "imm2": "0x6d3ddc5a", "rot": 2, "bit": 29, "mask": 1, "width": 1}, + {"i": 22, "op": "mulhi", "dst": 2, "src": 5, "src2": 0, "imm": "0xc7e9887a", "imm2": "0x19ec898f", "rot": 14, "bit": 9, "mask": 1, "width": 1}, + {"i": 23, "op": "scratch", "dst": 3, "src": 7, "src2": 2, "imm": "0xc7fcfc8f", "imm2": "0x8528b94f", "rot": 17, "bit": 13, "mask": 4, "width": 1}, + {"i": 24, "op": "mulhi", "dst": 7, "src": 3, "src2": 5, "imm": "0xd91641e8", "imm2": "0xaf77faf2", "rot": 22, "bit": 21, "mask": 1, "width": 1}, + {"i": 25, "op": "or", "dst": 5, "src": 4, "src2": 0, "imm": "0x84c03868", "imm2": "0xf6c691b7", "rot": 29, "bit": 14, "mask": 8, "width": 1}, + {"i": 26, "op": "mad", "dst": 4, "src": 5, "src2": 2, "imm": "0x3bb2b6ba", "imm2": "0x49d95fd5", "rot": 1, "bit": 5, "mask": 8, "width": 1}, + {"i": 27, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4, "width": 1}, + {"i": 28, "op": "mulhi", "dst": 6, "src": 7, "src2": 6, "imm": "0xd69c4715", "imm2": "0xe0ebc4ce", "rot": 29, "bit": 2, "mask": 8, "width": 1}, + {"i": 29, "op": "add", "dst": 6, "src": 1, "src2": 2, "imm": "0x3b2d2124", "imm2": "0x187a9128", "rot": 1, "bit": 9, "mask": 16, "width": 1}, + {"i": 30, "op": "rotr", "dst": 6, "src": 7, "src2": 0, "imm": "0x5c64a589", "imm2": "0x61c9a38d", "rot": 17, "bit": 21, "mask": 16, "width": 1}, + {"i": 31, "op": "load", "dst": 3, "src": 1, "src2": 7, "imm": "0xc37723fa", "imm2": "0xf3b024da", "rot": 16, "bit": 27, "mask": 16, "width": 1}, + {"i": 32, "op": "scratch", "dst": 1, "src": 0, "src2": 7, "imm": "0xcc7972c4", "imm2": "0xad098d15", "rot": 30, "bit": 21, "mask": 8, "width": 1}, + {"i": 33, "op": "add", "dst": 0, "src": 4, "src2": 4, "imm": "0x2c35699f", "imm2": "0x351dde38", "rot": 21, "bit": 18, "mask": 4, "width": 1}, + {"i": 34, "op": "scratch", "dst": 0, "src": 2, "src2": 3, "imm": "0xfae8902b", "imm2": "0x5cd8306f", "rot": 5, "bit": 28, "mask": 16, "width": 1}, + {"i": 35, "op": "mul", "dst": 0, "src": 3, "src2": 1, "imm": "0x4fa3f3db", "imm2": "0xdbf37e75", "rot": 7, "bit": 18, "mask": 4, "width": 1}, + {"i": 36, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2, "width": 1}, + {"i": 37, "op": "scratch", "dst": 4, "src": 0, "src2": 0, "imm": "0x04cc1d55", "imm2": "0x35c52d04", "rot": 11, "bit": 14, "mask": 2, "width": 1}, + {"i": 38, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0x3958f280", "imm2": "0x8713c7e1", "rot": 5, "bit": 23, "mask": 16, "width": 1}, + {"i": 39, "op": "add", "dst": 0, "src": 3, "src2": 3, "imm": "0xa907b90b", "imm2": "0x1b053acf", "rot": 30, "bit": 25, "mask": 16, "width": 1}, + {"i": 40, "op": "rotr", "dst": 2, "src": 5, "src2": 4, "imm": "0xf8662282", "imm2": "0x10bb9e30", "rot": 8, "bit": 6, "mask": 2, "width": 1}, + {"i": 41, "op": "mul", "dst": 3, "src": 2, "src2": 4, "imm": "0x49087d74", "imm2": "0x6348b489", "rot": 17, "bit": 9, "mask": 16, "width": 1}, + {"i": 42, "op": "add", "dst": 1, "src": 5, "src2": 1, "imm": "0xa32e000c", "imm2": "0x6058c2e3", "rot": 25, "bit": 20, "mask": 8, "width": 1}, + {"i": 43, "op": "xor", "dst": 3, "src": 4, "src2": 2, "imm": "0x3dad0eb6", "imm2": "0xb97578cb", "rot": 3, "bit": 27, "mask": 1, "width": 1}, + {"i": 44, "op": "scratch", "dst": 3, "src": 5, "src2": 7, "imm": "0x374aec92", "imm2": "0x626f11df", "rot": 20, "bit": 18, "mask": 8, "width": 1}, + {"i": 45, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4, "width": 1}, + {"i": 46, "op": "xor", "dst": 7, "src": 1, "src2": 0, "imm": "0xef6ac348", "imm2": "0x963bb7e6", "rot": 26, "bit": 3, "mask": 8, "width": 1}, + {"i": 47, "op": "add", "dst": 0, "src": 3, "src2": 0, "imm": "0x838b5065", "imm2": "0x36360066", "rot": 3, "bit": 31, "mask": 4, "width": 1}, + {"i": 48, "op": "mulhi", "dst": 7, "src": 5, "src2": 0, "imm": "0x8458f7ac", "imm2": "0xc1c15026", "rot": 27, "bit": 15, "mask": 8, "width": 1}, + {"i": 49, "op": "load", "dst": 0, "src": 2, "src2": 4, "imm": "0x636a9dc4", "imm2": "0xac023d9b", "rot": 22, "bit": 29, "mask": 1, "width": 1}, + {"i": 50, "op": "sub", "dst": 2, "src": 6, "src2": 0, "imm": "0x2baec8c9", "imm2": "0x4390f156", "rot": 3, "bit": 12, "mask": 8, "width": 1}, + {"i": 51, "op": "sub", "dst": 7, "src": 5, "src2": 7, "imm": "0x19234061", "imm2": "0xe84dfade", "rot": 4, "bit": 19, "mask": 1, "width": 1}, + {"i": 52, "op": "xor", "dst": 2, "src": 3, "src2": 5, "imm": "0xdc2cd71e", "imm2": "0x1b5d334b", "rot": 9, "bit": 8, "mask": 8, "width": 1}, + {"i": 53, "op": "sub", "dst": 7, "src": 0, "src2": 4, "imm": "0x605c31ec", "imm2": "0x9923ff88", "rot": 28, "bit": 25, "mask": 4, "width": 1}, + {"i": 54, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4, "width": 1}, + {"i": 55, "op": "xor", "dst": 7, "src": 5, "src2": 5, "imm": "0xad7493e7", "imm2": "0x3e400372", "rot": 13, "bit": 8, "mask": 1, "width": 1}, + {"i": 56, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0x87e933c9", "imm2": "0x8c854c1b", "rot": 17, "bit": 3, "mask": 8, "width": 1}, + {"i": 57, "op": "sub", "dst": 5, "src": 6, "src2": 5, "imm": "0x11be3bc9", "imm2": "0xbbaa8e24", "rot": 6, "bit": 5, "mask": 16, "width": 1}, + {"i": 58, "op": "load", "dst": 1, "src": 3, "src2": 2, "imm": "0xa732351a", "imm2": "0xc01349cd", "rot": 14, "bit": 17, "mask": 16, "width": 1}, + {"i": 59, "op": "scratch", "dst": 1, "src": 4, "src2": 0, "imm": "0xb20547b2", "imm2": "0xc94655de", "rot": 27, "bit": 30, "mask": 1, "width": 1}, + {"i": 60, "op": "sub", "dst": 4, "src": 6, "src2": 7, "imm": "0x67cf904c", "imm2": "0x6873b216", "rot": 27, "bit": 7, "mask": 16, "width": 1}, + {"i": 61, "op": "mul", "dst": 1, "src": 2, "src2": 7, "imm": "0x93ab0bf4", "imm2": "0x96158375", "rot": 14, "bit": 0, "mask": 16, "width": 1}, + {"i": 62, "op": "mad", "dst": 3, "src": 6, "src2": 0, "imm": "0x41a443a3", "imm2": "0xe69d7919", "rot": 9, "bit": 0, "mask": 16, "width": 1}, + {"i": 63, "op": "add", "dst": 0, "src": 1, "src2": 3, "imm": "0x2fe0e98b", "imm2": "0xc88e2942", "rot": 5, "bit": 16, "mask": 16, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/scr8/program.metal b/proto-cuda/packs-readwidth/scr8/program.metal new file mode 100644 index 000000000..73532c454 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8/program.metal @@ -0,0 +1,126 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + device uint* scratch [[buffer(3)]], + constant uint& groups [[buffer(4)]], + constant uint& salt [[buffer(5)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint s_ = r6 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + { uint s_ = r7 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + { uint s_ = r0 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + { uint s_ = r0 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + { uint s_ = r5 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + { uint s_ = r4 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr8/program_bound.metal b/proto-cuda/packs-readwidth/scr8/program_bound.metal new file mode 100644 index 000000000..d260c34fe --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8/program_bound.metal @@ -0,0 +1,128 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + device uint* scratch [[buffer(4)]], + constant uint& groups [[buffer(5)]], + constant uint& salt [[buffer(6)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint s_ = r6 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + { uint s_ = r7 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + { uint s_ = r0 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + { uint s_ = r0 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + { uint s_ = r5 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + { uint s_ = r4 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr8/vectors.h b/proto-cuda/packs-readwidth/scr8/vectors.h new file mode 100644 index 000000000..43d22af98 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0xd0f846c3cd57ae09ull, 0x4d5e1caf41761a5cull, 0x7fc77bb7fa221a51ull, 0xe45b6317d4f52ae7ull, 0x09704a39107a3150ull, 0xd5166620f9dc49abull, 0xeaf1fa69e4075e55ull, 0x38e23d449b7616aeull, + 0xf4593424c8427320ull, 0x9359a44a5149bd60ull, 0x2df93b5bc274f083ull, 0xfd80eb32016f8659ull, 0xdeb10b8a37fc3ee7ull, 0xc18e2ed77ab77f34ull, 0x7f39b618cc81c1b4ull, 0xdc1b7299961ad5abull, + 0x8f9c03d151013ec4ull, 0x668927a4a267d76full, 0x233a50c329caf635ull, 0x7c80406541ae3f15ull, 0x18fb38606c48af4aull, 0x0452a913df5f112bull, 0x59065cb670ba6d7dull, 0xa2998e2ccd726df8ull, + 0xef690e221af37893ull, 0x4e87968e5c45903eull, 0x6c6e5d4c1b35d7a9ull, 0x54d7e322426dcae3ull, 0x0d3ed6c08deda320ull, 0x3d6a4301fbb06a5aull, 0x9eb78b04d2366566ull, 0x7806f8c64d2d09ffull + }, + { // base nonce 4096 + 0xe2c97c384a85c687ull, 0xcf905b005ab01ebcull, 0x4526799fae210c6bull, 0xd4daf72ed75e8a16ull, 0xb3703a9e7c6820a8ull, 0xb6be9393fbb920bdull, 0xe2fff04fb7816eb8ull, 0xed76ad5dbf400f91ull, + 0xea7d4b58ab0eb8d1ull, 0xf67a73b01f030532ull, 0x35ab823037510099ull, 0x5de1b1b0b3a26c69ull, 0xbe5dd90dbb7e632bull, 0x1475918a9237e24dull, 0x177fd53c45634d71ull, 0xa7cf00759ba28ce0ull, + 0xe51be0584ac3fbb4ull, 0x427049cc778aab35ull, 0x826bab125577d172ull, 0xd705b891b16237f5ull, 0xdf622fd44b180a87ull, 0x359398ecb79ec2deull, 0x2a1e075fb078da66ull, 0xd480ddd8e66d26e8ull, + 0x07e3e86f517de466ull, 0xa98b2f7423557445ull, 0x6bab15b36bb142feull, 0xf87d147bf2cc5c0bull, 0xf0294ea2b2820e03ull, 0xf219ac95e823d794ull, 0x9fa3fea85bc54264ull, 0xc4af3fafd2ff5201ull + }, + { // base nonce 1000000 + 0x501f772483fac0a3ull, 0x461363da2c1539e0ull, 0x750050674de592afull, 0x14ed105042cd912cull, 0xc6477310878614ebull, 0xe916b44e32e90e56ull, 0x531c92e69c2bdd73ull, 0x5bb0129ef7c3bd51ull, + 0x967ed91f7cbc7c79ull, 0x06fcc6a895b58b4aull, 0x3bdb29fbf93cbff3ull, 0xbed385172d2e6abeull, 0x921fc99ff4efac5eull, 0x6b0090bedd9f69c7ull, 0x0b105c18aaa53ab0ull, 0xf0c4203561f3b94dull, + 0x64abe33adccf7807ull, 0xd8f3a3b7e0242c04ull, 0x397da462f69fac5bull, 0xe6b6b32e467b71beull, 0x5318ac9e56278d04ull, 0xf7c5e348a4e1f5dbull, 0x79c994e6109646dfull, 0x9cdc42b0f6e82231ull, + 0x3f1160f02fd96ac7ull, 0xa69a23fc7058be21ull, 0xccde0a19bfdc25b6ull, 0xd3684a5b966f7497ull, 0x929db00b97a624f9ull, 0xfd14a40882d395a6ull, 0x49b0c1ecc514d6adull, 0x95d994c8349be12dull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/scr8/vectors.json b/proto-cuda/packs-readwidth/scr8/vectors.json new file mode 100644 index 000000000..64e628d6e --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-genesis", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0xd0f846c3cd57ae09", "0x4d5e1caf41761a5c", "0x7fc77bb7fa221a51", "0xe45b6317d4f52ae7", "0x09704a39107a3150", "0xd5166620f9dc49ab", "0xeaf1fa69e4075e55", "0x38e23d449b7616ae", + "0xf4593424c8427320", "0x9359a44a5149bd60", "0x2df93b5bc274f083", "0xfd80eb32016f8659", "0xdeb10b8a37fc3ee7", "0xc18e2ed77ab77f34", "0x7f39b618cc81c1b4", "0xdc1b7299961ad5ab", + "0x8f9c03d151013ec4", "0x668927a4a267d76f", "0x233a50c329caf635", "0x7c80406541ae3f15", "0x18fb38606c48af4a", "0x0452a913df5f112b", "0x59065cb670ba6d7d", "0xa2998e2ccd726df8", + "0xef690e221af37893", "0x4e87968e5c45903e", "0x6c6e5d4c1b35d7a9", "0x54d7e322426dcae3", "0x0d3ed6c08deda320", "0x3d6a4301fbb06a5a", "0x9eb78b04d2366566", "0x7806f8c64d2d09ff" + ]}, + {"base_nonce": 4096, "expected": [ + "0xe2c97c384a85c687", "0xcf905b005ab01ebc", "0x4526799fae210c6b", "0xd4daf72ed75e8a16", "0xb3703a9e7c6820a8", "0xb6be9393fbb920bd", "0xe2fff04fb7816eb8", "0xed76ad5dbf400f91", + "0xea7d4b58ab0eb8d1", "0xf67a73b01f030532", "0x35ab823037510099", "0x5de1b1b0b3a26c69", "0xbe5dd90dbb7e632b", "0x1475918a9237e24d", "0x177fd53c45634d71", "0xa7cf00759ba28ce0", + "0xe51be0584ac3fbb4", "0x427049cc778aab35", "0x826bab125577d172", "0xd705b891b16237f5", "0xdf622fd44b180a87", "0x359398ecb79ec2de", "0x2a1e075fb078da66", "0xd480ddd8e66d26e8", + "0x07e3e86f517de466", "0xa98b2f7423557445", "0x6bab15b36bb142fe", "0xf87d147bf2cc5c0b", "0xf0294ea2b2820e03", "0xf219ac95e823d794", "0x9fa3fea85bc54264", "0xc4af3fafd2ff5201" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x501f772483fac0a3", "0x461363da2c1539e0", "0x750050674de592af", "0x14ed105042cd912c", "0xc6477310878614eb", "0xe916b44e32e90e56", "0x531c92e69c2bdd73", "0x5bb0129ef7c3bd51", + "0x967ed91f7cbc7c79", "0x06fcc6a895b58b4a", "0x3bdb29fbf93cbff3", "0xbed385172d2e6abe", "0x921fc99ff4efac5e", "0x6b0090bedd9f69c7", "0x0b105c18aaa53ab0", "0xf0c4203561f3b94d", + "0x64abe33adccf7807", "0xd8f3a3b7e0242c04", "0x397da462f69fac5b", "0xe6b6b32e467b71be", "0x5318ac9e56278d04", "0xf7c5e348a4e1f5db", "0x79c994e6109646df", "0x9cdc42b0f6e82231", + "0x3f1160f02fd96ac7", "0xa69a23fc7058be21", "0xccde0a19bfdc25b6", "0xd3684a5b966f7497", "0x929db00b97a624f9", "0xfd14a40882d395a6", "0x49b0c1ecc514d6ad", "0x95d994c8349be12d" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/w16/kernel.cl b/proto-cuda/packs-readwidth/w16/kernel.cl new file mode 100644 index 000000000..fba96ead5 --- /dev/null +++ b/proto-cuda/packs-readwidth/w16/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 4 load + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 load + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 31 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 56 load + r5 = r5 - r6; // 57 sub + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 58 load + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/w16/kernel.cu b/proto-cuda/packs-readwidth/w16/kernel.cu new file mode 100644 index 000000000..db3806481 --- /dev/null +++ b/proto-cuda/packs-readwidth/w16/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 4 load + { uint32_t b_ = (r1 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 load + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 23 load + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + { uint32_t b_ = (r1 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 31 load + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 56 load + r5 = r5 - r6; // 57 sub + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 58 load + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/w16/kernel_bound.cl b/proto-cuda/packs-readwidth/w16/kernel_bound.cl new file mode 100644 index 000000000..b9c21f8ba --- /dev/null +++ b/proto-cuda/packs-readwidth/w16/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 4 load + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 load + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 31 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 56 load + r5 = r5 - r6; // 57 sub + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 58 load + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 4 load + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint b_ = (r6 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 load + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + { uint b_ = (r1 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 31 load + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint b_ = (r0 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint b_ = (r5 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + { uint b_ = (r2 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + { uint b_ = (r7 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 56 load + r5 = r5 - r6; // 57 sub + { uint b_ = (r3 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 58 load + { uint b_ = (r4 & mask) & ~3u; uint4 v0_ = vload4(0u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/w16/kernel_bound.cu b/proto-cuda/packs-readwidth/w16/kernel_bound.cu new file mode 100644 index 000000000..24a99193b --- /dev/null +++ b/proto-cuda/packs-readwidth/w16/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 4 load + { uint32_t b_ = (r1 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t b_ = (r6 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 load + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 23 load + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + { uint32_t b_ = (r1 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 31 load + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint32_t b_ = (r0 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint32_t b_ = (r5 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + { uint32_t b_ = (r2 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + { uint32_t b_ = (r7 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 56 load + r5 = r5 - r6; // 57 sub + { uint32_t b_ = (r3 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 58 load + { uint32_t b_ = (r4 & mask) & ~3u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/w16/memhard.h b/proto-cuda/packs-readwidth/w16/memhard.h new file mode 100644 index 000000000..4803d8e40 --- /dev/null +++ b/proto-cuda/packs-readwidth/w16/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/w16/memhard.metal b/proto-cuda/packs-readwidth/w16/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/w16/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/w16/program.h b/proto-cuda/packs-readwidth/w16/program.h new file mode 100644 index 000000000..2cb1f66d2 --- /dev/null +++ b/proto-cuda/packs-readwidth/w16/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-genesis" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x6a183165e23fd2bcull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=9 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 shfl=4 rotr=2 rotl=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "w16" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 0, 100, 0 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 0, 16, 0 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 2048 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/w16/program.json b/proto-cuda/packs-readwidth/w16/program.json new file mode 100644 index 000000000..8f563a1b7 --- /dev/null +++ b/proto-cuda/packs-readwidth/w16/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x6a183165e23fd2bc", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-genesis", + "seed_bytes": "69676e65756d2d67656e65736973", + "seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "w16", + "load_slots": 16, + "load_mix_percent_4_16_64": [0, 100, 0], + "load_width_counts_4_16_64": [0, 16, 0], + "bytes_per_hash": 2048, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "add": 9, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "shfl": 4, "rotr": 2, "rotl": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2, "width": 1}, + {"i": 1, "op": "add", "dst": 1, "src": 7, "src2": 2, "imm": "0x42da7657", "imm2": "0xc3bd2355", "rot": 25, "bit": 4, "mask": 16, "width": 1}, + {"i": 2, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x61f0b51c", "imm2": "0x2735a174", "rot": 4, "bit": 26, "mask": 2, "width": 1}, + {"i": 3, "op": "mad", "dst": 4, "src": 0, "src2": 6, "imm": "0x679648a8", "imm2": "0x3044ba32", "rot": 31, "bit": 31, "mask": 4, "width": 1}, + {"i": 4, "op": "load", "dst": 7, "src": 2, "src2": 6, "imm": "0x5d1ca2a2", "imm2": "0xe2481807", "rot": 24, "bit": 3, "mask": 1, "width": 4}, + {"i": 5, "op": "load", "dst": 4, "src": 1, "src2": 2, "imm": "0x987c017a", "imm2": "0xf4d60559", "rot": 2, "bit": 0, "mask": 4, "width": 4}, + {"i": 6, "op": "shfl", "dst": 6, "src": 3, "src2": 7, "imm": "0x6ea7b2df", "imm2": "0x9fce5071", "rot": 7, "bit": 15, "mask": 4, "width": 1}, + {"i": 7, "op": "shfl", "dst": 1, "src": 5, "src2": 1, "imm": "0x26a2ecde", "imm2": "0xfec6ad22", "rot": 15, "bit": 11, "mask": 8, "width": 1}, + {"i": 8, "op": "xor", "dst": 7, "src": 5, "src2": 2, "imm": "0xbe4b445c", "imm2": "0x17a5a9c7", "rot": 8, "bit": 8, "mask": 1, "width": 1}, + {"i": 9, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4, "width": 1}, + {"i": 10, "op": "or", "dst": 1, "src": 2, "src2": 3, "imm": "0x4e7dc10d", "imm2": "0x196d165c", "rot": 14, "bit": 27, "mask": 16, "width": 1}, + {"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 1, "imm": "0xc5c3b55d", "imm2": "0xec061424", "rot": 26, "bit": 27, "mask": 8, "width": 4}, + {"i": 12, "op": "or", "dst": 6, "src": 2, "src2": 3, "imm": "0x306542fe", "imm2": "0x1bb1b429", "rot": 31, "bit": 0, "mask": 2, "width": 1}, + {"i": 13, "op": "mul", "dst": 2, "src": 5, "src2": 6, "imm": "0xa672cdd3", "imm2": "0x59a4829c", "rot": 22, "bit": 13, "mask": 16, "width": 1}, + {"i": 14, "op": "load", "dst": 1, "src": 2, "src2": 5, "imm": "0x028b4d37", "imm2": "0x7bbd78ea", "rot": 15, "bit": 2, "mask": 8, "width": 4}, + {"i": 15, "op": "rotl", "dst": 7, "src": 6, "src2": 6, "imm": "0x5c88a1a7", "imm2": "0x5c628769", "rot": 1, "bit": 3, "mask": 8, "width": 1}, + {"i": 16, "op": "load", "dst": 3, "src": 6, "src2": 7, "imm": "0xbac2ae81", "imm2": "0xcbbc7bdb", "rot": 18, "bit": 8, "mask": 8, "width": 4}, + {"i": 17, "op": "load", "dst": 7, "src": 4, "src2": 2, "imm": "0xe8ab93e9", "imm2": "0xa00de107", "rot": 2, "bit": 1, "mask": 16, "width": 4}, + {"i": 18, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2, "width": 1}, + {"i": 19, "op": "mad", "dst": 4, "src": 0, "src2": 2, "imm": "0x5fba7bc2", "imm2": "0xdf099cfb", "rot": 4, "bit": 15, "mask": 16, "width": 1}, + {"i": 20, "op": "shfl", "dst": 0, "src": 6, "src2": 3, "imm": "0x0a3056de", "imm2": "0x7f0c25c3", "rot": 27, "bit": 13, "mask": 8, "width": 1}, + {"i": 21, "op": "xor", "dst": 5, "src": 7, "src2": 4, "imm": "0xbd066e1d", "imm2": "0x6d3ddc5a", "rot": 2, "bit": 29, "mask": 1, "width": 1}, + {"i": 22, "op": "mulhi", "dst": 2, "src": 5, "src2": 0, "imm": "0xc7e9887a", "imm2": "0x19ec898f", "rot": 14, "bit": 9, "mask": 1, "width": 1}, + {"i": 23, "op": "load", "dst": 3, "src": 7, "src2": 2, "imm": "0xc7fcfc8f", "imm2": "0x8528b94f", "rot": 17, "bit": 13, "mask": 4, "width": 4}, + {"i": 24, "op": "mulhi", "dst": 7, "src": 3, "src2": 5, "imm": "0xd91641e8", "imm2": "0xaf77faf2", "rot": 22, "bit": 21, "mask": 1, "width": 1}, + {"i": 25, "op": "or", "dst": 5, "src": 4, "src2": 0, "imm": "0x84c03868", "imm2": "0xf6c691b7", "rot": 29, "bit": 14, "mask": 8, "width": 1}, + {"i": 26, "op": "mad", "dst": 4, "src": 5, "src2": 2, "imm": "0x3bb2b6ba", "imm2": "0x49d95fd5", "rot": 1, "bit": 5, "mask": 8, "width": 1}, + {"i": 27, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4, "width": 1}, + {"i": 28, "op": "mulhi", "dst": 6, "src": 7, "src2": 6, "imm": "0xd69c4715", "imm2": "0xe0ebc4ce", "rot": 29, "bit": 2, "mask": 8, "width": 1}, + {"i": 29, "op": "add", "dst": 6, "src": 1, "src2": 2, "imm": "0x3b2d2124", "imm2": "0x187a9128", "rot": 1, "bit": 9, "mask": 16, "width": 1}, + {"i": 30, "op": "rotr", "dst": 6, "src": 7, "src2": 0, "imm": "0x5c64a589", "imm2": "0x61c9a38d", "rot": 17, "bit": 21, "mask": 16, "width": 1}, + {"i": 31, "op": "load", "dst": 3, "src": 1, "src2": 7, "imm": "0xc37723fa", "imm2": "0xf3b024da", "rot": 16, "bit": 27, "mask": 16, "width": 4}, + {"i": 32, "op": "load", "dst": 1, "src": 0, "src2": 7, "imm": "0xcc7972c4", "imm2": "0xad098d15", "rot": 30, "bit": 21, "mask": 8, "width": 4}, + {"i": 33, "op": "add", "dst": 0, "src": 4, "src2": 4, "imm": "0x2c35699f", "imm2": "0x351dde38", "rot": 21, "bit": 18, "mask": 4, "width": 1}, + {"i": 34, "op": "load", "dst": 0, "src": 2, "src2": 3, "imm": "0xfae8902b", "imm2": "0x5cd8306f", "rot": 5, "bit": 28, "mask": 16, "width": 4}, + {"i": 35, "op": "mul", "dst": 0, "src": 3, "src2": 1, "imm": "0x4fa3f3db", "imm2": "0xdbf37e75", "rot": 7, "bit": 18, "mask": 4, "width": 1}, + {"i": 36, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2, "width": 1}, + {"i": 37, "op": "load", "dst": 4, "src": 0, "src2": 0, "imm": "0x04cc1d55", "imm2": "0x35c52d04", "rot": 11, "bit": 14, "mask": 2, "width": 4}, + {"i": 38, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0x3958f280", "imm2": "0x8713c7e1", "rot": 5, "bit": 23, "mask": 16, "width": 1}, + {"i": 39, "op": "add", "dst": 0, "src": 3, "src2": 3, "imm": "0xa907b90b", "imm2": "0x1b053acf", "rot": 30, "bit": 25, "mask": 16, "width": 1}, + {"i": 40, "op": "rotr", "dst": 2, "src": 5, "src2": 4, "imm": "0xf8662282", "imm2": "0x10bb9e30", "rot": 8, "bit": 6, "mask": 2, "width": 1}, + {"i": 41, "op": "mul", "dst": 3, "src": 2, "src2": 4, "imm": "0x49087d74", "imm2": "0x6348b489", "rot": 17, "bit": 9, "mask": 16, "width": 1}, + {"i": 42, "op": "add", "dst": 1, "src": 5, "src2": 1, "imm": "0xa32e000c", "imm2": "0x6058c2e3", "rot": 25, "bit": 20, "mask": 8, "width": 1}, + {"i": 43, "op": "xor", "dst": 3, "src": 4, "src2": 2, "imm": "0x3dad0eb6", "imm2": "0xb97578cb", "rot": 3, "bit": 27, "mask": 1, "width": 1}, + {"i": 44, "op": "load", "dst": 3, "src": 5, "src2": 7, "imm": "0x374aec92", "imm2": "0x626f11df", "rot": 20, "bit": 18, "mask": 8, "width": 4}, + {"i": 45, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4, "width": 1}, + {"i": 46, "op": "xor", "dst": 7, "src": 1, "src2": 0, "imm": "0xef6ac348", "imm2": "0x963bb7e6", "rot": 26, "bit": 3, "mask": 8, "width": 1}, + {"i": 47, "op": "add", "dst": 0, "src": 3, "src2": 0, "imm": "0x838b5065", "imm2": "0x36360066", "rot": 3, "bit": 31, "mask": 4, "width": 1}, + {"i": 48, "op": "mulhi", "dst": 7, "src": 5, "src2": 0, "imm": "0x8458f7ac", "imm2": "0xc1c15026", "rot": 27, "bit": 15, "mask": 8, "width": 1}, + {"i": 49, "op": "load", "dst": 0, "src": 2, "src2": 4, "imm": "0x636a9dc4", "imm2": "0xac023d9b", "rot": 22, "bit": 29, "mask": 1, "width": 4}, + {"i": 50, "op": "sub", "dst": 2, "src": 6, "src2": 0, "imm": "0x2baec8c9", "imm2": "0x4390f156", "rot": 3, "bit": 12, "mask": 8, "width": 1}, + {"i": 51, "op": "sub", "dst": 7, "src": 5, "src2": 7, "imm": "0x19234061", "imm2": "0xe84dfade", "rot": 4, "bit": 19, "mask": 1, "width": 1}, + {"i": 52, "op": "xor", "dst": 2, "src": 3, "src2": 5, "imm": "0xdc2cd71e", "imm2": "0x1b5d334b", "rot": 9, "bit": 8, "mask": 8, "width": 1}, + {"i": 53, "op": "sub", "dst": 7, "src": 0, "src2": 4, "imm": "0x605c31ec", "imm2": "0x9923ff88", "rot": 28, "bit": 25, "mask": 4, "width": 1}, + {"i": 54, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4, "width": 1}, + {"i": 55, "op": "xor", "dst": 7, "src": 5, "src2": 5, "imm": "0xad7493e7", "imm2": "0x3e400372", "rot": 13, "bit": 8, "mask": 1, "width": 1}, + {"i": 56, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0x87e933c9", "imm2": "0x8c854c1b", "rot": 17, "bit": 3, "mask": 8, "width": 4}, + {"i": 57, "op": "sub", "dst": 5, "src": 6, "src2": 5, "imm": "0x11be3bc9", "imm2": "0xbbaa8e24", "rot": 6, "bit": 5, "mask": 16, "width": 1}, + {"i": 58, "op": "load", "dst": 1, "src": 3, "src2": 2, "imm": "0xa732351a", "imm2": "0xc01349cd", "rot": 14, "bit": 17, "mask": 16, "width": 4}, + {"i": 59, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0xb20547b2", "imm2": "0xc94655de", "rot": 27, "bit": 30, "mask": 1, "width": 4}, + {"i": 60, "op": "sub", "dst": 4, "src": 6, "src2": 7, "imm": "0x67cf904c", "imm2": "0x6873b216", "rot": 27, "bit": 7, "mask": 16, "width": 1}, + {"i": 61, "op": "mul", "dst": 1, "src": 2, "src2": 7, "imm": "0x93ab0bf4", "imm2": "0x96158375", "rot": 14, "bit": 0, "mask": 16, "width": 1}, + {"i": 62, "op": "mad", "dst": 3, "src": 6, "src2": 0, "imm": "0x41a443a3", "imm2": "0xe69d7919", "rot": 9, "bit": 0, "mask": 16, "width": 1}, + {"i": 63, "op": "add", "dst": 0, "src": 1, "src2": 3, "imm": "0x2fe0e98b", "imm2": "0xc88e2942", "rot": 5, "bit": 16, "mask": 16, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/w16/program.metal b/proto-cuda/packs-readwidth/w16/program.metal new file mode 100644 index 000000000..937a8732c --- /dev/null +++ b/proto-cuda/packs-readwidth/w16/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 4 + { uint b_ = (r1 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + { uint b_ = (r1 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 31 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 56 + r5 = r5 - r6; // 57 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 58 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/w16/program_bound.metal b/proto-cuda/packs-readwidth/w16/program_bound.metal new file mode 100644 index 000000000..19d9dd80c --- /dev/null +++ b/proto-cuda/packs-readwidth/w16/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 4 + { uint b_ = (r1 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint b_ = (r6 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 16 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r7 = x_; } // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + { uint b_ = (r1 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 31 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + { uint b_ = (r0 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r4 = x_; } // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + { uint b_ = (r5 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r3 = x_; } // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + { uint b_ = (r2 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r0 = x_; } // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + { uint b_ = (r7 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r2 = x_; } // 56 + r5 = r5 - r6; // 57 + { uint b_ = (r3 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 58 + { uint b_ = (r4 & MASK) & ~3u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; r1 = x_; } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/w16/vectors.h b/proto-cuda/packs-readwidth/w16/vectors.h new file mode 100644 index 000000000..d0a6ac1ce --- /dev/null +++ b/proto-cuda/packs-readwidth/w16/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x8990d9ea9d842e47ull, 0x2e56d7a457c9b33cull, 0x7b7bf5095c5c8059ull, 0xb08eb13ad2b9fbb5ull, 0x09e38890a5e9fa39ull, 0x0e3a26c63b62ca35ull, 0x9fc6750de38838e2ull, 0xc11876c5fcabbcbaull, + 0x06c63dff72642d31ull, 0xae90e17028aff18dull, 0xb58b4e2fcf46b9deull, 0x6ac8c16aa8b40172ull, 0x249e93f57b9d67beull, 0xc736176690f8f7fcull, 0x6c59adccfdd04ce8ull, 0x63b27f1a8c87dfeaull, + 0x5fd17f7a71c71f4eull, 0x21b1f7f17f720c0eull, 0xcc4dddda4ab67b1eull, 0xe23e3528dc7a2590ull, 0xf1ddcfec6bc96214ull, 0xc9fe9a77d8bc8014ull, 0xb548c0a1c8e5b959ull, 0x0bd53aedd898991aull, + 0xa54a0c823fe11e6cull, 0x52242302b541704aull, 0x7fd871e0001cc3aeull, 0xda282aaa671cca04ull, 0xcce7d4a10f014c06ull, 0xcfe66d4da558f5f9ull, 0xe3205bc5628642daull, 0x5fc7e557beb9a47aull + }, + { // base nonce 4096 + 0x6af0908472947c8cull, 0xf690c04c339737daull, 0xd38c16fbd8b52797ull, 0xb4ab81b6aaa438dfull, 0x6c2feeb76557c2e8ull, 0xd5b034b393a2cbf7ull, 0x1b9bf865c564e5bcull, 0x54d1b2a27561bbd1ull, + 0xd2857d24c6d80489ull, 0xce59adb76feeaeeaull, 0xed3729c4c8d6ce36ull, 0xb7d7561fca026702ull, 0xd8bddf4bcb4c58f4ull, 0x87c4d854da6e7a87ull, 0x737adba1ee9a305cull, 0xe6b71508f3f50dc9ull, + 0x055f17b64d00f306ull, 0x1db21c5aa4867c51ull, 0x1bd65024010dd73eull, 0x07b8d976a1924860ull, 0x67ccbb1048613736ull, 0xed4a373e966b49e0ull, 0xeec2b3d18a32770aull, 0xe64da383108c1a82ull, + 0xbe7f6bbd40f656c6ull, 0xac46734b9df94f58ull, 0x66be3e5ad9c37dd4ull, 0x2e850961ecbb06a2ull, 0xfa0af285be67e947ull, 0xef160e045a32de4eull, 0xc01364041ac1f95cull, 0xf596d4fe7d8e383bull + }, + { // base nonce 1000000 + 0x7ef5e946fdf4cb46ull, 0x7037295d891effddull, 0xf0a765215fc4bdd4ull, 0x587a49fd3a69e71cull, 0x1327c43ce05eb049ull, 0xaa54732bece4fe61ull, 0x7245d7af07e602a0ull, 0x776efdc3b32fa074ull, + 0xb065c5e8bd1d9271ull, 0xe221c490138eca38ull, 0xbc308d3844dfdf00ull, 0x920a26a62d5c15e6ull, 0x6e643837a3068fffull, 0x86f3d5ab365d4b44ull, 0xb381a96a08bc2999ull, 0x498eebc6212b50b6ull, + 0xfec7059c93ccba08ull, 0x406963ddb2119903ull, 0xb2e50bf92bf9ef25ull, 0x05367327b1c4b448ull, 0x22ee0f56614ef9dfull, 0xdbe02e603db4fb8full, 0xd14301549bbad2aeull, 0xe24b884d4dccdd0aull, + 0xf58d446ad4122b21ull, 0x6078881289b599c4ull, 0xe03390fe9914ebb2ull, 0x3b9294b3891e7ed1ull, 0x4304325fd80e3bdaull, 0x0ae7281886301b04ull, 0x08795a399354da80ull, 0x00d90c98a85c30f9ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/w16/vectors.json b/proto-cuda/packs-readwidth/w16/vectors.json new file mode 100644 index 000000000..6c4acef02 --- /dev/null +++ b/proto-cuda/packs-readwidth/w16/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-genesis", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x8990d9ea9d842e47", "0x2e56d7a457c9b33c", "0x7b7bf5095c5c8059", "0xb08eb13ad2b9fbb5", "0x09e38890a5e9fa39", "0x0e3a26c63b62ca35", "0x9fc6750de38838e2", "0xc11876c5fcabbcba", + "0x06c63dff72642d31", "0xae90e17028aff18d", "0xb58b4e2fcf46b9de", "0x6ac8c16aa8b40172", "0x249e93f57b9d67be", "0xc736176690f8f7fc", "0x6c59adccfdd04ce8", "0x63b27f1a8c87dfea", + "0x5fd17f7a71c71f4e", "0x21b1f7f17f720c0e", "0xcc4dddda4ab67b1e", "0xe23e3528dc7a2590", "0xf1ddcfec6bc96214", "0xc9fe9a77d8bc8014", "0xb548c0a1c8e5b959", "0x0bd53aedd898991a", + "0xa54a0c823fe11e6c", "0x52242302b541704a", "0x7fd871e0001cc3ae", "0xda282aaa671cca04", "0xcce7d4a10f014c06", "0xcfe66d4da558f5f9", "0xe3205bc5628642da", "0x5fc7e557beb9a47a" + ]}, + {"base_nonce": 4096, "expected": [ + "0x6af0908472947c8c", "0xf690c04c339737da", "0xd38c16fbd8b52797", "0xb4ab81b6aaa438df", "0x6c2feeb76557c2e8", "0xd5b034b393a2cbf7", "0x1b9bf865c564e5bc", "0x54d1b2a27561bbd1", + "0xd2857d24c6d80489", "0xce59adb76feeaeea", "0xed3729c4c8d6ce36", "0xb7d7561fca026702", "0xd8bddf4bcb4c58f4", "0x87c4d854da6e7a87", "0x737adba1ee9a305c", "0xe6b71508f3f50dc9", + "0x055f17b64d00f306", "0x1db21c5aa4867c51", "0x1bd65024010dd73e", "0x07b8d976a1924860", "0x67ccbb1048613736", "0xed4a373e966b49e0", "0xeec2b3d18a32770a", "0xe64da383108c1a82", + "0xbe7f6bbd40f656c6", "0xac46734b9df94f58", "0x66be3e5ad9c37dd4", "0x2e850961ecbb06a2", "0xfa0af285be67e947", "0xef160e045a32de4e", "0xc01364041ac1f95c", "0xf596d4fe7d8e383b" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x7ef5e946fdf4cb46", "0x7037295d891effdd", "0xf0a765215fc4bdd4", "0x587a49fd3a69e71c", "0x1327c43ce05eb049", "0xaa54732bece4fe61", "0x7245d7af07e602a0", "0x776efdc3b32fa074", + "0xb065c5e8bd1d9271", "0xe221c490138eca38", "0xbc308d3844dfdf00", "0x920a26a62d5c15e6", "0x6e643837a3068fff", "0x86f3d5ab365d4b44", "0xb381a96a08bc2999", "0x498eebc6212b50b6", + "0xfec7059c93ccba08", "0x406963ddb2119903", "0xb2e50bf92bf9ef25", "0x05367327b1c4b448", "0x22ee0f56614ef9df", "0xdbe02e603db4fb8f", "0xd14301549bbad2ae", "0xe24b884d4dccdd0a", + "0xf58d446ad4122b21", "0x6078881289b599c4", "0xe03390fe9914ebb2", "0x3b9294b3891e7ed1", "0x4304325fd80e3bda", "0x0ae7281886301b04", "0x08795a399354da80", "0x00d90c98a85c30f9" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/w4/kernel.cl b/proto-cuda/packs-readwidth/w4/kernel.cl new file mode 100644 index 000000000..3736067d3 --- /dev/null +++ b/proto-cuda/packs-readwidth/w4/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r2 = r1 * r1 + r2; // 1 mad + r2 = r3 * r2 + r2; // 2 mad + r3 = r3 ^ r5; // 3 xor + r7 = r7 ^ ds[r2 & mask]; // 4 load + r5 = r5 ^ ds[r7 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r1 = r1 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r7 = r7 ^ t_; } // 7 shfl + r1 = mul_hi(r1, r5); // 8 mulhi + r6 = rotr_var(r6, r3); // 9 rotr + r3 = r3 | r4; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r0 = mul_hi(r0, r4); // 12 mulhi + r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add + r0 = r0 ^ ds[r4 & mask]; // 14 load + r2 = r2 - r4; // 15 sub + r2 = r2 ^ ds[r0 & mask]; // 16 load + r7 = r7 ^ ds[r2 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl + r5 = r5 * r0; // 19 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl + r6 = mul_hi(r6, r2); // 22 mulhi + r6 = r6 ^ ds[r1 & mask]; // 23 load + r5 = r5 * r0; // 24 mul + r5 = rotl_imm(r5, 19u); // 25 rotl + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl + r0 = r0 ^ r5; // 27 xor + r0 = r0 ^ r4; // 28 xor + r3 = r3 - r0; // 29 sub + r5 = r5 * r1; // 30 mul + r7 = r7 ^ ds[r2 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r5 = r5 ^ r6; // 33 xor + r5 = r5 ^ ds[r1 & mask]; // 34 load + r0 = mul_hi(r0, r5); // 35 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl + r7 = r7 ^ ds[r0 & mask]; // 37 load + r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl + r2 = r2 ^ r5; // 40 xor + r3 = r6 * r3 + r3; // 41 mad + r6 = r6 - r7; // 42 sub + r7 = r7 ^ r0; // 43 xor + r1 = r1 ^ ds[r7 & mask]; // 44 load + r2 = r2 * r3; // 45 mul + r1 = mul_hi(r1, r5); // 46 mulhi + r4 = r4 - r3; // 47 sub + r2 = rotr_var(r2, r6); // 48 rotr + r3 = r3 ^ ds[r5 & mask]; // 49 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add + r0 = r0 * r2; // 51 mul + r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add + r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add + r7 = rotl_imm(r7, 14u); // 54 rotl + r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add + r6 = r6 ^ ds[r7 & mask]; // 56 load + r1 = rotr_var(r1, r5); // 57 rotr + r5 = r5 ^ ds[r4 & mask]; // 58 load + r6 = r6 ^ ds[r2 & mask]; // 59 load + r3 = r5 * r0 + r3; // 60 mad + r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add + r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add + r5 = rotl_imm(r5, 19u); // 63 rotl + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/w4/kernel.cu b/proto-cuda/packs-readwidth/w4/kernel.cu new file mode 100644 index 000000000..e22472314 --- /dev/null +++ b/proto-cuda/packs-readwidth/w4/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r2 = r1 * r1 + r2; // 1 mad + r2 = r3 * r2 + r2; // 2 mad + r3 = r3 ^ r5; // 3 xor + r7 = r7 ^ ds[r2 & mask]; // 4 load + r5 = r5 ^ ds[r7 & mask]; // 5 load + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl + r1 = __umulhi(r1, r5); // 8 mulhi + r6 = rotr_var(r6, r3); // 9 rotr + r3 = r3 | r4; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r0 = __umulhi(r0, r4); // 12 mulhi + r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add + r0 = r0 ^ ds[r4 & mask]; // 14 load + r2 = r2 - r4; // 15 sub + r2 = r2 ^ ds[r0 & mask]; // 16 load + r7 = r7 ^ ds[r2 & mask]; // 17 load + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl + r5 = r5 * r0; // 19 mul + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl + r6 = __umulhi(r6, r2); // 22 mulhi + r6 = r6 ^ ds[r1 & mask]; // 23 load + r5 = r5 * r0; // 24 mul + r5 = rotl_imm(r5, 19u); // 25 rotl + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl + r0 = r0 ^ r5; // 27 xor + r0 = r0 ^ r4; // 28 xor + r3 = r3 - r0; // 29 sub + r5 = r5 * r1; // 30 mul + r7 = r7 ^ ds[r2 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r5 = r5 ^ r6; // 33 xor + r5 = r5 ^ ds[r1 & mask]; // 34 load + r0 = __umulhi(r0, r5); // 35 mulhi + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl + r7 = r7 ^ ds[r0 & mask]; // 37 load + r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl + r2 = r2 ^ r5; // 40 xor + r3 = r6 * r3 + r3; // 41 mad + r6 = r6 - r7; // 42 sub + r7 = r7 ^ r0; // 43 xor + r1 = r1 ^ ds[r7 & mask]; // 44 load + r2 = r2 * r3; // 45 mul + r1 = __umulhi(r1, r5); // 46 mulhi + r4 = r4 - r3; // 47 sub + r2 = rotr_var(r2, r6); // 48 rotr + r3 = r3 ^ ds[r5 & mask]; // 49 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add + r0 = r0 * r2; // 51 mul + r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add + r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add + r7 = rotl_imm(r7, 14u); // 54 rotl + r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add + r6 = r6 ^ ds[r7 & mask]; // 56 load + r1 = rotr_var(r1, r5); // 57 rotr + r5 = r5 ^ ds[r4 & mask]; // 58 load + r6 = r6 ^ ds[r2 & mask]; // 59 load + r3 = r5 * r0 + r3; // 60 mad + r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add + r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add + r5 = rotl_imm(r5, 19u); // 63 rotl + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/w4/kernel_bound.cl b/proto-cuda/packs-readwidth/w4/kernel_bound.cl new file mode 100644 index 000000000..16893c594 --- /dev/null +++ b/proto-cuda/packs-readwidth/w4/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r2 = r1 * r1 + r2; // 1 mad + r2 = r3 * r2 + r2; // 2 mad + r3 = r3 ^ r5; // 3 xor + r7 = r7 ^ ds[r2 & mask]; // 4 load + r5 = r5 ^ ds[r7 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r1 = r1 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r7 = r7 ^ t_; } // 7 shfl + r1 = mul_hi(r1, r5); // 8 mulhi + r6 = rotr_var(r6, r3); // 9 rotr + r3 = r3 | r4; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r0 = mul_hi(r0, r4); // 12 mulhi + r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add + r0 = r0 ^ ds[r4 & mask]; // 14 load + r2 = r2 - r4; // 15 sub + r2 = r2 ^ ds[r0 & mask]; // 16 load + r7 = r7 ^ ds[r2 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl + r5 = r5 * r0; // 19 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl + r6 = mul_hi(r6, r2); // 22 mulhi + r6 = r6 ^ ds[r1 & mask]; // 23 load + r5 = r5 * r0; // 24 mul + r5 = rotl_imm(r5, 19u); // 25 rotl + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl + r0 = r0 ^ r5; // 27 xor + r0 = r0 ^ r4; // 28 xor + r3 = r3 - r0; // 29 sub + r5 = r5 * r1; // 30 mul + r7 = r7 ^ ds[r2 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r5 = r5 ^ r6; // 33 xor + r5 = r5 ^ ds[r1 & mask]; // 34 load + r0 = mul_hi(r0, r5); // 35 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl + r7 = r7 ^ ds[r0 & mask]; // 37 load + r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl + r2 = r2 ^ r5; // 40 xor + r3 = r6 * r3 + r3; // 41 mad + r6 = r6 - r7; // 42 sub + r7 = r7 ^ r0; // 43 xor + r1 = r1 ^ ds[r7 & mask]; // 44 load + r2 = r2 * r3; // 45 mul + r1 = mul_hi(r1, r5); // 46 mulhi + r4 = r4 - r3; // 47 sub + r2 = rotr_var(r2, r6); // 48 rotr + r3 = r3 ^ ds[r5 & mask]; // 49 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add + r0 = r0 * r2; // 51 mul + r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add + r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add + r7 = rotl_imm(r7, 14u); // 54 rotl + r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add + r6 = r6 ^ ds[r7 & mask]; // 56 load + r1 = rotr_var(r1, r5); // 57 rotr + r5 = r5 ^ ds[r4 & mask]; // 58 load + r6 = r6 ^ ds[r2 & mask]; // 59 load + r3 = r5 * r0 + r3; // 60 mad + r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add + r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add + r5 = rotl_imm(r5, 19u); // 63 rotl + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r2 = r1 * r1 + r2; // 1 mad + r2 = r3 * r2 + r2; // 2 mad + r3 = r3 ^ r5; // 3 xor + r7 = r7 ^ ds[r2 & mask]; // 4 load + r5 = r5 ^ ds[r7 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r1 = r1 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r7 = r7 ^ t_; } // 7 shfl + r1 = mul_hi(r1, r5); // 8 mulhi + r6 = rotr_var(r6, r3); // 9 rotr + r3 = r3 | r4; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r0 = mul_hi(r0, r4); // 12 mulhi + r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add + r0 = r0 ^ ds[r4 & mask]; // 14 load + r2 = r2 - r4; // 15 sub + r2 = r2 ^ ds[r0 & mask]; // 16 load + r7 = r7 ^ ds[r2 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl + r5 = r5 * r0; // 19 mul + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl + r6 = mul_hi(r6, r2); // 22 mulhi + r6 = r6 ^ ds[r1 & mask]; // 23 load + r5 = r5 * r0; // 24 mul + r5 = rotl_imm(r5, 19u); // 25 rotl + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl + r0 = r0 ^ r5; // 27 xor + r0 = r0 ^ r4; // 28 xor + r3 = r3 - r0; // 29 sub + r5 = r5 * r1; // 30 mul + r7 = r7 ^ ds[r2 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r5 = r5 ^ r6; // 33 xor + r5 = r5 ^ ds[r1 & mask]; // 34 load + r0 = mul_hi(r0, r5); // 35 mulhi + { uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl + r7 = r7 ^ ds[r0 & mask]; // 37 load + r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl + r2 = r2 ^ r5; // 40 xor + r3 = r6 * r3 + r3; // 41 mad + r6 = r6 - r7; // 42 sub + r7 = r7 ^ r0; // 43 xor + r1 = r1 ^ ds[r7 & mask]; // 44 load + r2 = r2 * r3; // 45 mul + r1 = mul_hi(r1, r5); // 46 mulhi + r4 = r4 - r3; // 47 sub + r2 = rotr_var(r2, r6); // 48 rotr + r3 = r3 ^ ds[r5 & mask]; // 49 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add + r0 = r0 * r2; // 51 mul + r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add + r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add + r7 = rotl_imm(r7, 14u); // 54 rotl + r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add + r6 = r6 ^ ds[r7 & mask]; // 56 load + r1 = rotr_var(r1, r5); // 57 rotr + r5 = r5 ^ ds[r4 & mask]; // 58 load + r6 = r6 ^ ds[r2 & mask]; // 59 load + r3 = r5 * r0 + r3; // 60 mad + r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add + r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add + r5 = rotl_imm(r5, 19u); // 63 rotl + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/w4/kernel_bound.cu b/proto-cuda/packs-readwidth/w4/kernel_bound.cu new file mode 100644 index 000000000..e8f1a9a79 --- /dev/null +++ b/proto-cuda/packs-readwidth/w4/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r2 = r1 * r1 + r2; // 1 mad + r2 = r3 * r2 + r2; // 2 mad + r3 = r3 ^ r5; // 3 xor + r7 = r7 ^ ds[r2 & mask]; // 4 load + r5 = r5 ^ ds[r7 & mask]; // 5 load + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl + r1 = __umulhi(r1, r5); // 8 mulhi + r6 = rotr_var(r6, r3); // 9 rotr + r3 = r3 | r4; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r0 = __umulhi(r0, r4); // 12 mulhi + r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add + r0 = r0 ^ ds[r4 & mask]; // 14 load + r2 = r2 - r4; // 15 sub + r2 = r2 ^ ds[r0 & mask]; // 16 load + r7 = r7 ^ ds[r2 & mask]; // 17 load + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl + r5 = r5 * r0; // 19 mul + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl + r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl + r6 = __umulhi(r6, r2); // 22 mulhi + r6 = r6 ^ ds[r1 & mask]; // 23 load + r5 = r5 * r0; // 24 mul + r5 = rotl_imm(r5, 19u); // 25 rotl + r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl + r0 = r0 ^ r5; // 27 xor + r0 = r0 ^ r4; // 28 xor + r3 = r3 - r0; // 29 sub + r5 = r5 * r1; // 30 mul + r7 = r7 ^ ds[r2 & mask]; // 31 load + r1 = r1 ^ ds[r0 & mask]; // 32 load + r5 = r5 ^ r6; // 33 xor + r5 = r5 ^ ds[r1 & mask]; // 34 load + r0 = __umulhi(r0, r5); // 35 mulhi + r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl + r7 = r7 ^ ds[r0 & mask]; // 37 load + r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl + r2 = r2 ^ r5; // 40 xor + r3 = r6 * r3 + r3; // 41 mad + r6 = r6 - r7; // 42 sub + r7 = r7 ^ r0; // 43 xor + r1 = r1 ^ ds[r7 & mask]; // 44 load + r2 = r2 * r3; // 45 mul + r1 = __umulhi(r1, r5); // 46 mulhi + r4 = r4 - r3; // 47 sub + r2 = rotr_var(r2, r6); // 48 rotr + r3 = r3 ^ ds[r5 & mask]; // 49 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add + r0 = r0 * r2; // 51 mul + r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add + r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add + r7 = rotl_imm(r7, 14u); // 54 rotl + r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add + r6 = r6 ^ ds[r7 & mask]; // 56 load + r1 = rotr_var(r1, r5); // 57 rotr + r5 = r5 ^ ds[r4 & mask]; // 58 load + r6 = r6 ^ ds[r2 & mask]; // 59 load + r3 = r5 * r0 + r3; // 60 mad + r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add + r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add + r5 = rotl_imm(r5, 19u); // 63 rotl + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/w4/memhard.h b/proto-cuda/packs-readwidth/w4/memhard.h new file mode 100644 index 000000000..4803d8e40 --- /dev/null +++ b/proto-cuda/packs-readwidth/w4/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/w4/memhard.metal b/proto-cuda/packs-readwidth/w4/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/w4/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/w4/program.h b/proto-cuda/packs-readwidth/w4/program.h new file mode 100644 index 000000000..3c4d47df0 --- /dev/null +++ b/proto-cuda/packs-readwidth/w4/program.h @@ -0,0 +1,51 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-genesis" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0xbcc1248b10cc90f2ull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1" +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/w4/program.json b/proto-cuda/packs-readwidth/w4/program.json new file mode 100644 index 000000000..e6f74a72b --- /dev/null +++ b/proto-cuda/packs-readwidth/w4/program.json @@ -0,0 +1,121 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0xbcc1248b10cc90f2", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-genesis", + "seed_bytes": "69676e65756d2d67656e65736973", + "seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "op_mix": {"load": 16, "add": 8, "shfl": 8, "xor": 6, "mad": 5, "mul": 5, "mulhi": 5, "sub": 4, "rotl": 3, "rotr": 3, "or": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2}, + {"i": 1, "op": "mad", "dst": 2, "src": 1, "src2": 1, "imm": "0xdd04a5da", "imm2": "0x42da7657", "rot": 15, "bit": 30, "mask": 16}, + {"i": 2, "op": "mad", "dst": 2, "src": 3, "src2": 2, "imm": "0x734003fa", "imm2": "0x5bb67700", "rot": 3, "bit": 20, "mask": 1}, + {"i": 3, "op": "xor", "dst": 3, "src": 5, "src2": 5, "imm": "0xc55a1b1c", "imm2": "0xa19720f3", "rot": 7, "bit": 8, "mask": 1}, + {"i": 4, "op": "load", "dst": 7, "src": 2, "src2": 5, "imm": "0xad572dd7", "imm2": "0x9ceb3ea7", "rot": 18, "bit": 30, "mask": 2}, + {"i": 5, "op": "load", "dst": 5, "src": 7, "src2": 3, "imm": "0x769a53be", "imm2": "0x80f9067e", "rot": 12, "bit": 22, "mask": 1}, + {"i": 6, "op": "shfl", "dst": 1, "src": 4, "src2": 0, "imm": "0xd3613d88", "imm2": "0x262fb219", "rot": 10, "bit": 30, "mask": 8}, + {"i": 7, "op": "shfl", "dst": 7, "src": 3, "src2": 4, "imm": "0xce38e42f", "imm2": "0xb868b818", "rot": 11, "bit": 8, "mask": 8}, + {"i": 8, "op": "mulhi", "dst": 1, "src": 5, "src2": 2, "imm": "0xa5eebca5", "imm2": "0x5703a72b", "rot": 13, "bit": 13, "mask": 16}, + {"i": 9, "op": "rotr", "dst": 6, "src": 3, "src2": 4, "imm": "0x17a5a9c7", "imm2": "0xdcfb93a1", "rot": 20, "bit": 27, "mask": 2}, + {"i": 10, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4}, + {"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 2, "imm": "0x88cb9af3", "imm2": "0x4e7dc10d", "rot": 24, "bit": 17, "mask": 4}, + {"i": 12, "op": "mulhi", "dst": 0, "src": 4, "src2": 4, "imm": "0x45374321", "imm2": "0x3cd91989", "rot": 11, "bit": 4, "mask": 2}, + {"i": 13, "op": "add", "dst": 5, "src": 1, "src2": 2, "imm": "0xc7934706", "imm2": "0xd3177981", "rot": 16, "bit": 30, "mask": 2}, + {"i": 14, "op": "load", "dst": 0, "src": 4, "src2": 3, "imm": "0xfd7f56bb", "imm2": "0x65e14f52", "rot": 13, "bit": 22, "mask": 2}, + {"i": 15, "op": "sub", "dst": 2, "src": 4, "src2": 4, "imm": "0x35a80b49", "imm2": "0x060f2d13", "rot": 16, "bit": 20, "mask": 16}, + {"i": 16, "op": "load", "dst": 2, "src": 0, "src2": 2, "imm": "0xae0a32c2", "imm2": "0x4c2a4cfe", "rot": 8, "bit": 31, "mask": 16}, + {"i": 17, "op": "load", "dst": 7, "src": 2, "src2": 6, "imm": "0x82a84cc3", "imm2": "0x21a38d68", "rot": 15, "bit": 21, "mask": 2}, + {"i": 18, "op": "shfl", "dst": 7, "src": 3, "src2": 3, "imm": "0xa3818806", "imm2": "0x8f66b5c8", "rot": 14, "bit": 6, "mask": 4}, + {"i": 19, "op": "mul", "dst": 5, "src": 0, "src2": 1, "imm": "0xa00de107", "imm2": "0x77bfcaa5", "rot": 3, "bit": 10, "mask": 2}, + {"i": 20, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2}, + {"i": 21, "op": "shfl", "dst": 2, "src": 4, "src2": 5, "imm": "0x3ac915d2", "imm2": "0x5fba7bc2", "rot": 16, "bit": 1, "mask": 16}, + {"i": 22, "op": "mulhi", "dst": 6, "src": 2, "src2": 0, "imm": "0xdc3ec8fd", "imm2": "0x599e2fa3", "rot": 22, "bit": 3, "mask": 2}, + {"i": 23, "op": "load", "dst": 6, "src": 1, "src2": 5, "imm": "0x2a6b16d5", "imm2": "0xd73e396f", "rot": 28, "bit": 29, "mask": 2}, + {"i": 24, "op": "mul", "dst": 5, "src": 0, "src2": 7, "imm": "0x376d0223", "imm2": "0xe1c2169a", "rot": 4, "bit": 16, "mask": 16}, + {"i": 25, "op": "rotl", "dst": 5, "src": 7, "src2": 3, "imm": "0x78ad8c60", "imm2": "0x6f5b77d5", "rot": 19, "bit": 11, "mask": 16}, + {"i": 26, "op": "shfl", "dst": 7, "src": 6, "src2": 5, "imm": "0x93915b9f", "imm2": "0x1e61fb6b", "rot": 28, "bit": 23, "mask": 2}, + {"i": 27, "op": "xor", "dst": 0, "src": 5, "src2": 7, "imm": "0x6378fe15", "imm2": "0x66c78f42", "rot": 12, "bit": 31, "mask": 8}, + {"i": 28, "op": "xor", "dst": 0, "src": 4, "src2": 7, "imm": "0x20a57fda", "imm2": "0x088c848e", "rot": 16, "bit": 13, "mask": 4}, + {"i": 29, "op": "sub", "dst": 3, "src": 0, "src2": 2, "imm": "0x49d95fd5", "imm2": "0x1a5f946a", "rot": 6, "bit": 12, "mask": 1}, + {"i": 30, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4}, + {"i": 31, "op": "load", "dst": 7, "src": 2, "src2": 2, "imm": "0x09ed045e", "imm2": "0xd69c4715", "rot": 5, "bit": 9, "mask": 2}, + {"i": 32, "op": "load", "dst": 1, "src": 0, "src2": 6, "imm": "0xeb79ea49", "imm2": "0xcc587f5a", "rot": 6, "bit": 8, "mask": 16}, + {"i": 33, "op": "xor", "dst": 5, "src": 6, "src2": 1, "imm": "0x3027401e", "imm2": "0x5f20c27e", "rot": 18, "bit": 9, "mask": 2}, + {"i": 34, "op": "load", "dst": 5, "src": 1, "src2": 3, "imm": "0x0e1cab07", "imm2": "0x09356c5b", "rot": 19, "bit": 31, "mask": 1}, + {"i": 35, "op": "mulhi", "dst": 0, "src": 5, "src2": 2, "imm": "0x90e31357", "imm2": "0xabd32484", "rot": 26, "bit": 5, "mask": 8}, + {"i": 36, "op": "shfl", "dst": 5, "src": 2, "src2": 5, "imm": "0xee9a955f", "imm2": "0x31b3faed", "rot": 8, "bit": 24, "mask": 4}, + {"i": 37, "op": "load", "dst": 7, "src": 0, "src2": 7, "imm": "0x3ba2f832", "imm2": "0x1160dcd3", "rot": 4, "bit": 29, "mask": 1}, + {"i": 38, "op": "add", "dst": 3, "src": 1, "src2": 7, "imm": "0x75ba2fad", "imm2": "0x230c005c", "rot": 4, "bit": 27, "mask": 1}, + {"i": 39, "op": "shfl", "dst": 1, "src": 5, "src2": 3, "imm": "0xdbf37e75", "imm2": "0xb5ac1969", "rot": 30, "bit": 13, "mask": 4}, + {"i": 40, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2}, + {"i": 41, "op": "mad", "dst": 3, "src": 6, "src2": 3, "imm": "0xce13eff8", "imm2": "0x04cc1d55", "rot": 3, "bit": 1, "mask": 4}, + {"i": 42, "op": "sub", "dst": 6, "src": 7, "src2": 1, "imm": "0x6a65ab71", "imm2": "0x8fbc1bcd", "rot": 4, "bit": 1, "mask": 8}, + {"i": 43, "op": "xor", "dst": 7, "src": 0, "src2": 7, "imm": "0xdaeb4928", "imm2": "0xc0423027", "rot": 24, "bit": 11, "mask": 8}, + {"i": 44, "op": "load", "dst": 1, "src": 7, "src2": 7, "imm": "0x778f01c9", "imm2": "0x28cedcea", "rot": 12, "bit": 4, "mask": 16}, + {"i": 45, "op": "mul", "dst": 2, "src": 3, "src2": 2, "imm": "0xf4264f1b", "imm2": "0x0f627d56", "rot": 5, "bit": 28, "mask": 8}, + {"i": 46, "op": "mulhi", "dst": 1, "src": 5, "src2": 1, "imm": "0xffb2147a", "imm2": "0xccde9b05", "rot": 13, "bit": 9, "mask": 2}, + {"i": 47, "op": "sub", "dst": 4, "src": 3, "src2": 5, "imm": "0x73b36234", "imm2": "0x3f5d5997", "rot": 7, "bit": 18, "mask": 2}, + {"i": 48, "op": "rotr", "dst": 2, "src": 6, "src2": 3, "imm": "0x3a4d9aa9", "imm2": "0x212bec7b", "rot": 4, "bit": 29, "mask": 16}, + {"i": 49, "op": "load", "dst": 3, "src": 5, "src2": 2, "imm": "0x626f11df", "imm2": "0x56cd5bfd", "rot": 7, "bit": 1, "mask": 1}, + {"i": 50, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4}, + {"i": 51, "op": "mul", "dst": 0, "src": 2, "src2": 2, "imm": "0xa8848b30", "imm2": "0xef6ac348", "rot": 9, "bit": 15, "mask": 8}, + {"i": 52, "op": "add", "dst": 0, "src": 2, "src2": 0, "imm": "0x4f92b968", "imm2": "0x699fd448", "rot": 22, "bit": 6, "mask": 4}, + {"i": 53, "op": "add", "dst": 1, "src": 0, "src2": 2, "imm": "0x2bb965af", "imm2": "0x77b1520d", "rot": 2, "bit": 12, "mask": 8}, + {"i": 54, "op": "rotl", "dst": 7, "src": 1, "src2": 0, "imm": "0x553e678b", "imm2": "0x3cc8eae0", "rot": 14, "bit": 20, "mask": 2}, + {"i": 55, "op": "add", "dst": 3, "src": 7, "src2": 2, "imm": "0x7b0fe07a", "imm2": "0xa54c55a0", "rot": 10, "bit": 1, "mask": 1}, + {"i": 56, "op": "load", "dst": 6, "src": 7, "src2": 4, "imm": "0x01eba9aa", "imm2": "0x2758c0f7", "rot": 14, "bit": 15, "mask": 4}, + {"i": 57, "op": "rotr", "dst": 1, "src": 5, "src2": 2, "imm": "0x1f5267b3", "imm2": "0x236f5a27", "rot": 2, "bit": 31, "mask": 16}, + {"i": 58, "op": "load", "dst": 5, "src": 4, "src2": 3, "imm": "0xa9954a9b", "imm2": "0x6a54d4e8", "rot": 11, "bit": 10, "mask": 16}, + {"i": 59, "op": "load", "dst": 6, "src": 2, "src2": 4, "imm": "0x9923ff88", "imm2": "0x9357254e", "rot": 16, "bit": 1, "mask": 16}, + {"i": 60, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4}, + {"i": 61, "op": "add", "dst": 5, "src": 7, "src2": 7, "imm": "0xaf9dd72d", "imm2": "0xad7493e7", "rot": 7, "bit": 31, "mask": 16}, + {"i": 62, "op": "add", "dst": 4, "src": 6, "src2": 2, "imm": "0x89841d87", "imm2": "0x1e07c3d9", "rot": 6, "bit": 27, "mask": 1}, + {"i": 63, "op": "rotl", "dst": 5, "src": 4, "src2": 2, "imm": "0xae210f8d", "imm2": "0x8e499ba4", "rot": 19, "bit": 9, "mask": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/w4/program.metal b/proto-cuda/packs-readwidth/w4/program.metal new file mode 100644 index 000000000..d90be2f77 --- /dev/null +++ b/proto-cuda/packs-readwidth/w4/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r2 = r1 * r1 + r2; // 1 + r2 = r3 * r2 + r2; // 2 + r3 = r3 ^ r5; // 3 + r7 = r7 ^ dataset[r2 & MASK]; // 4 + r5 = r5 ^ dataset[r7 & MASK]; // 5 + r1 = r1 ^ simd_shuffle_xor(r4, (ushort)8); // 6 + r7 = r7 ^ simd_shuffle_xor(r3, (ushort)8); // 7 + r1 = mulhi(r1, r5); // 8 + r6 = rotr_var(r6, r3); // 9 + r3 = r3 | r4; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r0 = mulhi(r0, r4); // 12 + r5 = r5 + r1 + select(0xc7934706u, 0xd3177981u, ((sel >> 30u) & 1u) != 0u); // 13 + r0 = r0 ^ dataset[r4 & MASK]; // 14 + r2 = r2 - r4; // 15 + r2 = r2 ^ dataset[r0 & MASK]; // 16 + r7 = r7 ^ dataset[r2 & MASK]; // 17 + r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 18 + r5 = r5 * r0; // 19 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 20 + r2 = r2 ^ simd_shuffle_xor(r4, (ushort)16); // 21 + r6 = mulhi(r6, r2); // 22 + r6 = r6 ^ dataset[r1 & MASK]; // 23 + r5 = r5 * r0; // 24 + r5 = rotl_imm(r5, 19u); // 25 + r7 = r7 ^ simd_shuffle_xor(r6, (ushort)2); // 26 + r0 = r0 ^ r5; // 27 + r0 = r0 ^ r4; // 28 + r3 = r3 - r0; // 29 + r5 = r5 * r1; // 30 + r7 = r7 ^ dataset[r2 & MASK]; // 31 + r1 = r1 ^ dataset[r0 & MASK]; // 32 + r5 = r5 ^ r6; // 33 + r5 = r5 ^ dataset[r1 & MASK]; // 34 + r0 = mulhi(r0, r5); // 35 + r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 36 + r7 = r7 ^ dataset[r0 & MASK]; // 37 + r3 = r3 + r1 + select(0x75ba2fadu, 0x230c005cu, ((sel >> 27u) & 1u) != 0u); // 38 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)4); // 39 + r2 = r2 ^ r5; // 40 + r3 = r6 * r3 + r3; // 41 + r6 = r6 - r7; // 42 + r7 = r7 ^ r0; // 43 + r1 = r1 ^ dataset[r7 & MASK]; // 44 + r2 = r2 * r3; // 45 + r1 = mulhi(r1, r5); // 46 + r4 = r4 - r3; // 47 + r2 = rotr_var(r2, r6); // 48 + r3 = r3 ^ dataset[r5 & MASK]; // 49 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 50 + r0 = r0 * r2; // 51 + r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 52 + r1 = r1 + r0 + select(0x2bb965afu, 0x77b1520du, ((sel >> 12u) & 1u) != 0u); // 53 + r7 = rotl_imm(r7, 14u); // 54 + r3 = r3 + r7 + select(0x7b0fe07au, 0xa54c55a0u, ((sel >> 1u) & 1u) != 0u); // 55 + r6 = r6 ^ dataset[r7 & MASK]; // 56 + r1 = rotr_var(r1, r5); // 57 + r5 = r5 ^ dataset[r4 & MASK]; // 58 + r6 = r6 ^ dataset[r2 & MASK]; // 59 + r3 = r5 * r0 + r3; // 60 + r5 = r5 + r7 + select(0xaf9dd72du, 0xad7493e7u, ((sel >> 31u) & 1u) != 0u); // 61 + r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 62 + r5 = rotl_imm(r5, 19u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/w4/program_bound.metal b/proto-cuda/packs-readwidth/w4/program_bound.metal new file mode 100644 index 000000000..8fafce718 --- /dev/null +++ b/proto-cuda/packs-readwidth/w4/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r2 = r1 * r1 + r2; // 1 + r2 = r3 * r2 + r2; // 2 + r3 = r3 ^ r5; // 3 + r7 = r7 ^ dataset[r2 & MASK]; // 4 + r5 = r5 ^ dataset[r7 & MASK]; // 5 + r1 = r1 ^ simd_shuffle_xor(r4, (ushort)8); // 6 + r7 = r7 ^ simd_shuffle_xor(r3, (ushort)8); // 7 + r1 = mulhi(r1, r5); // 8 + r6 = rotr_var(r6, r3); // 9 + r3 = r3 | r4; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r0 = mulhi(r0, r4); // 12 + r5 = r5 + r1 + select(0xc7934706u, 0xd3177981u, ((sel >> 30u) & 1u) != 0u); // 13 + r0 = r0 ^ dataset[r4 & MASK]; // 14 + r2 = r2 - r4; // 15 + r2 = r2 ^ dataset[r0 & MASK]; // 16 + r7 = r7 ^ dataset[r2 & MASK]; // 17 + r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 18 + r5 = r5 * r0; // 19 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 20 + r2 = r2 ^ simd_shuffle_xor(r4, (ushort)16); // 21 + r6 = mulhi(r6, r2); // 22 + r6 = r6 ^ dataset[r1 & MASK]; // 23 + r5 = r5 * r0; // 24 + r5 = rotl_imm(r5, 19u); // 25 + r7 = r7 ^ simd_shuffle_xor(r6, (ushort)2); // 26 + r0 = r0 ^ r5; // 27 + r0 = r0 ^ r4; // 28 + r3 = r3 - r0; // 29 + r5 = r5 * r1; // 30 + r7 = r7 ^ dataset[r2 & MASK]; // 31 + r1 = r1 ^ dataset[r0 & MASK]; // 32 + r5 = r5 ^ r6; // 33 + r5 = r5 ^ dataset[r1 & MASK]; // 34 + r0 = mulhi(r0, r5); // 35 + r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 36 + r7 = r7 ^ dataset[r0 & MASK]; // 37 + r3 = r3 + r1 + select(0x75ba2fadu, 0x230c005cu, ((sel >> 27u) & 1u) != 0u); // 38 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)4); // 39 + r2 = r2 ^ r5; // 40 + r3 = r6 * r3 + r3; // 41 + r6 = r6 - r7; // 42 + r7 = r7 ^ r0; // 43 + r1 = r1 ^ dataset[r7 & MASK]; // 44 + r2 = r2 * r3; // 45 + r1 = mulhi(r1, r5); // 46 + r4 = r4 - r3; // 47 + r2 = rotr_var(r2, r6); // 48 + r3 = r3 ^ dataset[r5 & MASK]; // 49 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 50 + r0 = r0 * r2; // 51 + r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 52 + r1 = r1 + r0 + select(0x2bb965afu, 0x77b1520du, ((sel >> 12u) & 1u) != 0u); // 53 + r7 = rotl_imm(r7, 14u); // 54 + r3 = r3 + r7 + select(0x7b0fe07au, 0xa54c55a0u, ((sel >> 1u) & 1u) != 0u); // 55 + r6 = r6 ^ dataset[r7 & MASK]; // 56 + r1 = rotr_var(r1, r5); // 57 + r5 = r5 ^ dataset[r4 & MASK]; // 58 + r6 = r6 ^ dataset[r2 & MASK]; // 59 + r3 = r5 * r0 + r3; // 60 + r5 = r5 + r7 + select(0xaf9dd72du, 0xad7493e7u, ((sel >> 31u) & 1u) != 0u); // 61 + r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 62 + r5 = rotl_imm(r5, 19u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/w4/vectors.h b/proto-cuda/packs-readwidth/w4/vectors.h new file mode 100644 index 000000000..658504dc8 --- /dev/null +++ b/proto-cuda/packs-readwidth/w4/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x42246ba99fc58e4full, 0x19561eec0db4f9f4ull, 0x11a4fb70ea7b688full, 0x8872bfadec960949ull, 0x085e6cf2ff6b8303ull, 0x780e0d76e504fe6cull, 0x7bb05a713f8283faull, 0xfef7beeffd5ab1d1ull, + 0x1b83afb12e78d65bull, 0x2576c5384a2e1caeull, 0x549d620a0735d6d8ull, 0x2eb4e32f32e8d5f1ull, 0x81bb414492f69586ull, 0x092d17324b465e01ull, 0x8bde350ee2354b5aull, 0xd64b47e9c9e0ec07ull, + 0xdbbf1b78b9979a3full, 0xbee382abde89c111ull, 0x6b598d99c16a8c70ull, 0x69d72da59fb3b6e9ull, 0x1e309d2a549632faull, 0x98fa8255ab65f005ull, 0x2f48ab1bb516110cull, 0x7d2af17cadd18beaull, + 0x547c5978bd005e02ull, 0x25ea2e21e88ea9d3ull, 0x1c575ec6e43efc58ull, 0x077b80cb079958b4ull, 0xa96eebd0c8634981ull, 0x44b18bdf19eb7838ull, 0xc3ccb9fe9ef5953cull, 0xb08446b1f2de7793ull + }, + { // base nonce 4096 + 0x3d3903e310ca038full, 0x9e72b9a86ebb29e4ull, 0x1dca847362709effull, 0x4d0b021e131d6a4full, 0x88de07bd1539fa72ull, 0x042a436373b96ffeull, 0x3fcb1ef447b976fdull, 0x0237159348bd27fbull, + 0x282fb8b521df5e60ull, 0xc16d95f79fd99903ull, 0xd75a20a4d69df082ull, 0x809381bf5581fd22ull, 0xea6e9ed106945107ull, 0x3ef5fadeb7ac79d9ull, 0x75a4c882c4d02e44ull, 0x66b95d3105a88bf4ull, + 0xca7e020f5d415c29ull, 0x32dc62b7072cd5e4ull, 0x253fd43c6e7c2d57ull, 0x228235093e06d1afull, 0x1085653fe1d382a4ull, 0xe6bcd4d4693fd10dull, 0x065058e340b134a5ull, 0x4febdd41ea5e1913ull, + 0x8a9fba060f8c1d9full, 0x2be8ec07bf80f713ull, 0x476855e9a37e1b80ull, 0x5830538ad9e75e0aull, 0xb32a3ac396dc34bcull, 0xf521612e59ee414eull, 0xd5687c5ae184a180ull, 0x61c242509efdccddull + }, + { // base nonce 1000000 + 0xf218c1bd58e6dfe0ull, 0x8c5a362ee98971c2ull, 0x06394d22d03aa4eeull, 0x68166b204b0cf6ddull, 0x6e8e505080b2de7aull, 0xe922e3a6684d8c0aull, 0x89d0c35410c44bc9ull, 0x5a3df3edbf4d62a8ull, + 0xecf6b45e5a90ce00ull, 0x02221b7088fafc7full, 0x5d23dd2a7a4d3e16ull, 0x04e8e12e17a940f3ull, 0x4f671f11554578a6ull, 0x4e49f1b1117a97abull, 0xec0fb9eac4520ae9ull, 0x4ca8b3859d421f9bull, + 0xbf593793d02b0309ull, 0xcb11dc0a40088de6ull, 0xaa93a48f712fb2adull, 0xda314ca08359d2baull, 0xfba7742340db6445ull, 0xf0138acf64f37cb2ull, 0xfc307731b74f1d99ull, 0x660553bd0b06bd63ull, + 0xbecb2b151282f602ull, 0x373bead8040e3eb8ull, 0xbe39a4642591f3dfull, 0xb6c5a400b6586ff3ull, 0xd614efa47d7118e3ull, 0x1be1d5c73feceb6eull, 0xca6e33dd721350c4ull, 0x6c3b2c11adfbfcacull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/w4/vectors.json b/proto-cuda/packs-readwidth/w4/vectors.json new file mode 100644 index 000000000..82030ca66 --- /dev/null +++ b/proto-cuda/packs-readwidth/w4/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-genesis", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x42246ba99fc58e4f", "0x19561eec0db4f9f4", "0x11a4fb70ea7b688f", "0x8872bfadec960949", "0x085e6cf2ff6b8303", "0x780e0d76e504fe6c", "0x7bb05a713f8283fa", "0xfef7beeffd5ab1d1", + "0x1b83afb12e78d65b", "0x2576c5384a2e1cae", "0x549d620a0735d6d8", "0x2eb4e32f32e8d5f1", "0x81bb414492f69586", "0x092d17324b465e01", "0x8bde350ee2354b5a", "0xd64b47e9c9e0ec07", + "0xdbbf1b78b9979a3f", "0xbee382abde89c111", "0x6b598d99c16a8c70", "0x69d72da59fb3b6e9", "0x1e309d2a549632fa", "0x98fa8255ab65f005", "0x2f48ab1bb516110c", "0x7d2af17cadd18bea", + "0x547c5978bd005e02", "0x25ea2e21e88ea9d3", "0x1c575ec6e43efc58", "0x077b80cb079958b4", "0xa96eebd0c8634981", "0x44b18bdf19eb7838", "0xc3ccb9fe9ef5953c", "0xb08446b1f2de7793" + ]}, + {"base_nonce": 4096, "expected": [ + "0x3d3903e310ca038f", "0x9e72b9a86ebb29e4", "0x1dca847362709eff", "0x4d0b021e131d6a4f", "0x88de07bd1539fa72", "0x042a436373b96ffe", "0x3fcb1ef447b976fd", "0x0237159348bd27fb", + "0x282fb8b521df5e60", "0xc16d95f79fd99903", "0xd75a20a4d69df082", "0x809381bf5581fd22", "0xea6e9ed106945107", "0x3ef5fadeb7ac79d9", "0x75a4c882c4d02e44", "0x66b95d3105a88bf4", + "0xca7e020f5d415c29", "0x32dc62b7072cd5e4", "0x253fd43c6e7c2d57", "0x228235093e06d1af", "0x1085653fe1d382a4", "0xe6bcd4d4693fd10d", "0x065058e340b134a5", "0x4febdd41ea5e1913", + "0x8a9fba060f8c1d9f", "0x2be8ec07bf80f713", "0x476855e9a37e1b80", "0x5830538ad9e75e0a", "0xb32a3ac396dc34bc", "0xf521612e59ee414e", "0xd5687c5ae184a180", "0x61c242509efdccdd" + ]}, + {"base_nonce": 1000000, "expected": [ + "0xf218c1bd58e6dfe0", "0x8c5a362ee98971c2", "0x06394d22d03aa4ee", "0x68166b204b0cf6dd", "0x6e8e505080b2de7a", "0xe922e3a6684d8c0a", "0x89d0c35410c44bc9", "0x5a3df3edbf4d62a8", + "0xecf6b45e5a90ce00", "0x02221b7088fafc7f", "0x5d23dd2a7a4d3e16", "0x04e8e12e17a940f3", "0x4f671f11554578a6", "0x4e49f1b1117a97ab", "0xec0fb9eac4520ae9", "0x4ca8b3859d421f9b", + "0xbf593793d02b0309", "0xcb11dc0a40088de6", "0xaa93a48f712fb2ad", "0xda314ca08359d2ba", "0xfba7742340db6445", "0xf0138acf64f37cb2", "0xfc307731b74f1d99", "0x660553bd0b06bd63", + "0xbecb2b151282f602", "0x373bead8040e3eb8", "0xbe39a4642591f3df", "0xb6c5a400b6586ff3", "0xd614efa47d7118e3", "0x1be1d5c73feceb6e", "0xca6e33dd721350c4", "0x6c3b2c11adfbfcac" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/w64/kernel.cl b/proto-cuda/packs-readwidth/w64/kernel.cl new file mode 100644 index 000000000..373b3e3c4 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 4 load + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 16 load + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 31 load + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 56 load + r5 = r5 - r6; // 57 sub + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 58 load + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/w64/kernel.cu b/proto-cuda/packs-readwidth/w64/kernel.cu new file mode 100644 index 000000000..756160e3f --- /dev/null +++ b/proto-cuda/packs-readwidth/w64/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 4 load + { uint32_t b_ = (r1 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t b_ = (r6 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 16 load + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 23 load + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + { uint32_t b_ = (r1 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 31 load + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 56 load + r5 = r5 - r6; // 57 sub + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 58 load + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/w64/kernel_bound.cl b/proto-cuda/packs-readwidth/w64/kernel_bound.cl new file mode 100644 index 000000000..5968125ab --- /dev/null +++ b/proto-cuda/packs-readwidth/w64/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 4 load + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 16 load + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 31 load + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 56 load + r5 = r5 - r6; // 57 sub + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 58 load + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 4 load + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint b_ = (r6 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 16 load + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + { uint b_ = (r1 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 31 load + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint b_ = (r5 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + { uint b_ = (r2 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + { uint b_ = (r7 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 56 load + r5 = r5 - r6; // 57 sub + { uint b_ = (r3 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 58 load + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/w64/kernel_bound.cu b/proto-cuda/packs-readwidth/w64/kernel_bound.cu new file mode 100644 index 000000000..baf4d8cfb --- /dev/null +++ b/proto-cuda/packs-readwidth/w64/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 4 load + { uint32_t b_ = (r1 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t b_ = (r6 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 16 load + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 23 load + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + { uint32_t b_ = (r1 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 31 load + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 32 load + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 34 load + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint32_t b_ = (r5 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + { uint32_t b_ = (r2 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + { uint32_t b_ = (r7 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 56 load + r5 = r5 - r6; // 57 sub + { uint32_t b_ = (r3 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 58 load + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/w64/memhard.h b/proto-cuda/packs-readwidth/w64/memhard.h new file mode 100644 index 000000000..4803d8e40 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/w64/memhard.metal b/proto-cuda/packs-readwidth/w64/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/w64/program.h b/proto-cuda/packs-readwidth/w64/program.h new file mode 100644 index 000000000..c80f5a88a --- /dev/null +++ b/proto-cuda/packs-readwidth/w64/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-genesis" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x5c427d666b4ed4f4ull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=16 add=9 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 shfl=4 rotr=2 rotl=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "w64" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 0, 0, 100 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 0, 0, 16 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 8192 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/w64/program.json b/proto-cuda/packs-readwidth/w64/program.json new file mode 100644 index 000000000..48e3cfe4b --- /dev/null +++ b/proto-cuda/packs-readwidth/w64/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x5c427d666b4ed4f4", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-genesis", + "seed_bytes": "69676e65756d2d67656e65736973", + "seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "w64", + "load_slots": 16, + "load_mix_percent_4_16_64": [0, 0, 100], + "load_width_counts_4_16_64": [0, 0, 16], + "bytes_per_hash": 8192, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 16, "add": 9, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "shfl": 4, "rotr": 2, "rotl": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2, "width": 1}, + {"i": 1, "op": "add", "dst": 1, "src": 7, "src2": 2, "imm": "0x42da7657", "imm2": "0xc3bd2355", "rot": 25, "bit": 4, "mask": 16, "width": 1}, + {"i": 2, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x61f0b51c", "imm2": "0x2735a174", "rot": 4, "bit": 26, "mask": 2, "width": 1}, + {"i": 3, "op": "mad", "dst": 4, "src": 0, "src2": 6, "imm": "0x679648a8", "imm2": "0x3044ba32", "rot": 31, "bit": 31, "mask": 4, "width": 1}, + {"i": 4, "op": "load", "dst": 7, "src": 2, "src2": 6, "imm": "0x5d1ca2a2", "imm2": "0xe2481807", "rot": 24, "bit": 3, "mask": 1, "width": 16}, + {"i": 5, "op": "load", "dst": 4, "src": 1, "src2": 2, "imm": "0x987c017a", "imm2": "0xf4d60559", "rot": 2, "bit": 0, "mask": 4, "width": 16}, + {"i": 6, "op": "shfl", "dst": 6, "src": 3, "src2": 7, "imm": "0x6ea7b2df", "imm2": "0x9fce5071", "rot": 7, "bit": 15, "mask": 4, "width": 1}, + {"i": 7, "op": "shfl", "dst": 1, "src": 5, "src2": 1, "imm": "0x26a2ecde", "imm2": "0xfec6ad22", "rot": 15, "bit": 11, "mask": 8, "width": 1}, + {"i": 8, "op": "xor", "dst": 7, "src": 5, "src2": 2, "imm": "0xbe4b445c", "imm2": "0x17a5a9c7", "rot": 8, "bit": 8, "mask": 1, "width": 1}, + {"i": 9, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4, "width": 1}, + {"i": 10, "op": "or", "dst": 1, "src": 2, "src2": 3, "imm": "0x4e7dc10d", "imm2": "0x196d165c", "rot": 14, "bit": 27, "mask": 16, "width": 1}, + {"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 1, "imm": "0xc5c3b55d", "imm2": "0xec061424", "rot": 26, "bit": 27, "mask": 8, "width": 16}, + {"i": 12, "op": "or", "dst": 6, "src": 2, "src2": 3, "imm": "0x306542fe", "imm2": "0x1bb1b429", "rot": 31, "bit": 0, "mask": 2, "width": 1}, + {"i": 13, "op": "mul", "dst": 2, "src": 5, "src2": 6, "imm": "0xa672cdd3", "imm2": "0x59a4829c", "rot": 22, "bit": 13, "mask": 16, "width": 1}, + {"i": 14, "op": "load", "dst": 1, "src": 2, "src2": 5, "imm": "0x028b4d37", "imm2": "0x7bbd78ea", "rot": 15, "bit": 2, "mask": 8, "width": 16}, + {"i": 15, "op": "rotl", "dst": 7, "src": 6, "src2": 6, "imm": "0x5c88a1a7", "imm2": "0x5c628769", "rot": 1, "bit": 3, "mask": 8, "width": 1}, + {"i": 16, "op": "load", "dst": 3, "src": 6, "src2": 7, "imm": "0xbac2ae81", "imm2": "0xcbbc7bdb", "rot": 18, "bit": 8, "mask": 8, "width": 16}, + {"i": 17, "op": "load", "dst": 7, "src": 4, "src2": 2, "imm": "0xe8ab93e9", "imm2": "0xa00de107", "rot": 2, "bit": 1, "mask": 16, "width": 16}, + {"i": 18, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2, "width": 1}, + {"i": 19, "op": "mad", "dst": 4, "src": 0, "src2": 2, "imm": "0x5fba7bc2", "imm2": "0xdf099cfb", "rot": 4, "bit": 15, "mask": 16, "width": 1}, + {"i": 20, "op": "shfl", "dst": 0, "src": 6, "src2": 3, "imm": "0x0a3056de", "imm2": "0x7f0c25c3", "rot": 27, "bit": 13, "mask": 8, "width": 1}, + {"i": 21, "op": "xor", "dst": 5, "src": 7, "src2": 4, "imm": "0xbd066e1d", "imm2": "0x6d3ddc5a", "rot": 2, "bit": 29, "mask": 1, "width": 1}, + {"i": 22, "op": "mulhi", "dst": 2, "src": 5, "src2": 0, "imm": "0xc7e9887a", "imm2": "0x19ec898f", "rot": 14, "bit": 9, "mask": 1, "width": 1}, + {"i": 23, "op": "load", "dst": 3, "src": 7, "src2": 2, "imm": "0xc7fcfc8f", "imm2": "0x8528b94f", "rot": 17, "bit": 13, "mask": 4, "width": 16}, + {"i": 24, "op": "mulhi", "dst": 7, "src": 3, "src2": 5, "imm": "0xd91641e8", "imm2": "0xaf77faf2", "rot": 22, "bit": 21, "mask": 1, "width": 1}, + {"i": 25, "op": "or", "dst": 5, "src": 4, "src2": 0, "imm": "0x84c03868", "imm2": "0xf6c691b7", "rot": 29, "bit": 14, "mask": 8, "width": 1}, + {"i": 26, "op": "mad", "dst": 4, "src": 5, "src2": 2, "imm": "0x3bb2b6ba", "imm2": "0x49d95fd5", "rot": 1, "bit": 5, "mask": 8, "width": 1}, + {"i": 27, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4, "width": 1}, + {"i": 28, "op": "mulhi", "dst": 6, "src": 7, "src2": 6, "imm": "0xd69c4715", "imm2": "0xe0ebc4ce", "rot": 29, "bit": 2, "mask": 8, "width": 1}, + {"i": 29, "op": "add", "dst": 6, "src": 1, "src2": 2, "imm": "0x3b2d2124", "imm2": "0x187a9128", "rot": 1, "bit": 9, "mask": 16, "width": 1}, + {"i": 30, "op": "rotr", "dst": 6, "src": 7, "src2": 0, "imm": "0x5c64a589", "imm2": "0x61c9a38d", "rot": 17, "bit": 21, "mask": 16, "width": 1}, + {"i": 31, "op": "load", "dst": 3, "src": 1, "src2": 7, "imm": "0xc37723fa", "imm2": "0xf3b024da", "rot": 16, "bit": 27, "mask": 16, "width": 16}, + {"i": 32, "op": "load", "dst": 1, "src": 0, "src2": 7, "imm": "0xcc7972c4", "imm2": "0xad098d15", "rot": 30, "bit": 21, "mask": 8, "width": 16}, + {"i": 33, "op": "add", "dst": 0, "src": 4, "src2": 4, "imm": "0x2c35699f", "imm2": "0x351dde38", "rot": 21, "bit": 18, "mask": 4, "width": 1}, + {"i": 34, "op": "load", "dst": 0, "src": 2, "src2": 3, "imm": "0xfae8902b", "imm2": "0x5cd8306f", "rot": 5, "bit": 28, "mask": 16, "width": 16}, + {"i": 35, "op": "mul", "dst": 0, "src": 3, "src2": 1, "imm": "0x4fa3f3db", "imm2": "0xdbf37e75", "rot": 7, "bit": 18, "mask": 4, "width": 1}, + {"i": 36, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2, "width": 1}, + {"i": 37, "op": "load", "dst": 4, "src": 0, "src2": 0, "imm": "0x04cc1d55", "imm2": "0x35c52d04", "rot": 11, "bit": 14, "mask": 2, "width": 16}, + {"i": 38, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0x3958f280", "imm2": "0x8713c7e1", "rot": 5, "bit": 23, "mask": 16, "width": 1}, + {"i": 39, "op": "add", "dst": 0, "src": 3, "src2": 3, "imm": "0xa907b90b", "imm2": "0x1b053acf", "rot": 30, "bit": 25, "mask": 16, "width": 1}, + {"i": 40, "op": "rotr", "dst": 2, "src": 5, "src2": 4, "imm": "0xf8662282", "imm2": "0x10bb9e30", "rot": 8, "bit": 6, "mask": 2, "width": 1}, + {"i": 41, "op": "mul", "dst": 3, "src": 2, "src2": 4, "imm": "0x49087d74", "imm2": "0x6348b489", "rot": 17, "bit": 9, "mask": 16, "width": 1}, + {"i": 42, "op": "add", "dst": 1, "src": 5, "src2": 1, "imm": "0xa32e000c", "imm2": "0x6058c2e3", "rot": 25, "bit": 20, "mask": 8, "width": 1}, + {"i": 43, "op": "xor", "dst": 3, "src": 4, "src2": 2, "imm": "0x3dad0eb6", "imm2": "0xb97578cb", "rot": 3, "bit": 27, "mask": 1, "width": 1}, + {"i": 44, "op": "load", "dst": 3, "src": 5, "src2": 7, "imm": "0x374aec92", "imm2": "0x626f11df", "rot": 20, "bit": 18, "mask": 8, "width": 16}, + {"i": 45, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4, "width": 1}, + {"i": 46, "op": "xor", "dst": 7, "src": 1, "src2": 0, "imm": "0xef6ac348", "imm2": "0x963bb7e6", "rot": 26, "bit": 3, "mask": 8, "width": 1}, + {"i": 47, "op": "add", "dst": 0, "src": 3, "src2": 0, "imm": "0x838b5065", "imm2": "0x36360066", "rot": 3, "bit": 31, "mask": 4, "width": 1}, + {"i": 48, "op": "mulhi", "dst": 7, "src": 5, "src2": 0, "imm": "0x8458f7ac", "imm2": "0xc1c15026", "rot": 27, "bit": 15, "mask": 8, "width": 1}, + {"i": 49, "op": "load", "dst": 0, "src": 2, "src2": 4, "imm": "0x636a9dc4", "imm2": "0xac023d9b", "rot": 22, "bit": 29, "mask": 1, "width": 16}, + {"i": 50, "op": "sub", "dst": 2, "src": 6, "src2": 0, "imm": "0x2baec8c9", "imm2": "0x4390f156", "rot": 3, "bit": 12, "mask": 8, "width": 1}, + {"i": 51, "op": "sub", "dst": 7, "src": 5, "src2": 7, "imm": "0x19234061", "imm2": "0xe84dfade", "rot": 4, "bit": 19, "mask": 1, "width": 1}, + {"i": 52, "op": "xor", "dst": 2, "src": 3, "src2": 5, "imm": "0xdc2cd71e", "imm2": "0x1b5d334b", "rot": 9, "bit": 8, "mask": 8, "width": 1}, + {"i": 53, "op": "sub", "dst": 7, "src": 0, "src2": 4, "imm": "0x605c31ec", "imm2": "0x9923ff88", "rot": 28, "bit": 25, "mask": 4, "width": 1}, + {"i": 54, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4, "width": 1}, + {"i": 55, "op": "xor", "dst": 7, "src": 5, "src2": 5, "imm": "0xad7493e7", "imm2": "0x3e400372", "rot": 13, "bit": 8, "mask": 1, "width": 1}, + {"i": 56, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0x87e933c9", "imm2": "0x8c854c1b", "rot": 17, "bit": 3, "mask": 8, "width": 16}, + {"i": 57, "op": "sub", "dst": 5, "src": 6, "src2": 5, "imm": "0x11be3bc9", "imm2": "0xbbaa8e24", "rot": 6, "bit": 5, "mask": 16, "width": 1}, + {"i": 58, "op": "load", "dst": 1, "src": 3, "src2": 2, "imm": "0xa732351a", "imm2": "0xc01349cd", "rot": 14, "bit": 17, "mask": 16, "width": 16}, + {"i": 59, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0xb20547b2", "imm2": "0xc94655de", "rot": 27, "bit": 30, "mask": 1, "width": 16}, + {"i": 60, "op": "sub", "dst": 4, "src": 6, "src2": 7, "imm": "0x67cf904c", "imm2": "0x6873b216", "rot": 27, "bit": 7, "mask": 16, "width": 1}, + {"i": 61, "op": "mul", "dst": 1, "src": 2, "src2": 7, "imm": "0x93ab0bf4", "imm2": "0x96158375", "rot": 14, "bit": 0, "mask": 16, "width": 1}, + {"i": 62, "op": "mad", "dst": 3, "src": 6, "src2": 0, "imm": "0x41a443a3", "imm2": "0xe69d7919", "rot": 9, "bit": 0, "mask": 16, "width": 1}, + {"i": 63, "op": "add", "dst": 0, "src": 1, "src2": 3, "imm": "0x2fe0e98b", "imm2": "0xc88e2942", "rot": 5, "bit": 16, "mask": 16, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/w64/program.metal b/proto-cuda/packs-readwidth/w64/program.metal new file mode 100644 index 000000000..4003412be --- /dev/null +++ b/proto-cuda/packs-readwidth/w64/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 4 + { uint b_ = (r1 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint b_ = (r6 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 16 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + { uint b_ = (r1 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 31 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 56 + r5 = r5 - r6; // 57 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 58 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/w64/program_bound.metal b/proto-cuda/packs-readwidth/w64/program_bound.metal new file mode 100644 index 000000000..c3657c2b9 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 4 + { uint b_ = (r1 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint b_ = (r6 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 16 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r7 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r7 = x_; } // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + { uint b_ = (r1 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 31 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r4 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r4 = x_; } // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + { uint b_ = (r5 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + { uint b_ = (r2 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r0 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r0 = x_; } // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + { uint b_ = (r7 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r2 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r2 = x_; } // 56 + r5 = r5 - r6; // 57 + { uint b_ = (r3 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 58 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/w64/vectors.h b/proto-cuda/packs-readwidth/w64/vectors.h new file mode 100644 index 000000000..7f90abbef --- /dev/null +++ b/proto-cuda/packs-readwidth/w64/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x038d4ba51c8b5953ull, 0x6ed4390f858d607full, 0xcc5934eb413f2163ull, 0x9ef8bc61e4ff9010ull, 0x8510a782df64cb09ull, 0x3b2bd10e19f09e83ull, 0x0ce2c9f328a529e8ull, 0x4a25b1d2a9945c34ull, + 0x3ce471d313fbd1ddull, 0xb5d04b8faa5e6e1bull, 0x15adcbf13addc14full, 0xebaa63c35d931843ull, 0x88db9de190eaedd8ull, 0x2a52ec487180574bull, 0x15378c68afa45880ull, 0xa8da5311698b773eull, + 0xd04bdc631432b590ull, 0xa36028c8a5cfc9afull, 0x16c89c3870d01c06ull, 0xd9a490f00fa2cf51ull, 0x7129cc4a8eeb54b7ull, 0xb4c39136602bec32ull, 0xf8a038bd63f7c448ull, 0xc2b32a5818c514edull, + 0xabde76b03dcab2d6ull, 0xbfd624b667c93c55ull, 0x37ca3d66bb0377feull, 0xbc7c83c977c0f01cull, 0xf61e43fbe952e3eaull, 0x7cb9c8934f4988e2ull, 0xc0c8f8d9a0fb9616ull, 0x5fe5ee068519dcfeull + }, + { // base nonce 4096 + 0x84f5e8f5233979e7ull, 0xa46e3d6b347cf831ull, 0x9a5d0ffb01876b4dull, 0xba99e208c76b68f4ull, 0x2d01b99f77dd39faull, 0x219cd5dfa69bc5aeull, 0x4f4f6e2cdb06e61cull, 0x21c8ab91d8e8463dull, + 0x7ecca803668442c9ull, 0x2dc6ecbd3c04b0d5ull, 0xd2f6e4103d491cc5ull, 0xb5ad5cfe40ae1772ull, 0xde58df753fc12345ull, 0x822c3d93fafd58a6ull, 0x9fc2d2546d9c7aa0ull, 0x2ba76af772824996ull, + 0x4ae331b88e70ed63ull, 0xd0a6af2a0a6bd8c5ull, 0x758e048965050f70ull, 0x2cda739f6bafe6a8ull, 0x705eaef1634ab7f4ull, 0xf90c278dc92731e2ull, 0x8d7b2a6658337606ull, 0x1cedc3457e07d088ull, + 0xe9f0f6a87b0f3679ull, 0x731063764bb6a45full, 0x20409c2b6a66e2efull, 0x6f28df86bb1f54e8ull, 0xc5da6238f94bc667ull, 0x94051ff525799ec9ull, 0x453eadb75259cafeull, 0x2bd083e460e4bcd6ull + }, + { // base nonce 1000000 + 0x3565dd1aaaa422d7ull, 0xb05b6631a587cf65ull, 0xc046f81e30f37c06ull, 0x944a743c547cf24eull, 0x1c4139640e9fa77dull, 0x5e33b64f25ac7de7ull, 0xe56c627b53920fc8ull, 0x8ee024785bc0dfbbull, + 0x7e58372211331279ull, 0xbe283cc1d9ea71ecull, 0x19739c1a7bcb61baull, 0x113f3d16ee971f15ull, 0xa237d9dd737f3caeull, 0x1a11da3d009cb2aeull, 0x5f76c93ecc3d7920ull, 0xdc49a33be8fccfe8ull, + 0xfb0f9b8c218de3d6ull, 0x2ef5242b0959f5e2ull, 0x9b128efb0a7bacc6ull, 0xe51dd9b113057975ull, 0x1ad55b3294de8658ull, 0xfbfba77b85ceecacull, 0x292d5263c3eeeae8ull, 0x1bb33eeef9545786ull, + 0xedd5a01ed1ab6a02ull, 0x477a5a0810c70f59ull, 0xec98953e4343378eull, 0xa89312d1b0a5a924ull, 0x6643584abaf1f211ull, 0x95c8dc3ee8265688ull, 0x83b7f1f8823c2033ull, 0x6a04c1f48449706eull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/w64/vectors.json b/proto-cuda/packs-readwidth/w64/vectors.json new file mode 100644 index 000000000..426540b79 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-genesis", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x038d4ba51c8b5953", "0x6ed4390f858d607f", "0xcc5934eb413f2163", "0x9ef8bc61e4ff9010", "0x8510a782df64cb09", "0x3b2bd10e19f09e83", "0x0ce2c9f328a529e8", "0x4a25b1d2a9945c34", + "0x3ce471d313fbd1dd", "0xb5d04b8faa5e6e1b", "0x15adcbf13addc14f", "0xebaa63c35d931843", "0x88db9de190eaedd8", "0x2a52ec487180574b", "0x15378c68afa45880", "0xa8da5311698b773e", + "0xd04bdc631432b590", "0xa36028c8a5cfc9af", "0x16c89c3870d01c06", "0xd9a490f00fa2cf51", "0x7129cc4a8eeb54b7", "0xb4c39136602bec32", "0xf8a038bd63f7c448", "0xc2b32a5818c514ed", + "0xabde76b03dcab2d6", "0xbfd624b667c93c55", "0x37ca3d66bb0377fe", "0xbc7c83c977c0f01c", "0xf61e43fbe952e3ea", "0x7cb9c8934f4988e2", "0xc0c8f8d9a0fb9616", "0x5fe5ee068519dcfe" + ]}, + {"base_nonce": 4096, "expected": [ + "0x84f5e8f5233979e7", "0xa46e3d6b347cf831", "0x9a5d0ffb01876b4d", "0xba99e208c76b68f4", "0x2d01b99f77dd39fa", "0x219cd5dfa69bc5ae", "0x4f4f6e2cdb06e61c", "0x21c8ab91d8e8463d", + "0x7ecca803668442c9", "0x2dc6ecbd3c04b0d5", "0xd2f6e4103d491cc5", "0xb5ad5cfe40ae1772", "0xde58df753fc12345", "0x822c3d93fafd58a6", "0x9fc2d2546d9c7aa0", "0x2ba76af772824996", + "0x4ae331b88e70ed63", "0xd0a6af2a0a6bd8c5", "0x758e048965050f70", "0x2cda739f6bafe6a8", "0x705eaef1634ab7f4", "0xf90c278dc92731e2", "0x8d7b2a6658337606", "0x1cedc3457e07d088", + "0xe9f0f6a87b0f3679", "0x731063764bb6a45f", "0x20409c2b6a66e2ef", "0x6f28df86bb1f54e8", "0xc5da6238f94bc667", "0x94051ff525799ec9", "0x453eadb75259cafe", "0x2bd083e460e4bcd6" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x3565dd1aaaa422d7", "0xb05b6631a587cf65", "0xc046f81e30f37c06", "0x944a743c547cf24e", "0x1c4139640e9fa77d", "0x5e33b64f25ac7de7", "0xe56c627b53920fc8", "0x8ee024785bc0dfbb", + "0x7e58372211331279", "0xbe283cc1d9ea71ec", "0x19739c1a7bcb61ba", "0x113f3d16ee971f15", "0xa237d9dd737f3cae", "0x1a11da3d009cb2ae", "0x5f76c93ecc3d7920", "0xdc49a33be8fccfe8", + "0xfb0f9b8c218de3d6", "0x2ef5242b0959f5e2", "0x9b128efb0a7bacc6", "0xe51dd9b113057975", "0x1ad55b3294de8658", "0xfbfba77b85ceecac", "0x292d5263c3eeeae8", "0x1bb33eeef9545786", + "0xedd5a01ed1ab6a02", "0x477a5a0810c70f59", "0xec98953e4343378e", "0xa89312d1b0a5a924", "0x6643584abaf1f211", "0x95c8dc3ee8265688", "0x83b7f1f8823c2033", "0x6a04c1f48449706e" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/w64x4/kernel.cl b/proto-cuda/packs-readwidth/w64x4/kernel.cl new file mode 100644 index 000000000..953160d15 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64x4/kernel.cl @@ -0,0 +1,278 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r0 = r0 ^ t_; } // 0 shfl + r7 = mul_hi(r7, r6); // 1 mulhi + r5 = r5 + r7 + ((((sel >> 21u) & 1u) != 0u) ? 0xdd04a5dau : 0xe9239829u); // 2 add + r2 = r3 * r2 + r2; // 3 mad + r0 = mul_hi(r0, r7); // 4 mulhi + r5 = r5 + r7 + ((((sel >> 7u) & 1u) != 0u) ? 0x9043323eu : 0x265677dcu); // 5 add + r6 = rotr_var(r6, r5); // 6 rotr + r1 = r1 * r4; // 7 mul + r4 = r4 ^ r0; // 8 xor + r5 = r5 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x504c2692u : 0xc0aca976u); // 9 add + r7 = mul_hi(r7, r5); // 10 mulhi + r4 = r4 ^ r2; // 11 xor + r0 = mul_hi(r0, r4); // 12 mulhi + r2 = mul_hi(r2, r5); // 13 mulhi + r3 = mul_hi(r3, r0); // 14 mulhi + r1 = rotl_imm(r1, 22u); // 15 rotl + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 16 load + r0 = rotl_imm(r0, 11u); // 17 rotl + r6 = r6 + r1 + ((((sel >> 7u) & 1u) != 0u) ? 0x0d110a7au : 0xa377d905u); // 18 add + r4 = r4 + r6 + ((((sel >> 2u) & 1u) != 0u) ? 0x9b05eed7u : 0x56170e75u); // 19 add + r2 = r2 + r4 + ((((sel >> 27u) & 1u) != 0u) ? 0x3ac915d2u : 0xdc8d29d5u); // 20 add + r6 = mul_hi(r6, r2); // 21 mulhi + r6 = r6 - r2; // 22 sub + r7 = rotr_var(r7, r1); // 23 rotr + r0 = rotr_var(r0, r2); // 24 rotr + r3 = r2 * r7 + r3; // 25 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 26 shfl + r5 = rotl_imm(r5, 15u); // 27 rotl + r6 = r6 * r4; // 28 mul + r2 = r2 + r0 + ((((sel >> 14u) & 1u) != 0u) ? 0x09ed045eu : 0x2d1d020au); // 29 add + r1 = r1 - r6; // 30 sub + r0 = mul_hi(r0, r5); // 31 mulhi + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 32 load + r7 = r7 ^ r5; // 33 xor + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 load + r4 = rotr_var(r4, r0); // 35 rotr + r3 = r3 + r0 + ((((sel >> 21u) & 1u) != 0u) ? 0xd1b7c4e1u : 0x4c115f69u); // 36 add + r4 = r4 + r3 + ((((sel >> 19u) & 1u) != 0u) ? 0xaca40f85u : 0x4e8977ceu); // 37 add + r4 = r1 * r4 + r4; // 38 mad + r6 = r6 - r7; // 39 sub + r2 = r2 * r0; // 40 mul + r7 = r7 * r2; // 41 mul + r3 = r3 + r0 + ((((sel >> 9u) & 1u) != 0u) ? 0x87d9ef84u : 0x028aa63cu); // 42 add + r5 = r5 ^ r3; // 43 xor + r6 = rotl_imm(r6, 20u); // 44 rotl + r5 = r5 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0xff9a2d77u : 0x89948c4bu); // 45 add + r2 = r5 * r1 + r2; // 46 mad + r0 = r0 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0xa8848b30u : 0x0be29eaau); // 47 add + r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 48 add + r3 = r3 + r5 + ((((sel >> 6u) & 1u) != 0u) ? 0xb17aad78u : 0x77b1520du); // 49 add + r0 = r6 * r0 + r0; // 50 mad + r2 = r2 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0x5c933bd0u : 0x14bde8e1u); // 51 add + r7 = r7 - r0; // 52 sub + r7 = r7 + r4 + ((((sel >> 11u) & 1u) != 0u) ? 0x546f4095u : 0x7149e3c9u); // 53 add + r2 = r2 * r5; // 54 mul + r6 = r6 ^ r2; // 55 xor + r4 = r4 ^ r3; // 56 xor + r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 57 add + r0 = r0 | r6; // 58 or + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r5 = r5 ^ r2; // 60 xor + r0 = r0 + r3 + ((((sel >> 22u) & 1u) != 0u) ? 0xe686f917u : 0x83245ee9u); // 61 add + r1 = mul_hi(r1, r4); // 62 mulhi + r7 = rotl_imm(r7, 12u); // 63 rotl + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/w64x4/kernel.cu b/proto-cuda/packs-readwidth/w64x4/kernel.cu new file mode 100644 index 000000000..75e9726be --- /dev/null +++ b/proto-cuda/packs-readwidth/w64x4/kernel.cu @@ -0,0 +1,164 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 4); // 0 shfl + r7 = __umulhi(r7, r6); // 1 mulhi + r5 = r5 + r7 + ((((sel >> 21u) & 1u) != 0u) ? 0xdd04a5dau : 0xe9239829u); // 2 add + r2 = r3 * r2 + r2; // 3 mad + r0 = __umulhi(r0, r7); // 4 mulhi + r5 = r5 + r7 + ((((sel >> 7u) & 1u) != 0u) ? 0x9043323eu : 0x265677dcu); // 5 add + r6 = rotr_var(r6, r5); // 6 rotr + r1 = r1 * r4; // 7 mul + r4 = r4 ^ r0; // 8 xor + r5 = r5 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x504c2692u : 0xc0aca976u); // 9 add + r7 = __umulhi(r7, r5); // 10 mulhi + r4 = r4 ^ r2; // 11 xor + r0 = __umulhi(r0, r4); // 12 mulhi + r2 = __umulhi(r2, r5); // 13 mulhi + r3 = __umulhi(r3, r0); // 14 mulhi + r1 = rotl_imm(r1, 22u); // 15 rotl + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 16 load + r0 = rotl_imm(r0, 11u); // 17 rotl + r6 = r6 + r1 + ((((sel >> 7u) & 1u) != 0u) ? 0x0d110a7au : 0xa377d905u); // 18 add + r4 = r4 + r6 + ((((sel >> 2u) & 1u) != 0u) ? 0x9b05eed7u : 0x56170e75u); // 19 add + r2 = r2 + r4 + ((((sel >> 27u) & 1u) != 0u) ? 0x3ac915d2u : 0xdc8d29d5u); // 20 add + r6 = __umulhi(r6, r2); // 21 mulhi + r6 = r6 - r2; // 22 sub + r7 = rotr_var(r7, r1); // 23 rotr + r0 = rotr_var(r0, r2); // 24 rotr + r3 = r2 * r7 + r3; // 25 mad + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 26 shfl + r5 = rotl_imm(r5, 15u); // 27 rotl + r6 = r6 * r4; // 28 mul + r2 = r2 + r0 + ((((sel >> 14u) & 1u) != 0u) ? 0x09ed045eu : 0x2d1d020au); // 29 add + r1 = r1 - r6; // 30 sub + r0 = __umulhi(r0, r5); // 31 mulhi + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 32 load + r7 = r7 ^ r5; // 33 xor + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 load + r4 = rotr_var(r4, r0); // 35 rotr + r3 = r3 + r0 + ((((sel >> 21u) & 1u) != 0u) ? 0xd1b7c4e1u : 0x4c115f69u); // 36 add + r4 = r4 + r3 + ((((sel >> 19u) & 1u) != 0u) ? 0xaca40f85u : 0x4e8977ceu); // 37 add + r4 = r1 * r4 + r4; // 38 mad + r6 = r6 - r7; // 39 sub + r2 = r2 * r0; // 40 mul + r7 = r7 * r2; // 41 mul + r3 = r3 + r0 + ((((sel >> 9u) & 1u) != 0u) ? 0x87d9ef84u : 0x028aa63cu); // 42 add + r5 = r5 ^ r3; // 43 xor + r6 = rotl_imm(r6, 20u); // 44 rotl + r5 = r5 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0xff9a2d77u : 0x89948c4bu); // 45 add + r2 = r5 * r1 + r2; // 46 mad + r0 = r0 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0xa8848b30u : 0x0be29eaau); // 47 add + r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 48 add + r3 = r3 + r5 + ((((sel >> 6u) & 1u) != 0u) ? 0xb17aad78u : 0x77b1520du); // 49 add + r0 = r6 * r0 + r0; // 50 mad + r2 = r2 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0x5c933bd0u : 0x14bde8e1u); // 51 add + r7 = r7 - r0; // 52 sub + r7 = r7 + r4 + ((((sel >> 11u) & 1u) != 0u) ? 0x546f4095u : 0x7149e3c9u); // 53 add + r2 = r2 * r5; // 54 mul + r6 = r6 ^ r2; // 55 xor + r4 = r4 ^ r3; // 56 xor + r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 57 add + r0 = r0 | r6; // 58 or + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r5 = r5 ^ r2; // 60 xor + r0 = r0 + r3 + ((((sel >> 22u) & 1u) != 0u) ? 0xe686f917u : 0x83245ee9u); // 61 add + r1 = __umulhi(r1, r4); // 62 mulhi + r7 = rotl_imm(r7, 12u); // 63 rotl + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/w64x4/kernel_bound.cl b/proto-cuda/packs-readwidth/w64x4/kernel_bound.cl new file mode 100644 index 000000000..4e7687f30 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64x4/kernel_bound.cl @@ -0,0 +1,372 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r0 = r0 ^ t_; } // 0 shfl + r7 = mul_hi(r7, r6); // 1 mulhi + r5 = r5 + r7 + ((((sel >> 21u) & 1u) != 0u) ? 0xdd04a5dau : 0xe9239829u); // 2 add + r2 = r3 * r2 + r2; // 3 mad + r0 = mul_hi(r0, r7); // 4 mulhi + r5 = r5 + r7 + ((((sel >> 7u) & 1u) != 0u) ? 0x9043323eu : 0x265677dcu); // 5 add + r6 = rotr_var(r6, r5); // 6 rotr + r1 = r1 * r4; // 7 mul + r4 = r4 ^ r0; // 8 xor + r5 = r5 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x504c2692u : 0xc0aca976u); // 9 add + r7 = mul_hi(r7, r5); // 10 mulhi + r4 = r4 ^ r2; // 11 xor + r0 = mul_hi(r0, r4); // 12 mulhi + r2 = mul_hi(r2, r5); // 13 mulhi + r3 = mul_hi(r3, r0); // 14 mulhi + r1 = rotl_imm(r1, 22u); // 15 rotl + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 16 load + r0 = rotl_imm(r0, 11u); // 17 rotl + r6 = r6 + r1 + ((((sel >> 7u) & 1u) != 0u) ? 0x0d110a7au : 0xa377d905u); // 18 add + r4 = r4 + r6 + ((((sel >> 2u) & 1u) != 0u) ? 0x9b05eed7u : 0x56170e75u); // 19 add + r2 = r2 + r4 + ((((sel >> 27u) & 1u) != 0u) ? 0x3ac915d2u : 0xdc8d29d5u); // 20 add + r6 = mul_hi(r6, r2); // 21 mulhi + r6 = r6 - r2; // 22 sub + r7 = rotr_var(r7, r1); // 23 rotr + r0 = rotr_var(r0, r2); // 24 rotr + r3 = r2 * r7 + r3; // 25 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 26 shfl + r5 = rotl_imm(r5, 15u); // 27 rotl + r6 = r6 * r4; // 28 mul + r2 = r2 + r0 + ((((sel >> 14u) & 1u) != 0u) ? 0x09ed045eu : 0x2d1d020au); // 29 add + r1 = r1 - r6; // 30 sub + r0 = mul_hi(r0, r5); // 31 mulhi + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 32 load + r7 = r7 ^ r5; // 33 xor + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 load + r4 = rotr_var(r4, r0); // 35 rotr + r3 = r3 + r0 + ((((sel >> 21u) & 1u) != 0u) ? 0xd1b7c4e1u : 0x4c115f69u); // 36 add + r4 = r4 + r3 + ((((sel >> 19u) & 1u) != 0u) ? 0xaca40f85u : 0x4e8977ceu); // 37 add + r4 = r1 * r4 + r4; // 38 mad + r6 = r6 - r7; // 39 sub + r2 = r2 * r0; // 40 mul + r7 = r7 * r2; // 41 mul + r3 = r3 + r0 + ((((sel >> 9u) & 1u) != 0u) ? 0x87d9ef84u : 0x028aa63cu); // 42 add + r5 = r5 ^ r3; // 43 xor + r6 = rotl_imm(r6, 20u); // 44 rotl + r5 = r5 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0xff9a2d77u : 0x89948c4bu); // 45 add + r2 = r5 * r1 + r2; // 46 mad + r0 = r0 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0xa8848b30u : 0x0be29eaau); // 47 add + r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 48 add + r3 = r3 + r5 + ((((sel >> 6u) & 1u) != 0u) ? 0xb17aad78u : 0x77b1520du); // 49 add + r0 = r6 * r0 + r0; // 50 mad + r2 = r2 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0x5c933bd0u : 0x14bde8e1u); // 51 add + r7 = r7 - r0; // 52 sub + r7 = r7 + r4 + ((((sel >> 11u) & 1u) != 0u) ? 0x546f4095u : 0x7149e3c9u); // 53 add + r2 = r2 * r5; // 54 mul + r6 = r6 ^ r2; // 55 xor + r4 = r4 ^ r3; // 56 xor + r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 57 add + r0 = r0 | r6; // 58 or + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r5 = r5 ^ r2; // 60 xor + r0 = r0 + r3 + ((((sel >> 22u) & 1u) != 0u) ? 0xe686f917u : 0x83245ee9u); // 61 add + r1 = mul_hi(r1, r4); // 62 mulhi + r7 = rotl_imm(r7, 12u); // 63 rotl + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) { + uint gid = (uint)get_global_id(0); + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r0 = r0 ^ t_; } // 0 shfl + r7 = mul_hi(r7, r6); // 1 mulhi + r5 = r5 + r7 + ((((sel >> 21u) & 1u) != 0u) ? 0xdd04a5dau : 0xe9239829u); // 2 add + r2 = r3 * r2 + r2; // 3 mad + r0 = mul_hi(r0, r7); // 4 mulhi + r5 = r5 + r7 + ((((sel >> 7u) & 1u) != 0u) ? 0x9043323eu : 0x265677dcu); // 5 add + r6 = rotr_var(r6, r5); // 6 rotr + r1 = r1 * r4; // 7 mul + r4 = r4 ^ r0; // 8 xor + r5 = r5 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x504c2692u : 0xc0aca976u); // 9 add + r7 = mul_hi(r7, r5); // 10 mulhi + r4 = r4 ^ r2; // 11 xor + r0 = mul_hi(r0, r4); // 12 mulhi + r2 = mul_hi(r2, r5); // 13 mulhi + r3 = mul_hi(r3, r0); // 14 mulhi + r1 = rotl_imm(r1, 22u); // 15 rotl + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 16 load + r0 = rotl_imm(r0, 11u); // 17 rotl + r6 = r6 + r1 + ((((sel >> 7u) & 1u) != 0u) ? 0x0d110a7au : 0xa377d905u); // 18 add + r4 = r4 + r6 + ((((sel >> 2u) & 1u) != 0u) ? 0x9b05eed7u : 0x56170e75u); // 19 add + r2 = r2 + r4 + ((((sel >> 27u) & 1u) != 0u) ? 0x3ac915d2u : 0xdc8d29d5u); // 20 add + r6 = mul_hi(r6, r2); // 21 mulhi + r6 = r6 - r2; // 22 sub + r7 = rotr_var(r7, r1); // 23 rotr + r0 = rotr_var(r0, r2); // 24 rotr + r3 = r2 * r7 + r3; // 25 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 26 shfl + r5 = rotl_imm(r5, 15u); // 27 rotl + r6 = r6 * r4; // 28 mul + r2 = r2 + r0 + ((((sel >> 14u) & 1u) != 0u) ? 0x09ed045eu : 0x2d1d020au); // 29 add + r1 = r1 - r6; // 30 sub + r0 = mul_hi(r0, r5); // 31 mulhi + { uint b_ = (r0 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 32 load + r7 = r7 ^ r5; // 33 xor + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 load + r4 = rotr_var(r4, r0); // 35 rotr + r3 = r3 + r0 + ((((sel >> 21u) & 1u) != 0u) ? 0xd1b7c4e1u : 0x4c115f69u); // 36 add + r4 = r4 + r3 + ((((sel >> 19u) & 1u) != 0u) ? 0xaca40f85u : 0x4e8977ceu); // 37 add + r4 = r1 * r4 + r4; // 38 mad + r6 = r6 - r7; // 39 sub + r2 = r2 * r0; // 40 mul + r7 = r7 * r2; // 41 mul + r3 = r3 + r0 + ((((sel >> 9u) & 1u) != 0u) ? 0x87d9ef84u : 0x028aa63cu); // 42 add + r5 = r5 ^ r3; // 43 xor + r6 = rotl_imm(r6, 20u); // 44 rotl + r5 = r5 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0xff9a2d77u : 0x89948c4bu); // 45 add + r2 = r5 * r1 + r2; // 46 mad + r0 = r0 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0xa8848b30u : 0x0be29eaau); // 47 add + r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 48 add + r3 = r3 + r5 + ((((sel >> 6u) & 1u) != 0u) ? 0xb17aad78u : 0x77b1520du); // 49 add + r0 = r6 * r0 + r0; // 50 mad + r2 = r2 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0x5c933bd0u : 0x14bde8e1u); // 51 add + r7 = r7 - r0; // 52 sub + r7 = r7 + r4 + ((((sel >> 11u) & 1u) != 0u) ? 0x546f4095u : 0x7149e3c9u); // 53 add + r2 = r2 * r5; // 54 mul + r6 = r6 ^ r2; // 55 xor + r4 = r4 ^ r3; // 56 xor + r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 57 add + r0 = r0 | r6; // 58 or + { uint b_ = (r4 & mask) & ~15u; uint4 v0_ = vload4(0u, ds + b_); uint4 v1_ = vload4(1u, ds + b_); uint4 v2_ = vload4(2u, ds + b_); uint4 v3_ = vload4(3u, ds + b_); uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r5 = r5 ^ r2; // 60 xor + r0 = r0 + r3 + ((((sel >> 22u) & 1u) != 0u) ? 0xe686f917u : 0x83245ee9u); // 61 add + r1 = mul_hi(r1, r4); // 62 mulhi + r7 = rotl_imm(r7, 12u); // 63 rotl + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/w64x4/kernel_bound.cu b/proto-cuda/packs-readwidth/w64x4/kernel_bound.cu new file mode 100644 index 000000000..d28486ae0 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64x4/kernel_bound.cu @@ -0,0 +1,123 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) { + uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 4); // 0 shfl + r7 = __umulhi(r7, r6); // 1 mulhi + r5 = r5 + r7 + ((((sel >> 21u) & 1u) != 0u) ? 0xdd04a5dau : 0xe9239829u); // 2 add + r2 = r3 * r2 + r2; // 3 mad + r0 = __umulhi(r0, r7); // 4 mulhi + r5 = r5 + r7 + ((((sel >> 7u) & 1u) != 0u) ? 0x9043323eu : 0x265677dcu); // 5 add + r6 = rotr_var(r6, r5); // 6 rotr + r1 = r1 * r4; // 7 mul + r4 = r4 ^ r0; // 8 xor + r5 = r5 + r2 + ((((sel >> 7u) & 1u) != 0u) ? 0x504c2692u : 0xc0aca976u); // 9 add + r7 = __umulhi(r7, r5); // 10 mulhi + r4 = r4 ^ r2; // 11 xor + r0 = __umulhi(r0, r4); // 12 mulhi + r2 = __umulhi(r2, r5); // 13 mulhi + r3 = __umulhi(r3, r0); // 14 mulhi + r1 = rotl_imm(r1, 22u); // 15 rotl + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 16 load + r0 = rotl_imm(r0, 11u); // 17 rotl + r6 = r6 + r1 + ((((sel >> 7u) & 1u) != 0u) ? 0x0d110a7au : 0xa377d905u); // 18 add + r4 = r4 + r6 + ((((sel >> 2u) & 1u) != 0u) ? 0x9b05eed7u : 0x56170e75u); // 19 add + r2 = r2 + r4 + ((((sel >> 27u) & 1u) != 0u) ? 0x3ac915d2u : 0xdc8d29d5u); // 20 add + r6 = __umulhi(r6, r2); // 21 mulhi + r6 = r6 - r2; // 22 sub + r7 = rotr_var(r7, r1); // 23 rotr + r0 = rotr_var(r0, r2); // 24 rotr + r3 = r2 * r7 + r3; // 25 mad + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 26 shfl + r5 = rotl_imm(r5, 15u); // 27 rotl + r6 = r6 * r4; // 28 mul + r2 = r2 + r0 + ((((sel >> 14u) & 1u) != 0u) ? 0x09ed045eu : 0x2d1d020au); // 29 add + r1 = r1 - r6; // 30 sub + r0 = __umulhi(r0, r5); // 31 mulhi + { uint32_t b_ = (r0 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 32 load + r7 = r7 ^ r5; // 33 xor + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 load + r4 = rotr_var(r4, r0); // 35 rotr + r3 = r3 + r0 + ((((sel >> 21u) & 1u) != 0u) ? 0xd1b7c4e1u : 0x4c115f69u); // 36 add + r4 = r4 + r3 + ((((sel >> 19u) & 1u) != 0u) ? 0xaca40f85u : 0x4e8977ceu); // 37 add + r4 = r1 * r4 + r4; // 38 mad + r6 = r6 - r7; // 39 sub + r2 = r2 * r0; // 40 mul + r7 = r7 * r2; // 41 mul + r3 = r3 + r0 + ((((sel >> 9u) & 1u) != 0u) ? 0x87d9ef84u : 0x028aa63cu); // 42 add + r5 = r5 ^ r3; // 43 xor + r6 = rotl_imm(r6, 20u); // 44 rotl + r5 = r5 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0xff9a2d77u : 0x89948c4bu); // 45 add + r2 = r5 * r1 + r2; // 46 mad + r0 = r0 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0xa8848b30u : 0x0be29eaau); // 47 add + r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 48 add + r3 = r3 + r5 + ((((sel >> 6u) & 1u) != 0u) ? 0xb17aad78u : 0x77b1520du); // 49 add + r0 = r6 * r0 + r0; // 50 mad + r2 = r2 + r0 + ((((sel >> 22u) & 1u) != 0u) ? 0x5c933bd0u : 0x14bde8e1u); // 51 add + r7 = r7 - r0; // 52 sub + r7 = r7 + r4 + ((((sel >> 11u) & 1u) != 0u) ? 0x546f4095u : 0x7149e3c9u); // 53 add + r2 = r2 * r5; // 54 mul + r6 = r6 ^ r2; // 55 xor + r4 = r4 ^ r3; // 56 xor + r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 57 add + r0 = r0 | r6; // 58 or + { uint32_t b_ = (r4 & mask) & ~15u; const uint4* l_ = (const uint4*)(ds + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint32_t x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 load + r5 = r5 ^ r2; // 60 xor + r0 = r0 + r3 + ((((sel >> 22u) & 1u) != 0u) ? 0xe686f917u : 0x83245ee9u); // 61 add + r1 = __umulhi(r1, r4); // 62 mulhi + r7 = rotl_imm(r7, 12u); // 63 rotl + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; +} + +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) { + if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/w64x4/memhard.h b/proto-cuda/packs-readwidth/w64x4/memhard.h new file mode 100644 index 000000000..4803d8e40 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64x4/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/w64x4/memhard.metal b/proto-cuda/packs-readwidth/w64x4/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64x4/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/w64x4/program.h b/proto-cuda/packs-readwidth/w64x4/program.h new file mode 100644 index 000000000..7c0b10273 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64x4/program.h @@ -0,0 +1,60 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-genesis" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0x5c4269666b4eb2f8ull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 32 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "add=18 mulhi=9 xor=7 mad=5 mul=5 rotl=5 load=4 rotr=4 sub=4 shfl=2 or=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "w64x4" +#define IGNEUM_LOAD_SLOTS 4 +#define IGNEUM_LOAD_MIX { 0, 0, 100 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 0, 0, 4 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 2048 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/w64x4/program.json b/proto-cuda/packs-readwidth/w64x4/program.json new file mode 100644 index 000000000..de9bee3fa --- /dev/null +++ b/proto-cuda/packs-readwidth/w64x4/program.json @@ -0,0 +1,127 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0x5c4269666b4eb2f8", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-genesis", + "seed_bytes": "69676e65756d2d67656e65736973", + "seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 32, + "load_class": "w64x4", + "load_slots": 4, + "load_mix_percent_4_16_64": [0, 0, 100], + "load_width_counts_4_16_64": [0, 0, 4], + "bytes_per_hash": 2048, + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"add": 18, "mulhi": 9, "xor": 7, "mad": 5, "mul": 5, "rotl": 5, "load": 4, "rotr": 4, "sub": 4, "shfl": 2, "or": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "shfl", "dst": 0, "src": 6, "src2": 7, "imm": "0x6ec894a2", "imm2": "0xa1e74575", "rot": 19, "bit": 1, "mask": 4, "width": 1}, + {"i": 1, "op": "mulhi", "dst": 7, "src": 6, "src2": 2, "imm": "0x70a5b5dd", "imm2": "0xa7a66264", "rot": 21, "bit": 14, "mask": 16, "width": 1}, + {"i": 2, "op": "add", "dst": 5, "src": 7, "src2": 1, "imm": "0xe9239829", "imm2": "0xdd04a5da", "rot": 30, "bit": 21, "mask": 1, "width": 1}, + {"i": 3, "op": "mad", "dst": 2, "src": 3, "src2": 2, "imm": "0x734003fa", "imm2": "0x5bb67700", "rot": 3, "bit": 20, "mask": 1, "width": 1}, + {"i": 4, "op": "mulhi", "dst": 0, "src": 7, "src2": 4, "imm": "0xa19720f3", "imm2": "0x4d183796", "rot": 2, "bit": 18, "mask": 4, "width": 1}, + {"i": 5, "op": "add", "dst": 5, "src": 7, "src2": 7, "imm": "0x265677dc", "imm2": "0x9043323e", "rot": 30, "bit": 7, "mask": 4, "width": 1}, + {"i": 6, "op": "rotr", "dst": 6, "src": 5, "src2": 4, "imm": "0x5a069596", "imm2": "0xaba3dbaa", "rot": 15, "bit": 25, "mask": 2, "width": 1}, + {"i": 7, "op": "mul", "dst": 1, "src": 4, "src2": 6, "imm": "0x12a93d05", "imm2": "0x10761047", "rot": 17, "bit": 17, "mask": 4, "width": 1}, + {"i": 8, "op": "xor", "dst": 4, "src": 0, "src2": 1, "imm": "0x31009b67", "imm2": "0xbc24f7b9", "rot": 3, "bit": 2, "mask": 1, "width": 1}, + {"i": 9, "op": "add", "dst": 5, "src": 2, "src2": 7, "imm": "0xc0aca976", "imm2": "0x504c2692", "rot": 17, "bit": 7, "mask": 1, "width": 1}, + {"i": 10, "op": "mulhi", "dst": 7, "src": 5, "src2": 3, "imm": "0xe630c672", "imm2": "0x08c46d31", "rot": 13, "bit": 12, "mask": 2, "width": 1}, + {"i": 11, "op": "xor", "dst": 4, "src": 2, "src2": 1, "imm": "0xb0e9df02", "imm2": "0x88cb9af3", "rot": 24, "bit": 28, "mask": 4, "width": 1}, + {"i": 12, "op": "mulhi", "dst": 0, "src": 4, "src2": 4, "imm": "0x45374321", "imm2": "0x3cd91989", "rot": 11, "bit": 4, "mask": 2, "width": 1}, + {"i": 13, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0xd3177981", "imm2": "0x17c9c95b", "rot": 30, "bit": 9, "mask": 8, "width": 1}, + {"i": 14, "op": "mulhi", "dst": 3, "src": 0, "src2": 2, "imm": "0x590f9e11", "imm2": "0xa4c13036", "rot": 3, "bit": 28, "mask": 16, "width": 1}, + {"i": 15, "op": "rotl", "dst": 1, "src": 7, "src2": 1, "imm": "0xf069b834", "imm2": "0xe3c4bf4d", "rot": 22, "bit": 10, "mask": 1, "width": 1}, + {"i": 16, "op": "load", "dst": 6, "src": 4, "src2": 7, "imm": "0x71fbe6f2", "imm2": "0x61b6261e", "rot": 15, "bit": 9, "mask": 1, "width": 16}, + {"i": 17, "op": "rotl", "dst": 0, "src": 6, "src2": 3, "imm": "0x2aa480c2", "imm2": "0x90a248f7", "rot": 11, "bit": 27, "mask": 16, "width": 1}, + {"i": 18, "op": "add", "dst": 6, "src": 1, "src2": 7, "imm": "0xa377d905", "imm2": "0x0d110a7a", "rot": 17, "bit": 7, "mask": 4, "width": 1}, + {"i": 19, "op": "add", "dst": 4, "src": 6, "src2": 3, "imm": "0x56170e75", "imm2": "0x9b05eed7", "rot": 6, "bit": 2, "mask": 4, "width": 1}, + {"i": 20, "op": "add", "dst": 2, "src": 4, "src2": 4, "imm": "0xdc8d29d5", "imm2": "0x3ac915d2", "rot": 6, "bit": 27, "mask": 2, "width": 1}, + {"i": 21, "op": "mulhi", "dst": 6, "src": 2, "src2": 0, "imm": "0xdc3ec8fd", "imm2": "0x599e2fa3", "rot": 22, "bit": 3, "mask": 2, "width": 1}, + {"i": 22, "op": "sub", "dst": 6, "src": 2, "src2": 5, "imm": "0xd73e396f", "imm2": "0x089f5404", "rot": 3, "bit": 26, "mask": 1, "width": 1}, + {"i": 23, "op": "rotr", "dst": 7, "src": 1, "src2": 2, "imm": "0x3e26afea", "imm2": "0xba573970", "rot": 11, "bit": 15, "mask": 8, "width": 1}, + {"i": 24, "op": "rotr", "dst": 0, "src": 2, "src2": 3, "imm": "0x30db112b", "imm2": "0x97acf9f2", "rot": 21, "bit": 15, "mask": 16, "width": 1}, + {"i": 25, "op": "mad", "dst": 3, "src": 2, "src2": 7, "imm": "0x400383b6", "imm2": "0xe9bab735", "rot": 16, "bit": 18, "mask": 1, "width": 1}, + {"i": 26, "op": "shfl", "dst": 3, "src": 6, "src2": 5, "imm": "0x55b67f9f", "imm2": "0xe74626f8", "rot": 29, "bit": 23, "mask": 1, "width": 1}, + {"i": 27, "op": "rotl", "dst": 5, "src": 3, "src2": 4, "imm": "0xecdd7d43", "imm2": "0x574f3d02", "rot": 15, "bit": 21, "mask": 4, "width": 1}, + {"i": 28, "op": "mul", "dst": 6, "src": 4, "src2": 5, "imm": "0x88750d70", "imm2": "0x3ee75acf", "rot": 11, "bit": 19, "mask": 1, "width": 1}, + {"i": 29, "op": "add", "dst": 2, "src": 0, "src2": 6, "imm": "0x2d1d020a", "imm2": "0x09ed045e", "rot": 26, "bit": 14, "mask": 16, "width": 1}, + {"i": 30, "op": "sub", "dst": 1, "src": 6, "src2": 6, "imm": "0xeb79ea49", "imm2": "0xcc587f5a", "rot": 6, "bit": 8, "mask": 16, "width": 1}, + {"i": 31, "op": "mulhi", "dst": 0, "src": 5, "src2": 6, "imm": "0x5f20c27e", "imm2": "0xe3f24158", "rot": 30, "bit": 13, "mask": 8, "width": 1}, + {"i": 32, "op": "load", "dst": 3, "src": 0, "src2": 3, "imm": "0x63578bc1", "imm2": "0xbf64a89f", "rot": 16, "bit": 26, "mask": 1, "width": 16}, + {"i": 33, "op": "xor", "dst": 7, "src": 5, "src2": 1, "imm": "0x0389cf85", "imm2": "0x84cad367", "rot": 6, "bit": 21, "mask": 1, "width": 1}, + {"i": 34, "op": "load", "dst": 5, "src": 4, "src2": 0, "imm": "0xc24a7d70", "imm2": "0x6591d24c", "rot": 21, "bit": 24, "mask": 1, "width": 16}, + {"i": 35, "op": "rotr", "dst": 4, "src": 0, "src2": 0, "imm": "0x144ca538", "imm2": "0xf68f348b", "rot": 11, "bit": 15, "mask": 4, "width": 1}, + {"i": 36, "op": "add", "dst": 3, "src": 0, "src2": 0, "imm": "0x4c115f69", "imm2": "0xd1b7c4e1", "rot": 25, "bit": 21, "mask": 1, "width": 1}, + {"i": 37, "op": "add", "dst": 4, "src": 3, "src2": 2, "imm": "0x4e8977ce", "imm2": "0xaca40f85", "rot": 3, "bit": 19, "mask": 8, "width": 1}, + {"i": 38, "op": "mad", "dst": 4, "src": 1, "src2": 4, "imm": "0x7de7e64b", "imm2": "0xce13eff8", "rot": 13, "bit": 4, "mask": 4, "width": 1}, + {"i": 39, "op": "sub", "dst": 6, "src": 7, "src2": 1, "imm": "0x6a65ab71", "imm2": "0x8fbc1bcd", "rot": 4, "bit": 1, "mask": 8, "width": 1}, + {"i": 40, "op": "mul", "dst": 2, "src": 0, "src2": 0, "imm": "0xc0423027", "imm2": "0xc1ae8d3b", "rot": 4, "bit": 15, "mask": 16, "width": 1}, + {"i": 41, "op": "mul", "dst": 7, "src": 2, "src2": 2, "imm": "0xc4f18bec", "imm2": "0x8a3e2464", "rot": 30, "bit": 16, "mask": 2, "width": 1}, + {"i": 42, "op": "add", "dst": 3, "src": 0, "src2": 3, "imm": "0x028aa63c", "imm2": "0x87d9ef84", "rot": 1, "bit": 9, "mask": 8, "width": 1}, + {"i": 43, "op": "xor", "dst": 5, "src": 3, "src2": 1, "imm": "0x74a59b7d", "imm2": "0xc4766ff1", "rot": 30, "bit": 3, "mask": 16, "width": 1}, + {"i": 44, "op": "rotl", "dst": 6, "src": 0, "src2": 3, "imm": "0x9342df0d", "imm2": "0xd78ff4fa", "rot": 20, "bit": 11, "mask": 1, "width": 1}, + {"i": 45, "op": "add", "dst": 5, "src": 3, "src2": 3, "imm": "0x89948c4b", "imm2": "0xff9a2d77", "rot": 20, "bit": 31, "mask": 16, "width": 1}, + {"i": 46, "op": "mad", "dst": 2, "src": 5, "src2": 1, "imm": "0xb26e5d9d", "imm2": "0xdf72f45e", "rot": 22, "bit": 12, "mask": 8, "width": 1}, + {"i": 47, "op": "add", "dst": 0, "src": 1, "src2": 7, "imm": "0x0be29eaa", "imm2": "0xa8848b30", "rot": 2, "bit": 6, "mask": 16, "width": 1}, + {"i": 48, "op": "add", "dst": 0, "src": 2, "src2": 0, "imm": "0x4f92b968", "imm2": "0x699fd448", "rot": 22, "bit": 6, "mask": 4, "width": 1}, + {"i": 49, "op": "add", "dst": 3, "src": 5, "src2": 7, "imm": "0x77b1520d", "imm2": "0xb17aad78", "rot": 4, "bit": 6, "mask": 16, "width": 1}, + {"i": 50, "op": "mad", "dst": 0, "src": 6, "src2": 0, "imm": "0x25955401", "imm2": "0xf4689674", "rot": 14, "bit": 27, "mask": 4, "width": 1}, + {"i": 51, "op": "add", "dst": 2, "src": 0, "src2": 2, "imm": "0x14bde8e1", "imm2": "0x5c933bd0", "rot": 9, "bit": 22, "mask": 4, "width": 1}, + {"i": 52, "op": "sub", "dst": 7, "src": 0, "src2": 7, "imm": "0xf62515d5", "imm2": "0x4a164f9f", "rot": 28, "bit": 30, "mask": 2, "width": 1}, + {"i": 53, "op": "add", "dst": 7, "src": 4, "src2": 2, "imm": "0x7149e3c9", "imm2": "0x546f4095", "rot": 5, "bit": 11, "mask": 4, "width": 1}, + {"i": 54, "op": "mul", "dst": 2, "src": 5, "src2": 7, "imm": "0x9ab072be", "imm2": "0xa2bcf23c", "rot": 9, "bit": 8, "mask": 4, "width": 1}, + {"i": 55, "op": "xor", "dst": 6, "src": 2, "src2": 3, "imm": "0xbfb337f3", "imm2": "0x65660398", "rot": 11, "bit": 27, "mask": 4, "width": 1}, + {"i": 56, "op": "xor", "dst": 4, "src": 3, "src2": 7, "imm": "0xd76e14d7", "imm2": "0xaf9dd72d", "rot": 12, "bit": 18, "mask": 8, "width": 1}, + {"i": 57, "op": "add", "dst": 4, "src": 6, "src2": 2, "imm": "0x89841d87", "imm2": "0x1e07c3d9", "rot": 6, "bit": 27, "mask": 1, "width": 1}, + {"i": 58, "op": "or", "dst": 0, "src": 6, "src2": 5, "imm": "0x8e499ba4", "imm2": "0x1bf306dd", "rot": 27, "bit": 4, "mask": 16, "width": 1}, + {"i": 59, "op": "load", "dst": 1, "src": 4, "src2": 1, "imm": "0x7cc2ce84", "imm2": "0xcdd8b1f2", "rot": 30, "bit": 13, "mask": 4, "width": 16}, + {"i": 60, "op": "xor", "dst": 5, "src": 2, "src2": 1, "imm": "0x5742c15a", "imm2": "0x7d088b30", "rot": 2, "bit": 30, "mask": 4, "width": 1}, + {"i": 61, "op": "add", "dst": 0, "src": 3, "src2": 4, "imm": "0x83245ee9", "imm2": "0xe686f917", "rot": 16, "bit": 22, "mask": 2, "width": 1}, + {"i": 62, "op": "mulhi", "dst": 1, "src": 4, "src2": 1, "imm": "0x1b9441fb", "imm2": "0xb52a7a0f", "rot": 20, "bit": 21, "mask": 1, "width": 1}, + {"i": 63, "op": "rotl", "dst": 7, "src": 6, "src2": 3, "imm": "0xf509c1a8", "imm2": "0x69360550", "rot": 12, "bit": 25, "mask": 16, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/w64x4/program.metal b/proto-cuda/packs-readwidth/w64x4/program.metal new file mode 100644 index 000000000..232f921f1 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64x4/program.metal @@ -0,0 +1,109 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)4); // 0 + r7 = mulhi(r7, r6); // 1 + r5 = r5 + r7 + select(0xe9239829u, 0xdd04a5dau, ((sel >> 21u) & 1u) != 0u); // 2 + r2 = r3 * r2 + r2; // 3 + r0 = mulhi(r0, r7); // 4 + r5 = r5 + r7 + select(0x265677dcu, 0x9043323eu, ((sel >> 7u) & 1u) != 0u); // 5 + r6 = rotr_var(r6, r5); // 6 + r1 = r1 * r4; // 7 + r4 = r4 ^ r0; // 8 + r5 = r5 + r2 + select(0xc0aca976u, 0x504c2692u, ((sel >> 7u) & 1u) != 0u); // 9 + r7 = mulhi(r7, r5); // 10 + r4 = r4 ^ r2; // 11 + r0 = mulhi(r0, r4); // 12 + r2 = mulhi(r2, r5); // 13 + r3 = mulhi(r3, r0); // 14 + r1 = rotl_imm(r1, 22u); // 15 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 16 + r0 = rotl_imm(r0, 11u); // 17 + r6 = r6 + r1 + select(0xa377d905u, 0x0d110a7au, ((sel >> 7u) & 1u) != 0u); // 18 + r4 = r4 + r6 + select(0x56170e75u, 0x9b05eed7u, ((sel >> 2u) & 1u) != 0u); // 19 + r2 = r2 + r4 + select(0xdc8d29d5u, 0x3ac915d2u, ((sel >> 27u) & 1u) != 0u); // 20 + r6 = mulhi(r6, r2); // 21 + r6 = r6 - r2; // 22 + r7 = rotr_var(r7, r1); // 23 + r0 = rotr_var(r0, r2); // 24 + r3 = r2 * r7 + r3; // 25 + r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 26 + r5 = rotl_imm(r5, 15u); // 27 + r6 = r6 * r4; // 28 + r2 = r2 + r0 + select(0x2d1d020au, 0x09ed045eu, ((sel >> 14u) & 1u) != 0u); // 29 + r1 = r1 - r6; // 30 + r0 = mulhi(r0, r5); // 31 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 32 + r7 = r7 ^ r5; // 33 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 + r4 = rotr_var(r4, r0); // 35 + r3 = r3 + r0 + select(0x4c115f69u, 0xd1b7c4e1u, ((sel >> 21u) & 1u) != 0u); // 36 + r4 = r4 + r3 + select(0x4e8977ceu, 0xaca40f85u, ((sel >> 19u) & 1u) != 0u); // 37 + r4 = r1 * r4 + r4; // 38 + r6 = r6 - r7; // 39 + r2 = r2 * r0; // 40 + r7 = r7 * r2; // 41 + r3 = r3 + r0 + select(0x028aa63cu, 0x87d9ef84u, ((sel >> 9u) & 1u) != 0u); // 42 + r5 = r5 ^ r3; // 43 + r6 = rotl_imm(r6, 20u); // 44 + r5 = r5 + r3 + select(0x89948c4bu, 0xff9a2d77u, ((sel >> 31u) & 1u) != 0u); // 45 + r2 = r5 * r1 + r2; // 46 + r0 = r0 + r1 + select(0x0be29eaau, 0xa8848b30u, ((sel >> 6u) & 1u) != 0u); // 47 + r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 48 + r3 = r3 + r5 + select(0x77b1520du, 0xb17aad78u, ((sel >> 6u) & 1u) != 0u); // 49 + r0 = r6 * r0 + r0; // 50 + r2 = r2 + r0 + select(0x14bde8e1u, 0x5c933bd0u, ((sel >> 22u) & 1u) != 0u); // 51 + r7 = r7 - r0; // 52 + r7 = r7 + r4 + select(0x7149e3c9u, 0x546f4095u, ((sel >> 11u) & 1u) != 0u); // 53 + r2 = r2 * r5; // 54 + r6 = r6 ^ r2; // 55 + r4 = r4 ^ r3; // 56 + r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 57 + r0 = r0 | r6; // 58 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 + r5 = r5 ^ r2; // 60 + r0 = r0 + r3 + select(0x83245ee9u, 0xe686f917u, ((sel >> 22u) & 1u) != 0u); // 61 + r1 = mulhi(r1, r4); // 62 + r7 = rotl_imm(r7, 12u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/w64x4/program_bound.metal b/proto-cuda/packs-readwidth/w64x4/program_bound.metal new file mode 100644 index 000000000..a7e992ea1 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64x4/program_bound.metal @@ -0,0 +1,111 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + uint gid [[thread_position_in_grid]]) { + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)4); // 0 + r7 = mulhi(r7, r6); // 1 + r5 = r5 + r7 + select(0xe9239829u, 0xdd04a5dau, ((sel >> 21u) & 1u) != 0u); // 2 + r2 = r3 * r2 + r2; // 3 + r0 = mulhi(r0, r7); // 4 + r5 = r5 + r7 + select(0x265677dcu, 0x9043323eu, ((sel >> 7u) & 1u) != 0u); // 5 + r6 = rotr_var(r6, r5); // 6 + r1 = r1 * r4; // 7 + r4 = r4 ^ r0; // 8 + r5 = r5 + r2 + select(0xc0aca976u, 0x504c2692u, ((sel >> 7u) & 1u) != 0u); // 9 + r7 = mulhi(r7, r5); // 10 + r4 = r4 ^ r2; // 11 + r0 = mulhi(r0, r4); // 12 + r2 = mulhi(r2, r5); // 13 + r3 = mulhi(r3, r0); // 14 + r1 = rotl_imm(r1, 22u); // 15 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r6 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r6 = x_; } // 16 + r0 = rotl_imm(r0, 11u); // 17 + r6 = r6 + r1 + select(0xa377d905u, 0x0d110a7au, ((sel >> 7u) & 1u) != 0u); // 18 + r4 = r4 + r6 + select(0x56170e75u, 0x9b05eed7u, ((sel >> 2u) & 1u) != 0u); // 19 + r2 = r2 + r4 + select(0xdc8d29d5u, 0x3ac915d2u, ((sel >> 27u) & 1u) != 0u); // 20 + r6 = mulhi(r6, r2); // 21 + r6 = r6 - r2; // 22 + r7 = rotr_var(r7, r1); // 23 + r0 = rotr_var(r0, r2); // 24 + r3 = r2 * r7 + r3; // 25 + r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 26 + r5 = rotl_imm(r5, 15u); // 27 + r6 = r6 * r4; // 28 + r2 = r2 + r0 + select(0x2d1d020au, 0x09ed045eu, ((sel >> 14u) & 1u) != 0u); // 29 + r1 = r1 - r6; // 30 + r0 = mulhi(r0, r5); // 31 + { uint b_ = (r0 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r3 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r3 = x_; } // 32 + r7 = r7 ^ r5; // 33 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r5 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r5 = x_; } // 34 + r4 = rotr_var(r4, r0); // 35 + r3 = r3 + r0 + select(0x4c115f69u, 0xd1b7c4e1u, ((sel >> 21u) & 1u) != 0u); // 36 + r4 = r4 + r3 + select(0x4e8977ceu, 0xaca40f85u, ((sel >> 19u) & 1u) != 0u); // 37 + r4 = r1 * r4 + r4; // 38 + r6 = r6 - r7; // 39 + r2 = r2 * r0; // 40 + r7 = r7 * r2; // 41 + r3 = r3 + r0 + select(0x028aa63cu, 0x87d9ef84u, ((sel >> 9u) & 1u) != 0u); // 42 + r5 = r5 ^ r3; // 43 + r6 = rotl_imm(r6, 20u); // 44 + r5 = r5 + r3 + select(0x89948c4bu, 0xff9a2d77u, ((sel >> 31u) & 1u) != 0u); // 45 + r2 = r5 * r1 + r2; // 46 + r0 = r0 + r1 + select(0x0be29eaau, 0xa8848b30u, ((sel >> 6u) & 1u) != 0u); // 47 + r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 48 + r3 = r3 + r5 + select(0x77b1520du, 0xb17aad78u, ((sel >> 6u) & 1u) != 0u); // 49 + r0 = r6 * r0 + r0; // 50 + r2 = r2 + r0 + select(0x14bde8e1u, 0x5c933bd0u, ((sel >> 22u) & 1u) != 0u); // 51 + r7 = r7 - r0; // 52 + r7 = r7 + r4 + select(0x7149e3c9u, 0x546f4095u, ((sel >> 11u) & 1u) != 0u); // 53 + r2 = r2 * r5; // 54 + r6 = r6 ^ r2; // 55 + r4 = r4 ^ r3; // 56 + r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 57 + r0 = r0 | r6; // 58 + { uint b_ = (r4 & MASK) & ~15u; device const uint4* l_ = (device const uint4*)(dataset + b_); uint4 v0_ = l_[0]; uint4 v1_ = l_[1]; uint4 v2_ = l_[2]; uint4 v3_ = l_[3]; uint x_ = r1 ^ v0_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v0_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v1_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v2_.w; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.x; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.y; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.z; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ v3_.w; r1 = x_; } // 59 + r5 = r5 ^ r2; // 60 + r0 = r0 + r3 + select(0x83245ee9u, 0xe686f917u, ((sel >> 22u) & 1u) != 0u); // 61 + r1 = mulhi(r1, r4); // 62 + r7 = rotl_imm(r7, 12u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; +} diff --git a/proto-cuda/packs-readwidth/w64x4/vectors.h b/proto-cuda/packs-readwidth/w64x4/vectors.h new file mode 100644 index 000000000..03f707649 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64x4/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0xfe5dcc9efa189e98ull, 0xe6f859e133cd6343ull, 0xdd8b2875dbf2f985ull, 0xaefe3d5ae18af3d9ull, 0xc0f7e3eb67142930ull, 0x13d0cb06527b0e95ull, 0x38e603d82ce40f3dull, 0xeb7bcfbaacfe5c62ull, + 0x6a2b7ed18742fb69ull, 0x59d62b3e9f122cacull, 0x3de59dbba159aec0ull, 0xd6ec0490527cf675ull, 0x3f96eeb47a84c841ull, 0xcc21fd74435b1ab0ull, 0x5f2da48a3e1bfcffull, 0xecd64d86f0a86b0cull, + 0x98a7303577de5cc1ull, 0x5b294f3e199b2e6dull, 0x94586e8efaf67937ull, 0xc7eab0e4d6bbf095ull, 0x6772dbd145e9475aull, 0x8b9dfbb27f69cf0aull, 0xdeebbe23891dffd4ull, 0xf0498d6b32124bf2ull, + 0x8452f24ed7c6fed3ull, 0x460eaa2e0f43b7dcull, 0x3b2fa47f9c42a219ull, 0xb7786b2da5f1ba64ull, 0x60221dc073be8b17ull, 0x2a57d2a15f7f371eull, 0xbbbbcddf1617a962ull, 0x5d3b053c6de2e1adull + }, + { // base nonce 4096 + 0x90fddfc1b0249983ull, 0x85a62ae03a1804a4ull, 0x9f85f2a06ae6e26full, 0xd7a8a6365fd537f1ull, 0x20e85fa620d208c2ull, 0x80d05d7505f6e7ceull, 0x183a782a7131dd47ull, 0x4177e969c24b05e4ull, + 0xf1f1d7113769933full, 0x0089b7e3f0853afaull, 0x35f758aca76dea97ull, 0x1d2ad56f11493ca1ull, 0xd9712fc3b5a00286ull, 0x783a3ab87c504a40ull, 0x80d7d7b2f3cb1ceaull, 0x962a16e431210a49ull, + 0xc42eb62a15638090ull, 0xf42b81d37d72ad05ull, 0x215af96aef57b7c6ull, 0x748b45bfb0548a82ull, 0x1dc85bef6aa4a786ull, 0x93ab6f69be7078cfull, 0xc86266a70e5803f9ull, 0x349bb7aa412b88dfull, + 0xa59722e64627d640ull, 0x51038cb78970a996ull, 0xee4ff7b0226aa1f4ull, 0xef3a4bcafc51c1e1ull, 0xe203013799946902ull, 0xf56887ebbf1ff7e4ull, 0xcf655dcd00c1f4c1ull, 0xee083a544155f1e0ull + }, + { // base nonce 1000000 + 0xbe2ca17ab95f8c06ull, 0xb2c803d9957ca796ull, 0x628cf6ae2d808b4full, 0x4fecf19b1873f7dbull, 0xfc8fded7f3a302bbull, 0x159f10fa74b6c048ull, 0xe12c1e1fc12e3b8aull, 0x5bbece7c4f9c0b09ull, + 0x19763f2abb4536f7ull, 0x55d67305ac79bf72ull, 0xebcdf574af866db8ull, 0xa8a19b58a710d4f3ull, 0x7552086f820e59e9ull, 0x9dad73b580b84c47ull, 0xb77143e3d8bd225cull, 0x9507fdce2708f827ull, + 0xa4d3cbbeed4c25cfull, 0x3b139f4f815a97faull, 0xdf0c711ce8767304ull, 0xb5eff3e5b4e14b7eull, 0xaf22dbb2d670530dull, 0xf4af61141861dbcaull, 0x28054c01a0a22bfeull, 0x79c7fce297ff2047ull, + 0x0477fb35b663e383ull, 0x65120611d29ac152ull, 0xd22d930224d5c38aull, 0xfdb346491e405983ull, 0x10b3726282474f80ull, 0xbb14dee47980ea55ull, 0x9e5c024bbd2ef6e9ull, 0xb99261595c3d2dd2ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/w64x4/vectors.json b/proto-cuda/packs-readwidth/w64x4/vectors.json new file mode 100644 index 000000000..83041ee51 --- /dev/null +++ b/proto-cuda/packs-readwidth/w64x4/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-genesis", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0xfe5dcc9efa189e98", "0xe6f859e133cd6343", "0xdd8b2875dbf2f985", "0xaefe3d5ae18af3d9", "0xc0f7e3eb67142930", "0x13d0cb06527b0e95", "0x38e603d82ce40f3d", "0xeb7bcfbaacfe5c62", + "0x6a2b7ed18742fb69", "0x59d62b3e9f122cac", "0x3de59dbba159aec0", "0xd6ec0490527cf675", "0x3f96eeb47a84c841", "0xcc21fd74435b1ab0", "0x5f2da48a3e1bfcff", "0xecd64d86f0a86b0c", + "0x98a7303577de5cc1", "0x5b294f3e199b2e6d", "0x94586e8efaf67937", "0xc7eab0e4d6bbf095", "0x6772dbd145e9475a", "0x8b9dfbb27f69cf0a", "0xdeebbe23891dffd4", "0xf0498d6b32124bf2", + "0x8452f24ed7c6fed3", "0x460eaa2e0f43b7dc", "0x3b2fa47f9c42a219", "0xb7786b2da5f1ba64", "0x60221dc073be8b17", "0x2a57d2a15f7f371e", "0xbbbbcddf1617a962", "0x5d3b053c6de2e1ad" + ]}, + {"base_nonce": 4096, "expected": [ + "0x90fddfc1b0249983", "0x85a62ae03a1804a4", "0x9f85f2a06ae6e26f", "0xd7a8a6365fd537f1", "0x20e85fa620d208c2", "0x80d05d7505f6e7ce", "0x183a782a7131dd47", "0x4177e969c24b05e4", + "0xf1f1d7113769933f", "0x0089b7e3f0853afa", "0x35f758aca76dea97", "0x1d2ad56f11493ca1", "0xd9712fc3b5a00286", "0x783a3ab87c504a40", "0x80d7d7b2f3cb1cea", "0x962a16e431210a49", + "0xc42eb62a15638090", "0xf42b81d37d72ad05", "0x215af96aef57b7c6", "0x748b45bfb0548a82", "0x1dc85bef6aa4a786", "0x93ab6f69be7078cf", "0xc86266a70e5803f9", "0x349bb7aa412b88df", + "0xa59722e64627d640", "0x51038cb78970a996", "0xee4ff7b0226aa1f4", "0xef3a4bcafc51c1e1", "0xe203013799946902", "0xf56887ebbf1ff7e4", "0xcf655dcd00c1f4c1", "0xee083a544155f1e0" + ]}, + {"base_nonce": 1000000, "expected": [ + "0xbe2ca17ab95f8c06", "0xb2c803d9957ca796", "0x628cf6ae2d808b4f", "0x4fecf19b1873f7db", "0xfc8fded7f3a302bb", "0x159f10fa74b6c048", "0xe12c1e1fc12e3b8a", "0x5bbece7c4f9c0b09", + "0x19763f2abb4536f7", "0x55d67305ac79bf72", "0xebcdf574af866db8", "0xa8a19b58a710d4f3", "0x7552086f820e59e9", "0x9dad73b580b84c47", "0xb77143e3d8bd225c", "0x9507fdce2708f827", + "0xa4d3cbbeed4c25cf", "0x3b139f4f815a97fa", "0xdf0c711ce8767304", "0xb5eff3e5b4e14b7e", "0xaf22dbb2d670530d", "0xf4af61141861dbca", "0x28054c01a0a22bfe", "0x79c7fce297ff2047", + "0x0477fb35b663e383", "0x65120611d29ac152", "0xd22d930224d5c38a", "0xfdb346491e405983", "0x10b3726282474f80", "0xbb14dee47980ea55", "0x9e5c024bbd2ef6e9", "0xb99261595c3d2dd2" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-opencl/README.md b/proto-opencl/README.md index fc6a37832..d058796af 100644 --- a/proto-opencl/README.md +++ b/proto-opencl/README.md @@ -32,6 +32,25 @@ the driver) and the Khronos headers fetched by `fetch-redist.sh`. `test-generic. OpenCL with the two packs `proto-cuda/nvrtc/emu/test.sh` writes (PASS on 4 October 2026: 192 found lines, prepare and swap, 15 sampled hashes equal to `igneum-pow hash-bound`). The bench and the compiled-in serve mode are unchanged. +## 5 October 2026: the RX 9070 XT (gfx1201, RDNA 4) on PC 1 + +Measured in `docs/bench-log.md` ("the 9070 XT on the eGPU"). What changed in `host.c`: + +- `--list` folds the same card listed by two platforms of one vendor (an old driver's OpenCL registration left behind + after an update: PC 1 had 3652.0 and 3683.0) into one entry; the older one prints as ` dup [N] ...` and the app's + parser (`detect.rs`) never makes a card of it. Before, the app ran two workers on one 9070 XT at half rate each. + Indices stay flat, so `--device N` still reaches the hidden entry for a comparison. +- Every path prints the kernel as the driver compiled it: `kernel: ... preferred multiple, local and private memory, + sub-group size`. Private memory above 0 means spilled registers. The sub-group size is queried on the local-memory + path too (the gfx1201 answers 32: wave32). +- `--serve` reads back the hits and 34 sentinel words through a GPU-side select pass instead of 8 bytes per nonce + (`--readback full` or `IGNEUM_READBACK=full` keeps the old path). The stats line every 200 jobs carries the bytes up + and down per chunk and the mean device time of the hash kernel, the select pass, the read-back and the host scan. +- `--memprobe [--probe-mib N]`: no pack. Dependent random 4-byte loads against lanes in flight (latency and the + random-read ceiling), eight independent loads per lane, random 64-byte lines, a coalesced stream and an integer + chain, at 4, 64 and 1024 MiB. The hash is 128 dependent random 4-byte loads, so the 1024 MiB chase ceiling divided + by 128 is the card's hash-rate ceiling for this program class. + ## Layout ``` @@ -39,6 +58,7 @@ proto-opencl/ host.c C99 host: device list, runtime kernel build, cache + dataset fill, self-tests, vectors, bench, sweep, --serve, --pack cl_dynamic.h Windows one-click build: OpenCL.dll loaded at run time (IGNEUM_CL_DYNAMIC) test-generic.sh the --pack mode checked here through Apple OpenCL (needs proto-cuda/nvrtc/emu/test.sh's packs) + test_host.c device-free unit tests of host.c's rules (the duplicate-platform fold); run with test-host.sh build.sh macOS (-framework OpenCL, or the Khronos ICD loader) and Linux (-lOpenCL) build.bat Windows (MSVC cl.exe + OpenCL.lib) WAVEFRONT.md wave32 vs wave64 on AMD, and why the kernel cannot tell the difference diff --git a/proto-opencl/WAVEFRONT.md b/proto-opencl/WAVEFRONT.md index 6bfb2e8e7..2e05e438b 100644 --- a/proto-opencl/WAVEFRONT.md +++ b/proto-opencl/WAVEFRONT.md @@ -86,6 +86,15 @@ exchange alone would survive a wave64 sub-group (point 1 above); host.c still re device because of point 2 and because of `sub_group_broadcast`. The conservative rule costs nothing in correctness and, on a wave64 card, the local-memory path is what runs. +## RDNA 4, measured (5 October 2026) + +An RX 9070 XT (gfx1201, Adrenalin 26.9.2, OpenCL driver 3683.0 PAL,LC) on PC 1 reports `AMD wavefront width 32`, +lists `cl_khr_subgroups` without a shuffle extension, and `clGetKernelSubGroupInfoKHR` on `igneum_hash` answers a +sub-group of 32 for a 32-item work-group: wave32, so a work-group of 32 is one full wave and path 0 (local memory) +runs with its barriers elided by the compiler. `--group-warps 1, 2, 4, 8` give 18.02, 18.06, 18.07 and 18.04 MH/s, the +same number: the exchange path and the work-group shape cost nothing there. The rate is set by the card's random-read +throughput at the dataset size (`docs/bench-log.md`, "the 9070 XT on the eGPU"). + ## What is not proven here - No AMD compiler has compiled `kernel.cl`, and no AMD device has run it. The emulator is clang; pocl is LLVM on a diff --git a/proto-opencl/emu/emu_opencl.h b/proto-opencl/emu/emu_opencl.h index cc5edf02e..ef00263a5 100644 --- a/proto-opencl/emu/emu_opencl.h +++ b/proto-opencl/emu/emu_opencl.h @@ -43,6 +43,12 @@ uint intel_sub_group_shuffle_xor(uint v, uint mask); // same semantics uint sub_group_broadcast(uint v, uint lane); uint* emu_local_words(uint n); +// 16-byte vectors (read-width experiment, 5 October 2026): vload4/vstore4 over uint*, element-aligned as the spec says. +struct uint4 { uint x, y, z, w; }; +static inline uint4 vload4(size_t offset, const uint* p) { uint4 v; v.x = p[offset * 4]; v.y = p[offset * 4 + 1]; v.z = p[offset * 4 + 2]; v.w = p[offset * 4 + 3]; return v; } +static inline void vstore4(uint4 v, size_t offset, uint* p) { p[offset * 4] = v.x; p[offset * 4 + 1] = v.y; p[offset * 4 + 2] = v.z; p[offset * 4 + 3] = v.w; } +static inline uint4 IGNEUM_U4(uint a, uint b, uint c, uint d) { uint4 v; v.x = a; v.y = b; v.z = c; v.w = d; return v; } + // Integer built-ins with OpenCL semantics. static inline uint mul_hi(uint a, uint b) { return (uint)(((uint64_t)a * (uint64_t)b) >> 32); } // rotate(v, i): bits shifted left by i modulo the bit width (OpenCL C spec 6.3 and 6.12.3). diff --git a/proto-opencl/host.c b/proto-opencl/host.c index 457ec2090..31be99b18 100644 --- a/proto-opencl/host.c +++ b/proto-opencl/host.c @@ -202,6 +202,11 @@ typedef struct { const char* vendor; // --vendor S: pick the first GPU whose vendor string contains S (default: first GPU of any vendor) const char* packDir; // --pack D (serve only): generic mode, the pack is read from D at run time; the compiled-in pack is // then only the build-time placeholder of the prebuilt exe (4 October 2026) + int readback; // --readback: 0 select (a GPU-side pass reads back only the hits and 34 sentinel words), 1 full + // (every output word comes back, 8 bytes per nonce, the path before 5 October 2026) + int memprobe; // --memprobe: dependent-load latency and throughput, independent-load throughput and an ALU + // chain on the chosen device, no pack needed (5 October 2026, the 9070 XT on the eGPU) + int probeMib; // --probe-mib N: --memprobe at that one buffer size only (default 0 = 4, 64 and 1024 MiB) } Options; static int packMib(void) { return (int)(((1ull << IGNEUM_DATASET_LOG2) * 4ull) >> 20); } @@ -228,7 +233,11 @@ static void usage(void) { " --no-prepare with --serve: no prepare support (the miner then falls back to exit 42 at a seed change).\n" " Builds the pack's kernel_bound.cl (next to the compiled-in kernel.cl) unless --kernel says otherwise.\n" " --pack D with --serve: serve the pack in directory D (program.h, seeds.txt, vectors.h, kernel_bound.cl), whatever\n" - " pack this exe was built against; it is self-tested against its vectors.h first (the one-click worker)\n", packMib(), IGNEUM_KERNEL_PATH); + " pack this exe was built against; it is self-tested against its vectors.h first (the one-click worker)\n" + " --readback M with --serve: select (default) reads back only the hits and 34 sentinel words of each dispatch through a\n" + " GPU-side pass; full reads back every output (8 bytes per nonce). IGNEUM_READBACK=full does the same.\n" + " --memprobe no pack: dependent random loads (latency and throughput against lanes in flight), independent random\n" + " loads and an ALU chain on the chosen device, at 4, 64 and 1024 MiB (--probe-mib N for one size)\n", packMib(), IGNEUM_KERNEL_PATH); } static int isPow2(long long v) { return v > 0 && (v & (v - 1)) == 0; } @@ -239,6 +248,7 @@ static Options parseArgs(int argc, char** argv) { int i; o.datasetMib = 1024; o.batchLog2 = 24; o.batches = 5; o.groupWarps = 1; o.sweep = 0; o.device = -1; o.exchange = 0; o.list = 0; o.timeWall = -1; o.kernelPath = IGNEUM_KERNEL_PATH; o.extraOpts = ""; o.serve = 0; o.noPrepare = 0; o.kernelGiven = 0; o.vendor = NULL; o.packDir = NULL; + o.readback = (getenv("IGNEUM_READBACK") && strcmp(getenv("IGNEUM_READBACK"), "full") == 0) ? 1 : 0; o.memprobe = 0; o.probeMib = 0; for (i = 1; i < argc; ++i) { const char* a = argv[i]; int needs = (strcmp(a, "--dataset-mib") == 0 || strcmp(a, "--batch-log2") == 0 || strcmp(a, "--batches") == 0 || @@ -255,6 +265,16 @@ static Options parseArgs(int argc, char** argv) { else if (strcmp(a, "--no-prepare") == 0) o.noPrepare = 1; else if (strcmp(a, "--vendor") == 0) { if (i + 1 >= argc) { usage(); exit(2); } o.vendor = argv[++i]; } else if (strcmp(a, "--pack") == 0) { if (i + 1 >= argc) { usage(); exit(2); } o.packDir = argv[++i]; } + else if (strcmp(a, "--readback") == 0) { + const char* m; + if (i + 1 >= argc) { usage(); exit(2); } + m = argv[++i]; + if (strcmp(m, "select") == 0) o.readback = 0; + else if (strcmp(m, "full") == 0) o.readback = 1; + else { printf("--readback must be select or full\n"); exit(2); } + } + else if (strcmp(a, "--memprobe") == 0) o.memprobe = 1; + else if (strcmp(a, "--probe-mib") == 0) { if (i + 1 >= argc) { usage(); exit(2); } o.probeMib = atoi(argv[++i]); } else if (strcmp(a, "--build-opts") == 0) o.extraOpts = argv[++i]; else if (strcmp(a, "--time") == 0) { const char* m = argv[++i]; @@ -297,8 +317,49 @@ typedef struct { int cMajor, cMinor; // OpenCL C version int dMajor, dMinor; // device (platform profile) version cl_uint amdWavefront, nvWarp; // 0 if not reported + int dupOf; // index of the same card on a newer platform of the same vendor, -1 if none (5 October 2026) } DeviceInfo; +/* The driver version as a number for ordering ("3683.0 (PAL,LC)" -> 3683.0; "617.14" -> 617.14; 0 when unreadable). */ +static double driverNumber(const char* driver) { + const char* p = driver; + while (*p && (*p < '0' || *p > '9')) ++p; + return *p ? strtod(p, NULL) : 0.0; +} + +/* Two AMD platforms are registered after a driver update on Windows (PC 1, 5 October 2026: 32.0.21042 and 32.0.32015, + * OpenCL driver strings 3652.0 and 3683.0): every card is listed twice, the app ran two workers on one 9070 XT, and + * each got half. The same card on the same-named platform with a different platform version is the one card; the + * entry whose driver number is lower is the duplicate. Two real cards of one model sit on the SAME platform and are + * never folded. Indices stay flat (the app passes them back as --device). Returns how many duplicates were marked. */ +static int markDuplicates(DeviceInfo* list, int n) { + int i, j, marked = 0; + for (i = 0; i < n; ++i) list[i].dupOf = -1; + for (i = 0; i < n; ++i) { + if (list[i].dupOf >= 0) continue; + for (j = i + 1; j < n; ++j) { + int older; + if (list[j].dupOf >= 0) continue; + if (strcmp(list[i].platformName, list[j].platformName) != 0) continue; + if (strcmp(list[i].platformVersion, list[j].platformVersion) == 0) continue; + if (strcmp(list[i].name, list[j].name) != 0 || strcmp(list[i].vendor, list[j].vendor) != 0) continue; + if (list[i].globalMem != list[j].globalMem || list[i].computeUnits != list[j].computeUnits) continue; + older = driverNumber(list[j].driver) < driverNumber(list[i].driver) ? j : i; + if (older == j) { list[j].dupOf = i; } + else { list[i].dupOf = j; } + ++marked; + if (older == i) break; /* i itself is the duplicate; j stays the real one */ + } + } + /* three registrations of one card: every duplicate points at the one that stays, not at another duplicate */ + for (i = 0; i < n; ++i) { + int k = list[i].dupOf, hops = 0; + while (k >= 0 && list[k].dupOf >= 0 && hops++ < n) k = list[k].dupOf; + if (list[i].dupOf >= 0) list[i].dupOf = k; + } + return marked; +} + static void devStr(cl_device_id d, cl_device_info what, char* out, size_t n) { out[0] = 0; clGetDeviceInfo(d, what, n - 1, out, NULL); @@ -357,6 +418,7 @@ static int enumerateDevices(DeviceInfo** outList) { list[n++] = di; } } + if (n) markDuplicates(list, n); *outList = list; return n; } @@ -365,6 +427,13 @@ static void printDevice(int idx, const DeviceInfo* d, int chosen) { const char* subExt = strstr(d->extensions, "cl_khr_subgroup_shuffle") ? "cl_khr_subgroup_shuffle" : strstr(d->extensions, "cl_intel_subgroups") ? "cl_intel_subgroups" : strstr(d->extensions, "cl_khr_subgroups") ? "cl_khr_subgroups (no shuffle extension)" : "none"; + if (d->dupOf >= 0) { + /* Hidden from the app's card list: its parser takes only lines that start with "[" (detect.rs). The index + * is still valid for --device, so a run on the older platform stays possible for a comparison. */ + printf(" dup [%d] %s | %s (%s): the same card as [%d] on an older platform (driver %s); hidden, use [%d]\n", + idx, d->name, d->platformName, d->platformVersion, d->dupOf, d->driver, d->dupOf); + return; + } printf("%s[%d] %s | %s (%s)\n", chosen ? "*" : " ", idx, d->name, d->platformName, d->platformVersion); printf(" %s, vendor %s, driver %s, %s, %u compute units, %u MHz\n", typeName(d->type), d->vendor, d->driver, d->cVersion, d->computeUnits, d->clockMHz); printf(" global %llu MiB, max alloc %llu MiB, local %llu KiB, max work-group %llu, sub-group extension: %s", @@ -457,19 +526,45 @@ static void releaseProgram(Device* dv) { dv->kHash = dv->kCacheFill = dv->kBuild = dv->kFill = dv->kHashBound = NULL; dv->prog = NULL; } -// Sub-group size of igneum_hash for a work-group of `local` items. 0 if the query is unavailable (reason in *why). -static size_t querySubGroupSize(const Device* dv, const DeviceInfo* di, size_t local, char* why, size_t whyLen) { +// Sub-group size of kernel k for a work-group of `local` items. 0 if the query is unavailable (reason in *why). +static size_t querySubGroupSizeOf(cl_kernel k, const DeviceInfo* di, size_t local, char* why, size_t whyLen) { ig_pfn_subgroup_info fn = (ig_pfn_subgroup_info)clGetExtensionFunctionAddressForPlatform(di->platform, "clGetKernelSubGroupInfoKHR"); const char* via = "clGetKernelSubGroupInfoKHR"; size_t sg = 0; cl_int e; if (!fn) { fn = (ig_pfn_subgroup_info)loadSym("clGetKernelSubGroupInfo"); via = "clGetKernelSubGroupInfo (OpenCL 2.1 core, through the loader)"; } if (!fn) { snprintf(why, whyLen, "the sub-group size could not be queried (neither clGetKernelSubGroupInfoKHR nor clGetKernelSubGroupInfo is available)"); return 0; } - e = fn(dv->kHash, di->device, IG_CL_KERNEL_MAX_SUB_GROUP_SIZE_FOR_NDRANGE, sizeof(local), &local, sizeof(sg), &sg, NULL); + e = fn(k, di->device, IG_CL_KERNEL_MAX_SUB_GROUP_SIZE_FOR_NDRANGE, sizeof(local), &local, sizeof(sg), &sg, NULL); if (e != CL_SUCCESS) { snprintf(why, whyLen, "the sub-group size query failed: %s returned %s (%d)", via, clErrName(e), (int)e); return 0; } snprintf(why, whyLen, "queried through %s", via); return sg; } +static size_t querySubGroupSize(const Device* dv, const DeviceInfo* di, size_t local, char* why, size_t whyLen) { + return querySubGroupSizeOf(dv->kHash, di, local, why, whyLen); +} + +/* The compiled kernel as the driver sees it, on every exchange path (5 October 2026, the 9070 XT question): the + * work-group limit, the preferred multiple (the wave width the compiler chose), local memory, private memory (scratch: + * anything above 0 means spilled registers, which on AMD costs a memory round trip per spill), and the sub-group size + * for the built work-group size (wave32 or wave64 on RDNA; the local-memory path never queried it before). */ +#define IG_CL_KERNEL_PREFERRED_WORK_GROUP_SIZE_MULTIPLE 0x11B3 +#define IG_CL_KERNEL_PRIVATE_MEM_SIZE 0x11B4 +static void printKernelInfo(const DeviceInfo* di, cl_kernel k, const char* kernelName, int groupSize, const char* prefix) { + size_t wg = 0, mult = 0; + cl_ulong lmem = 0, pmem = 0; + char why[256]; + size_t sg; + clGetKernelWorkGroupInfo(k, di->device, CL_KERNEL_WORK_GROUP_SIZE, sizeof(wg), &wg, NULL); + clGetKernelWorkGroupInfo(k, di->device, IG_CL_KERNEL_PREFERRED_WORK_GROUP_SIZE_MULTIPLE, sizeof(mult), &mult, NULL); + clGetKernelWorkGroupInfo(k, di->device, CL_KERNEL_LOCAL_MEM_SIZE, sizeof(lmem), &lmem, NULL); + clGetKernelWorkGroupInfo(k, di->device, IG_CL_KERNEL_PRIVATE_MEM_SIZE, sizeof(pmem), &pmem, NULL); + sg = querySubGroupSizeOf(k, di, (size_t)groupSize, why, sizeof(why)); + printf("%skernel: %s max work-group %llu, preferred multiple %llu, local memory %llu bytes, private memory %llu bytes%s, work-group %d, sub-group size %llu (%s)%s\n", + prefix, kernelName, (unsigned long long)wg, (unsigned long long)mult, (unsigned long long)lmem, (unsigned long long)pmem, + pmem ? " (SPILLED: registers in scratch memory)" : "", groupSize, (unsigned long long)sg, why, + (sg > (size_t)groupSize) ? " (the work-group fills only part of a wave: see --group-warps)" : ""); + fflush(stdout); +} // Decide the exchange implementation and build. See WAVEFRONT.md for the rule. static void setupProgram(Device* dv, const DeviceInfo* di, const Options* o, const char* src, size_t srcLen) { @@ -1166,6 +1261,61 @@ static int runServe(Device* dv, const DeviceInfo* di, const Options* o) { hOut = (uint64_t*)malloc((size_t)batch * sizeof(uint64_t)); strncpy(devName, di->name, 255); devName[255] = 0; for (k = 0; devName[k]; ++k) if (devName[k] == ' ') devName[k] = '_'; + printKernelInfo(di, cur->kHashBound, "igneum_hash_bound", (int)groupSize, "info "); + /* The select pass (5 October 2026). Before it every dispatch read back 8 bytes per nonce (16 MiB for a 2^21-nonce + * job) over the bus and scanned them on the host; on an eGPU over USB4 that is a measurable part of every job. + * Now a tiny kernel built here (no pack involved) writes the hits (index, hash) behind an atomic counter plus 34 + * sentinel words (the first 32 outputs, the middle and the last), and the host reads back a few hundred bytes. + * The fault detectors read the sentinels; the found lines are printed in nonce order from the sorted hits. If a + * chunk has more hits than the table holds (a target that loose is a test, not a block), the chunk falls back to the + * full read. --readback full or IGNEUM_READBACK=full keeps the old path for a comparison. */ +#define IG_MAX_HITS 256u + cl_program selProg = NULL; + cl_kernel kSelect = NULL; + cl_mem dCount = NULL, dHits = NULL, dSentinel = NULL; + size_t selLocal = di->maxWorkGroup < 256 ? di->maxWorkGroup : 256; + uint32_t hCount = 0; + uint64_t hHits[IG_MAX_HITS * 2]; + uint64_t hSentinel[34]; + unsigned long long bytesUp = 0, bytesDown = 0; + double kernelMsSum = 0, selectMsSum = 0, readMsSum = 0, scanMsSum = 0; + unsigned long fullFallbacks = 0; + if (!o->readback) { + static const char* SELECT_SRC = + "__kernel void igneum_select(__global const ulong* out, uint n, ulong target, volatile __global uint* count,\n" + " __global ulong* hits, uint maxHits, __global ulong* sentinel) {\n" + " uint i = (uint)get_global_id(0);\n" + " if (i < n) {\n" + " ulong h = out[i];\n" + " if (h <= target) { uint k = atomic_inc(count); if (k < maxHits) { hits[2u * k] = (ulong)i; hits[2u * k + 1u] = h; } }\n" + " if (i < 32u) sentinel[i] = h;\n" + " if (i == 0u) { sentinel[32] = out[n / 2u]; sentinel[33] = out[n - 1u]; }\n" + " }\n" + "}\n"; + size_t selLen = strlen(SELECT_SRC); + selProg = clCreateProgramWithSource(dv->ctx, 1, &SELECT_SRC, &selLen, &err); CL_CHECK_ERR(err, "clCreateProgramWithSource select"); + err = clBuildProgram(selProg, 1, &di->device, "-cl-std=CL1.2", NULL, NULL); + if (err != CL_SUCCESS) { + size_t logLen = 0; char* log; + clGetProgramBuildInfo(selProg, di->device, CL_PROGRAM_BUILD_LOG, 0, NULL, &logLen); + log = (char*)calloc(logLen + 1, 1); + if (logLen) clGetProgramBuildInfo(selProg, di->device, CL_PROGRAM_BUILD_LOG, logLen, log, NULL); + printf("info the select pass did not build (%s): %.300s; using the full read-back\n", clErrName(err), log); + free(log); clReleaseProgram(selProg); selProg = NULL; ((Options*)o)->readback = 1; + } else { + kSelect = clCreateKernel(selProg, "igneum_select", &err); CL_CHECK_ERR(err, "clCreateKernel igneum_select"); + dCount = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, 4, NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer count"); + dHits = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, IG_MAX_HITS * 2 * sizeof(uint64_t), NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer hits"); + dSentinel = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, 34 * sizeof(uint64_t), NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer sentinel"); + gMemCreated += 3; + { + size_t wg = 0; + if (clGetKernelWorkGroupInfo(kSelect, di->device, CL_KERNEL_WORK_GROUP_SIZE, sizeof(wg), &wg, NULL) == CL_SUCCESS && wg && wg < selLocal) selLocal = wg; + } + } + } + printf("info readback %s (per dispatch of %u nonces: %s)\n", o->readback ? "full" : "select", + batch, o->readback ? "8 bytes per nonce come back and the host scans them" : "the hits and 34 sentinel words come back; the GPU scans"); printf("ready opencl %s platform %s pack %s dataset-log2 %d batch %u exchange %d prepare %d path %s\n", devName, di->platformName, gGeneric ? gPack.seedString : IGNEUM_SEED_STRING, log2u32(words), batch, dv->exchange, o->noPrepare ? 0 : 1, gGeneric ? "prebuilt-generic" : "compiled-in"); fflush(stdout); @@ -1178,6 +1328,7 @@ static int runServe(Device* dv, const DeviceInfo* di, const Options* o) { long gFaultTestChunk = getenv("IGNEUM_FAULT_TEST") ? atol(getenv("IGNEUM_FAULT_TEST")) : -1; double meanNsPerNonce = 0; /* running mean of wall ns per nonce over the chunks so far */ unsigned long chunksSeen = 0, jobsSeen = 0; + unsigned long long noncesSeen = 0; /* for the mean chunk in the stats line */ uint64_t prevSig = 0; int havePrevSig = 0; double jobMsSum = 0; #define SERVE_FATAL(jobIdStr, fmt, ...) do { printf("error %s worker fault: " fmt "; exiting 3 so the miner restarts the worker\n", jobIdStr, __VA_ARGS__); fflush(stdout); exit(3); } while (0) @@ -1278,11 +1429,15 @@ static int runServe(Device* dv, const DeviceInfo* di, const Options* o) { b[45] = (uint8_t)hi; b[46] = (uint8_t)(hi >> 8); b[47] = (uint8_t)(hi >> 16); b[48] = (uint8_t)(hi >> 24); seedWordsFromBytes(b, 49, iw); { - double c0 = wallMs(), cms; + double c0 = wallMs(), cms, kernelMs = -1.0, selectMs = 0.0, readMs = 0.0, scanMs = 0.0, r0; cl_int status = 0; size_t g = ((chunk + groupSize - 1) / groupSize) * groupSize; uint64_t sig; + int useSelect = (kSelect != NULL); + uint32_t nHits = 0; SERVE_CHECK(jobId, clEnqueueWriteBuffer(dv->q, dInit, CL_TRUE, 0, 32, iw, 0, NULL, NULL)); + bytesUp += 32; + if (useSelect) { hCount = 0; SERVE_CHECK(jobId, clEnqueueWriteBuffer(dv->q, dCount, CL_TRUE, 0, 4, &hCount, 0, NULL, NULL)); bytesUp += 4; } baseNonce = lo; SERVE_CHECK(jobId, clSetKernelArg(cur->kHashBound, 0, sizeof(cl_mem), &cur->ds)); SERVE_CHECK(jobId, clSetKernelArg(cur->kHashBound, 1, sizeof(cl_mem), &dOut)); @@ -1302,11 +1457,50 @@ static int runServe(Device* dv, const DeviceInfo* di, const Options* o) { err = clWaitForEvents(1, &ev); if (err != CL_SUCCESS) { countRelease(ev); SERVE_FATAL(jobId, "clWaitForEvents on the dispatch returned %s (%d)", clErrName(err), (int)err); } err = clGetEventInfo(ev, CL_EVENT_COMMAND_EXECUTION_STATUS, sizeof(status), &status, NULL); + kernelMs = eventMs(ev); countRelease(ev); if (err != CL_SUCCESS) SERVE_FATAL(jobId, "clGetEventInfo on the dispatch returned %s (%d)", clErrName(err), (int)err); if (status != CL_COMPLETE) SERVE_FATAL(jobId, "the dispatch event ended with status %d, not CL_COMPLETE (a device reset or a lost context)", (int)status); + if (useSelect) { + cl_event evs = NULL; + cl_uint nArg = chunk, maxArg = IG_MAX_HITS; + cl_ulong tArg = (cl_ulong)target; + size_t gs = ((chunk + selLocal - 1) / selLocal) * selLocal; + SERVE_CHECK(jobId, clSetKernelArg(kSelect, 0, sizeof(cl_mem), &dOut)); + SERVE_CHECK(jobId, clSetKernelArg(kSelect, 1, sizeof(cl_uint), &nArg)); + SERVE_CHECK(jobId, clSetKernelArg(kSelect, 2, sizeof(cl_ulong), &tArg)); + SERVE_CHECK(jobId, clSetKernelArg(kSelect, 3, sizeof(cl_mem), &dCount)); + SERVE_CHECK(jobId, clSetKernelArg(kSelect, 4, sizeof(cl_mem), &dHits)); + SERVE_CHECK(jobId, clSetKernelArg(kSelect, 5, sizeof(cl_uint), &maxArg)); + SERVE_CHECK(jobId, clSetKernelArg(kSelect, 6, sizeof(cl_mem), &dSentinel)); + SERVE_CHECK(jobId, clEnqueueNDRangeKernel(dv->q, kSelect, 1, NULL, &gs, &selLocal, 0, NULL, &evs)); + ++gEvCreated; + err = clWaitForEvents(1, &evs); + if (err != CL_SUCCESS) { countRelease(evs); SERVE_FATAL(jobId, "clWaitForEvents on the select pass returned %s (%d)", clErrName(err), (int)err); } + selectMs = eventMs(evs); + countRelease(evs); } - SERVE_CHECK(jobId, clEnqueueReadBuffer(dv->q, dOut, CL_TRUE, 0, (size_t)chunk * sizeof(uint64_t), hOut, 0, NULL, NULL)); + } + r0 = wallMs(); + if (useSelect) { + SERVE_CHECK(jobId, clEnqueueReadBuffer(dv->q, dCount, CL_TRUE, 0, 4, &hCount, 0, NULL, NULL)); + bytesDown += 4; + nHits = hCount; + if (nHits > IG_MAX_HITS) { + /* more hits than the table holds: this chunk takes the full path (the test target case) */ + ++fullFallbacks; + useSelect = 0; + } else { + if (nHits) { SERVE_CHECK(jobId, clEnqueueReadBuffer(dv->q, dHits, CL_TRUE, 0, (size_t)nHits * 2 * sizeof(uint64_t), hHits, 0, NULL, NULL)); bytesDown += (unsigned long long)nHits * 16; } + SERVE_CHECK(jobId, clEnqueueReadBuffer(dv->q, dSentinel, CL_TRUE, 0, 34 * sizeof(uint64_t), hSentinel, 0, NULL, NULL)); + bytesDown += 34 * 8; + } + } + if (!useSelect) { + SERVE_CHECK(jobId, clEnqueueReadBuffer(dv->q, dOut, CL_TRUE, 0, (size_t)chunk * sizeof(uint64_t), hOut, 0, NULL, NULL)); + bytesDown += (unsigned long long)chunk * 8; + } + readMs = wallMs() - r0; cms = wallMs() - c0; /* Plausibility: wall time per nonce against the running mean (the first chunk sets it; a chunk is a full * batch except at the 32-bit boundary, so per nonce is the comparable unit). 20x faster = the kernel did not run. */ @@ -1315,18 +1509,39 @@ static int runServe(Device* dv, const DeviceInfo* di, const Options* o) { if (chunksSeen >= 4 && meanNsPerNonce > 0 && ns * 20.0 < meanNsPerNonce) SERVE_FATAL(jobId, "%u nonces reported complete in %.3f ms, %.0fx faster than the running mean of %.2f ms per million (the runtime is not running the kernel)", chunk, cms, meanNsPerNonce / ns, meanNsPerNonce / 1e3); meanNsPerNonce = chunksSeen == 0 ? ns : meanNsPerNonce + (ns - meanNsPerNonce) / (double)(chunksSeen + 1 < 64 ? chunksSeen + 1 : 64); - ++chunksSeen; + ++chunksSeen; noncesSeen += chunk; } - /* Stale output: the first, middle and last words plus an FNV of the first 32. Two chunks never share it. */ - sig = fnv1a64(hOut, 32 * sizeof(uint64_t)) ^ hOut[chunk / 2] ^ hOut[chunk - 1]; + /* Stale output: the first, middle and last words plus an FNV of the first 32. Two chunks never share it. + * On the select path the same 34 words come from the sentinel buffer the select pass wrote. */ + if (useSelect) sig = fnv1a64(hSentinel, 32 * sizeof(uint64_t)) ^ hSentinel[32] ^ hSentinel[33]; + else sig = fnv1a64(hOut, 32 * sizeof(uint64_t)) ^ hOut[chunk / 2] ^ hOut[chunk - 1]; if (chunk >= 64) { if (havePrevSig && sig == prevSig) SERVE_FATAL(jobId, "the output buffer is unchanged since the previous dispatch (signature %016llx): the kernel did not run", (unsigned long long)sig); prevSig = sig; havePrevSig = 1; } + r0 = wallMs(); + if (useSelect) { + /* nonce order, as the full scan printed them (the miner submits the first found of a job) */ + uint32_t a, b; + for (a = 1; a < nHits; ++a) { + uint64_t ki = hHits[2 * a], kh = hHits[2 * a + 1]; + for (b = a; b > 0 && hHits[2 * (b - 1)] > ki; --b) { hHits[2 * b] = hHits[2 * (b - 1)]; hHits[2 * b + 1] = hHits[2 * (b - 1) + 1]; } + hHits[2 * b] = ki; hHits[2 * b + 1] = kh; + } + for (a = 0; a < nHits; ++a) { + uint32_t idx = (uint32_t)hHits[2 * a]; + unsigned long long nonce = ((unsigned long long)hi << 32) | (unsigned long long)(uint32_t)(lo + idx); + printf("found %s %llu %016llx\n", jobId, nonce, (unsigned long long)hHits[2 * a + 1]); + } + } else { + for (i = 0; i < chunk; ++i) if (hOut[i] <= target) { + unsigned long long nonce = ((unsigned long long)hi << 32) | (unsigned long long)(uint32_t)(lo + i); + printf("found %s %llu %016llx\n", jobId, nonce, (unsigned long long)hOut[i]); + } } - for (i = 0; i < chunk; ++i) if (hOut[i] <= target) { - unsigned long long nonce = ((unsigned long long)hi << 32) | (unsigned long long)(uint32_t)(lo + i); - printf("found %s %llu %016llx\n", jobId, nonce, (unsigned long long)hOut[i]); + scanMs = wallMs() - r0; + if (kernelMs > 0) kernelMsSum += kernelMs; + selectMsSum += selectMs; readMsSum += readMs; scanMsSum += scanMs; } fflush(stdout); hashes += chunk; @@ -1340,6 +1555,12 @@ static int runServe(Device* dv, const DeviceInfo* di, const Options* o) { if (jobsSeen % 200 == 0) { printf("info stats jobs %lu chunks %lu mean job %.1f ms mean %.2f ms per million nonces; events created %lu released %lu live %lu; buffers created %lu released %lu live %lu\n", jobsSeen, chunksSeen, jobMsSum / (double)jobsSeen, meanNsPerNonce / 1e3, gEvCreated, gEvReleased, gEvCreated - gEvReleased, gMemCreated, gMemReleased, gMemCreated - gMemReleased); + /* The transfer and time budget per chunk (5 October 2026): bytes up (init words, the count reset), bytes + * down (hits and sentinels, or 8 bytes per nonce on the full path), and the mean device time of the hash + * kernel (event profiling), the select pass, the blocking read-back and the host scan. */ + printf("info transfers per chunk: up %.0f B, down %.0f B (%s, %lu full fallbacks); mean per chunk: kernel %.2f ms (device), select %.3f ms, read-back %.2f ms, scan %.2f ms, chunk wall %.2f ms\n", + (double)bytesUp / (double)chunksSeen, (double)bytesDown / (double)chunksSeen, kSelect ? "select" : "full", fullFallbacks, + kernelMsSum / (double)chunksSeen, selectMsSum / (double)chunksSeen, readMsSum / (double)chunksSeen, scanMsSum / (double)chunksSeen, meanNsPerNonce * ((double)noncesSeen / (double)chunksSeen) / 1e6); if (gEvCreated - gEvReleased > 16 || gMemCreated - gMemReleased > 12) SERVE_FATAL(jobId, "object leak: %lu events and %lu buffers live after %lu jobs", gEvCreated - gEvReleased, gMemCreated - gMemReleased, jobsSeen); } if (switched && old) { releasePair(old); old = NULL; printf("info dropped the previous pair (its program, cache and dataset)\n"); } @@ -1349,6 +1570,9 @@ static int runServe(Device* dv, const DeviceInfo* di, const Options* o) { clReleaseMemObject(dInit); clReleaseMemObject(dOut); gMemReleased += 2; + if (kSelect) { clReleaseKernel(kSelect); clReleaseMemObject(dCount); clReleaseMemObject(dHits); clReleaseMemObject(dSentinel); gMemReleased += 3; } + if (selProg) clReleaseProgram(selProg); + if (chunksSeen) printf("info transfers total: %lu chunks, up %llu B, down %llu B (%s, %lu full fallbacks)\n", chunksSeen, bytesUp, bytesDown, kSelect ? "select" : "full", fullFallbacks); if (old) releasePair(old); if (prepared) releasePair(prepared); releasePair(cur); @@ -1356,6 +1580,224 @@ static int runServe(Device* dv, const DeviceInfo* di, const Options* o) { #endif } +/* --------------------------------------------------------------------------------------------- + * --memprobe (5 October 2026): what the device itself can do with the access pattern of the hash, with no pack and no + * program. Three kernels built from the text below: + * chase one dependent random 4-byte load per step per lane (the next address comes from the loaded word), so the + * time per step at a small lane count is the loaded-latency of one random read, and the loads per second at a + * large lane count is the device's random-read throughput for a dependent chain (the hash is 128 of these) + * indep eight independent chains per lane: the throughput when latency is hidden inside one lane + * alu an integer multiply-add-rotate chain with no memory: the achieved integer rate, which moves with the clock + * Sizes 4 MiB (cache-resident), 64 MiB (the last-level cache on RDNA 3 and 4 is 64 MiB or more, approximate) and + * 1024 MiB (the dataset size: DRAM). Times from event profiling (wall on Apple). Every number is printed with the + * configuration that produced it. Rates are G loads/s = 1e9 loads per second; ns per load = time / steps. + */ +static const char* PROBE_SRC = + "static inline uint pm_mix(uint x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }\n" + "__kernel void probe_fill(__global uint* ds, uint n) { uint i = (uint)get_global_id(0); if (i < n) ds[i] = pm_mix(i ^ 0x9E3779B9u); }\n" + "__kernel void probe_chase(__global const uint* ds, uint mask, uint steps, uint seed, __global uint* out) {\n" + " uint x = pm_mix((uint)get_global_id(0) ^ seed);\n" + " for (uint s = 0u; s < steps; ++s) x = ds[x & mask] ^ (x * 0x9E3779B1u + s);\n" + " out[get_global_id(0)] = x;\n" + "}\n" + "__kernel void probe_indep(__global const uint* ds, uint mask, uint steps, uint seed, __global uint* out) {\n" + " uint g = (uint)get_global_id(0);\n" + " uint x0 = pm_mix(g * 8u ^ seed), x1 = pm_mix((g * 8u + 1u) ^ seed), x2 = pm_mix((g * 8u + 2u) ^ seed), x3 = pm_mix((g * 8u + 3u) ^ seed);\n" + " uint x4 = pm_mix((g * 8u + 4u) ^ seed), x5 = pm_mix((g * 8u + 5u) ^ seed), x6 = pm_mix((g * 8u + 6u) ^ seed), x7 = pm_mix((g * 8u + 7u) ^ seed);\n" + " for (uint s = 0u; s < steps; ++s) {\n" + " x0 = ds[x0 & mask] ^ (x0 * 0x9E3779B1u + s); x1 = ds[x1 & mask] ^ (x1 * 0x9E3779B1u + s);\n" + " x2 = ds[x2 & mask] ^ (x2 * 0x9E3779B1u + s); x3 = ds[x3 & mask] ^ (x3 * 0x9E3779B1u + s);\n" + " x4 = ds[x4 & mask] ^ (x4 * 0x9E3779B1u + s); x5 = ds[x5 & mask] ^ (x5 * 0x9E3779B1u + s);\n" + " x6 = ds[x6 & mask] ^ (x6 * 0x9E3779B1u + s); x7 = ds[x7 & mask] ^ (x7 * 0x9E3779B1u + s);\n" + " }\n" + " out[g] = x0 ^ x1 ^ x2 ^ x3 ^ x4 ^ x5 ^ x6 ^ x7;\n" + "}\n" + "__kernel void probe_line(__global const uint4* ds, uint lineMask, uint steps, uint seed, __global uint* out) {\n" + " uint x = pm_mix((uint)get_global_id(0) ^ seed);\n" + " for (uint s = 0u; s < steps; ++s) {\n" + " uint l = (x & lineMask) * 4u;\n" + " uint4 a = ds[l], b = ds[l + 1u], c = ds[l + 2u], d = ds[l + 3u];\n" + " x = (a.x ^ b.y ^ c.z ^ d.w) ^ (x * 0x9E3779B1u + s);\n" + " }\n" + " out[get_global_id(0)] = x;\n" + "}\n" + "__kernel void probe_stream(__global const uint4* ds, uint perLane, __global uint* out) {\n" + " uint g = (uint)get_global_id(0), n = (uint)get_global_size(0);\n" + " uint4 acc = (uint4)(0u, 0u, 0u, 0u);\n" + " for (uint s = 0u; s < perLane; ++s) acc ^= ds[s * n + g];\n" + " out[g] = acc.x ^ acc.y ^ acc.z ^ acc.w;\n" + "}\n" + "__kernel void probe_alu(uint steps, uint seed, __global uint* out) {\n" + " uint g = (uint)get_global_id(0); uint x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u;\n" + " for (uint s = 0u; s < steps; ++s) { x = x * 0x9E3779B1u + rotate(y, 7u); y = (y ^ x) + s; }\n" + " out[g] = x ^ y;\n" + "}\n"; + +/* One timed launch, best of `reps`, in ms (event time, or wall when the platform's events are unusable). */ +static double probeLaunch(Device* dv, const Options* o, cl_kernel k, size_t global, size_t local, int reps, int seedArg, cl_uint seed) { + double best = -1.0; + int r; + for (r = 0; r < reps; ++r) { + cl_event ev = NULL; + double w0 = wallMs(), ms; + /* a fresh seed per repetition: a replay of the same addresses would be served from the last-level cache + * (65,536 reads x 64 B = 4 MiB fits any of them) and read as DRAM latency (seen on the 9070 XT, 5 October) */ + if (seedArg >= 0) { cl_uint s = seed + (cl_uint)r * 0x9E3779B9u; CL_CHECK(clSetKernelArg(k, (cl_uint)seedArg, sizeof(cl_uint), &s)); } + CL_CHECK(clEnqueueNDRangeKernel(dv->q, k, 1, NULL, &global, &local, 0, NULL, &ev)); + CL_CHECK(clWaitForEvents(1, &ev)); + ms = o->timeWall ? wallMs() - w0 : eventMs(ev); + if (ms < 0) ms = wallMs() - w0; + clReleaseEvent(ev); + if (best < 0 || ms < best) best = ms; + } + return best; +} + +static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) { + cl_int err = 0; + cl_program prog; + cl_kernel kFill, kChase, kIndep, kAlu, kLine, kStream; + size_t srcLen = strlen(PROBE_SRC); + int sizes[3] = { 4, 64, 1024 }, nSizes = 3, si; + size_t lanesList[8] = { 256, 1024, 1u << 12, 1u << 14, 1u << 16, 1u << 18, 1u << 20, 1u << 22 }; + const int nLanes = 8; + size_t groups[2] = { 32, 256 }; + const cl_uint STEPS = 256u, ALU_STEPS = 4096u; + const size_t maxLanes = 1u << 22; + cl_mem dOut; + if (o->probeMib > 0) { sizes[0] = o->probeMib; nSizes = 1; } + prog = clCreateProgramWithSource(dv->ctx, 1, &PROBE_SRC, &srcLen, &err); CL_CHECK_ERR(err, "clCreateProgramWithSource probe"); + err = clBuildProgram(prog, 1, &di->device, "-cl-std=CL1.2", NULL, NULL); + if (err != CL_SUCCESS) { + size_t logLen = 0; char* log; + clGetProgramBuildInfo(prog, di->device, CL_PROGRAM_BUILD_LOG, 0, NULL, &logLen); + log = (char*)calloc(logLen + 1, 1); + if (logLen) clGetProgramBuildInfo(prog, di->device, CL_PROGRAM_BUILD_LOG, logLen, log, NULL); + printf("memprobe: build FAILED (%s)\n--- build log ---\n%s\n--- end of build log ---\n", clErrName(err), log); + free(log); + return 2; + } + kFill = clCreateKernel(prog, "probe_fill", &err); CL_CHECK_ERR(err, "probe_fill"); + kChase = clCreateKernel(prog, "probe_chase", &err); CL_CHECK_ERR(err, "probe_chase"); + kIndep = clCreateKernel(prog, "probe_indep", &err); CL_CHECK_ERR(err, "probe_indep"); + kAlu = clCreateKernel(prog, "probe_alu", &err); CL_CHECK_ERR(err, "probe_alu"); + kLine = clCreateKernel(prog, "probe_line", &err); CL_CHECK_ERR(err, "probe_line"); + kStream = clCreateKernel(prog, "probe_stream", &err); CL_CHECK_ERR(err, "probe_stream"); + dOut = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, maxLanes * 4u, NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer probe out"); + printf("memprobe on [%s] %s, driver %s, %u compute units, %u MHz, %s time\n", di->platformName, di->name, di->driver, di->computeUnits, di->clockMHz, o->timeWall ? "wall" : "device event"); + printKernelInfo(di, kChase, "probe_chase", 32, "memprobe "); + printKernelInfo(di, kChase, "probe_chase", 256, "memprobe "); + printf("| probe | MiB | work-group | lanes in flight | steps per lane | best ms | G loads/s | ns per dependent load |\n|---|---|---|---|---|---|---|---|\n"); + for (si = 0; si < nSizes; ++si) { + int mib = sizes[si]; + uint64_t bytes = (uint64_t)mib << 20; + cl_uint words = (cl_uint)(bytes / 4ull), mask = words - 1u, n = words; + cl_mem dDs; + size_t gi, li; + size_t fillLocal = di->maxWorkGroup < 256 ? di->maxWorkGroup : 256; + size_t fillGlobal = ((size_t)words + fillLocal - 1) / fillLocal * fillLocal; + if ((uint64_t)di->maxAlloc < bytes) { printf("| chase | %d | skipped: max alloc %llu MiB | | | | | |\n", mib, (unsigned long long)(di->maxAlloc >> 20)); continue; } + dDs = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, (size_t)bytes, NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer probe dataset"); + CL_CHECK(clSetKernelArg(kFill, 0, sizeof(cl_mem), &dDs)); + CL_CHECK(clSetKernelArg(kFill, 1, sizeof(cl_uint), &n)); + probeLaunch(dv, o, kFill, fillGlobal, fillLocal, 1, -1, 0u); + for (gi = 0; gi < 2; ++gi) { + size_t local = groups[gi]; + if (local > di->maxWorkGroup) continue; + for (li = 0; li < (size_t)nLanes; ++li) { + size_t lanes = lanesList[li]; + cl_uint seed = (cl_uint)(0x1234567u + (cl_uint)li * 977u); + double ms; + if (lanes < local) continue; + CL_CHECK(clSetKernelArg(kChase, 0, sizeof(cl_mem), &dDs)); + CL_CHECK(clSetKernelArg(kChase, 1, sizeof(cl_uint), &mask)); + CL_CHECK(clSetKernelArg(kChase, 2, sizeof(cl_uint), &STEPS)); + CL_CHECK(clSetKernelArg(kChase, 3, sizeof(cl_uint), &seed)); + CL_CHECK(clSetKernelArg(kChase, 4, sizeof(cl_mem), &dOut)); + ms = probeLaunch(dv, o, kChase, lanes, local, 3, 3, seed); + printf("| chase | %d | %llu | %llu | %u | %.3f | %.3f | %.0f |\n", mib, (unsigned long long)local, (unsigned long long)lanes, STEPS, ms, + (double)lanes * (double)STEPS / (ms / 1000.0) / 1e9, ms * 1e6 / (double)STEPS); + fflush(stdout); + } + } + { + size_t local = di->maxWorkGroup < 256 ? di->maxWorkGroup : 256; + size_t lanes; + for (lanes = 1u << 16; lanes <= maxLanes; lanes <<= 2) { + cl_uint seed = 0x7654321u; + double ms; + CL_CHECK(clSetKernelArg(kIndep, 0, sizeof(cl_mem), &dDs)); + CL_CHECK(clSetKernelArg(kIndep, 1, sizeof(cl_uint), &mask)); + CL_CHECK(clSetKernelArg(kIndep, 2, sizeof(cl_uint), &STEPS)); + CL_CHECK(clSetKernelArg(kIndep, 3, sizeof(cl_uint), &seed)); + CL_CHECK(clSetKernelArg(kIndep, 4, sizeof(cl_mem), &dOut)); + ms = probeLaunch(dv, o, kIndep, lanes, local, 3, 3, seed); + printf("| indep x8 | %d | %llu | %llu | %u | %.3f | %.3f | (8 loads in flight per lane) |\n", mib, (unsigned long long)local, (unsigned long long)lanes, STEPS, ms, + (double)lanes * 8.0 * (double)STEPS / (ms / 1000.0) / 1e9); + fflush(stdout); + } + } + { + /* Random 64-byte lines (16 words, four uint4 loads) in a dependent chain: lines per second against the + * 4-byte chase above says what one random 4-byte read costs the memory system. If the two rates are + * equal, every 4-byte read fetches a whole line; if lines/s is a quarter of loads/s, reads cost a 16-byte + * sector. */ + size_t local = di->maxWorkGroup < 256 ? di->maxWorkGroup : 256; + size_t lanes; + cl_uint lineMask = (words / 16u) - 1u; + for (lanes = 1u << 14; lanes <= maxLanes; lanes <<= 2) { + cl_uint seed = 0x3141592u; + double ms; + CL_CHECK(clSetKernelArg(kLine, 0, sizeof(cl_mem), &dDs)); + CL_CHECK(clSetKernelArg(kLine, 1, sizeof(cl_uint), &lineMask)); + CL_CHECK(clSetKernelArg(kLine, 2, sizeof(cl_uint), &STEPS)); + CL_CHECK(clSetKernelArg(kLine, 3, sizeof(cl_uint), &seed)); + CL_CHECK(clSetKernelArg(kLine, 4, sizeof(cl_mem), &dOut)); + ms = probeLaunch(dv, o, kLine, lanes, local, 3, 3, seed); + printf("| line 64 B | %d | %llu | %llu | %u | %.3f | %.3f G lines/s | %.1f GB/s in lines |\n", mib, (unsigned long long)local, (unsigned long long)lanes, STEPS, ms, + (double)lanes * (double)STEPS / (ms / 1000.0) / 1e9, (double)lanes * (double)STEPS * 64.0 / (ms / 1000.0) / 1e9); + fflush(stdout); + } + } + { + /* Coalesced read of the whole buffer (uint4 per lane per step, consecutive lanes consecutive addresses): + * the sequential bandwidth. Against the card's rated figure this says whether the memory clock is in its + * full state; a card parked in a middle memory state shows about half (approximate). */ + size_t local = di->maxWorkGroup < 256 ? di->maxWorkGroup : 256; + size_t lanes = 1u << 20; + cl_uint perLane = (cl_uint)((uint64_t)words / 4ull / (uint64_t)lanes); + double ms, bytes = (double)perLane * (double)lanes * 16.0; + if (perLane == 0) { perLane = 1; lanes = (size_t)words / 4u; bytes = (double)lanes * 16.0; } + CL_CHECK(clSetKernelArg(kStream, 0, sizeof(cl_mem), &dDs)); + CL_CHECK(clSetKernelArg(kStream, 1, sizeof(cl_uint), &perLane)); + CL_CHECK(clSetKernelArg(kStream, 2, sizeof(cl_mem), &dOut)); + ms = probeLaunch(dv, o, kStream, lanes, local, 3, -1, 0u); + printf("| stream | %d | %llu | %llu | %u | %.3f | %.1f GB/s coalesced | (%.0f MiB read once) |\n", mib, (unsigned long long)local, (unsigned long long)lanes, perLane, ms, bytes / (ms / 1000.0) / 1e9, bytes / 1048576.0); + fflush(stdout); + } + clReleaseMemObject(dDs); + } + { + size_t local = di->maxWorkGroup < 256 ? di->maxWorkGroup : 256; + size_t lanes = 1u << 20; + cl_uint seed = 0x2468aceu; + double ms, ops; + CL_CHECK(clSetKernelArg(kAlu, 0, sizeof(cl_uint), &ALU_STEPS)); + CL_CHECK(clSetKernelArg(kAlu, 1, sizeof(cl_uint), &seed)); + CL_CHECK(clSetKernelArg(kAlu, 2, sizeof(cl_mem), &dOut)); + ms = probeLaunch(dv, o, kAlu, lanes, local, 3, 1, seed); + ops = (double)lanes * (double)ALU_STEPS * 5.0; /* mul, add, rotate, xor, add per step */ + printf("| alu | 0 | %llu | %llu | %u | %.3f | %.1f G int ops/s | %.3f G steps/s per compute unit (approximate: 5 ops per step counted) |\n", + (unsigned long long)local, (unsigned long long)lanes, ALU_STEPS, ms, ops / (ms / 1000.0) / 1e9, + (double)lanes * (double)ALU_STEPS / (ms / 1000.0) / 1e9 / (double)(di->computeUnits ? di->computeUnits : 1)); + } + clReleaseMemObject(dOut); + clReleaseKernel(kFill); clReleaseKernel(kChase); clReleaseKernel(kIndep); clReleaseKernel(kAlu); clReleaseKernel(kLine); clReleaseKernel(kStream); + clReleaseProgram(prog); + printf("memprobe: done\n"); + return 0; +} + int main(int argc, char** argv) { Options o = parseArgs(argc, argv); DeviceInfo* devs = NULL; @@ -1397,7 +1839,7 @@ int main(int argc, char** argv) { if (o.device >= nDev) { printf("FAIL: --device %d out of range (%d devices)\n", o.device, nDev); return 2; } chosen = o.device; } else if (o.vendor) { - for (i = 0; i < nDev; ++i) if ((devs[i].type & CL_DEVICE_TYPE_GPU) && strstr(devs[i].vendor, o.vendor)) { chosen = i; break; } + for (i = 0; i < nDev; ++i) if ((devs[i].type & CL_DEVICE_TYPE_GPU) && devs[i].dupOf < 0 && strstr(devs[i].vendor, o.vendor)) { chosen = i; break; } if (chosen < 0) { printf("OpenCL devices (%d):\n", nDev); for (i = 0; i < nDev; ++i) printDevice(i, &devs[i], 0); @@ -1405,14 +1847,20 @@ int main(int argc, char** argv) { return 2; } } else { - for (i = 0; i < nDev; ++i) if (devs[i].type & CL_DEVICE_TYPE_GPU) { chosen = i; break; } + for (i = 0; i < nDev; ++i) if ((devs[i].type & CL_DEVICE_TYPE_GPU) && devs[i].dupOf < 0) { chosen = i; break; } if (chosen < 0) chosen = 0; } printf("OpenCL devices (%d):\n", nDev); for (i = 0; i < nDev; ++i) printDevice(i, &devs[i], i == chosen && !o.list); + { + int dups = 0; + for (i = 0; i < nDev; ++i) if (devs[i].dupOf >= 0) ++dups; + if (dups) printf("platforms: %d device(s) hidden as the same card on an older platform of the same vendor (an old driver's OpenCL registration is still present)\n", dups); + } if (o.list) return 0; di = &devs[chosen]; printf("using device [%d] %s\n", chosen, di->name); + if (di->dupOf >= 0) printf("NOTE: --device %d is the older platform's listing of the card [%d] (driver %s); results on it are for comparison only\n", chosen, di->dupOf, di->driver); if (o.timeWall < 0) o.timeWall = (strcmp(di->platformName, "Apple") == 0) ? 1 : 0; if (o.timeWall) printf("timing: host wall time (Apple's OpenCL event timestamps are not usable; the rate is still a device rate, see README.md)\n"); else printf("timing: device event profiling (CL_PROFILING_COMMAND_START/END), like cudaEvent elapsed time\n"); @@ -1423,6 +1871,12 @@ int main(int argc, char** argv) { dv.q = clCreateCommandQueue(dv.ctx, di->device, CL_QUEUE_PROFILING_ENABLE, &err); CL_CHECK_ERR(err, "clCreateCommandQueue"); + if (o.memprobe) { + int rc = runMemprobe(&dv, di, &o); + clReleaseCommandQueue(dv.q); + clReleaseContext(dv.ctx); + return rc; + } if (o.serve && !o.kernelGiven) { /* The bound kernel lives next to the compiled-in kernel.cl as kernel_bound.cl (packs from igneum-pow or igneum-miner export-pack). */ static char boundPath[1024]; @@ -1440,14 +1894,7 @@ int main(int argc, char** argv) { printf("build options: %s\n", dv.buildOptions); printf("exchange: %s\n", dv.exchangeNote); if (o.serve) return runServe(&dv, di, &o); - { - size_t wg = 0; - cl_ulong lmem = 0; - clGetKernelWorkGroupInfo(dv.kHash, di->device, CL_KERNEL_WORK_GROUP_SIZE, sizeof(wg), &wg, NULL); - clGetKernelWorkGroupInfo(dv.kHash, di->device, CL_KERNEL_LOCAL_MEM_SIZE, sizeof(lmem), &lmem, NULL); - printf("kernel: igneum_hash max work-group %llu, local memory %llu bytes, work-group %d x 32\n", - (unsigned long long)wg, (unsigned long long)lmem, o.groupWarps); - } + printKernelInfo(di, dv.kHash, "igneum_hash", dv.groupSize, ""); printf("program: %d instructions x %d iterations, loads/hash %d, op mix %s\n", IGNEUM_INSTR_COUNT, IGNEUM_ITERATIONS, IGNEUM_LOADS_PER_HASH, IGNEUM_OP_MIX); printf("seed words: %08x %08x %08x %08x %08x %08x %08x %08x\n", diff --git a/proto-opencl/test-host.sh b/proto-opencl/test-host.sh new file mode 100755 index 000000000..1c826833e --- /dev/null +++ b/proto-opencl/test-host.sh @@ -0,0 +1,10 @@ +#!/usr/bin/env bash +# Builds and runs test_host.c (the device-free rules of host.c) against the placeholder pack. No GPU is touched. +set -euo pipefail +cd "$(dirname "$0")" +P="../proto-cuda/packs/igneum-devnet-v4-epoch0" +case "$(uname -s)" in + Darwin) cc -std=c99 -O1 -Wall -Wextra -Wno-deprecated-declarations -Wno-unused-function -I "$P" -DIGNEUM_KERNEL_PATH='"kernel_bound.cl"' -o test_host test_host.c -framework OpenCL ;; + *) cc -std=c99 -O1 -Wall -Wextra -Wno-unused-function -I "$P" -DIGNEUM_KERNEL_PATH='"kernel_bound.cl"' -o test_host test_host.c -lOpenCL -ldl ;; +esac +./test_host diff --git a/proto-opencl/test_host.c b/proto-opencl/test_host.c new file mode 100644 index 000000000..9b271b3d3 --- /dev/null +++ b/proto-opencl/test_host.c @@ -0,0 +1,71 @@ +/* Unit tests for the host-side rules of host.c that need no device (5 October 2026): the duplicate-platform fold + * (markDuplicates, driverNumber). host.c is included with its main renamed. Build and run: ./test-host.sh */ +#define main host_main +#include "host.c" +#undef main + +static int failures = 0; +#define CHECK(cond, what) do { if (!(cond)) { printf("FAIL: %s (line %d)\n", what, __LINE__); ++failures; } else printf("ok: %s\n", what); } while (0) + +static DeviceInfo dev(const char* platform, const char* version, const char* name, const char* vendor, const char* driver, cl_ulong mem, cl_uint cus) { + DeviceInfo d; + memset(&d, 0, sizeof(d)); + strncpy(d.platformName, platform, 255); strncpy(d.platformVersion, version, 255); + strncpy(d.name, name, 255); strncpy(d.vendor, vendor, 255); strncpy(d.driver, driver, 255); + d.globalMem = mem; d.computeUnits = cus; d.type = CL_DEVICE_TYPE_GPU; d.dupOf = -1; + return d; +} + +int main(void) { + const char* AMD = "AMD Accelerated Parallel Processing"; + const char* V_NEW = "OpenCL 2.1 AMD-APP (3683.0)", *V_OLD = "OpenCL 2.1 AMD-APP (3652.0)", *V_OLDER = "OpenCL 2.1 AMD-APP (3600.0)"; + const char* AMDV = "Advanced Micro Devices, Inc."; + CHECK(driverNumber("3683.0 (PAL,LC)") == 3683.0, "driverNumber reads the AMD string"); + CHECK(driverNumber("617.14") > 617.1 && driverNumber("617.14") < 617.2, "driverNumber reads the NVIDIA string"); + CHECK(driverNumber("") == 0.0, "driverNumber of an empty string is 0"); + { + /* PC 1 on 5 October 2026: the new platform first, the old one second, NVIDIA last */ + DeviceInfo l[5]; + l[0] = dev(AMD, V_NEW, "gfx1036", AMDV, "3683.0 (PAL,LC)", 59589ull << 20, 1); + l[1] = dev(AMD, V_NEW, "gfx1201", AMDV, "3683.0 (PAL,LC)", 16304ull << 20, 32); + l[2] = dev(AMD, V_OLD, "gfx1036", AMDV, "3652.0 (PAL,LC)", 59589ull << 20, 1); + l[3] = dev(AMD, V_OLD, "gfx1201", AMDV, "3652.0 (PAL,LC)", 16304ull << 20, 32); + l[4] = dev("NVIDIA CUDA", "OpenCL 3.0 CUDA 13.4.96", "NVIDIA GeForce RTX 5090", "NVIDIA Corporation", "617.14", 32579ull << 20, 170); + CHECK(markDuplicates(l, 5) == 2, "PC 1: two duplicates marked"); + CHECK(l[0].dupOf == -1 && l[1].dupOf == -1 && l[4].dupOf == -1, "PC 1: the new platform's cards and the NVIDIA card stay"); + CHECK(l[2].dupOf == 0 && l[3].dupOf == 1, "PC 1: the old platform's entries point at the new ones"); + } + { + /* the old platform enumerated first: the newer entry must still be the one kept */ + DeviceInfo l[2]; + l[0] = dev(AMD, V_OLD, "gfx1201", AMDV, "3652.0 (PAL,LC)", 16304ull << 20, 32); + l[1] = dev(AMD, V_NEW, "gfx1201", AMDV, "3683.0 (PAL,LC)", 16304ull << 20, 32); + CHECK(markDuplicates(l, 2) == 1, "old first: one duplicate"); + CHECK(l[0].dupOf == 1 && l[1].dupOf == -1, "old first: the older entry is the duplicate"); + } + { + /* two real cards of one model on ONE platform: never folded */ + DeviceInfo l[2]; + l[0] = dev(AMD, V_NEW, "gfx1201", AMDV, "3683.0 (PAL,LC)", 16304ull << 20, 32); + l[1] = dev(AMD, V_NEW, "gfx1201", AMDV, "3683.0 (PAL,LC)", 16304ull << 20, 32); + CHECK(markDuplicates(l, 2) == 0 && l[0].dupOf == -1 && l[1].dupOf == -1, "two identical cards on one platform stay two cards"); + } + { + /* three registrations of one card: one stays */ + DeviceInfo l[3]; + l[0] = dev(AMD, V_OLD, "gfx1201", AMDV, "3652.0 (PAL,LC)", 16304ull << 20, 32); + l[1] = dev(AMD, V_OLDER, "gfx1201", AMDV, "3600.0 (PAL,LC)", 16304ull << 20, 32); + l[2] = dev(AMD, V_NEW, "gfx1201", AMDV, "3683.0 (PAL,LC)", 16304ull << 20, 32); + CHECK(markDuplicates(l, 3) == 2, "three platforms: two duplicates"); + CHECK(l[2].dupOf == -1 && l[0].dupOf == 2 && l[1].dupOf == 2, "three platforms: the newest stays and both duplicates point at it"); + } + { + /* a different card with the same name but other memory (an 8 GB and a 16 GB model) on two platforms: not folded */ + DeviceInfo l[2]; + l[0] = dev(AMD, V_NEW, "gfx1201", AMDV, "3683.0 (PAL,LC)", 16304ull << 20, 32); + l[1] = dev(AMD, V_OLD, "gfx1201", AMDV, "3652.0 (PAL,LC)", 8000ull << 20, 32); + CHECK(markDuplicates(l, 2) == 0, "different memory sizes are different cards"); + } + printf("%s: %d failure(s)\n", failures ? "FAIL" : "PASS", failures); + return failures ? 1 : 0; +} From b970ddade49990e0beee137a65f5aa8d78557099 Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:58:20 +0100 Subject: [PATCH 002/311] read-width: scratch per warp is a class parameter (32 or 128 KiB, under the 6 GB working-set cap), distinct-address rule bounds dataset loads only; Metal pack harness; OpenCL --bench-pack, scratch args and 16-byte probe; packfile class fields; OpenCL emulator persistent launch Co-Authored-By: Claude Fable 5.1 --- igneum-pow/src/accept.rs | 20 +- igneum-pow/src/emit.rs | 84 ++-- igneum-pow/src/generator.rs | 75 ++-- igneum-pow/src/verify.rs | 36 +- proto-cuda/nvrtc/packfile.h | 9 + .../{scr0 => scr0k32}/kernel.cl | 4 +- .../{scr0 => scr0k32}/kernel.cu | 4 +- .../{scr0 => scr0k32}/kernel_bound.cl | 6 +- .../{scr0 => scr0k32}/kernel_bound.cu | 4 +- .../{scr0 => scr0k32}/memhard.h | 0 .../{scr0 => scr0k32}/memhard.metal | 0 .../{scr0 => scr0k32}/program.h | 12 +- .../{scr0 => scr0k32}/program.json | 7 +- .../{scr0 => scr0k32}/program.metal | 4 +- .../{scr0 => scr0k32}/program_bound.metal | 4 +- .../{scr0 => scr0k32}/vectors.h | 0 .../{scr0 => scr0k32}/vectors.json | 0 .../{scr2 => scr2k128}/kernel.cl | 8 +- .../{scr2 => scr2k128}/kernel.cu | 8 +- .../{scr2 => scr2k128}/kernel_bound.cl | 14 +- .../{scr2 => scr2k128}/kernel_bound.cu | 8 +- .../{scr2 => scr2k128}/memhard.h | 0 .../{scr2 => scr2k128}/memhard.metal | 0 proto-cuda/packs-readwidth/scr2k128/program.h | 67 +++ .../{scr2 => scr2k128}/program.json | 7 +- .../{scr2 => scr2k128}/program.metal | 8 +- .../{scr2 => scr2k128}/program_bound.metal | 8 +- .../{scr4 => scr2k128}/vectors.h | 24 +- .../{scr8 => scr2k128}/vectors.json | 24 +- .../{scr4 => scr2k32}/kernel.cl | 12 +- .../{scr4 => scr2k32}/kernel.cu | 12 +- .../{scr4 => scr2k32}/kernel_bound.cl | 22 +- .../{scr4 => scr2k32}/kernel_bound.cu | 12 +- .../{scr4 => scr2k32}/memhard.h | 0 .../{scr4 => scr2k32}/memhard.metal | 0 .../{scr2 => scr2k32}/program.h | 12 +- .../packs-readwidth/scr2k32/program.json | 130 ++++++ .../{scr4 => scr2k32}/program.metal | 12 +- .../{scr4 => scr2k32}/program_bound.metal | 12 +- .../{scr8 => scr2k32}/vectors.h | 24 +- .../{scr2 => scr2k32}/vectors.json | 24 +- .../{scr8 => scr4k128}/kernel.cl | 20 +- .../{scr8 => scr4k128}/kernel.cu | 20 +- .../{scr8 => scr4k128}/kernel_bound.cl | 38 +- .../{scr8 => scr4k128}/kernel_bound.cu | 20 +- .../{scr8 => scr4k128}/memhard.h | 0 .../{scr8 => scr4k128}/memhard.metal | 0 proto-cuda/packs-readwidth/scr4k128/program.h | 67 +++ .../{scr4 => scr4k128}/program.json | 7 +- .../packs-readwidth/scr4k128/program.metal | 126 ++++++ .../scr4k128/program_bound.metal | 128 ++++++ .../{scr2 => scr4k128}/vectors.h | 24 +- .../{scr4 => scr4k128}/vectors.json | 24 +- proto-cuda/packs-readwidth/scr4k32/kernel.cl | 291 +++++++++++++ proto-cuda/packs-readwidth/scr4k32/kernel.cu | 177 ++++++++ .../packs-readwidth/scr4k32/kernel_bound.cl | 393 ++++++++++++++++++ .../packs-readwidth/scr4k32/kernel_bound.cu | 136 ++++++ proto-cuda/packs-readwidth/scr4k32/memhard.h | 108 +++++ .../packs-readwidth/scr4k32/memhard.metal | 106 +++++ .../{scr4 => scr4k32}/program.h | 12 +- .../packs-readwidth/scr4k32/program.json | 130 ++++++ .../packs-readwidth/scr4k32/program.metal | 126 ++++++ .../scr4k32/program_bound.metal | 128 ++++++ proto-cuda/packs-readwidth/scr4k32/vectors.h | 57 +++ .../packs-readwidth/scr4k32/vectors.json | 36 ++ proto-cuda/packs-readwidth/scr8k128/kernel.cl | 291 +++++++++++++ proto-cuda/packs-readwidth/scr8k128/kernel.cu | 177 ++++++++ .../packs-readwidth/scr8k128/kernel_bound.cl | 393 ++++++++++++++++++ .../packs-readwidth/scr8k128/kernel_bound.cu | 136 ++++++ proto-cuda/packs-readwidth/scr8k128/memhard.h | 108 +++++ .../packs-readwidth/scr8k128/memhard.metal | 106 +++++ proto-cuda/packs-readwidth/scr8k128/program.h | 67 +++ .../{scr8 => scr8k128}/program.json | 7 +- .../{scr8 => scr8k128}/program.metal | 20 +- .../{scr8 => scr8k128}/program_bound.metal | 20 +- proto-cuda/packs-readwidth/scr8k128/vectors.h | 57 +++ .../packs-readwidth/scr8k128/vectors.json | 36 ++ proto-cuda/packs-readwidth/scr8k32/kernel.cl | 291 +++++++++++++ proto-cuda/packs-readwidth/scr8k32/kernel.cu | 177 ++++++++ .../packs-readwidth/scr8k32/kernel_bound.cl | 393 ++++++++++++++++++ .../packs-readwidth/scr8k32/kernel_bound.cu | 136 ++++++ proto-cuda/packs-readwidth/scr8k32/memhard.h | 108 +++++ .../packs-readwidth/scr8k32/memhard.metal | 106 +++++ .../{scr8 => scr8k32}/program.h | 12 +- .../packs-readwidth/scr8k32/program.json | 130 ++++++ .../packs-readwidth/scr8k32/program.metal | 126 ++++++ .../scr8k32/program_bound.metal | 128 ++++++ proto-cuda/packs-readwidth/scr8k32/vectors.h | 57 +++ .../packs-readwidth/scr8k32/vectors.json | 36 ++ proto-metal/packbench.swift | 209 ++++++++++ proto-opencl/emu/emu_main.cpp | 34 +- proto-opencl/host.c | 146 ++++++- 92 files changed, 6031 insertions(+), 367 deletions(-) rename proto-cuda/packs-readwidth/{scr0 => scr0k32}/kernel.cl (99%) rename proto-cuda/packs-readwidth/{scr0 => scr0k32}/kernel.cu (98%) rename proto-cuda/packs-readwidth/{scr0 => scr0k32}/kernel_bound.cl (99%) rename proto-cuda/packs-readwidth/{scr0 => scr0k32}/kernel_bound.cu (98%) rename proto-cuda/packs-readwidth/{scr0 => scr0k32}/memhard.h (100%) rename proto-cuda/packs-readwidth/{scr0 => scr0k32}/memhard.metal (100%) rename proto-cuda/packs-readwidth/{scr0 => scr0k32}/program.h (91%) rename proto-cuda/packs-readwidth/{scr0 => scr0k32}/program.json (96%) rename proto-cuda/packs-readwidth/{scr0 => scr0k32}/program.metal (99%) rename proto-cuda/packs-readwidth/{scr0 => scr0k32}/program_bound.metal (99%) rename proto-cuda/packs-readwidth/{scr0 => scr0k32}/vectors.h (100%) rename proto-cuda/packs-readwidth/{scr0 => scr0k32}/vectors.json (100%) rename proto-cuda/packs-readwidth/{scr2 => scr2k128}/kernel.cl (93%) rename proto-cuda/packs-readwidth/{scr2 => scr2k128}/kernel.cu (88%) rename proto-cuda/packs-readwidth/{scr2 => scr2k128}/kernel_bound.cl (90%) rename proto-cuda/packs-readwidth/{scr2 => scr2k128}/kernel_bound.cu (85%) rename proto-cuda/packs-readwidth/{scr2 => scr2k128}/memhard.h (100%) rename proto-cuda/packs-readwidth/{scr2 => scr2k128}/memhard.metal (100%) create mode 100644 proto-cuda/packs-readwidth/scr2k128/program.h rename proto-cuda/packs-readwidth/{scr2 => scr2k128}/program.json (98%) rename proto-cuda/packs-readwidth/{scr2 => scr2k128}/program.metal (83%) rename proto-cuda/packs-readwidth/{scr2 => scr2k128}/program_bound.metal (84%) rename proto-cuda/packs-readwidth/{scr4 => scr2k128}/vectors.h (60%) rename proto-cuda/packs-readwidth/{scr8 => scr2k128}/vectors.json (65%) rename proto-cuda/packs-readwidth/{scr4 => scr2k32}/kernel.cl (88%) rename proto-cuda/packs-readwidth/{scr4 => scr2k32}/kernel.cu (79%) rename proto-cuda/packs-readwidth/{scr4 => scr2k32}/kernel_bound.cl (83%) rename proto-cuda/packs-readwidth/{scr4 => scr2k32}/kernel_bound.cu (75%) rename proto-cuda/packs-readwidth/{scr4 => scr2k32}/memhard.h (100%) rename proto-cuda/packs-readwidth/{scr4 => scr2k32}/memhard.metal (100%) rename proto-cuda/packs-readwidth/{scr2 => scr2k32}/program.h (91%) create mode 100644 proto-cuda/packs-readwidth/scr2k32/program.json rename proto-cuda/packs-readwidth/{scr4 => scr2k32}/program.metal (72%) rename proto-cuda/packs-readwidth/{scr4 => scr2k32}/program_bound.metal (73%) rename proto-cuda/packs-readwidth/{scr8 => scr2k32}/vectors.h (60%) rename proto-cuda/packs-readwidth/{scr2 => scr2k32}/vectors.json (65%) rename proto-cuda/packs-readwidth/{scr8 => scr4k128}/kernel.cl (78%) rename proto-cuda/packs-readwidth/{scr8 => scr4k128}/kernel.cu (66%) rename proto-cuda/packs-readwidth/{scr8 => scr4k128}/kernel_bound.cl (70%) rename proto-cuda/packs-readwidth/{scr8 => scr4k128}/kernel_bound.cu (60%) rename proto-cuda/packs-readwidth/{scr8 => scr4k128}/memhard.h (100%) rename proto-cuda/packs-readwidth/{scr8 => scr4k128}/memhard.metal (100%) create mode 100644 proto-cuda/packs-readwidth/scr4k128/program.h rename proto-cuda/packs-readwidth/{scr4 => scr4k128}/program.json (98%) create mode 100644 proto-cuda/packs-readwidth/scr4k128/program.metal create mode 100644 proto-cuda/packs-readwidth/scr4k128/program_bound.metal rename proto-cuda/packs-readwidth/{scr2 => scr4k128}/vectors.h (60%) rename proto-cuda/packs-readwidth/{scr4 => scr4k128}/vectors.json (65%) create mode 100644 proto-cuda/packs-readwidth/scr4k32/kernel.cl create mode 100644 proto-cuda/packs-readwidth/scr4k32/kernel.cu create mode 100644 proto-cuda/packs-readwidth/scr4k32/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/scr4k32/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/scr4k32/memhard.h create mode 100644 proto-cuda/packs-readwidth/scr4k32/memhard.metal rename proto-cuda/packs-readwidth/{scr4 => scr4k32}/program.h (91%) create mode 100644 proto-cuda/packs-readwidth/scr4k32/program.json create mode 100644 proto-cuda/packs-readwidth/scr4k32/program.metal create mode 100644 proto-cuda/packs-readwidth/scr4k32/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/scr4k32/vectors.h create mode 100644 proto-cuda/packs-readwidth/scr4k32/vectors.json create mode 100644 proto-cuda/packs-readwidth/scr8k128/kernel.cl create mode 100644 proto-cuda/packs-readwidth/scr8k128/kernel.cu create mode 100644 proto-cuda/packs-readwidth/scr8k128/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/scr8k128/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/scr8k128/memhard.h create mode 100644 proto-cuda/packs-readwidth/scr8k128/memhard.metal create mode 100644 proto-cuda/packs-readwidth/scr8k128/program.h rename proto-cuda/packs-readwidth/{scr8 => scr8k128}/program.json (98%) rename proto-cuda/packs-readwidth/{scr8 => scr8k128}/program.metal (56%) rename proto-cuda/packs-readwidth/{scr8 => scr8k128}/program_bound.metal (57%) create mode 100644 proto-cuda/packs-readwidth/scr8k128/vectors.h create mode 100644 proto-cuda/packs-readwidth/scr8k128/vectors.json create mode 100644 proto-cuda/packs-readwidth/scr8k32/kernel.cl create mode 100644 proto-cuda/packs-readwidth/scr8k32/kernel.cu create mode 100644 proto-cuda/packs-readwidth/scr8k32/kernel_bound.cl create mode 100644 proto-cuda/packs-readwidth/scr8k32/kernel_bound.cu create mode 100644 proto-cuda/packs-readwidth/scr8k32/memhard.h create mode 100644 proto-cuda/packs-readwidth/scr8k32/memhard.metal rename proto-cuda/packs-readwidth/{scr8 => scr8k32}/program.h (91%) create mode 100644 proto-cuda/packs-readwidth/scr8k32/program.json create mode 100644 proto-cuda/packs-readwidth/scr8k32/program.metal create mode 100644 proto-cuda/packs-readwidth/scr8k32/program_bound.metal create mode 100644 proto-cuda/packs-readwidth/scr8k32/vectors.h create mode 100644 proto-cuda/packs-readwidth/scr8k32/vectors.json create mode 100644 proto-metal/packbench.swift diff --git a/igneum-pow/src/accept.rs b/igneum-pow/src/accept.rs index 7f8c4438a..aca1702ea 100644 --- a/igneum-pow/src/accept.rs +++ b/igneum-pow/src/accept.rs @@ -14,7 +14,7 @@ //! costs about a millisecond on one core. The census (section 7.3) checked on 100,000 programs that the //! closed-form verdict agrees with the memory-hard one on all but 39 threshold-edge cases. -use crate::generator::{Instr, Op, Program, INSTR_COUNT, ITERATIONS, LANES, SCRATCH_SLOT_MASK}; +use crate::generator::{Instr, Op, Program, INSTR_COUNT, ITERATIONS, LANES}; use crate::seed::{fnv1a64, SplitMix64}; use crate::verify::{dataset_elem, fold_words, splitmix32, ScratchModel}; @@ -33,8 +33,10 @@ pub const BIAS_TOLERANCE: u32 = 136; /// Distinct addresses per lane per evaluation, summed over 2,048 evaluations, must exceed this (mean above 120). pub const MIN_DISTINCT_SUM: u64 = 245_760; -/// The distinct-address bound for a program with `loads` loads per hash: the same 120 of 128 ratio, so +/// The distinct-address bound for a program with `loads` dataset loads per hash: the same 120 of 128 ratio, so /// [`MIN_DISTINCT_SUM`] for the lottery hash and `loads x 1,920` for the read-width classes with other counts. +/// Variant 5's scratch read-modify-writes are not dataset loads: their slots repeat by design (a later +/// read-modify-write sees an earlier write), so they are neither counted nor bounded here. pub fn min_distinct_sum(loads: usize) -> u64 { loads as u64 * ACCEPT_HASHES as u64 * 120 / 128 } @@ -73,7 +75,7 @@ impl std::fmt::Display for Reject { Reject::Saturated { count } => write!(f, "(c) {count} of 16384 final register values saturated (limit 163)"), Reject::OutputBias { bit, ones } => write!(f, "(c) output bit {bit} set in {ones} of 2048 hashes"), Reject::DistinctAddresses { sum } => { - write!(f, "(c) distinct addresses {sum} over 2048 hashes (mean {:.2}, needs above 120 of 128 loads)", *sum as f64 / 2048.0) + write!(f, "(c) distinct dataset addresses {sum} over 2048 hashes (mean {:.2}, needs above 120 of 128 of the dataset loads)", *sum as f64 / 2048.0) } } } @@ -189,7 +191,8 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut } let mut idx = [0u32; LANES]; let mut nload = 0usize; - let mut scratch = if p.has_scratch() { Some(ScratchModel::new()) } else { None }; + let mut scratch = if p.has_scratch() { Some(ScratchModel::new(p.class.scratch_slots_per_lane())) } else { None }; + let slot_mask = p.class.scratch_slot_mask(); for it in 0..ITERATIONS { let sel = r[0]; for (k, ins) in p.instrs.iter().enumerate() { @@ -200,7 +203,7 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut // Variant 5: the slot stands in for the address (bit 31 set so it never aliases a dataset word). let m = scratch.as_mut().expect("a scratch op needs a scratch class"); for lane in 0..LANES { - idx[lane] = r[a][lane] & SCRATCH_SLOT_MASK; + idx[lane] = r[a][lane] & slot_mask; } if idx.iter().all(|&x| x == idx[0]) { return Err(Reject::LaneConstantSite { iteration: it as u8, instr: k as u8, unit: unit as u8 }); @@ -332,7 +335,8 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut sl.sort_unstable(); let mut distinct = 0u64; for k in 0..loads { - if k == 0 || sl[k] != sl[k - 1] { + // scratch slots carry bit 31 (variant 5) and are not dataset addresses + if sl[k] & 0x8000_0000 == 0 && (k == 0 || sl[k] != sl[k - 1]) { distinct += 1; } } @@ -367,7 +371,7 @@ pub fn check_dynamic(p: &Program) -> Result { } bias_max = bias_max.max(d); } - if acc.distinct_sum <= min_distinct_sum(loads) { + if acc.distinct_sum <= min_distinct_sum(loads - p.scratch_ops_per_hash()) { return Err(Reject::DistinctAddresses { sum: acc.distinct_sum }); } Ok(AcceptReport { distinct_sum: acc.distinct_sum, saturated: acc.saturated, bias_max }) @@ -395,7 +399,7 @@ mod tests { /// with `verify.rs` on every class (the fold is shared, the addresses are aligned the same way). #[test] fn classes_pass_and_match_verify() { - for name in ["w16", "w64", "w64x4", "50,35,15", "25,50,25", "scr2", "scr8"] { + for name in ["w16", "w64", "w64x4", "50,35,15", "25,50,25", "scr2k32", "scr8k128"] { let c = LoadClass::parse(name).unwrap(); let p = generate_class("igneum-genesis", c); assert!(check(&p).is_ok(), "{name}"); diff --git a/igneum-pow/src/emit.rs b/igneum-pow/src/emit.rs index 4ea7a9f03..77029f280 100644 --- a/igneum-pow/src/emit.rs +++ b/igneum-pow/src/emit.rs @@ -15,7 +15,6 @@ use crate::memhard::{ CACHE_TAG, CACHE_WORDS, CHACHA_ROUNDS, CHACHA_SIGMA, ITEM_ROUNDS, }; use crate::seed::SplitMix64; -use crate::generator::{SCRATCH_BYTES_PER_WARP, SCRATCH_SLOTS, SCRATCH_SLOT_MASK, SCRATCH_WORDS_PER_LANE}; use crate::verify::{DatasetMode, DatasetSource, Epoch, FOLD_MUL, FOLD_ROT}; /// Where the words of a wide load come from (read-width experiment). @@ -127,15 +126,11 @@ fn scratch_prelude(p: &Program, dialect: CoreDialect) -> String { CoreDialect::OpenCl => ("uint", "static inline"), }; let mut s = String::new(); - s.push_str(&format!("// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, {} slots of -", SCRATCH_SLOTS)); - s.push_str("// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not -"); - s.push_str("// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. -"); + s.push_str(&format!("// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a {} KiB scratch per warp, {} slots of\n", p.class.scratch_kb, p.class.scratch_slots_per_lane())); + s.push_str("// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not\n"); + s.push_str("// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill.\n"); s.push_str(&format!( - "{fn_} {u} scr_fill({u} gbase, {u} lane, {u} slot, {u} j) {{ {u} sw = (j == 0u) ? {} : ((j == 1u) ? {} : {}); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); }} -", + "{fn_} {u} scr_fill({u} gbase, {u} lane, {u} slot, {u} j) {{ {u} sw = (j == 0u) ? {} : ((j == 1u) ? {} : {}); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); }}\n", hex(p.seed[0]), hex(p.seed[1]), hex(p.seed[2]) @@ -145,15 +140,14 @@ fn scratch_prelude(p: &Program, dialect: CoreDialect) -> String { /// Variant 5: one scratch read-modify-write as a statement block. `arena`, `tag`, `gbase` and `lane` are in scope /// (the persistent prologue). Reads 16 bytes, folds the three data words into dst, rewrites the slot behind the tag. -fn scratch_stmt(dialect: CoreDialect, d: &str, a: &str) -> String { +fn scratch_stmt(dialect: CoreDialect, d: &str, a: &str, slot_mask: u32) -> String { let (u, load, store) = match dialect { CoreDialect::Metal => ("uint", "uint4 v_ = *(device const uint4*)(arena + s_ * 4u);", "*(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_);"), CoreDialect::Cuda => ("uint32_t", "uint4 v_ = *(const uint4*)(arena + s_ * 4u);", "*(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_);"), CoreDialect::OpenCl => ("uint", "uint4 v_ = vload4(s_, arena);", "vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena);"), }; format!( - "{{ {u} s_ = {a} & {}u; {load} {u} m_ = (v_.x == tag) ? 0xffffffffu : 0u; {u} w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); {u} w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); {u} w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); {u} x_ = {d} ^ w0_; x_ = (rotl_imm(x_, {FOLD_ROT}u) * {k}) ^ w1_; x_ = (rotl_imm(x_, {FOLD_ROT}u) * {k}) ^ w2_; {d} = x_; {store} }}", - SCRATCH_SLOT_MASK, + "{{ {u} s_ = {a} & {slot_mask}u; {load} {u} m_ = (v_.x == tag) ? 0xffffffffu : 0u; {u} w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); {u} w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); {u} w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); {u} x_ = {d} ^ w0_; x_ = (rotl_imm(x_, {FOLD_ROT}u) * {k}) ^ w1_; x_ = (rotl_imm(x_, {FOLD_ROT}u) * {k}) ^ w2_; {d} = x_; {store} }}", k = hex(FOLD_MUL) ) } @@ -162,29 +156,21 @@ fn scratch_stmt(dialect: CoreDialect, d: &str, a: &str) -> String { /// choice); warp `w` owns arena `w` and runs the units `w, w + N, w + 2N, ...` of the launch. Inside the loop the /// lottery hash's text is unchanged: `gid` is the unit's first output index plus the lane. The host MUST launch /// `groups` as a multiple of N (a uniform trip count: the OpenCL local-memory exchange carries a barrier). -fn persistent_prologue(dialect: CoreDialect) -> String { +fn persistent_prologue(dialect: CoreDialect, words_per_lane: usize) -> String { let (u, tid, nthreads, ptr) = match dialect { CoreDialect::Metal => ("uint", "tid", "nthreads", "device uint*"), CoreDialect::Cuda => ("uint32_t", "(blockIdx.x * blockDim.x + threadIdx.x)", "(gridDim.x * blockDim.x)", "uint32_t*"), CoreDialect::OpenCl => ("uint", "(uint)get_global_id(0)", "(uint)get_global_size(0)", "__global uint*"), }; let mut s = String::new(); - s.push_str(&format!(" {u} lane = {tid} & 31u; -")); - s.push_str(&format!(" {u} warp_ = {tid} >> 5; -")); - s.push_str(&format!(" {u} nwarps_ = {nthreads} >> 5; -")); - s.push_str(&format!(" {ptr} arena = scratch + ((size_t)warp_ * 32u + lane) * {}u; -", SCRATCH_WORDS_PER_LANE)); - s.push_str(&format!(" for ({u} g_ = warp_; g_ < groups; g_ += nwarps_) {{ -")); - s.push_str(&format!(" {u} gid = g_ * 32u + lane; -")); - s.push_str(&format!(" {u} gbase = baseNonce + g_ * 32u; -")); - s.push_str(&format!(" {u} tag = salt + g_; -")); + s.push_str(&format!(" {u} lane = {tid} & 31u;\n")); + s.push_str(&format!(" {u} warp_ = {tid} >> 5;\n")); + s.push_str(&format!(" {u} nwarps_ = {nthreads} >> 5;\n")); + s.push_str(&format!(" {ptr} arena = scratch + ((size_t)warp_ * 32u + lane) * {words_per_lane}u;\n")); + s.push_str(&format!(" for ({u} g_ = warp_; g_ < groups; g_ += nwarps_) {{\n")); + s.push_str(&format!(" {u} gid = g_ * 32u + lane;\n")); + s.push_str(&format!(" {u} gbase = baseNonce + g_ * 32u;\n")); + s.push_str(&format!(" {u} tag = salt + g_;\n")); s } @@ -194,20 +180,13 @@ fn scratch_header_lines(p: &Program) -> String { return String::new(); } let mut s = String::new(); - s.push_str("// Variant 5: persistent warps, a 1 MiB scratch per launched warp (the host launches N warps and passes scratch, -"); - s.push_str("// groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). -"); - s.push_str("#define IGNEUM_PERSISTENT_WARPS 1 -"); - s.push_str(&format!("#define IGNEUM_SCRATCH_OPS {} // scratch read-modify-writes per program ({} per hash) -", p.class.scratch_slots(), p.scratch_ops_per_hash())); - s.push_str(&format!("#define IGNEUM_SCRATCH_SLOTS {SCRATCH_SLOTS}u -")); - s.push_str(&format!("#define IGNEUM_SCRATCH_WORDS_PER_LANE {SCRATCH_WORDS_PER_LANE}u -")); - s.push_str(&format!("#define IGNEUM_SCRATCH_BYTES_PER_WARP {SCRATCH_BYTES_PER_WARP}u -")); + s.push_str(&format!("// Variant 5: persistent warps, a {} KiB scratch per launched warp (the host launches N warps and passes scratch,\n", p.class.scratch_kb)); + s.push_str("// groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit).\n"); + s.push_str("#define IGNEUM_PERSISTENT_WARPS 1\n"); + s.push_str(&format!("#define IGNEUM_SCRATCH_OPS {} // scratch read-modify-writes per program ({} per hash)\n", p.class.scratch_slots(), p.scratch_ops_per_hash())); + s.push_str(&format!("#define IGNEUM_SCRATCH_SLOTS {}u\n", p.class.scratch_slots_per_lane())); + s.push_str(&format!("#define IGNEUM_SCRATCH_WORDS_PER_LANE {}u\n", p.class.scratch_words_per_lane())); + s.push_str(&format!("#define IGNEUM_SCRATCH_BYTES_PER_WARP {}u\n", p.class.scratch_bytes_per_warp())); s } @@ -472,7 +451,7 @@ fn metal_program_impl(p: &Program, dataset_log2: u32, source: LoadSource, bound: s.push_str(&format!(" constant uint& salt [[buffer({})]],\n", b + 2)); s.push_str(" uint tid [[thread_position_in_grid]],\n"); s.push_str(" uint nthreads [[threads_per_grid]]) {\n"); - s.push_str(&persistent_prologue(CoreDialect::Metal)); + s.push_str(&persistent_prologue(CoreDialect::Metal, p.class.scratch_words_per_lane())); } else { s.push_str(" uint gid [[thread_position_in_grid]]) {\n"); } @@ -534,7 +513,7 @@ fn metal_program_impl(p: &Program, dataset_log2: u32, source: LoadSource, bound: } Op::Load => format!("{d} = {d} ^ {};", fetch(word_index(&a, false))), Op::WLoad => format!("{d} = {d} ^ {};", fetch(word_index(&a, true))), - Op::Scratch => scratch_stmt(CoreDialect::Metal, &d, &a), + Op::Scratch => scratch_stmt(CoreDialect::Metal, &d, &a, p.class.scratch_slot_mask()), }; s.push_str(&format!(" {line} // {k}\n")); } @@ -599,7 +578,7 @@ fn cuda_instr_lines(p: &Program) -> String { Op::Load if load_width(ins) > 1 => wide_load_stmt(CoreDialect::Cuda, &d, &a, ins.width, WideSource::Stored, None), Op::Load => format!("{d} = {d} ^ ds[{a} & mask];"), Op::WLoad => format!("{d} = {d} ^ ds[(__shfl_sync(0xffffffffu, {a}, 0) & wmask) + lane];"), - Op::Scratch => scratch_stmt(CoreDialect::Cuda, &d, &a), + Op::Scratch => scratch_stmt(CoreDialect::Cuda, &d, &a, p.class.scratch_slot_mask()), }; s.push_str(&format!(" {line} // {k} {}\n", ins.op.name())); } @@ -671,7 +650,7 @@ pub fn cuda_kernel(p: &Program, memhard: Option<&MixParams>) -> String { let scratch_args = if p.has_scratch() { ", uint32_t* scratch, uint32_t groups, uint32_t salt" } else { "" }; s.push_str(&format!("__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask{scratch_args}) {{\n")); if p.has_scratch() { - s.push_str(&persistent_prologue(CoreDialect::Cuda)); + s.push_str(&persistent_prologue(CoreDialect::Cuda, p.class.scratch_words_per_lane())); } else { s.push_str(" uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;\n"); } @@ -792,7 +771,7 @@ pub fn cuda_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String { let scratch_args = if p.has_scratch() { ", uint32_t* scratch, uint32_t groups, uint32_t salt" } else { "" }; s.push_str(&format!("__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw{scratch_args}) {{\n")); if p.has_scratch() { - s.push_str(&persistent_prologue(CoreDialect::Cuda)); + s.push_str(&persistent_prologue(CoreDialect::Cuda, p.class.scratch_words_per_lane())); } else { s.push_str(" uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;\n"); } @@ -878,7 +857,7 @@ fn opencl_instr_lines(p: &Program) -> String { Op::Load if load_width(ins) > 1 => wide_load_stmt(CoreDialect::OpenCl, &d, &a, ins.width, WideSource::Stored, None), Op::Load => format!("{d} = {d} ^ ds[{a} & mask];"), Op::WLoad => format!("{{ uint t_; IGNEUM_BCAST0(t_, {a}); {d} = {d} ^ ds[(t_ & wmask) + lane]; }}"), - Op::Scratch => scratch_stmt(CoreDialect::OpenCl, &d, &a), + Op::Scratch => scratch_stmt(CoreDialect::OpenCl, &d, &a, p.class.scratch_slot_mask()), }; s.push_str(&format!(" {line} // {k} {}\n", ins.op.name())); } @@ -897,7 +876,7 @@ pub fn opencl_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String { let scratch_args = if p.has_scratch() { ", __global uint* scratch, uint groups, uint salt" } else { "" }; s.push_str(&format!("IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw{scratch_args}) {{\n")); if p.has_scratch() { - s.push_str(&persistent_prologue(CoreDialect::OpenCl)); + s.push_str(&persistent_prologue(CoreDialect::OpenCl, p.class.scratch_words_per_lane())); } else { s.push_str(" uint gid = (uint)get_global_id(0);\n"); } @@ -1029,7 +1008,7 @@ pub fn opencl_kernel(p: &Program, memhard: Option<&MixParams>) -> String { let scratch_args = if p.has_scratch() { ", __global uint* scratch, uint groups, uint salt" } else { "" }; s.push_str(&format!("IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask{scratch_args}) {{\n")); if p.has_scratch() { - s.push_str(&persistent_prologue(CoreDialect::OpenCl)); + s.push_str(&persistent_prologue(CoreDialect::OpenCl, p.class.scratch_words_per_lane())); } else { s.push_str(" uint gid = (uint)get_global_id(0);\n"); } @@ -1303,7 +1282,8 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String { s.push_str(&format!(" \"bytes_per_hash\": {},\n", p.bytes_per_hash())); if p.has_scratch() { s.push_str(&format!(" \"scratch_ops_per_hash\": {},\n", p.scratch_ops_per_hash())); - s.push_str(&format!(" \"scratch\": \"variant 5 (measurement only): persistent warps; a 1 MiB scratch per warp of {SCRATCH_SLOTS} 16-byte slots per lane (lane-major); slot = src & 0x{SCRATCH_SLOT_MASK:x}; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)\",\n")); + s.push_str(&format!(" \"scratch_kib_per_warp\": {},\n", p.class.scratch_kb)); + s.push_str(&format!(" \"scratch\": \"variant 5 (measurement only): persistent warps; a {kb} KiB scratch per warp of {slots} 16-byte slots per lane (lane-major); slot = src & 0x{smask:x}; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)\",\n", kb = p.class.scratch_kb, slots = p.class.scratch_slots_per_lane(), smask = p.class.scratch_slot_mask())); } s.push_str(&format!(" \"wide_load\": \"read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, {FOLD_ROT}) * 0x{FOLD_MUL:08x}) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots\",\n")); } diff --git a/igneum-pow/src/generator.rs b/igneum-pow/src/generator.rs index 0092818e9..bcd8bec25 100644 --- a/igneum-pow/src/generator.rs +++ b/igneum-pow/src/generator.rs @@ -170,50 +170,70 @@ pub const WIDTH_WORDS: [u8; 3] = [1, 4, 16]; pub struct LoadClass { pub mix: [u8; 3], pub load_slots: u8, - /// Variant 5: `Some(k)` gives the program a 1 MiB per-warp scratch (the kernels run persistent warps) and - /// turns `k` of the load slots into scratch read-modify-writes. `None` for every other class. + /// Variant 5: `Some(k)` gives the program a per-warp scratch (the kernels run persistent warps) and turns `k` + /// of the load slots into scratch read-modify-writes. `None` for every other class. pub scratch: Option, + /// Variant 5: the scratch per warp in KiB (32 or 128; the whole working set of a card at full occupancy must + /// stay under 6 GB, coordinator's cap of 5 October 2026). 0 for every other class. + pub scratch_kb: u8, } -/// Scratch geometry (variant 5): 2^11 slots of 16 bytes per lane (32 KiB), 32 lanes per warp (1 MiB), lane-major. -pub const SCRATCH_SLOT_BITS: u32 = 11; -pub const SCRATCH_SLOTS: usize = 1 << SCRATCH_SLOT_BITS; -pub const SCRATCH_SLOT_MASK: u32 = SCRATCH_SLOTS as u32 - 1; -pub const SCRATCH_WORDS_PER_LANE: usize = SCRATCH_SLOTS * 4; -pub const SCRATCH_BYTES_PER_WARP: usize = SCRATCH_WORDS_PER_LANE * 4 * LANES; +/// Scratch geometry (variant 5): 16-byte slots, lane-major, 32 lanes per warp; `scratch_kb` KiB per warp gives +/// `scratch_kb x 2` slots per lane (32 KiB: 64 slots, 128 KiB: 256 slots). +pub const SCRATCH_SLOT_BYTES: usize = 16; + +impl LoadClass { + /// Slots per lane of the scratch (0 without one). + pub fn scratch_slots_per_lane(&self) -> usize { + self.scratch_kb as usize * 1024 / LANES / SCRATCH_SLOT_BYTES + } + pub fn scratch_slot_mask(&self) -> u32 { + self.scratch_slots_per_lane().saturating_sub(1) as u32 + } + pub fn scratch_words_per_lane(&self) -> usize { + self.scratch_slots_per_lane() * 4 + } + pub fn scratch_bytes_per_warp(&self) -> usize { + self.scratch_kb as usize * 1024 + } +} impl LoadClass { /// Generator version 2 as adopted on 4 October 2026: 16 loads of one word. The lottery hash. - pub const V2: LoadClass = LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None }; + pub const V2: LoadClass = LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0 }; /// A fixed width (1, 4 or 16 words) with `load_slots` loads per program. pub fn fixed(width_words: u8, load_slots: u8) -> LoadClass { let mut mix = [0u8; 3]; let i = WIDTH_WORDS.iter().position(|&w| w == width_words).expect("width must be 1, 4 or 16 words"); mix[i] = 100; - LoadClass { mix, load_slots, scratch: None } + LoadClass { mix, load_slots, scratch: None, scratch_kb: 0 } } /// Per-load width drawn from `mix` (percent for 4, 16, 64 bytes), 16 loads per program. pub fn mixed(mix: [u8; 3]) -> LoadClass { assert_eq!(mix.iter().map(|&m| m as u32).sum::(), 100, "the mix must sum to 100"); - LoadClass { mix, load_slots: LOAD_SLOTS as u8, scratch: None } + LoadClass { mix, load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0 } } - /// Variant 5: version 2 widths, 16 memory operations of which `k` are scratch read-modify-writes. - pub fn scratch(k: u8) -> LoadClass { + /// Variant 5: version 2 widths, 16 memory operations of which `k` are scratch read-modify-writes into a + /// scratch of `kb` KiB per warp (a power of two, 1 to 128: at least one slot per lane, under the 6 GB cap). + pub fn scratch(k: u8, kb: u8) -> LoadClass { assert!(k as usize <= LOAD_SLOTS); - LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: Some(k) } + assert!(kb.is_power_of_two() && kb <= 128, "scratch per warp must be a power of two up to 128 KiB"); + LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: Some(k), scratch_kb: kb } } - /// Parse "p4,p16,p64" or one of the names of [`LoadClass::name`]. + /// Parse "p4,p16,p64" or one of the names of [`LoadClass::name`] ("scr4k32": 4 scratch ops, 32 KiB per warp). pub fn parse(s: &str) -> Option { - if let Some(k) = s.strip_prefix("scr") { + if let Some(rest) = s.strip_prefix("scr") { + let (k, kb) = rest.split_once('k')?; let k: u8 = k.parse().ok()?; - if k as usize > LOAD_SLOTS { + let kb: u8 = kb.parse().ok()?; + if k as usize > LOAD_SLOTS || !kb.is_power_of_two() || kb > 128 { return None; } - return Some(LoadClass::scratch(k)); + return Some(LoadClass::scratch(k, kb)); } let (mix_s, slots) = match s.split_once("x") { Some((m, n)) if !m.contains(',') => (m, n.parse::().ok()?), @@ -235,7 +255,7 @@ impl LoadClass { if slots == 0 || slots as usize >= INSTR_COUNT { return None; } - Some(LoadClass { mix, load_slots: slots, scratch: None }) + Some(LoadClass { mix, load_slots: slots, scratch: None, scratch_kb: 0 }) } /// Scratch read-modify-writes per program (0 without a scratch). @@ -253,7 +273,7 @@ impl LoadClass { return "v2".to_string(); } if let Some(k) = self.scratch { - return format!("scr{k}"); + return format!("scr{k}k{}", self.scratch_kb); } let base = match self.mix { [100, 0, 0] => "w4".to_string(), @@ -387,6 +407,7 @@ pub fn program_id_class(generator: u32, seed: &[u32; 8], attempt: u32, class: &L if let Some(k) = class.scratch { b.extend_from_slice(b"scratch/"); b.push(k); + b.push(class.scratch_kb); } fnv1a64(&b) } @@ -855,10 +876,12 @@ mod tests { assert!(ids.insert(p.program_id()), "{name}: program id collides"); } // variant 5: k scratch ops among the 16 memory operations, the rest one-word loads - for k in [0u8, 2, 4, 8] { - let c = LoadClass::parse(&format!("scr{k}")).unwrap(); - assert_eq!(c, LoadClass::scratch(k)); - assert_eq!(c.name(), format!("scr{k}")); + for (k, kb) in [(0u8, 32u8), (2, 32), (4, 128), (8, 128)] { + let c = LoadClass::parse(&format!("scr{k}k{kb}")).unwrap(); + assert_eq!(c, LoadClass::scratch(k, kb)); + assert_eq!(c.name(), format!("scr{k}k{kb}")); + assert_eq!(c.scratch_bytes_per_warp(), kb as usize * 1024); + assert_eq!(c.scratch_slots_per_lane(), kb as usize * 2); let p = candidate_class("igneum-genesis", b"igneum-genesis", 0, c); assert_eq!(p.loads_per_hash(), 128); assert_eq!(p.scratch_ops_per_hash(), 8 * k as usize); @@ -866,7 +889,9 @@ mod tests { assert!(p.instrs.iter().all(|i| i.width == 1)); assert!(ids.insert(p.program_id()), "scr{k}: program id collides"); } - assert!(!LoadClass::scratch(0).is_v2()); + assert!(!LoadClass::scratch(0, 32).is_v2()); + assert_ne!(LoadClass::scratch(4, 32).name(), LoadClass::scratch(4, 128).name()); + assert_eq!(LoadClass::parse("scr4"), None); // a class with the version 2 widths but another slot count takes the extra roll: a different stream let w4x8 = candidate_class("igneum-genesis", b"igneum-genesis", 0, LoadClass::fixed(1, 8)); assert_ne!(w4x8.instrs, v2.instrs); diff --git a/igneum-pow/src/verify.rs b/igneum-pow/src/verify.rs index 45808be94..fedfe1b91 100644 --- a/igneum-pow/src/verify.rs +++ b/igneum-pow/src/verify.rs @@ -1,9 +1,7 @@ //! The CPU reference interpreter for one 32-lane warp (`cpuWarpTraced` in the Swift) and the API the node //! calls. Dataset words come from the memory-hard cache (default) or from the closed form (old packs). -use crate::generator::{ - generate, generate_class, Instr, LoadClass, Op, Program, ITERATIONS, LANES, SCRATCH_SLOTS, SCRATCH_SLOT_MASK, -}; +use crate::generator::{generate, generate_class, Instr, LoadClass, Op, Program, ITERATIONS, LANES}; use crate::memhard::MemhardCpu; use crate::seed::day_key; @@ -45,9 +43,10 @@ pub fn scratch_rewrite(x: u32, w: &[u32; 3]) -> [u32; 3] { } /// The CPU model of one unit's scratch (variant 5): per lane, the written slots and their words. Unwritten slots -/// read as [`scratch_fill`]. A unit touches at most `scratch ops x 32` slots, so the model is small whatever the -/// nominal 1 MiB; a GPU keeps the real 1 MiB per resident warp with a per-unit tag per slot. +/// read as [`scratch_fill`]. A unit touches at most `scratch ops x 32` slots; a GPU keeps the real scratch per +/// resident warp with a per-unit tag per slot. pub struct ScratchModel { + slots: usize, written: Vec, data: Vec<[u32; 3]>, pub reads: usize, @@ -55,13 +54,19 @@ pub struct ScratchModel { } impl ScratchModel { - pub fn new() -> Self { - Self { written: vec![false; LANES * SCRATCH_SLOTS], data: vec![[0; 3]; LANES * SCRATCH_SLOTS], reads: 0, writes: 0 } + pub fn new(slots_per_lane: usize) -> Self { + Self { + slots: slots_per_lane, + written: vec![false; LANES * slots_per_lane], + data: vec![[0; 3]; LANES * slots_per_lane], + reads: 0, + writes: 0, + } } /// Read slot `slot` of `lane`, then rewrite it from the fold result `x`. Returns the three words read. #[inline] pub fn rmw(&mut self, seed: &[u32; 8], base: u32, lane: usize, slot: u32, dst: u32) -> u32 { - let i = lane * SCRATCH_SLOTS + slot as usize; + let i = lane * self.slots + slot as usize; let w = if self.written[i] { self.data[i] } else { @@ -80,12 +85,6 @@ impl ScratchModel { } } -impl Default for ScratchModel { - fn default() -> Self { - Self::new() - } -} - /// Dataset element, closed form of (day words, index). The original prototype's six-operation element. #[inline(always)] pub fn dataset_elem(i: u32, d0: u32, d1: u32) -> u32 { @@ -258,7 +257,8 @@ pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32, let mut items_derived = 0usize; let mut idx = [0u32; LANES]; let mut val = [0u32; LANES]; - let mut scratch = if program.has_scratch() { Some(ScratchModel::new()) } else { None }; + let mut scratch = if program.has_scratch() { Some(ScratchModel::new(program.class.scratch_slots_per_lane())) } else { None }; + let slot_mask = program.class.scratch_slot_mask(); for _ in 0..ITERATIONS { let sel = r[0]; for ins in &program.instrs { @@ -267,7 +267,7 @@ pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32, let m = scratch.as_mut().expect("a scratch op needs a scratch class"); let (d, a) = (ins.dst as usize, ins.src as usize); for lane in 0..LANES { - let slot = r[a][lane] & SCRATCH_SLOT_MASK; + let slot = r[a][lane] & slot_mask; r[d][lane] = m.rmw(&program.seed, base_nonce, lane, slot, r[d][lane]); } } @@ -538,14 +538,14 @@ mod tests { assert_eq!(scratch_fill(&seed, 32, 3, 100, 1), scratch_fill(&seed, 32, 3, 100, 1)); assert_ne!(scratch_fill(&seed, 32, 3, 100, 1), scratch_fill(&seed, 32, 3, 100, 2)); assert_ne!(scratch_fill(&seed, 32, 3, 100, 1), scratch_fill(&seed, 64, 3, 100, 1)); - let mut m = ScratchModel::new(); + let mut m = ScratchModel::new(256); let w = [scratch_fill(&seed, 32, 3, 100, 0), scratch_fill(&seed, 32, 3, 100, 1), scratch_fill(&seed, 32, 3, 100, 2)]; let x = m.rmw(&seed, 32, 3, 100, 0xabcd); assert_eq!(x, fold_words(0xabcd, &w)); let x2 = m.rmw(&seed, 32, 3, 100, 0xabcd); assert_eq!(x2, fold_words(0xabcd, &scratch_rewrite(x, &w))); assert_eq!(m.reads, 2); - let e = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::ClosedForm, 20, LoadClass::scratch(4)); + let e = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::ClosedForm, 20, LoadClass::scratch(4, 128)); assert_eq!(e.program.scratch_ops_per_hash(), 32); assert_eq!(e.hash_warp(0), e.hash_warp(0)); } diff --git a/proto-cuda/nvrtc/packfile.h b/proto-cuda/nvrtc/packfile.h index fa407f93c..ba1719e8e 100644 --- a/proto-cuda/nvrtc/packfile.h +++ b/proto-cuda/nvrtc/packfile.h @@ -25,6 +25,9 @@ typedef struct { uint32_t datasetLog2, cacheLog2Words, cacheSegments, datasetMode, generator; uint32_t seedw[8], keyw[8]; char seedString[600]; + // read-width experiment (5 October 2026): the load class (0 when absent), bytes per hash, variant 5's scratch + uint32_t loadsPerHash, bytesPerHash, scratchOps, persistent, scratchWordsPerLane; + char loadClass[64]; // seeds.txt (or program.h): the seeds as the worker protocol carries them char epochHex[65]; char dayHex[PF_HEX_CAP]; @@ -254,6 +257,12 @@ static int pf_load(const char* dir, PfPack* pk, char* err, size_t cap) { if (!pf_define_u32(prog, "IGNEUM_CACHE_SEGMENTS", &pk->cacheSegments)) { free(prog); return pf_fail(err, cap, "program.h has no IGNEUM_CACHE_SEGMENTS"); } if (!pf_define_u32(prog, "IGNEUM_GENERATOR", &pk->generator)) pk->generator = 1; if (pf_define_words(prog, "IGNEUM_SEEDW_INIT", pk->seedw, 8) != 8) { free(prog); return pf_fail(err, cap, "program.h has no IGNEUM_SEEDW_INIT with 8 words"); } + pk->loadsPerHash = 128; pf_define_u32(prog, "IGNEUM_LOADS_PER_HASH", &pk->loadsPerHash); + pk->bytesPerHash = pk->loadsPerHash * 4u; pf_define_u32(prog, "IGNEUM_BYTES_PER_HASH", &pk->bytesPerHash); + pk->scratchOps = 0; pf_define_u32(prog, "IGNEUM_SCRATCH_OPS", &pk->scratchOps); + pk->persistent = 0; pf_define_u32(prog, "IGNEUM_PERSISTENT_WARPS", &pk->persistent); + pk->scratchWordsPerLane = 8192; pf_define_u32(prog, "IGNEUM_SCRATCH_WORDS_PER_LANE", &pk->scratchWordsPerLane); + strcpy(pk->loadClass, "v2"); pf_define_str(prog, "IGNEUM_LOAD_CLASS", pk->loadClass, sizeof(pk->loadClass)); if (pf_define_words(prog, "IGNEUM_KEY_INIT", pk->keyw, 8) != 8) { free(prog); return pf_fail(err, cap, "program.h has no IGNEUM_KEY_INIT with 8 words"); } if (!pf_define_str(prog, "IGNEUM_SEED_STRING", pk->seedString, sizeof(pk->seedString))) strncpy(pk->seedString, "(no IGNEUM_SEED_STRING)", sizeof(pk->seedString) - 1); pf_define_str(prog, "IGNEUM_SEED_BYTES_HEX", ehex, sizeof(ehex)); diff --git a/proto-cuda/packs-readwidth/scr0/kernel.cl b/proto-cuda/packs-readwidth/scr0k32/kernel.cl similarity index 99% rename from proto-cuda/packs-readwidth/scr0/kernel.cl rename to proto-cuda/packs-readwidth/scr0k32/kernel.cl index 5ab647860..040d671a2 100644 --- a/proto-cuda/packs-readwidth/scr0/kernel.cl +++ b/proto-cuda/packs-readwidth/scr0k32/kernel.cl @@ -177,7 +177,7 @@ __kernel void igneum_build(__global uint* ds, __global const uint* cache, uint n // One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the // lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and // __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -185,7 +185,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint lane = (uint)get_global_id(0) & 31u; uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; - __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; diff --git a/proto-cuda/packs-readwidth/scr0/kernel.cu b/proto-cuda/packs-readwidth/scr0k32/kernel.cu similarity index 98% rename from proto-cuda/packs-readwidth/scr0/kernel.cu rename to proto-cuda/packs-readwidth/scr0k32/kernel.cu index f7ac2c32d..14709616c 100644 --- a/proto-cuda/packs-readwidth/scr0/kernel.cu +++ b/proto-cuda/packs-readwidth/scr0k32/kernel.cu @@ -44,7 +44,7 @@ __global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItem // One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every // __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a // 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. __device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -52,7 +52,7 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; - uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { uint32_t gid = g_ * 32u + lane; uint32_t gbase = baseNonce + g_ * 32u; diff --git a/proto-cuda/packs-readwidth/scr0/kernel_bound.cl b/proto-cuda/packs-readwidth/scr0k32/kernel_bound.cl similarity index 99% rename from proto-cuda/packs-readwidth/scr0/kernel_bound.cl rename to proto-cuda/packs-readwidth/scr0k32/kernel_bound.cl index eaba1546c..057f228d0 100644 --- a/proto-cuda/packs-readwidth/scr0/kernel_bound.cl +++ b/proto-cuda/packs-readwidth/scr0k32/kernel_bound.cl @@ -177,7 +177,7 @@ __kernel void igneum_build(__global uint* ds, __global const uint* cache, uint n // One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the // lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and // __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -185,7 +185,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint lane = (uint)get_global_id(0) & 31u; uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; - __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -295,7 +295,7 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon uint lane = (uint)get_global_id(0) & 31u; uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; - __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; diff --git a/proto-cuda/packs-readwidth/scr0/kernel_bound.cu b/proto-cuda/packs-readwidth/scr0k32/kernel_bound.cu similarity index 98% rename from proto-cuda/packs-readwidth/scr0/kernel_bound.cu rename to proto-cuda/packs-readwidth/scr0k32/kernel_bound.cu index 5f1b9aa92..47bd8b04d 100644 --- a/proto-cuda/packs-readwidth/scr0/kernel_bound.cu +++ b/proto-cuda/packs-readwidth/scr0k32/kernel_bound.cu @@ -20,7 +20,7 @@ __device__ __forceinline__ uint32_t splitmix32(uint32_t x) { __device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } __device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. __device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -28,7 +28,7 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; - uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { uint32_t gid = g_ * 32u + lane; uint32_t gbase = baseNonce + g_ * 32u; diff --git a/proto-cuda/packs-readwidth/scr0/memhard.h b/proto-cuda/packs-readwidth/scr0k32/memhard.h similarity index 100% rename from proto-cuda/packs-readwidth/scr0/memhard.h rename to proto-cuda/packs-readwidth/scr0k32/memhard.h diff --git a/proto-cuda/packs-readwidth/scr0/memhard.metal b/proto-cuda/packs-readwidth/scr0k32/memhard.metal similarity index 100% rename from proto-cuda/packs-readwidth/scr0/memhard.metal rename to proto-cuda/packs-readwidth/scr0k32/memhard.metal diff --git a/proto-cuda/packs-readwidth/scr0/program.h b/proto-cuda/packs-readwidth/scr0k32/program.h similarity index 91% rename from proto-cuda/packs-readwidth/scr0/program.h rename to proto-cuda/packs-readwidth/scr0k32/program.h index 2409c10b0..04c678b6e 100644 --- a/proto-cuda/packs-readwidth/scr0/program.h +++ b/proto-cuda/packs-readwidth/scr0k32/program.h @@ -15,7 +15,7 @@ #define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" #define IGNEUM_GENERATOR 2 #define IGNEUM_PROGRAM_ATTEMPT 0 -#define IGNEUM_PROGRAM_ID 0x2f098ae568f38029ull +#define IGNEUM_PROGRAM_ID 0xe0b70cd155c28f4bull #define IGNEUM_DAY_STRING "2026-10-03" #define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" #define IGNEUM_DAY0 0x3067619fu @@ -30,20 +30,20 @@ #define IGNEUM_OP_MIX "load=16 add=9 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 shfl=4 rotr=2 rotl=1" // Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads // the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. -#define IGNEUM_LOAD_CLASS "scr0" +#define IGNEUM_LOAD_CLASS "scr0k32" #define IGNEUM_LOAD_SLOTS 16 #define IGNEUM_LOAD_MIX { 100, 0, 0 } #define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program #define IGNEUM_BYTES_PER_HASH 512 #define IGNEUM_FOLD_ROT 11 #define IGNEUM_FOLD_MUL 0x9e3779b1u -// Variant 5: persistent warps, a 1 MiB scratch per launched warp (the host launches N warps and passes scratch, +// Variant 5: persistent warps, a 32 KiB scratch per launched warp (the host launches N warps and passes scratch, // groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). #define IGNEUM_PERSISTENT_WARPS 1 #define IGNEUM_SCRATCH_OPS 0 // scratch read-modify-writes per program (0 per hash) -#define IGNEUM_SCRATCH_SLOTS 2048u -#define IGNEUM_SCRATCH_WORDS_PER_LANE 8192u -#define IGNEUM_SCRATCH_BYTES_PER_WARP 1048576u +#define IGNEUM_SCRATCH_SLOTS 64u +#define IGNEUM_SCRATCH_WORDS_PER_LANE 256u +#define IGNEUM_SCRATCH_BYTES_PER_WARP 32768u // 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) #define IGNEUM_DATASET_MODE 1 diff --git a/proto-cuda/packs-readwidth/scr0/program.json b/proto-cuda/packs-readwidth/scr0k32/program.json similarity index 96% rename from proto-cuda/packs-readwidth/scr0/program.json rename to proto-cuda/packs-readwidth/scr0k32/program.json index ad5bdb8c0..2d28cab4a 100644 --- a/proto-cuda/packs-readwidth/scr0/program.json +++ b/proto-cuda/packs-readwidth/scr0k32/program.json @@ -2,7 +2,7 @@ "format": "igneum-program-pack-3", "generator": 2, "attempt": 0, - "program_id": "0x2f098ae568f38029", + "program_id": "0xe0b70cd155c28f4b", "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", "dataset_mode": "memory-hard", "seed": "igneum-genesis", @@ -15,13 +15,14 @@ "iterations": 8, "instruction_count": 64, "loads_per_hash": 128, - "load_class": "scr0", + "load_class": "scr0k32", "load_slots": 16, "load_mix_percent_4_16_64": [100, 0, 0], "load_width_counts_4_16_64": [16, 0, 0], "bytes_per_hash": 512, "scratch_ops_per_hash": 0, - "scratch": "variant 5 (measurement only): persistent warps; a 1 MiB scratch per warp of 2048 16-byte slots per lane (lane-major); slot = src & 0x7ff; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", + "scratch_kib_per_warp": 32, + "scratch": "variant 5 (measurement only): persistent warps; a 32 KiB scratch per warp of 64 16-byte slots per lane (lane-major); slot = src & 0x3f; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", "op_mix": {"load": 16, "add": 9, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "shfl": 4, "rotr": 2, "rotl": 1}, "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", diff --git a/proto-cuda/packs-readwidth/scr0/program.metal b/proto-cuda/packs-readwidth/scr0k32/program.metal similarity index 99% rename from proto-cuda/packs-readwidth/scr0/program.metal rename to proto-cuda/packs-readwidth/scr0k32/program.metal index 9835a937b..a64d3b5ff 100644 --- a/proto-cuda/packs-readwidth/scr0/program.metal +++ b/proto-cuda/packs-readwidth/scr0k32/program.metal @@ -21,7 +21,7 @@ inline uint ds_elem(uint i, uint d0, uint d1) { return x; } -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -36,7 +36,7 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], uint lane = tid & 31u; uint warp_ = tid >> 5; uint nwarps_ = nthreads >> 5; - device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; diff --git a/proto-cuda/packs-readwidth/scr0/program_bound.metal b/proto-cuda/packs-readwidth/scr0k32/program_bound.metal similarity index 99% rename from proto-cuda/packs-readwidth/scr0/program_bound.metal rename to proto-cuda/packs-readwidth/scr0k32/program_bound.metal index fbfa7581b..9106aa169 100644 --- a/proto-cuda/packs-readwidth/scr0/program_bound.metal +++ b/proto-cuda/packs-readwidth/scr0k32/program_bound.metal @@ -21,7 +21,7 @@ inline uint ds_elem(uint i, uint d0, uint d1) { return x; } -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -38,7 +38,7 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], uint lane = tid & 31u; uint warp_ = tid >> 5; uint nwarps_ = nthreads >> 5; - device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; diff --git a/proto-cuda/packs-readwidth/scr0/vectors.h b/proto-cuda/packs-readwidth/scr0k32/vectors.h similarity index 100% rename from proto-cuda/packs-readwidth/scr0/vectors.h rename to proto-cuda/packs-readwidth/scr0k32/vectors.h diff --git a/proto-cuda/packs-readwidth/scr0/vectors.json b/proto-cuda/packs-readwidth/scr0k32/vectors.json similarity index 100% rename from proto-cuda/packs-readwidth/scr0/vectors.json rename to proto-cuda/packs-readwidth/scr0k32/vectors.json diff --git a/proto-cuda/packs-readwidth/scr2/kernel.cl b/proto-cuda/packs-readwidth/scr2k128/kernel.cl similarity index 93% rename from proto-cuda/packs-readwidth/scr2/kernel.cl rename to proto-cuda/packs-readwidth/scr2k128/kernel.cl index b7d17259d..8fc591da5 100644 --- a/proto-cuda/packs-readwidth/scr2/kernel.cl +++ b/proto-cuda/packs-readwidth/scr2k128/kernel.cl @@ -177,7 +177,7 @@ __kernel void igneum_build(__global uint* ds, __global const uint* cache, uint n // One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the // lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and // __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -185,7 +185,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint lane = (uint)get_global_id(0) & 31u; uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; - __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -226,7 +226,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + { uint s_ = r6 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl r4 = r0 * r2 + r4; // 19 mad @@ -244,7 +244,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r3 = r3 ^ ds[r1 & mask]; // 31 load r1 = r1 ^ ds[r0 & mask]; // 32 load r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + { uint s_ = r2 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor r4 = r4 ^ ds[r0 & mask]; // 37 load diff --git a/proto-cuda/packs-readwidth/scr2/kernel.cu b/proto-cuda/packs-readwidth/scr2k128/kernel.cu similarity index 88% rename from proto-cuda/packs-readwidth/scr2/kernel.cu rename to proto-cuda/packs-readwidth/scr2k128/kernel.cu index b1b566dea..6d8bf993d 100644 --- a/proto-cuda/packs-readwidth/scr2/kernel.cu +++ b/proto-cuda/packs-readwidth/scr2k128/kernel.cu @@ -44,7 +44,7 @@ __global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItem // One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every // __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a // 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. __device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -52,7 +52,7 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; - uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { uint32_t gid = g_ * 32u + lane; uint32_t gbase = baseNonce + g_ * 32u; @@ -86,7 +86,7 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint32_t s_ = r6 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + { uint32_t s_ = r6 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl r4 = r0 * r2 + r4; // 19 mad @@ -104,7 +104,7 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc r3 = r3 ^ ds[r1 & mask]; // 31 load r1 = r1 ^ ds[r0 & mask]; // 32 load r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + { uint32_t s_ = r2 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor r4 = r4 ^ ds[r0 & mask]; // 37 load diff --git a/proto-cuda/packs-readwidth/scr2/kernel_bound.cl b/proto-cuda/packs-readwidth/scr2k128/kernel_bound.cl similarity index 90% rename from proto-cuda/packs-readwidth/scr2/kernel_bound.cl rename to proto-cuda/packs-readwidth/scr2k128/kernel_bound.cl index 9889162fc..b08bc77a8 100644 --- a/proto-cuda/packs-readwidth/scr2/kernel_bound.cl +++ b/proto-cuda/packs-readwidth/scr2k128/kernel_bound.cl @@ -177,7 +177,7 @@ __kernel void igneum_build(__global uint* ds, __global const uint* cache, uint n // One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the // lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and // __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -185,7 +185,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint lane = (uint)get_global_id(0) & 31u; uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; - __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -226,7 +226,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + { uint s_ = r6 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl r4 = r0 * r2 + r4; // 19 mad @@ -244,7 +244,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r3 = r3 ^ ds[r1 & mask]; // 31 load r1 = r1 ^ ds[r0 & mask]; // 32 load r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + { uint s_ = r2 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor r4 = r4 ^ ds[r0 & mask]; // 37 load @@ -295,7 +295,7 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon uint lane = (uint)get_global_id(0) & 31u; uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; - __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -337,7 +337,7 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + { uint s_ = r6 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl r4 = r0 * r2 + r4; // 19 mad @@ -355,7 +355,7 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon r3 = r3 ^ ds[r1 & mask]; // 31 load r1 = r1 ^ ds[r0 & mask]; // 32 load r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + { uint s_ = r2 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor r4 = r4 ^ ds[r0 & mask]; // 37 load diff --git a/proto-cuda/packs-readwidth/scr2/kernel_bound.cu b/proto-cuda/packs-readwidth/scr2k128/kernel_bound.cu similarity index 85% rename from proto-cuda/packs-readwidth/scr2/kernel_bound.cu rename to proto-cuda/packs-readwidth/scr2k128/kernel_bound.cu index b012ad62e..81875c255 100644 --- a/proto-cuda/packs-readwidth/scr2/kernel_bound.cu +++ b/proto-cuda/packs-readwidth/scr2k128/kernel_bound.cu @@ -20,7 +20,7 @@ __device__ __forceinline__ uint32_t splitmix32(uint32_t x) { __device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } __device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. __device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -28,7 +28,7 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; - uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { uint32_t gid = g_ * 32u + lane; uint32_t gbase = baseNonce + g_ * 32u; @@ -62,7 +62,7 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint32_t s_ = r6 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + { uint32_t s_ = r6 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl r4 = r0 * r2 + r4; // 19 mad @@ -80,7 +80,7 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba r3 = r3 ^ ds[r1 & mask]; // 31 load r1 = r1 ^ ds[r0 & mask]; // 32 load r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + { uint32_t s_ = r2 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor r4 = r4 ^ ds[r0 & mask]; // 37 load diff --git a/proto-cuda/packs-readwidth/scr2/memhard.h b/proto-cuda/packs-readwidth/scr2k128/memhard.h similarity index 100% rename from proto-cuda/packs-readwidth/scr2/memhard.h rename to proto-cuda/packs-readwidth/scr2k128/memhard.h diff --git a/proto-cuda/packs-readwidth/scr2/memhard.metal b/proto-cuda/packs-readwidth/scr2k128/memhard.metal similarity index 100% rename from proto-cuda/packs-readwidth/scr2/memhard.metal rename to proto-cuda/packs-readwidth/scr2k128/memhard.metal diff --git a/proto-cuda/packs-readwidth/scr2k128/program.h b/proto-cuda/packs-readwidth/scr2k128/program.h new file mode 100644 index 000000000..b2b831632 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2k128/program.h @@ -0,0 +1,67 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-genesis" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0xe0afe0d155bc25d9ull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=14 add=9 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 shfl=4 rotr=2 scratch=2 rotl=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "scr2k128" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 100, 0, 0 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 14, 0, 0 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 448 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// Variant 5: persistent warps, a 128 KiB scratch per launched warp (the host launches N warps and passes scratch, +// groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). +#define IGNEUM_PERSISTENT_WARPS 1 +#define IGNEUM_SCRATCH_OPS 2 // scratch read-modify-writes per program (16 per hash) +#define IGNEUM_SCRATCH_SLOTS 256u +#define IGNEUM_SCRATCH_WORDS_PER_LANE 1024u +#define IGNEUM_SCRATCH_BYTES_PER_WARP 131072u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/scr2/program.json b/proto-cuda/packs-readwidth/scr2k128/program.json similarity index 98% rename from proto-cuda/packs-readwidth/scr2/program.json rename to proto-cuda/packs-readwidth/scr2k128/program.json index b9b6b935c..4b6d69fa5 100644 --- a/proto-cuda/packs-readwidth/scr2/program.json +++ b/proto-cuda/packs-readwidth/scr2k128/program.json @@ -2,7 +2,7 @@ "format": "igneum-program-pack-3", "generator": 2, "attempt": 0, - "program_id": "0x2f0988e568f37cc3", + "program_id": "0xe0afe0d155bc25d9", "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", "dataset_mode": "memory-hard", "seed": "igneum-genesis", @@ -15,13 +15,14 @@ "iterations": 8, "instruction_count": 64, "loads_per_hash": 128, - "load_class": "scr2", + "load_class": "scr2k128", "load_slots": 16, "load_mix_percent_4_16_64": [100, 0, 0], "load_width_counts_4_16_64": [14, 0, 0], "bytes_per_hash": 448, "scratch_ops_per_hash": 16, - "scratch": "variant 5 (measurement only): persistent warps; a 1 MiB scratch per warp of 2048 16-byte slots per lane (lane-major); slot = src & 0x7ff; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", + "scratch_kib_per_warp": 128, + "scratch": "variant 5 (measurement only): persistent warps; a 128 KiB scratch per warp of 256 16-byte slots per lane (lane-major); slot = src & 0xff; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", "op_mix": {"load": 14, "add": 9, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "shfl": 4, "rotr": 2, "scratch": 2, "rotl": 1}, "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", diff --git a/proto-cuda/packs-readwidth/scr2/program.metal b/proto-cuda/packs-readwidth/scr2k128/program.metal similarity index 83% rename from proto-cuda/packs-readwidth/scr2/program.metal rename to proto-cuda/packs-readwidth/scr2k128/program.metal index 5c8835ec2..cf5d87bc2 100644 --- a/proto-cuda/packs-readwidth/scr2/program.metal +++ b/proto-cuda/packs-readwidth/scr2k128/program.metal @@ -21,7 +21,7 @@ inline uint ds_elem(uint i, uint d0, uint d1) { return x; } -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -36,7 +36,7 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], uint lane = tid & 31u; uint warp_ = tid >> 5; uint nwarps_ = nthreads >> 5; - device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -70,7 +70,7 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], r2 = r2 * r5; // 13 r1 = r1 ^ dataset[r2 & MASK]; // 14 r7 = rotl_imm(r7, 1u); // 15 - { uint s_ = r6 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + { uint s_ = r6 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 r7 = r7 ^ dataset[r4 & MASK]; // 17 r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 r4 = r0 * r2 + r4; // 19 @@ -88,7 +88,7 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], r3 = r3 ^ dataset[r1 & MASK]; // 31 r1 = r1 ^ dataset[r0 & MASK]; // 32 r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 - { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + { uint s_ = r2 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 r0 = r0 * r3; // 35 r2 = r2 ^ r5; // 36 r4 = r4 ^ dataset[r0 & MASK]; // 37 diff --git a/proto-cuda/packs-readwidth/scr2/program_bound.metal b/proto-cuda/packs-readwidth/scr2k128/program_bound.metal similarity index 84% rename from proto-cuda/packs-readwidth/scr2/program_bound.metal rename to proto-cuda/packs-readwidth/scr2k128/program_bound.metal index b521473ef..ba42f27c1 100644 --- a/proto-cuda/packs-readwidth/scr2/program_bound.metal +++ b/proto-cuda/packs-readwidth/scr2k128/program_bound.metal @@ -21,7 +21,7 @@ inline uint ds_elem(uint i, uint d0, uint d1) { return x; } -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -38,7 +38,7 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], uint lane = tid & 31u; uint warp_ = tid >> 5; uint nwarps_ = nthreads >> 5; - device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -72,7 +72,7 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], r2 = r2 * r5; // 13 r1 = r1 ^ dataset[r2 & MASK]; // 14 r7 = rotl_imm(r7, 1u); // 15 - { uint s_ = r6 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + { uint s_ = r6 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 r7 = r7 ^ dataset[r4 & MASK]; // 17 r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 r4 = r0 * r2 + r4; // 19 @@ -90,7 +90,7 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], r3 = r3 ^ dataset[r1 & MASK]; // 31 r1 = r1 ^ dataset[r0 & MASK]; // 32 r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 - { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + { uint s_ = r2 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 r0 = r0 * r3; // 35 r2 = r2 ^ r5; // 36 r4 = r4 ^ dataset[r0 & MASK]; // 37 diff --git a/proto-cuda/packs-readwidth/scr4/vectors.h b/proto-cuda/packs-readwidth/scr2k128/vectors.h similarity index 60% rename from proto-cuda/packs-readwidth/scr4/vectors.h rename to proto-cuda/packs-readwidth/scr2k128/vectors.h index 88076f2e3..2a556f269 100644 --- a/proto-cuda/packs-readwidth/scr4/vectors.h +++ b/proto-cuda/packs-readwidth/scr2k128/vectors.h @@ -11,22 +11,22 @@ static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { { // base nonce 0 - 0x66cffcc97c46e625ull, 0xbc9019f8df50fbfdull, 0x65629c90dde6016eull, 0xd647a41effa03d3bull, 0x86da3b6bbd751b99ull, 0x6ccf4240a0fb2d19ull, 0xeb39a1e06f17378cull, 0x2ea6b349b289fb10ull, - 0x6067211e6c220500ull, 0x6e6095dfedd1360full, 0xbd1190d8b50e1b48ull, 0x216dc72a0c08d5b5ull, 0x5be1f8c080836b0cull, 0x2a32932a5953ed73ull, 0xcc2a3d68be83c802ull, 0xe46daca15338278full, - 0xb2de43b96761e459ull, 0x9004acd06588cbeaull, 0x6a9a3543cf93004full, 0xff956d859cb6e408ull, 0x4397ec6e3c5fb045ull, 0x521dea569cd481d5ull, 0x89832b34108759f0ull, 0xf66e393836ffe4eaull, - 0xb4e39af6c40ea2f4ull, 0x3adc22085dd8d648ull, 0x27efe270958bbfbbull, 0x6c80be0e8dca60d8ull, 0xa0afbc6a60260d59ull, 0x5d9a257fb9189537ull, 0xeb837aeef55dc3edull, 0xd381174dc14f8951ull + 0x870ae6d97d9e85d8ull, 0x82f91989add778d3ull, 0x127e79aa060861e3ull, 0x148dec51aec6ee46ull, 0xcb5b4b144055bebaull, 0x832ea2d8305f7177ull, 0x374d405f6f35141dull, 0xe83f0b25fcc5b98cull, + 0x8c6696ba39a3dfbcull, 0x5e588ceed27c0f20ull, 0x87382ec243a8208full, 0x831da1dd50dedfd7ull, 0x11af1b35d85d23faull, 0x3f203e64b5a6ae40ull, 0x5562b30447cf8941ull, 0xdbc5ddc8b06cab3cull, + 0x089351d365256721ull, 0xf4e67e3e0ce8ce3dull, 0x9a6581aba1e08812ull, 0xf8c1f2017f31d0b2ull, 0x00e282fb6d67ed1bull, 0x64ceb85ec97d5368ull, 0x0db3762c566cd35full, 0x9ecf6fb65b27a141ull, + 0xdfb6edb29d069ef0ull, 0xa3a2eb24fa67fb93ull, 0x2527bac1b676b544ull, 0x275704b557b5c1d5ull, 0x909fe0fbbb3d7e3eull, 0x6d3b7083f0b9dac1ull, 0xf7a58d446af0c3f0ull, 0x45a10eec5d0361ebull }, { // base nonce 4096 - 0xfb1f61aaeaeaef94ull, 0x184c2160963d8b57ull, 0x42a053c628625778ull, 0xeaad0e41c4770812ull, 0x1d6d389ceb462ce1ull, 0x4a639827672bdbd4ull, 0x857e42aa5a42dd6full, 0xb4ef399e5339979cull, - 0x497a29225b099233ull, 0x71d8b42862d81954ull, 0x0af995663313bf04ull, 0xf436fd126619d7a1ull, 0x199e4e3333cff269ull, 0x64077952f3775768ull, 0x51af1d126c5e8388ull, 0xffbaf44fe6b15cfdull, - 0xfc8fed86ecae34a7ull, 0x4cb548616f7a7d6bull, 0xc21d938c8b5bef35ull, 0x34789cbdd7088f71ull, 0xacb099a2c207d891ull, 0xfe1902d162374413ull, 0x26f7831c28f4020bull, 0xdf5192952b4af6b0ull, - 0xecab61fe88dbaff4ull, 0x941c491f7fdb86e5ull, 0x2b1900c53f746e77ull, 0x8c40507b1caffeb2ull, 0x7532a1ec2b9169efull, 0x1cf399b0c8bfb520ull, 0xdf003d2bb8a2cc0cull, 0x4da853307fc977a9ull + 0x5ba9a19be2ac506full, 0xf01aba1b9e1fbd4bull, 0x82576a1ada6a06aaull, 0x3cdfb035063961adull, 0x3b1c0146bee5cc0bull, 0xb9eb92e4388bb2edull, 0xfb3d93c5214abf98ull, 0xb624a0986997e24aull, + 0x6dc096da5e72a34aull, 0x598baa91443c82dbull, 0x689d8cef7afc8df5ull, 0x41ae32225004d576ull, 0x06adbced85e30f4dull, 0x1bef955028e11da8ull, 0x8e3f1fde00391e41ull, 0xf1e29599bd9c776eull, + 0xd85a42863321dfbcull, 0x2f220e8389179830ull, 0x658f1c559f2e3e28ull, 0x7ddc9adae1172cfaull, 0x493dad4ec6a7d467ull, 0x4f8cf35bbfa01901ull, 0x87971b6666cd0093ull, 0xbe71aa56ca4f56b0ull, + 0xe9cee6ec93582a96ull, 0xa10e3cc76912027cull, 0x0d3b49bb042a2033ull, 0x680fa00b5d161278ull, 0xdef83e00736c287aull, 0x354ed038aae23286ull, 0xd280ce9bd9a71c97ull, 0x4f48b2b22716cd62ull }, { // base nonce 1000000 - 0x3d094bd04694b96full, 0xaecbd76cecd1a20aull, 0xbcb86febe56b17feull, 0x98082b557ba97517ull, 0xbb5f94108888564bull, 0xea3284877a30fc87ull, 0xc608fa4d5a8bb2adull, 0x946c721c511e0729ull, - 0x46c6eba292083aedull, 0x936cb97231eb6795ull, 0xb1413c434c712cbbull, 0xedfd554d3948c1bdull, 0xa8a20cbef2faccd5ull, 0x5fe39d756cadbcadull, 0x208b2627380791feull, 0xf52f9374ce480218ull, - 0xb9db7cd8814eb29eull, 0xf32ed2192b5a8719ull, 0x4f1b06a054940aefull, 0x406df498e4365eb5ull, 0x1982075caad345efull, 0x590f725623dbbbbdull, 0xa26d9192dedfefa5ull, 0x36219ec00da18980ull, - 0x5361d0dcb0f8b1a3ull, 0x35bdefa2fbb5ffc3ull, 0xba4c2a4e473a9c80ull, 0x107d9d3030f8b9d3ull, 0xa8bb094266d6b987ull, 0x86164fdfbb1426e8ull, 0xa6e8cb895021cbbdull, 0xfe12809e9d99a243ull + 0x3220aa9dc0bca592ull, 0x5409a7301dfcc3b6ull, 0x31c9b27aad8ef845ull, 0x564ceb4647002ad9ull, 0xcb7dc5fb129b07a6ull, 0xd940b224720c0393ull, 0x0e7f07a345108096ull, 0xc55a7d439b46f6a4ull, + 0x0a0fab6343d5756aull, 0x397b5e178eeaa82cull, 0xe6a0adaa085838f8ull, 0x7389e3a61a07941bull, 0xeff9f511b46237bbull, 0xb35b6bf8e31ad14cull, 0x0b7d29ad2f411c1bull, 0x988e30319baf21daull, + 0x362ff74de128b05bull, 0x312357fc7857b317ull, 0xb006adf1d322c446ull, 0x4dcdf54a98b429eeull, 0x18a48dc7b38023bdull, 0xb2af25e04714b6ceull, 0x4d7b97fe2f335dc2ull, 0xc161c509e57b4616ull, + 0xb2b90c940eb64229ull, 0x2549b669b9e63f78ull, 0x1a1a4ff0e629079cull, 0x80ccafd83d359a54ull, 0x0232c0dfa9240d4dull, 0xc98a751590860b3full, 0x8e96a0aae858e53bull, 0x906b462113b3f109ull } }; diff --git a/proto-cuda/packs-readwidth/scr8/vectors.json b/proto-cuda/packs-readwidth/scr2k128/vectors.json similarity index 65% rename from proto-cuda/packs-readwidth/scr8/vectors.json rename to proto-cuda/packs-readwidth/scr2k128/vectors.json index 64e628d6e..cf0908113 100644 --- a/proto-cuda/packs-readwidth/scr8/vectors.json +++ b/proto-cuda/packs-readwidth/scr2k128/vectors.json @@ -8,22 +8,22 @@ "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", "warps": [ {"base_nonce": 0, "expected": [ - "0xd0f846c3cd57ae09", "0x4d5e1caf41761a5c", "0x7fc77bb7fa221a51", "0xe45b6317d4f52ae7", "0x09704a39107a3150", "0xd5166620f9dc49ab", "0xeaf1fa69e4075e55", "0x38e23d449b7616ae", - "0xf4593424c8427320", "0x9359a44a5149bd60", "0x2df93b5bc274f083", "0xfd80eb32016f8659", "0xdeb10b8a37fc3ee7", "0xc18e2ed77ab77f34", "0x7f39b618cc81c1b4", "0xdc1b7299961ad5ab", - "0x8f9c03d151013ec4", "0x668927a4a267d76f", "0x233a50c329caf635", "0x7c80406541ae3f15", "0x18fb38606c48af4a", "0x0452a913df5f112b", "0x59065cb670ba6d7d", "0xa2998e2ccd726df8", - "0xef690e221af37893", "0x4e87968e5c45903e", "0x6c6e5d4c1b35d7a9", "0x54d7e322426dcae3", "0x0d3ed6c08deda320", "0x3d6a4301fbb06a5a", "0x9eb78b04d2366566", "0x7806f8c64d2d09ff" + "0x870ae6d97d9e85d8", "0x82f91989add778d3", "0x127e79aa060861e3", "0x148dec51aec6ee46", "0xcb5b4b144055beba", "0x832ea2d8305f7177", "0x374d405f6f35141d", "0xe83f0b25fcc5b98c", + "0x8c6696ba39a3dfbc", "0x5e588ceed27c0f20", "0x87382ec243a8208f", "0x831da1dd50dedfd7", "0x11af1b35d85d23fa", "0x3f203e64b5a6ae40", "0x5562b30447cf8941", "0xdbc5ddc8b06cab3c", + "0x089351d365256721", "0xf4e67e3e0ce8ce3d", "0x9a6581aba1e08812", "0xf8c1f2017f31d0b2", "0x00e282fb6d67ed1b", "0x64ceb85ec97d5368", "0x0db3762c566cd35f", "0x9ecf6fb65b27a141", + "0xdfb6edb29d069ef0", "0xa3a2eb24fa67fb93", "0x2527bac1b676b544", "0x275704b557b5c1d5", "0x909fe0fbbb3d7e3e", "0x6d3b7083f0b9dac1", "0xf7a58d446af0c3f0", "0x45a10eec5d0361eb" ]}, {"base_nonce": 4096, "expected": [ - "0xe2c97c384a85c687", "0xcf905b005ab01ebc", "0x4526799fae210c6b", "0xd4daf72ed75e8a16", "0xb3703a9e7c6820a8", "0xb6be9393fbb920bd", "0xe2fff04fb7816eb8", "0xed76ad5dbf400f91", - "0xea7d4b58ab0eb8d1", "0xf67a73b01f030532", "0x35ab823037510099", "0x5de1b1b0b3a26c69", "0xbe5dd90dbb7e632b", "0x1475918a9237e24d", "0x177fd53c45634d71", "0xa7cf00759ba28ce0", - "0xe51be0584ac3fbb4", "0x427049cc778aab35", "0x826bab125577d172", "0xd705b891b16237f5", "0xdf622fd44b180a87", "0x359398ecb79ec2de", "0x2a1e075fb078da66", "0xd480ddd8e66d26e8", - "0x07e3e86f517de466", "0xa98b2f7423557445", "0x6bab15b36bb142fe", "0xf87d147bf2cc5c0b", "0xf0294ea2b2820e03", "0xf219ac95e823d794", "0x9fa3fea85bc54264", "0xc4af3fafd2ff5201" + "0x5ba9a19be2ac506f", "0xf01aba1b9e1fbd4b", "0x82576a1ada6a06aa", "0x3cdfb035063961ad", "0x3b1c0146bee5cc0b", "0xb9eb92e4388bb2ed", "0xfb3d93c5214abf98", "0xb624a0986997e24a", + "0x6dc096da5e72a34a", "0x598baa91443c82db", "0x689d8cef7afc8df5", "0x41ae32225004d576", "0x06adbced85e30f4d", "0x1bef955028e11da8", "0x8e3f1fde00391e41", "0xf1e29599bd9c776e", + "0xd85a42863321dfbc", "0x2f220e8389179830", "0x658f1c559f2e3e28", "0x7ddc9adae1172cfa", "0x493dad4ec6a7d467", "0x4f8cf35bbfa01901", "0x87971b6666cd0093", "0xbe71aa56ca4f56b0", + "0xe9cee6ec93582a96", "0xa10e3cc76912027c", "0x0d3b49bb042a2033", "0x680fa00b5d161278", "0xdef83e00736c287a", "0x354ed038aae23286", "0xd280ce9bd9a71c97", "0x4f48b2b22716cd62" ]}, {"base_nonce": 1000000, "expected": [ - "0x501f772483fac0a3", "0x461363da2c1539e0", "0x750050674de592af", "0x14ed105042cd912c", "0xc6477310878614eb", "0xe916b44e32e90e56", "0x531c92e69c2bdd73", "0x5bb0129ef7c3bd51", - "0x967ed91f7cbc7c79", "0x06fcc6a895b58b4a", "0x3bdb29fbf93cbff3", "0xbed385172d2e6abe", "0x921fc99ff4efac5e", "0x6b0090bedd9f69c7", "0x0b105c18aaa53ab0", "0xf0c4203561f3b94d", - "0x64abe33adccf7807", "0xd8f3a3b7e0242c04", "0x397da462f69fac5b", "0xe6b6b32e467b71be", "0x5318ac9e56278d04", "0xf7c5e348a4e1f5db", "0x79c994e6109646df", "0x9cdc42b0f6e82231", - "0x3f1160f02fd96ac7", "0xa69a23fc7058be21", "0xccde0a19bfdc25b6", "0xd3684a5b966f7497", "0x929db00b97a624f9", "0xfd14a40882d395a6", "0x49b0c1ecc514d6ad", "0x95d994c8349be12d" + "0x3220aa9dc0bca592", "0x5409a7301dfcc3b6", "0x31c9b27aad8ef845", "0x564ceb4647002ad9", "0xcb7dc5fb129b07a6", "0xd940b224720c0393", "0x0e7f07a345108096", "0xc55a7d439b46f6a4", + "0x0a0fab6343d5756a", "0x397b5e178eeaa82c", "0xe6a0adaa085838f8", "0x7389e3a61a07941b", "0xeff9f511b46237bb", "0xb35b6bf8e31ad14c", "0x0b7d29ad2f411c1b", "0x988e30319baf21da", + "0x362ff74de128b05b", "0x312357fc7857b317", "0xb006adf1d322c446", "0x4dcdf54a98b429ee", "0x18a48dc7b38023bd", "0xb2af25e04714b6ce", "0x4d7b97fe2f335dc2", "0xc161c509e57b4616", + "0xb2b90c940eb64229", "0x2549b669b9e63f78", "0x1a1a4ff0e629079c", "0x80ccafd83d359a54", "0x0232c0dfa9240d4d", "0xc98a751590860b3f", "0x8e96a0aae858e53b", "0x906b462113b3f109" ]} ], "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], diff --git a/proto-cuda/packs-readwidth/scr4/kernel.cl b/proto-cuda/packs-readwidth/scr2k32/kernel.cl similarity index 88% rename from proto-cuda/packs-readwidth/scr4/kernel.cl rename to proto-cuda/packs-readwidth/scr2k32/kernel.cl index a75793f6f..6befdde06 100644 --- a/proto-cuda/packs-readwidth/scr4/kernel.cl +++ b/proto-cuda/packs-readwidth/scr2k32/kernel.cl @@ -177,7 +177,7 @@ __kernel void igneum_build(__global uint* ds, __global const uint* cache, uint n // One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the // lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and // __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -185,7 +185,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint lane = (uint)get_global_id(0) & 31u; uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; - __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -226,7 +226,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + { uint s_ = r6 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl r4 = r0 * r2 + r4; // 19 mad @@ -242,9 +242,9 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add r6 = rotr_var(r6, r7); // 30 rotr r3 = r3 ^ ds[r1 & mask]; // 31 load - { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r1 = r1 ^ ds[r0 & mask]; // 32 load r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + { uint s_ = r2 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor r4 = r4 ^ ds[r0 & mask]; // 37 load @@ -269,7 +269,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r2 = r2 ^ ds[r7 & mask]; // 56 load r5 = r5 - r6; // 57 sub r1 = r1 ^ ds[r3 & mask]; // 58 load - { uint s_ = r4 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r1 = r1 ^ ds[r4 & mask]; // 59 load r4 = r4 - r6; // 60 sub r1 = r1 * r2; // 61 mul r3 = r6 * r0 + r3; // 62 mad diff --git a/proto-cuda/packs-readwidth/scr4/kernel.cu b/proto-cuda/packs-readwidth/scr2k32/kernel.cu similarity index 79% rename from proto-cuda/packs-readwidth/scr4/kernel.cu rename to proto-cuda/packs-readwidth/scr2k32/kernel.cu index fd9584208..5fd93b513 100644 --- a/proto-cuda/packs-readwidth/scr4/kernel.cu +++ b/proto-cuda/packs-readwidth/scr2k32/kernel.cu @@ -44,7 +44,7 @@ __global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItem // One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every // __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a // 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. __device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -52,7 +52,7 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; - uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { uint32_t gid = g_ * 32u + lane; uint32_t gbase = baseNonce + g_ * 32u; @@ -86,7 +86,7 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint32_t s_ = r6 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + { uint32_t s_ = r6 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl r4 = r0 * r2 + r4; // 19 mad @@ -102,9 +102,9 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add r6 = rotr_var(r6, r7); // 30 rotr r3 = r3 ^ ds[r1 & mask]; // 31 load - { uint32_t s_ = r0 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + r1 = r1 ^ ds[r0 & mask]; // 32 load r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + { uint32_t s_ = r2 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor r4 = r4 ^ ds[r0 & mask]; // 37 load @@ -129,7 +129,7 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc r2 = r2 ^ ds[r7 & mask]; // 56 load r5 = r5 - r6; // 57 sub r1 = r1 ^ ds[r3 & mask]; // 58 load - { uint32_t s_ = r4 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + r1 = r1 ^ ds[r4 & mask]; // 59 load r4 = r4 - r6; // 60 sub r1 = r1 * r2; // 61 mul r3 = r6 * r0 + r3; // 62 mad diff --git a/proto-cuda/packs-readwidth/scr4/kernel_bound.cl b/proto-cuda/packs-readwidth/scr2k32/kernel_bound.cl similarity index 83% rename from proto-cuda/packs-readwidth/scr4/kernel_bound.cl rename to proto-cuda/packs-readwidth/scr2k32/kernel_bound.cl index e33afd4bc..9fb87c712 100644 --- a/proto-cuda/packs-readwidth/scr4/kernel_bound.cl +++ b/proto-cuda/packs-readwidth/scr2k32/kernel_bound.cl @@ -177,7 +177,7 @@ __kernel void igneum_build(__global uint* ds, __global const uint* cache, uint n // One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the // lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and // __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -185,7 +185,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint lane = (uint)get_global_id(0) & 31u; uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; - __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -226,7 +226,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + { uint s_ = r6 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl r4 = r0 * r2 + r4; // 19 mad @@ -242,9 +242,9 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add r6 = rotr_var(r6, r7); // 30 rotr r3 = r3 ^ ds[r1 & mask]; // 31 load - { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r1 = r1 ^ ds[r0 & mask]; // 32 load r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + { uint s_ = r2 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor r4 = r4 ^ ds[r0 & mask]; // 37 load @@ -269,7 +269,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r2 = r2 ^ ds[r7 & mask]; // 56 load r5 = r5 - r6; // 57 sub r1 = r1 ^ ds[r3 & mask]; // 58 load - { uint s_ = r4 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r1 = r1 ^ ds[r4 & mask]; // 59 load r4 = r4 - r6; // 60 sub r1 = r1 * r2; // 61 mul r3 = r6 * r0 + r3; // 62 mad @@ -295,7 +295,7 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon uint lane = (uint)get_global_id(0) & 31u; uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; - __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -337,7 +337,7 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + { uint s_ = r6 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl r4 = r0 * r2 + r4; // 19 mad @@ -353,9 +353,9 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add r6 = rotr_var(r6, r7); // 30 rotr r3 = r3 ^ ds[r1 & mask]; // 31 load - { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r1 = r1 ^ ds[r0 & mask]; // 32 load r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + { uint s_ = r2 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor r4 = r4 ^ ds[r0 & mask]; // 37 load @@ -380,7 +380,7 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon r2 = r2 ^ ds[r7 & mask]; // 56 load r5 = r5 - r6; // 57 sub r1 = r1 ^ ds[r3 & mask]; // 58 load - { uint s_ = r4 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r1 = r1 ^ ds[r4 & mask]; // 59 load r4 = r4 - r6; // 60 sub r1 = r1 * r2; // 61 mul r3 = r6 * r0 + r3; // 62 mad diff --git a/proto-cuda/packs-readwidth/scr4/kernel_bound.cu b/proto-cuda/packs-readwidth/scr2k32/kernel_bound.cu similarity index 75% rename from proto-cuda/packs-readwidth/scr4/kernel_bound.cu rename to proto-cuda/packs-readwidth/scr2k32/kernel_bound.cu index a4d184a0a..94e353cf8 100644 --- a/proto-cuda/packs-readwidth/scr4/kernel_bound.cu +++ b/proto-cuda/packs-readwidth/scr2k32/kernel_bound.cu @@ -20,7 +20,7 @@ __device__ __forceinline__ uint32_t splitmix32(uint32_t x) { __device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } __device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. __device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -28,7 +28,7 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; - uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { uint32_t gid = g_ * 32u + lane; uint32_t gbase = baseNonce + g_ * 32u; @@ -62,7 +62,7 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint32_t s_ = r6 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + { uint32_t s_ = r6 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl r4 = r0 * r2 + r4; // 19 mad @@ -78,9 +78,9 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add r6 = rotr_var(r6, r7); // 30 rotr r3 = r3 ^ ds[r1 & mask]; // 31 load - { uint32_t s_ = r0 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + r1 = r1 ^ ds[r0 & mask]; // 32 load r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + { uint32_t s_ = r2 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor r4 = r4 ^ ds[r0 & mask]; // 37 load @@ -105,7 +105,7 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba r2 = r2 ^ ds[r7 & mask]; // 56 load r5 = r5 - r6; // 57 sub r1 = r1 ^ ds[r3 & mask]; // 58 load - { uint32_t s_ = r4 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + r1 = r1 ^ ds[r4 & mask]; // 59 load r4 = r4 - r6; // 60 sub r1 = r1 * r2; // 61 mul r3 = r6 * r0 + r3; // 62 mad diff --git a/proto-cuda/packs-readwidth/scr4/memhard.h b/proto-cuda/packs-readwidth/scr2k32/memhard.h similarity index 100% rename from proto-cuda/packs-readwidth/scr4/memhard.h rename to proto-cuda/packs-readwidth/scr2k32/memhard.h diff --git a/proto-cuda/packs-readwidth/scr4/memhard.metal b/proto-cuda/packs-readwidth/scr2k32/memhard.metal similarity index 100% rename from proto-cuda/packs-readwidth/scr4/memhard.metal rename to proto-cuda/packs-readwidth/scr2k32/memhard.metal diff --git a/proto-cuda/packs-readwidth/scr2/program.h b/proto-cuda/packs-readwidth/scr2k32/program.h similarity index 91% rename from proto-cuda/packs-readwidth/scr2/program.h rename to proto-cuda/packs-readwidth/scr2k32/program.h index 4892b7988..b7f0eb488 100644 --- a/proto-cuda/packs-readwidth/scr2/program.h +++ b/proto-cuda/packs-readwidth/scr2k32/program.h @@ -15,7 +15,7 @@ #define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" #define IGNEUM_GENERATOR 2 #define IGNEUM_PROGRAM_ATTEMPT 0 -#define IGNEUM_PROGRAM_ID 0x2f0988e568f37cc3ull +#define IGNEUM_PROGRAM_ID 0xe0b080d155bd35b9ull #define IGNEUM_DAY_STRING "2026-10-03" #define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" #define IGNEUM_DAY0 0x3067619fu @@ -30,20 +30,20 @@ #define IGNEUM_OP_MIX "load=14 add=9 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 shfl=4 rotr=2 scratch=2 rotl=1" // Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads // the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. -#define IGNEUM_LOAD_CLASS "scr2" +#define IGNEUM_LOAD_CLASS "scr2k32" #define IGNEUM_LOAD_SLOTS 16 #define IGNEUM_LOAD_MIX { 100, 0, 0 } #define IGNEUM_LOAD_WIDTH_COUNTS { 14, 0, 0 } // loads of 4, 16, 64 bytes per program #define IGNEUM_BYTES_PER_HASH 448 #define IGNEUM_FOLD_ROT 11 #define IGNEUM_FOLD_MUL 0x9e3779b1u -// Variant 5: persistent warps, a 1 MiB scratch per launched warp (the host launches N warps and passes scratch, +// Variant 5: persistent warps, a 32 KiB scratch per launched warp (the host launches N warps and passes scratch, // groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). #define IGNEUM_PERSISTENT_WARPS 1 #define IGNEUM_SCRATCH_OPS 2 // scratch read-modify-writes per program (16 per hash) -#define IGNEUM_SCRATCH_SLOTS 2048u -#define IGNEUM_SCRATCH_WORDS_PER_LANE 8192u -#define IGNEUM_SCRATCH_BYTES_PER_WARP 1048576u +#define IGNEUM_SCRATCH_SLOTS 64u +#define IGNEUM_SCRATCH_WORDS_PER_LANE 256u +#define IGNEUM_SCRATCH_BYTES_PER_WARP 32768u // 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) #define IGNEUM_DATASET_MODE 1 diff --git a/proto-cuda/packs-readwidth/scr2k32/program.json b/proto-cuda/packs-readwidth/scr2k32/program.json new file mode 100644 index 000000000..e9e8d9cbf --- /dev/null +++ b/proto-cuda/packs-readwidth/scr2k32/program.json @@ -0,0 +1,130 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0xe0b080d155bd35b9", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-genesis", + "seed_bytes": "69676e65756d2d67656e65736973", + "seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "scr2k32", + "load_slots": 16, + "load_mix_percent_4_16_64": [100, 0, 0], + "load_width_counts_4_16_64": [14, 0, 0], + "bytes_per_hash": 448, + "scratch_ops_per_hash": 16, + "scratch_kib_per_warp": 32, + "scratch": "variant 5 (measurement only): persistent warps; a 32 KiB scratch per warp of 64 16-byte slots per lane (lane-major); slot = src & 0x3f; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 14, "add": 9, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "shfl": 4, "rotr": 2, "scratch": 2, "rotl": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2, "width": 1}, + {"i": 1, "op": "add", "dst": 1, "src": 7, "src2": 2, "imm": "0x42da7657", "imm2": "0xc3bd2355", "rot": 25, "bit": 4, "mask": 16, "width": 1}, + {"i": 2, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x61f0b51c", "imm2": "0x2735a174", "rot": 4, "bit": 26, "mask": 2, "width": 1}, + {"i": 3, "op": "mad", "dst": 4, "src": 0, "src2": 6, "imm": "0x679648a8", "imm2": "0x3044ba32", "rot": 31, "bit": 31, "mask": 4, "width": 1}, + {"i": 4, "op": "load", "dst": 7, "src": 2, "src2": 6, "imm": "0x5d1ca2a2", "imm2": "0xe2481807", "rot": 24, "bit": 3, "mask": 1, "width": 1}, + {"i": 5, "op": "load", "dst": 4, "src": 1, "src2": 2, "imm": "0x987c017a", "imm2": "0xf4d60559", "rot": 2, "bit": 0, "mask": 4, "width": 1}, + {"i": 6, "op": "shfl", "dst": 6, "src": 3, "src2": 7, "imm": "0x6ea7b2df", "imm2": "0x9fce5071", "rot": 7, "bit": 15, "mask": 4, "width": 1}, + {"i": 7, "op": "shfl", "dst": 1, "src": 5, "src2": 1, "imm": "0x26a2ecde", "imm2": "0xfec6ad22", "rot": 15, "bit": 11, "mask": 8, "width": 1}, + {"i": 8, "op": "xor", "dst": 7, "src": 5, "src2": 2, "imm": "0xbe4b445c", "imm2": "0x17a5a9c7", "rot": 8, "bit": 8, "mask": 1, "width": 1}, + {"i": 9, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4, "width": 1}, + {"i": 10, "op": "or", "dst": 1, "src": 2, "src2": 3, "imm": "0x4e7dc10d", "imm2": "0x196d165c", "rot": 14, "bit": 27, "mask": 16, "width": 1}, + {"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 1, "imm": "0xc5c3b55d", "imm2": "0xec061424", "rot": 26, "bit": 27, "mask": 8, "width": 1}, + {"i": 12, "op": "or", "dst": 6, "src": 2, "src2": 3, "imm": "0x306542fe", "imm2": "0x1bb1b429", "rot": 31, "bit": 0, "mask": 2, "width": 1}, + {"i": 13, "op": "mul", "dst": 2, "src": 5, "src2": 6, "imm": "0xa672cdd3", "imm2": "0x59a4829c", "rot": 22, "bit": 13, "mask": 16, "width": 1}, + {"i": 14, "op": "load", "dst": 1, "src": 2, "src2": 5, "imm": "0x028b4d37", "imm2": "0x7bbd78ea", "rot": 15, "bit": 2, "mask": 8, "width": 1}, + {"i": 15, "op": "rotl", "dst": 7, "src": 6, "src2": 6, "imm": "0x5c88a1a7", "imm2": "0x5c628769", "rot": 1, "bit": 3, "mask": 8, "width": 1}, + {"i": 16, "op": "scratch", "dst": 3, "src": 6, "src2": 7, "imm": "0xbac2ae81", "imm2": "0xcbbc7bdb", "rot": 18, "bit": 8, "mask": 8, "width": 1}, + {"i": 17, "op": "load", "dst": 7, "src": 4, "src2": 2, "imm": "0xe8ab93e9", "imm2": "0xa00de107", "rot": 2, "bit": 1, "mask": 16, "width": 1}, + {"i": 18, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2, "width": 1}, + {"i": 19, "op": "mad", "dst": 4, "src": 0, "src2": 2, "imm": "0x5fba7bc2", "imm2": "0xdf099cfb", "rot": 4, "bit": 15, "mask": 16, "width": 1}, + {"i": 20, "op": "shfl", "dst": 0, "src": 6, "src2": 3, "imm": "0x0a3056de", "imm2": "0x7f0c25c3", "rot": 27, "bit": 13, "mask": 8, "width": 1}, + {"i": 21, "op": "xor", "dst": 5, "src": 7, "src2": 4, "imm": "0xbd066e1d", "imm2": "0x6d3ddc5a", "rot": 2, "bit": 29, "mask": 1, "width": 1}, + {"i": 22, "op": "mulhi", "dst": 2, "src": 5, "src2": 0, "imm": "0xc7e9887a", "imm2": "0x19ec898f", "rot": 14, "bit": 9, "mask": 1, "width": 1}, + {"i": 23, "op": "load", "dst": 3, "src": 7, "src2": 2, "imm": "0xc7fcfc8f", "imm2": "0x8528b94f", "rot": 17, "bit": 13, "mask": 4, "width": 1}, + {"i": 24, "op": "mulhi", "dst": 7, "src": 3, "src2": 5, "imm": "0xd91641e8", "imm2": "0xaf77faf2", "rot": 22, "bit": 21, "mask": 1, "width": 1}, + {"i": 25, "op": "or", "dst": 5, "src": 4, "src2": 0, "imm": "0x84c03868", "imm2": "0xf6c691b7", "rot": 29, "bit": 14, "mask": 8, "width": 1}, + {"i": 26, "op": "mad", "dst": 4, "src": 5, "src2": 2, "imm": "0x3bb2b6ba", "imm2": "0x49d95fd5", "rot": 1, "bit": 5, "mask": 8, "width": 1}, + {"i": 27, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4, "width": 1}, + {"i": 28, "op": "mulhi", "dst": 6, "src": 7, "src2": 6, "imm": "0xd69c4715", "imm2": "0xe0ebc4ce", "rot": 29, "bit": 2, "mask": 8, "width": 1}, + {"i": 29, "op": "add", "dst": 6, "src": 1, "src2": 2, "imm": "0x3b2d2124", "imm2": "0x187a9128", "rot": 1, "bit": 9, "mask": 16, "width": 1}, + {"i": 30, "op": "rotr", "dst": 6, "src": 7, "src2": 0, "imm": "0x5c64a589", "imm2": "0x61c9a38d", "rot": 17, "bit": 21, "mask": 16, "width": 1}, + {"i": 31, "op": "load", "dst": 3, "src": 1, "src2": 7, "imm": "0xc37723fa", "imm2": "0xf3b024da", "rot": 16, "bit": 27, "mask": 16, "width": 1}, + {"i": 32, "op": "load", "dst": 1, "src": 0, "src2": 7, "imm": "0xcc7972c4", "imm2": "0xad098d15", "rot": 30, "bit": 21, "mask": 8, "width": 1}, + {"i": 33, "op": "add", "dst": 0, "src": 4, "src2": 4, "imm": "0x2c35699f", "imm2": "0x351dde38", "rot": 21, "bit": 18, "mask": 4, "width": 1}, + {"i": 34, "op": "scratch", "dst": 0, "src": 2, "src2": 3, "imm": "0xfae8902b", "imm2": "0x5cd8306f", "rot": 5, "bit": 28, "mask": 16, "width": 1}, + {"i": 35, "op": "mul", "dst": 0, "src": 3, "src2": 1, "imm": "0x4fa3f3db", "imm2": "0xdbf37e75", "rot": 7, "bit": 18, "mask": 4, "width": 1}, + {"i": 36, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2, "width": 1}, + {"i": 37, "op": "load", "dst": 4, "src": 0, "src2": 0, "imm": "0x04cc1d55", "imm2": "0x35c52d04", "rot": 11, "bit": 14, "mask": 2, "width": 1}, + {"i": 38, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0x3958f280", "imm2": "0x8713c7e1", "rot": 5, "bit": 23, "mask": 16, "width": 1}, + {"i": 39, "op": "add", "dst": 0, "src": 3, "src2": 3, "imm": "0xa907b90b", "imm2": "0x1b053acf", "rot": 30, "bit": 25, "mask": 16, "width": 1}, + {"i": 40, "op": "rotr", "dst": 2, "src": 5, "src2": 4, "imm": "0xf8662282", "imm2": "0x10bb9e30", "rot": 8, "bit": 6, "mask": 2, "width": 1}, + {"i": 41, "op": "mul", "dst": 3, "src": 2, "src2": 4, "imm": "0x49087d74", "imm2": "0x6348b489", "rot": 17, "bit": 9, "mask": 16, "width": 1}, + {"i": 42, "op": "add", "dst": 1, "src": 5, "src2": 1, "imm": "0xa32e000c", "imm2": "0x6058c2e3", "rot": 25, "bit": 20, "mask": 8, "width": 1}, + {"i": 43, "op": "xor", "dst": 3, "src": 4, "src2": 2, "imm": "0x3dad0eb6", "imm2": "0xb97578cb", "rot": 3, "bit": 27, "mask": 1, "width": 1}, + {"i": 44, "op": "load", "dst": 3, "src": 5, "src2": 7, "imm": "0x374aec92", "imm2": "0x626f11df", "rot": 20, "bit": 18, "mask": 8, "width": 1}, + {"i": 45, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4, "width": 1}, + {"i": 46, "op": "xor", "dst": 7, "src": 1, "src2": 0, "imm": "0xef6ac348", "imm2": "0x963bb7e6", "rot": 26, "bit": 3, "mask": 8, "width": 1}, + {"i": 47, "op": "add", "dst": 0, "src": 3, "src2": 0, "imm": "0x838b5065", "imm2": "0x36360066", "rot": 3, "bit": 31, "mask": 4, "width": 1}, + {"i": 48, "op": "mulhi", "dst": 7, "src": 5, "src2": 0, "imm": "0x8458f7ac", "imm2": "0xc1c15026", "rot": 27, "bit": 15, "mask": 8, "width": 1}, + {"i": 49, "op": "load", "dst": 0, "src": 2, "src2": 4, "imm": "0x636a9dc4", "imm2": "0xac023d9b", "rot": 22, "bit": 29, "mask": 1, "width": 1}, + {"i": 50, "op": "sub", "dst": 2, "src": 6, "src2": 0, "imm": "0x2baec8c9", "imm2": "0x4390f156", "rot": 3, "bit": 12, "mask": 8, "width": 1}, + {"i": 51, "op": "sub", "dst": 7, "src": 5, "src2": 7, "imm": "0x19234061", "imm2": "0xe84dfade", "rot": 4, "bit": 19, "mask": 1, "width": 1}, + {"i": 52, "op": "xor", "dst": 2, "src": 3, "src2": 5, "imm": "0xdc2cd71e", "imm2": "0x1b5d334b", "rot": 9, "bit": 8, "mask": 8, "width": 1}, + {"i": 53, "op": "sub", "dst": 7, "src": 0, "src2": 4, "imm": "0x605c31ec", "imm2": "0x9923ff88", "rot": 28, "bit": 25, "mask": 4, "width": 1}, + {"i": 54, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4, "width": 1}, + {"i": 55, "op": "xor", "dst": 7, "src": 5, "src2": 5, "imm": "0xad7493e7", "imm2": "0x3e400372", "rot": 13, "bit": 8, "mask": 1, "width": 1}, + {"i": 56, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0x87e933c9", "imm2": "0x8c854c1b", "rot": 17, "bit": 3, "mask": 8, "width": 1}, + {"i": 57, "op": "sub", "dst": 5, "src": 6, "src2": 5, "imm": "0x11be3bc9", "imm2": "0xbbaa8e24", "rot": 6, "bit": 5, "mask": 16, "width": 1}, + {"i": 58, "op": "load", "dst": 1, "src": 3, "src2": 2, "imm": "0xa732351a", "imm2": "0xc01349cd", "rot": 14, "bit": 17, "mask": 16, "width": 1}, + {"i": 59, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0xb20547b2", "imm2": "0xc94655de", "rot": 27, "bit": 30, "mask": 1, "width": 1}, + {"i": 60, "op": "sub", "dst": 4, "src": 6, "src2": 7, "imm": "0x67cf904c", "imm2": "0x6873b216", "rot": 27, "bit": 7, "mask": 16, "width": 1}, + {"i": 61, "op": "mul", "dst": 1, "src": 2, "src2": 7, "imm": "0x93ab0bf4", "imm2": "0x96158375", "rot": 14, "bit": 0, "mask": 16, "width": 1}, + {"i": 62, "op": "mad", "dst": 3, "src": 6, "src2": 0, "imm": "0x41a443a3", "imm2": "0xe69d7919", "rot": 9, "bit": 0, "mask": 16, "width": 1}, + {"i": 63, "op": "add", "dst": 0, "src": 1, "src2": 3, "imm": "0x2fe0e98b", "imm2": "0xc88e2942", "rot": 5, "bit": 16, "mask": 16, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/scr4/program.metal b/proto-cuda/packs-readwidth/scr2k32/program.metal similarity index 72% rename from proto-cuda/packs-readwidth/scr4/program.metal rename to proto-cuda/packs-readwidth/scr2k32/program.metal index 4a8922643..453ad7237 100644 --- a/proto-cuda/packs-readwidth/scr4/program.metal +++ b/proto-cuda/packs-readwidth/scr2k32/program.metal @@ -21,7 +21,7 @@ inline uint ds_elem(uint i, uint d0, uint d1) { return x; } -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -36,7 +36,7 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], uint lane = tid & 31u; uint warp_ = tid >> 5; uint nwarps_ = nthreads >> 5; - device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -70,7 +70,7 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], r2 = r2 * r5; // 13 r1 = r1 ^ dataset[r2 & MASK]; // 14 r7 = rotl_imm(r7, 1u); // 15 - { uint s_ = r6 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + { uint s_ = r6 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 r7 = r7 ^ dataset[r4 & MASK]; // 17 r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 r4 = r0 * r2 + r4; // 19 @@ -86,9 +86,9 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 r6 = rotr_var(r6, r7); // 30 r3 = r3 ^ dataset[r1 & MASK]; // 31 - { uint s_ = r0 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + r1 = r1 ^ dataset[r0 & MASK]; // 32 r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 - { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + { uint s_ = r2 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 r0 = r0 * r3; // 35 r2 = r2 ^ r5; // 36 r4 = r4 ^ dataset[r0 & MASK]; // 37 @@ -113,7 +113,7 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], r2 = r2 ^ dataset[r7 & MASK]; // 56 r5 = r5 - r6; // 57 r1 = r1 ^ dataset[r3 & MASK]; // 58 - { uint s_ = r4 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + r1 = r1 ^ dataset[r4 & MASK]; // 59 r4 = r4 - r6; // 60 r1 = r1 * r2; // 61 r3 = r6 * r0 + r3; // 62 diff --git a/proto-cuda/packs-readwidth/scr4/program_bound.metal b/proto-cuda/packs-readwidth/scr2k32/program_bound.metal similarity index 73% rename from proto-cuda/packs-readwidth/scr4/program_bound.metal rename to proto-cuda/packs-readwidth/scr2k32/program_bound.metal index 6235bce84..24e7f8b4d 100644 --- a/proto-cuda/packs-readwidth/scr4/program_bound.metal +++ b/proto-cuda/packs-readwidth/scr2k32/program_bound.metal @@ -21,7 +21,7 @@ inline uint ds_elem(uint i, uint d0, uint d1) { return x; } -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -38,7 +38,7 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], uint lane = tid & 31u; uint warp_ = tid >> 5; uint nwarps_ = nthreads >> 5; - device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -72,7 +72,7 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], r2 = r2 * r5; // 13 r1 = r1 ^ dataset[r2 & MASK]; // 14 r7 = rotl_imm(r7, 1u); // 15 - { uint s_ = r6 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + { uint s_ = r6 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 r7 = r7 ^ dataset[r4 & MASK]; // 17 r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 r4 = r0 * r2 + r4; // 19 @@ -88,9 +88,9 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 r6 = rotr_var(r6, r7); // 30 r3 = r3 ^ dataset[r1 & MASK]; // 31 - { uint s_ = r0 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + r1 = r1 ^ dataset[r0 & MASK]; // 32 r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 - { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + { uint s_ = r2 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 r0 = r0 * r3; // 35 r2 = r2 ^ r5; // 36 r4 = r4 ^ dataset[r0 & MASK]; // 37 @@ -115,7 +115,7 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], r2 = r2 ^ dataset[r7 & MASK]; // 56 r5 = r5 - r6; // 57 r1 = r1 ^ dataset[r3 & MASK]; // 58 - { uint s_ = r4 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + r1 = r1 ^ dataset[r4 & MASK]; // 59 r4 = r4 - r6; // 60 r1 = r1 * r2; // 61 r3 = r6 * r0 + r3; // 62 diff --git a/proto-cuda/packs-readwidth/scr8/vectors.h b/proto-cuda/packs-readwidth/scr2k32/vectors.h similarity index 60% rename from proto-cuda/packs-readwidth/scr8/vectors.h rename to proto-cuda/packs-readwidth/scr2k32/vectors.h index 43d22af98..e65ab66b4 100644 --- a/proto-cuda/packs-readwidth/scr8/vectors.h +++ b/proto-cuda/packs-readwidth/scr2k32/vectors.h @@ -11,22 +11,22 @@ static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { { // base nonce 0 - 0xd0f846c3cd57ae09ull, 0x4d5e1caf41761a5cull, 0x7fc77bb7fa221a51ull, 0xe45b6317d4f52ae7ull, 0x09704a39107a3150ull, 0xd5166620f9dc49abull, 0xeaf1fa69e4075e55ull, 0x38e23d449b7616aeull, - 0xf4593424c8427320ull, 0x9359a44a5149bd60ull, 0x2df93b5bc274f083ull, 0xfd80eb32016f8659ull, 0xdeb10b8a37fc3ee7ull, 0xc18e2ed77ab77f34ull, 0x7f39b618cc81c1b4ull, 0xdc1b7299961ad5abull, - 0x8f9c03d151013ec4ull, 0x668927a4a267d76full, 0x233a50c329caf635ull, 0x7c80406541ae3f15ull, 0x18fb38606c48af4aull, 0x0452a913df5f112bull, 0x59065cb670ba6d7dull, 0xa2998e2ccd726df8ull, - 0xef690e221af37893ull, 0x4e87968e5c45903eull, 0x6c6e5d4c1b35d7a9ull, 0x54d7e322426dcae3ull, 0x0d3ed6c08deda320ull, 0x3d6a4301fbb06a5aull, 0x9eb78b04d2366566ull, 0x7806f8c64d2d09ffull + 0x589f62cd61c27dc1ull, 0xf6be7bab8b00a7c7ull, 0x349551c5979e330eull, 0x6d8156f5afaf0064ull, 0x81971552315d55b8ull, 0x20a3bb31ef5e202cull, 0x91818c4ec9fd5fecull, 0x35ca8cde74ac9715ull, + 0xc1db0a90f80c8801ull, 0xbc52521a33053a8aull, 0x89fda590bf946dd6ull, 0xc54fe4aeaf11975full, 0xffc712960d7f4022ull, 0x4da6b3de6f9a0abdull, 0xf70ff0468e9b595aull, 0xc2d1d4434c1eb8c9ull, + 0x70dc278b36920857ull, 0xfb20a2d85fa65b83ull, 0xd24316c07937542dull, 0xd505e1e3e85694bdull, 0x88c759248ae122a7ull, 0xf1c04d8ad71d5e8cull, 0x7e4d0e4e99717f12ull, 0x63d9e4b8f619ac98ull, + 0xcfef2a4243ba8608ull, 0xb2908c3da3593ba9ull, 0x48bee70b2ee45b84ull, 0x8f7a09d71be2c321ull, 0x0aadd48107e05bdbull, 0x8b73ff0f34ddaaaaull, 0x9ff3a4875ecf0fe3ull, 0xaf3d701cd1cbbd4bull }, { // base nonce 4096 - 0xe2c97c384a85c687ull, 0xcf905b005ab01ebcull, 0x4526799fae210c6bull, 0xd4daf72ed75e8a16ull, 0xb3703a9e7c6820a8ull, 0xb6be9393fbb920bdull, 0xe2fff04fb7816eb8ull, 0xed76ad5dbf400f91ull, - 0xea7d4b58ab0eb8d1ull, 0xf67a73b01f030532ull, 0x35ab823037510099ull, 0x5de1b1b0b3a26c69ull, 0xbe5dd90dbb7e632bull, 0x1475918a9237e24dull, 0x177fd53c45634d71ull, 0xa7cf00759ba28ce0ull, - 0xe51be0584ac3fbb4ull, 0x427049cc778aab35ull, 0x826bab125577d172ull, 0xd705b891b16237f5ull, 0xdf622fd44b180a87ull, 0x359398ecb79ec2deull, 0x2a1e075fb078da66ull, 0xd480ddd8e66d26e8ull, - 0x07e3e86f517de466ull, 0xa98b2f7423557445ull, 0x6bab15b36bb142feull, 0xf87d147bf2cc5c0bull, 0xf0294ea2b2820e03ull, 0xf219ac95e823d794ull, 0x9fa3fea85bc54264ull, 0xc4af3fafd2ff5201ull + 0xc93652d639480287ull, 0x2bff33a7119c0ec9ull, 0x7c676a1cf9f37474ull, 0xcb1dab3b216bdd52ull, 0x236552ac169e0f4aull, 0xffb7303a10cb2833ull, 0xdf1ca9f83cf0740cull, 0xe88fcc23bd9a0a9full, + 0x48118af81df77459ull, 0x49379fea0c36ec78ull, 0x931e8c0ca930cd35ull, 0xba3e0b2487710abfull, 0x56b607c6f3672398ull, 0xeb0215d7735482c3ull, 0x104b2405ae428d28ull, 0xa8c21e4eb7e2b744ull, + 0xbc41f10ffa4e4621ull, 0xc51ab216e6e0a339ull, 0x285cdde4bd22e712ull, 0x4f985d36d1302ebdull, 0x82563e5cd9a31b86ull, 0x9495c481dd399662ull, 0xec6111e88d79f207ull, 0x112bf6a954166121ull, + 0x1a6eda3d1a3846b6ull, 0xd25ebd17cbfaf07dull, 0xe552179181d4390full, 0xdd0129cf2d8db153ull, 0x0863f60becfbabedull, 0x1825c21e13698cecull, 0x20acd589e0408f6eull, 0x817a8413e24a68d4ull }, { // base nonce 1000000 - 0x501f772483fac0a3ull, 0x461363da2c1539e0ull, 0x750050674de592afull, 0x14ed105042cd912cull, 0xc6477310878614ebull, 0xe916b44e32e90e56ull, 0x531c92e69c2bdd73ull, 0x5bb0129ef7c3bd51ull, - 0x967ed91f7cbc7c79ull, 0x06fcc6a895b58b4aull, 0x3bdb29fbf93cbff3ull, 0xbed385172d2e6abeull, 0x921fc99ff4efac5eull, 0x6b0090bedd9f69c7ull, 0x0b105c18aaa53ab0ull, 0xf0c4203561f3b94dull, - 0x64abe33adccf7807ull, 0xd8f3a3b7e0242c04ull, 0x397da462f69fac5bull, 0xe6b6b32e467b71beull, 0x5318ac9e56278d04ull, 0xf7c5e348a4e1f5dbull, 0x79c994e6109646dfull, 0x9cdc42b0f6e82231ull, - 0x3f1160f02fd96ac7ull, 0xa69a23fc7058be21ull, 0xccde0a19bfdc25b6ull, 0xd3684a5b966f7497ull, 0x929db00b97a624f9ull, 0xfd14a40882d395a6ull, 0x49b0c1ecc514d6adull, 0x95d994c8349be12dull + 0x7535b29ea3e2823eull, 0xd84fe0e0281b6538ull, 0x14f4b14bf5185a26ull, 0x5fc4ec481d8c65bcull, 0x9ff6a4ec626c4cbcull, 0x52131ecd506117f0ull, 0x9a7db1822213f9b9ull, 0x025ac827f2f88c7cull, + 0x4076da8ca02131e1ull, 0xbc9956bc70d53e0bull, 0x077cf7d297357750ull, 0xb00c8db428fbccf2ull, 0x50f16eecd1fb65c4ull, 0x4daddb3cf455583dull, 0x952b3cca95e87c92ull, 0xa7b7af6eac1a0222ull, + 0xc58c0db5a99ced05ull, 0xa71a8ce697d65e94ull, 0xe29bab54459076d2ull, 0x5f613619a76c6400ull, 0xb43e9559e242a8d4ull, 0x4b5433e68aa1f302ull, 0x382b7f105840032cull, 0xbf402649fb8e9618ull, + 0xf45994bb29726a41ull, 0x8a11f358c24795cbull, 0xc2c4f8902007527cull, 0xe12f65396a832dcdull, 0x307d3f495790aff0ull, 0x5fc00c5eb0c3e81eull, 0xb46600e4685191eeull, 0xce65db0e1d36f875ull } }; diff --git a/proto-cuda/packs-readwidth/scr2/vectors.json b/proto-cuda/packs-readwidth/scr2k32/vectors.json similarity index 65% rename from proto-cuda/packs-readwidth/scr2/vectors.json rename to proto-cuda/packs-readwidth/scr2k32/vectors.json index d30ae7d93..c32644571 100644 --- a/proto-cuda/packs-readwidth/scr2/vectors.json +++ b/proto-cuda/packs-readwidth/scr2k32/vectors.json @@ -8,22 +8,22 @@ "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", "warps": [ {"base_nonce": 0, "expected": [ - "0x2273e2732203e32a", "0xa513d354bd107990", "0xe005b7515054c85f", "0x18a61b37b30cd1fb", "0xa21d5b98e8d07e9c", "0x5c24171a391d5ed0", "0x0f29e7583e1794b8", "0x7eca8a374d1a4f70", - "0x260ca011cd9ea10c", "0xce1050798fce3d43", "0x9560d939dca19041", "0x8480a8440b80ecc3", "0xfaa99aac459b739e", "0x7f083e72458e08ab", "0x78d876842f68672b", "0x3b9bcf6275d3575c", - "0x0256af61bdbf11b3", "0xefc6771cae646cbd", "0xbc44f1c9f9f54d87", "0x6caedd783487eb7d", "0x001b31fcb4fe0d4d", "0x947a7ba1057e25b6", "0xb9e5a0204d68c22a", "0x50489bed25d42661", - "0xb1019bff6d1057cd", "0xd1442990562ce940", "0xcd986a47f98801db", "0x9c8796b6df23300f", "0xbd53ad05d2c877a9", "0xc95e863774a15b0a", "0x132d8a91fb2fa67a", "0x53dd38e8eadc8a24" + "0x589f62cd61c27dc1", "0xf6be7bab8b00a7c7", "0x349551c5979e330e", "0x6d8156f5afaf0064", "0x81971552315d55b8", "0x20a3bb31ef5e202c", "0x91818c4ec9fd5fec", "0x35ca8cde74ac9715", + "0xc1db0a90f80c8801", "0xbc52521a33053a8a", "0x89fda590bf946dd6", "0xc54fe4aeaf11975f", "0xffc712960d7f4022", "0x4da6b3de6f9a0abd", "0xf70ff0468e9b595a", "0xc2d1d4434c1eb8c9", + "0x70dc278b36920857", "0xfb20a2d85fa65b83", "0xd24316c07937542d", "0xd505e1e3e85694bd", "0x88c759248ae122a7", "0xf1c04d8ad71d5e8c", "0x7e4d0e4e99717f12", "0x63d9e4b8f619ac98", + "0xcfef2a4243ba8608", "0xb2908c3da3593ba9", "0x48bee70b2ee45b84", "0x8f7a09d71be2c321", "0x0aadd48107e05bdb", "0x8b73ff0f34ddaaaa", "0x9ff3a4875ecf0fe3", "0xaf3d701cd1cbbd4b" ]}, {"base_nonce": 4096, "expected": [ - "0x78c93312a03fb0ee", "0x184fea638ec9b5fb", "0x5687e8dcc4301dbf", "0xed02c94f23681dfc", "0x326d70162241ff6d", "0x452017eb4ed2dfcf", "0xc10b0e016f1e28c9", "0x691ce0cecf2a99ba", - "0x6c9506f34e0e63ce", "0x447a98c2b7fdfa40", "0x07486b0e4055b2c9", "0x41781460bd47fd5c", "0x01db316e35198291", "0xccd7e727f139a880", "0xdd7bd9efd16bf21c", "0x8285d37966656366", - "0x383deade15fe0ecb", "0x5fd64f5873c8e324", "0xad584cb6839c5e1d", "0xbb842707fb5e9460", "0x4e8bc8f87978fcbd", "0x18eb56f4a1fae881", "0x4c3b731a6b0c47a1", "0xda52cf9d69b252eb", - "0xb5ff19b2b3eeb13e", "0xe2595cea2afe42dd", "0x3ff108424c9e6e38", "0x3a8a9e1995f359ca", "0x6a6b1da662cf2126", "0x54e684c127bb181f", "0x2018caa81f1a7d50", "0x33e94d2c92d9d148" + "0xc93652d639480287", "0x2bff33a7119c0ec9", "0x7c676a1cf9f37474", "0xcb1dab3b216bdd52", "0x236552ac169e0f4a", "0xffb7303a10cb2833", "0xdf1ca9f83cf0740c", "0xe88fcc23bd9a0a9f", + "0x48118af81df77459", "0x49379fea0c36ec78", "0x931e8c0ca930cd35", "0xba3e0b2487710abf", "0x56b607c6f3672398", "0xeb0215d7735482c3", "0x104b2405ae428d28", "0xa8c21e4eb7e2b744", + "0xbc41f10ffa4e4621", "0xc51ab216e6e0a339", "0x285cdde4bd22e712", "0x4f985d36d1302ebd", "0x82563e5cd9a31b86", "0x9495c481dd399662", "0xec6111e88d79f207", "0x112bf6a954166121", + "0x1a6eda3d1a3846b6", "0xd25ebd17cbfaf07d", "0xe552179181d4390f", "0xdd0129cf2d8db153", "0x0863f60becfbabed", "0x1825c21e13698cec", "0x20acd589e0408f6e", "0x817a8413e24a68d4" ]}, {"base_nonce": 1000000, "expected": [ - "0x04a41389bf3dfd3d", "0xc509164def9207df", "0x4a8ffdbdf46e429d", "0xff13bf0dc1b39aeb", "0xb852acc8e24133d7", "0x4bdd991ae56252ac", "0xa7739e74b3a054e9", "0xb4e36218d4b45fdc", - "0x8bbd323155f5edc5", "0xb7b56a90659e7fd2", "0xdff7c495b7027480", "0xffa8adb5c0302b06", "0xe97d7967d89a5672", "0x0d0c2d4e6493926e", "0xe9a5cda333cf2043", "0xdc95256d0986e5d8", - "0xd0dc211b811d6843", "0x68dfa3d0fb9a569b", "0xa9e0028dfd9178c0", "0x4a36ca1fc40b20a9", "0xe7c765c5a735294b", "0xf08954b015cb2628", "0xc69ee66ecf2740c5", "0xe3d01e899e46b089", - "0xc3558c74159c8603", "0x4c7aeb196bd01b04", "0x13c17119385f1910", "0xda7fca98e0989b8a", "0x6f95baf340817945", "0x1af52756fd3afcab", "0xb8eefc370bbe7e4b", "0x94a55055b48bd4db" + "0x7535b29ea3e2823e", "0xd84fe0e0281b6538", "0x14f4b14bf5185a26", "0x5fc4ec481d8c65bc", "0x9ff6a4ec626c4cbc", "0x52131ecd506117f0", "0x9a7db1822213f9b9", "0x025ac827f2f88c7c", + "0x4076da8ca02131e1", "0xbc9956bc70d53e0b", "0x077cf7d297357750", "0xb00c8db428fbccf2", "0x50f16eecd1fb65c4", "0x4daddb3cf455583d", "0x952b3cca95e87c92", "0xa7b7af6eac1a0222", + "0xc58c0db5a99ced05", "0xa71a8ce697d65e94", "0xe29bab54459076d2", "0x5f613619a76c6400", "0xb43e9559e242a8d4", "0x4b5433e68aa1f302", "0x382b7f105840032c", "0xbf402649fb8e9618", + "0xf45994bb29726a41", "0x8a11f358c24795cb", "0xc2c4f8902007527c", "0xe12f65396a832dcd", "0x307d3f495790aff0", "0x5fc00c5eb0c3e81e", "0xb46600e4685191ee", "0xce65db0e1d36f875" ]} ], "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], diff --git a/proto-cuda/packs-readwidth/scr8/kernel.cl b/proto-cuda/packs-readwidth/scr4k128/kernel.cl similarity index 78% rename from proto-cuda/packs-readwidth/scr8/kernel.cl rename to proto-cuda/packs-readwidth/scr4k128/kernel.cl index 50015ca9b..3917a927d 100644 --- a/proto-cuda/packs-readwidth/scr8/kernel.cl +++ b/proto-cuda/packs-readwidth/scr4k128/kernel.cl @@ -177,7 +177,7 @@ __kernel void igneum_build(__global uint* ds, __global const uint* cache, uint n // One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the // lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and // __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -185,7 +185,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint lane = (uint)get_global_id(0) & 31u; uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; - __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -214,7 +214,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add r4 = r0 * r6 + r4; // 3 mad - { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 4 scratch + r7 = r7 ^ ds[r2 & mask]; // 4 load r4 = r4 ^ ds[r1 & mask]; // 5 load { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl @@ -226,14 +226,14 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + { uint s_ = r6 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl r4 = r0 * r2 + r4; // 19 mad { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl r5 = r5 ^ r7; // 21 xor r2 = mul_hi(r2, r5); // 22 mulhi - { uint s_ = r7 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 23 scratch + r3 = r3 ^ ds[r7 & mask]; // 23 load r7 = mul_hi(r7, r3); // 24 mulhi r5 = r5 | r4; // 25 or r4 = r5 * r2 + r4; // 26 mad @@ -242,19 +242,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add r6 = rotr_var(r6, r7); // 30 rotr r3 = r3 ^ ds[r1 & mask]; // 31 load - { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + { uint s_ = r0 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + { uint s_ = r2 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor - { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 37 scratch + r4 = r4 ^ ds[r0 & mask]; // 37 load r1 = r3 * r5 + r1; // 38 mad r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add r2 = rotr_var(r2, r5); // 40 rotr r3 = r3 * r2; // 41 mul r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add r3 = r3 ^ r4; // 43 xor - { uint s_ = r5 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 44 scratch + r3 = r3 ^ ds[r5 & mask]; // 44 load r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add r7 = r7 ^ r1; // 46 xor r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add @@ -269,7 +269,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r2 = r2 ^ ds[r7 & mask]; // 56 load r5 = r5 - r6; // 57 sub r1 = r1 ^ ds[r3 & mask]; // 58 load - { uint s_ = r4 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + { uint s_ = r4 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch r4 = r4 - r6; // 60 sub r1 = r1 * r2; // 61 mul r3 = r6 * r0 + r3; // 62 mad diff --git a/proto-cuda/packs-readwidth/scr8/kernel.cu b/proto-cuda/packs-readwidth/scr4k128/kernel.cu similarity index 66% rename from proto-cuda/packs-readwidth/scr8/kernel.cu rename to proto-cuda/packs-readwidth/scr4k128/kernel.cu index 85212204b..67b3ff7d8 100644 --- a/proto-cuda/packs-readwidth/scr8/kernel.cu +++ b/proto-cuda/packs-readwidth/scr4k128/kernel.cu @@ -44,7 +44,7 @@ __global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItem // One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every // __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a // 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. __device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -52,7 +52,7 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; - uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { uint32_t gid = g_ * 32u + lane; uint32_t gbase = baseNonce + g_ * 32u; @@ -74,7 +74,7 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add r4 = r0 * r6 + r4; // 3 mad - { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 scratch + r7 = r7 ^ ds[r2 & mask]; // 4 load r4 = r4 ^ ds[r1 & mask]; // 5 load r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl @@ -86,14 +86,14 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint32_t s_ = r6 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + { uint32_t s_ = r6 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl r4 = r0 * r2 + r4; // 19 mad r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl r5 = r5 ^ r7; // 21 xor r2 = __umulhi(r2, r5); // 22 mulhi - { uint32_t s_ = r7 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 scratch + r3 = r3 ^ ds[r7 & mask]; // 23 load r7 = __umulhi(r7, r3); // 24 mulhi r5 = r5 | r4; // 25 or r4 = r5 * r2 + r4; // 26 mad @@ -102,19 +102,19 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add r6 = rotr_var(r6, r7); // 30 rotr r3 = r3 ^ ds[r1 & mask]; // 31 load - { uint32_t s_ = r0 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + { uint32_t s_ = r0 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + { uint32_t s_ = r2 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor - { uint32_t s_ = r0 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 scratch + r4 = r4 ^ ds[r0 & mask]; // 37 load r1 = r3 * r5 + r1; // 38 mad r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add r2 = rotr_var(r2, r5); // 40 rotr r3 = r3 * r2; // 41 mul r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add r3 = r3 ^ r4; // 43 xor - { uint32_t s_ = r5 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 scratch + r3 = r3 ^ ds[r5 & mask]; // 44 load r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add r7 = r7 ^ r1; // 46 xor r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add @@ -129,7 +129,7 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc r2 = r2 ^ ds[r7 & mask]; // 56 load r5 = r5 - r6; // 57 sub r1 = r1 ^ ds[r3 & mask]; // 58 load - { uint32_t s_ = r4 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + { uint32_t s_ = r4 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch r4 = r4 - r6; // 60 sub r1 = r1 * r2; // 61 mul r3 = r6 * r0 + r3; // 62 mad diff --git a/proto-cuda/packs-readwidth/scr8/kernel_bound.cl b/proto-cuda/packs-readwidth/scr4k128/kernel_bound.cl similarity index 70% rename from proto-cuda/packs-readwidth/scr8/kernel_bound.cl rename to proto-cuda/packs-readwidth/scr4k128/kernel_bound.cl index 73b2b9c9f..669eb1d51 100644 --- a/proto-cuda/packs-readwidth/scr8/kernel_bound.cl +++ b/proto-cuda/packs-readwidth/scr4k128/kernel_bound.cl @@ -177,7 +177,7 @@ __kernel void igneum_build(__global uint* ds, __global const uint* cache, uint n // One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the // lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and // __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -185,7 +185,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint lane = (uint)get_global_id(0) & 31u; uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; - __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -214,7 +214,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add r4 = r0 * r6 + r4; // 3 mad - { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 4 scratch + r7 = r7 ^ ds[r2 & mask]; // 4 load r4 = r4 ^ ds[r1 & mask]; // 5 load { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl @@ -226,14 +226,14 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + { uint s_ = r6 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl r4 = r0 * r2 + r4; // 19 mad { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl r5 = r5 ^ r7; // 21 xor r2 = mul_hi(r2, r5); // 22 mulhi - { uint s_ = r7 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 23 scratch + r3 = r3 ^ ds[r7 & mask]; // 23 load r7 = mul_hi(r7, r3); // 24 mulhi r5 = r5 | r4; // 25 or r4 = r5 * r2 + r4; // 26 mad @@ -242,19 +242,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add r6 = rotr_var(r6, r7); // 30 rotr r3 = r3 ^ ds[r1 & mask]; // 31 load - { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + { uint s_ = r0 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + { uint s_ = r2 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor - { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 37 scratch + r4 = r4 ^ ds[r0 & mask]; // 37 load r1 = r3 * r5 + r1; // 38 mad r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add r2 = rotr_var(r2, r5); // 40 rotr r3 = r3 * r2; // 41 mul r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add r3 = r3 ^ r4; // 43 xor - { uint s_ = r5 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 44 scratch + r3 = r3 ^ ds[r5 & mask]; // 44 load r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add r7 = r7 ^ r1; // 46 xor r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add @@ -269,7 +269,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out r2 = r2 ^ ds[r7 & mask]; // 56 load r5 = r5 - r6; // 57 sub r1 = r1 ^ ds[r3 & mask]; // 58 load - { uint s_ = r4 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + { uint s_ = r4 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch r4 = r4 - r6; // 60 sub r1 = r1 * r2; // 61 mul r3 = r6 * r0 + r3; // 62 mad @@ -295,7 +295,7 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon uint lane = (uint)get_global_id(0) & 31u; uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; - __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -325,7 +325,7 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add r4 = r0 * r6 + r4; // 3 mad - { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 4 scratch + r7 = r7 ^ ds[r2 & mask]; // 4 load r4 = r4 ^ ds[r1 & mask]; // 5 load { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl @@ -337,14 +337,14 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint s_ = r6 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + { uint s_ = r6 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl r4 = r0 * r2 + r4; // 19 mad { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl r5 = r5 ^ r7; // 21 xor r2 = mul_hi(r2, r5); // 22 mulhi - { uint s_ = r7 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 23 scratch + r3 = r3 ^ ds[r7 & mask]; // 23 load r7 = mul_hi(r7, r3); // 24 mulhi r5 = r5 | r4; // 25 or r4 = r5 * r2 + r4; // 26 mad @@ -353,19 +353,19 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add r6 = rotr_var(r6, r7); // 30 rotr r3 = r3 ^ ds[r1 & mask]; // 31 load - { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + { uint s_ = r0 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint s_ = r2 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + { uint s_ = r2 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor - { uint s_ = r0 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 37 scratch + r4 = r4 ^ ds[r0 & mask]; // 37 load r1 = r3 * r5 + r1; // 38 mad r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add r2 = rotr_var(r2, r5); // 40 rotr r3 = r3 * r2; // 41 mul r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add r3 = r3 ^ r4; // 43 xor - { uint s_ = r5 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 44 scratch + r3 = r3 ^ ds[r5 & mask]; // 44 load r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add r7 = r7 ^ r1; // 46 xor r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add @@ -380,7 +380,7 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon r2 = r2 ^ ds[r7 & mask]; // 56 load r5 = r5 - r6; // 57 sub r1 = r1 ^ ds[r3 & mask]; // 58 load - { uint s_ = r4 & 2047u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + { uint s_ = r4 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch r4 = r4 - r6; // 60 sub r1 = r1 * r2; // 61 mul r3 = r6 * r0 + r3; // 62 mad diff --git a/proto-cuda/packs-readwidth/scr8/kernel_bound.cu b/proto-cuda/packs-readwidth/scr4k128/kernel_bound.cu similarity index 60% rename from proto-cuda/packs-readwidth/scr8/kernel_bound.cu rename to proto-cuda/packs-readwidth/scr4k128/kernel_bound.cu index 13d40a9fc..9197cfce9 100644 --- a/proto-cuda/packs-readwidth/scr8/kernel_bound.cu +++ b/proto-cuda/packs-readwidth/scr4k128/kernel_bound.cu @@ -20,7 +20,7 @@ __device__ __forceinline__ uint32_t splitmix32(uint32_t x) { __device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } __device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. __device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -28,7 +28,7 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; - uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { uint32_t gid = g_ * 32u + lane; uint32_t gbase = baseNonce + g_ * 32u; @@ -50,7 +50,7 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add r4 = r0 * r6 + r4; // 3 mad - { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 scratch + r7 = r7 ^ ds[r2 & mask]; // 4 load r4 = r4 ^ ds[r1 & mask]; // 5 load r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl @@ -62,14 +62,14 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba r2 = r2 * r5; // 13 mul r1 = r1 ^ ds[r2 & mask]; // 14 load r7 = rotl_imm(r7, 1u); // 15 rotl - { uint32_t s_ = r6 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + { uint32_t s_ = r6 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch r7 = r7 ^ ds[r4 & mask]; // 17 load r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl r4 = r0 * r2 + r4; // 19 mad r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl r5 = r5 ^ r7; // 21 xor r2 = __umulhi(r2, r5); // 22 mulhi - { uint32_t s_ = r7 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 scratch + r3 = r3 ^ ds[r7 & mask]; // 23 load r7 = __umulhi(r7, r3); // 24 mulhi r5 = r5 | r4; // 25 or r4 = r5 * r2 + r4; // 26 mad @@ -78,19 +78,19 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add r6 = rotr_var(r6, r7); // 30 rotr r3 = r3 ^ ds[r1 & mask]; // 31 load - { uint32_t s_ = r0 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + { uint32_t s_ = r0 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add - { uint32_t s_ = r2 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + { uint32_t s_ = r2 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch r0 = r0 * r3; // 35 mul r2 = r2 ^ r5; // 36 xor - { uint32_t s_ = r0 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 scratch + r4 = r4 ^ ds[r0 & mask]; // 37 load r1 = r3 * r5 + r1; // 38 mad r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add r2 = rotr_var(r2, r5); // 40 rotr r3 = r3 * r2; // 41 mul r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add r3 = r3 ^ r4; // 43 xor - { uint32_t s_ = r5 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 scratch + r3 = r3 ^ ds[r5 & mask]; // 44 load r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add r7 = r7 ^ r1; // 46 xor r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add @@ -105,7 +105,7 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba r2 = r2 ^ ds[r7 & mask]; // 56 load r5 = r5 - r6; // 57 sub r1 = r1 ^ ds[r3 & mask]; // 58 load - { uint32_t s_ = r4 & 2047u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + { uint32_t s_ = r4 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch r4 = r4 - r6; // 60 sub r1 = r1 * r2; // 61 mul r3 = r6 * r0 + r3; // 62 mad diff --git a/proto-cuda/packs-readwidth/scr8/memhard.h b/proto-cuda/packs-readwidth/scr4k128/memhard.h similarity index 100% rename from proto-cuda/packs-readwidth/scr8/memhard.h rename to proto-cuda/packs-readwidth/scr4k128/memhard.h diff --git a/proto-cuda/packs-readwidth/scr8/memhard.metal b/proto-cuda/packs-readwidth/scr4k128/memhard.metal similarity index 100% rename from proto-cuda/packs-readwidth/scr8/memhard.metal rename to proto-cuda/packs-readwidth/scr4k128/memhard.metal diff --git a/proto-cuda/packs-readwidth/scr4k128/program.h b/proto-cuda/packs-readwidth/scr4k128/program.h new file mode 100644 index 000000000..ee7cc85d0 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k128/program.h @@ -0,0 +1,67 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-genesis" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0xe0c444d155cd78cfull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "load=12 add=9 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 scratch=4 shfl=4 rotr=2 rotl=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "scr4k128" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 100, 0, 0 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 12, 0, 0 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 384 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// Variant 5: persistent warps, a 128 KiB scratch per launched warp (the host launches N warps and passes scratch, +// groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). +#define IGNEUM_PERSISTENT_WARPS 1 +#define IGNEUM_SCRATCH_OPS 4 // scratch read-modify-writes per program (32 per hash) +#define IGNEUM_SCRATCH_SLOTS 256u +#define IGNEUM_SCRATCH_WORDS_PER_LANE 1024u +#define IGNEUM_SCRATCH_BYTES_PER_WARP 131072u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/scr4/program.json b/proto-cuda/packs-readwidth/scr4k128/program.json similarity index 98% rename from proto-cuda/packs-readwidth/scr4/program.json rename to proto-cuda/packs-readwidth/scr4k128/program.json index 61fe86581..4d0d0ad78 100644 --- a/proto-cuda/packs-readwidth/scr4/program.json +++ b/proto-cuda/packs-readwidth/scr4k128/program.json @@ -2,7 +2,7 @@ "format": "igneum-program-pack-3", "generator": 2, "attempt": 0, - "program_id": "0x2f098ee568f386f5", + "program_id": "0xe0c444d155cd78cf", "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", "dataset_mode": "memory-hard", "seed": "igneum-genesis", @@ -15,13 +15,14 @@ "iterations": 8, "instruction_count": 64, "loads_per_hash": 128, - "load_class": "scr4", + "load_class": "scr4k128", "load_slots": 16, "load_mix_percent_4_16_64": [100, 0, 0], "load_width_counts_4_16_64": [12, 0, 0], "bytes_per_hash": 384, "scratch_ops_per_hash": 32, - "scratch": "variant 5 (measurement only): persistent warps; a 1 MiB scratch per warp of 2048 16-byte slots per lane (lane-major); slot = src & 0x7ff; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", + "scratch_kib_per_warp": 128, + "scratch": "variant 5 (measurement only): persistent warps; a 128 KiB scratch per warp of 256 16-byte slots per lane (lane-major); slot = src & 0xff; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", "op_mix": {"load": 12, "add": 9, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "scratch": 4, "shfl": 4, "rotr": 2, "rotl": 1}, "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", diff --git a/proto-cuda/packs-readwidth/scr4k128/program.metal b/proto-cuda/packs-readwidth/scr4k128/program.metal new file mode 100644 index 000000000..f4f56728c --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k128/program.metal @@ -0,0 +1,126 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + device uint* scratch [[buffer(3)]], + constant uint& groups [[buffer(4)]], + constant uint& salt [[buffer(5)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + r7 = r7 ^ dataset[r2 & MASK]; // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint s_ = r6 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + r3 = r3 ^ dataset[r7 & MASK]; // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + { uint s_ = r0 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint s_ = r2 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + r4 = r4 ^ dataset[r0 & MASK]; // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + r3 = r3 ^ dataset[r5 & MASK]; // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + { uint s_ = r4 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr4k128/program_bound.metal b/proto-cuda/packs-readwidth/scr4k128/program_bound.metal new file mode 100644 index 000000000..3257d287e --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k128/program_bound.metal @@ -0,0 +1,128 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + device uint* scratch [[buffer(4)]], + constant uint& groups [[buffer(5)]], + constant uint& salt [[buffer(6)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + r7 = r7 ^ dataset[r2 & MASK]; // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint s_ = r6 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + r3 = r3 ^ dataset[r7 & MASK]; // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + { uint s_ = r0 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint s_ = r2 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + r4 = r4 ^ dataset[r0 & MASK]; // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + r3 = r3 ^ dataset[r5 & MASK]; // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + { uint s_ = r4 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr2/vectors.h b/proto-cuda/packs-readwidth/scr4k128/vectors.h similarity index 60% rename from proto-cuda/packs-readwidth/scr2/vectors.h rename to proto-cuda/packs-readwidth/scr4k128/vectors.h index 820c17e88..81688d9cc 100644 --- a/proto-cuda/packs-readwidth/scr2/vectors.h +++ b/proto-cuda/packs-readwidth/scr4k128/vectors.h @@ -11,22 +11,22 @@ static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { { // base nonce 0 - 0x2273e2732203e32aull, 0xa513d354bd107990ull, 0xe005b7515054c85full, 0x18a61b37b30cd1fbull, 0xa21d5b98e8d07e9cull, 0x5c24171a391d5ed0ull, 0x0f29e7583e1794b8ull, 0x7eca8a374d1a4f70ull, - 0x260ca011cd9ea10cull, 0xce1050798fce3d43ull, 0x9560d939dca19041ull, 0x8480a8440b80ecc3ull, 0xfaa99aac459b739eull, 0x7f083e72458e08abull, 0x78d876842f68672bull, 0x3b9bcf6275d3575cull, - 0x0256af61bdbf11b3ull, 0xefc6771cae646cbdull, 0xbc44f1c9f9f54d87ull, 0x6caedd783487eb7dull, 0x001b31fcb4fe0d4dull, 0x947a7ba1057e25b6ull, 0xb9e5a0204d68c22aull, 0x50489bed25d42661ull, - 0xb1019bff6d1057cdull, 0xd1442990562ce940ull, 0xcd986a47f98801dbull, 0x9c8796b6df23300full, 0xbd53ad05d2c877a9ull, 0xc95e863774a15b0aull, 0x132d8a91fb2fa67aull, 0x53dd38e8eadc8a24ull + 0xd48ade5043a1a440ull, 0x0236be04051c86abull, 0x4a368a9d6ce0d6eaull, 0x58518f225df7cd40ull, 0xa82cf7b3417df902ull, 0x4f2722d25e5aa401ull, 0xf94296c41a24bdb4ull, 0x4fca15537288398bull, + 0x9ba276d89178f795ull, 0x7205ccbf3772e4e5ull, 0x4149e0fbf1c35bb9ull, 0x91d484090f0b09e0ull, 0xb95010c64c2d27dbull, 0x2b9fc3c52f733771ull, 0x1f2f6dd44045e10cull, 0xb78170533d4b16a6ull, + 0x679056d6c824110full, 0x0f81ba5b1476c314ull, 0xeb3d83ce6ccd9cd2ull, 0x370575fe0b0e7119ull, 0x7b4ae5a0f119315full, 0x12ceac820c840cfcull, 0xd190a33169dd6d61ull, 0xf7f90f768fc7ec83ull, + 0x90f11a170fce1e73ull, 0xca168b60b8a41d61ull, 0xab9da4f155dc3d4cull, 0x884d93725ea8fc2full, 0x202bf861848ba637ull, 0x7d508d345e67589eull, 0x8dfcbd9365f96ddaull, 0x5e8315d79af30be3ull }, { // base nonce 4096 - 0x78c93312a03fb0eeull, 0x184fea638ec9b5fbull, 0x5687e8dcc4301dbfull, 0xed02c94f23681dfcull, 0x326d70162241ff6dull, 0x452017eb4ed2dfcfull, 0xc10b0e016f1e28c9ull, 0x691ce0cecf2a99baull, - 0x6c9506f34e0e63ceull, 0x447a98c2b7fdfa40ull, 0x07486b0e4055b2c9ull, 0x41781460bd47fd5cull, 0x01db316e35198291ull, 0xccd7e727f139a880ull, 0xdd7bd9efd16bf21cull, 0x8285d37966656366ull, - 0x383deade15fe0ecbull, 0x5fd64f5873c8e324ull, 0xad584cb6839c5e1dull, 0xbb842707fb5e9460ull, 0x4e8bc8f87978fcbdull, 0x18eb56f4a1fae881ull, 0x4c3b731a6b0c47a1ull, 0xda52cf9d69b252ebull, - 0xb5ff19b2b3eeb13eull, 0xe2595cea2afe42ddull, 0x3ff108424c9e6e38ull, 0x3a8a9e1995f359caull, 0x6a6b1da662cf2126ull, 0x54e684c127bb181full, 0x2018caa81f1a7d50ull, 0x33e94d2c92d9d148ull + 0x0925cd0a405af0f8ull, 0x9eeae6619738a6a0ull, 0x69d83343d36fa441ull, 0xef6e0dda67f22db7ull, 0xf5a5baddc99fc6e8ull, 0x048243c3a33d6313ull, 0xa2ed984433185d72ull, 0x5ce6444f5132231eull, + 0xf9c92e489ee1479bull, 0x13df97d418a1bb1dull, 0x54d8a4aa14eb5bf2ull, 0xc93ebe91c3aa1860ull, 0x12e1ca6f27af870bull, 0xa37cd8c938ec675bull, 0x0085b0d9144040d0ull, 0x389ce17c45d36ec5ull, + 0xc82d6694437e2f54ull, 0xf5a7b7357bc34eb6ull, 0x5e8e9c4cdaedb41dull, 0x18bff888b957603aull, 0x661b918790cf28f7ull, 0xf1411538bea6c80full, 0x0b3f28dd15dd2a0full, 0x076cd4e3230c6857ull, + 0x2cf5028d8fe8a18full, 0x270327a0333a8520ull, 0x53461f279f163486ull, 0x834a13f157378136ull, 0x5f8609aa7fbda1aeull, 0xd57be373dec55a73ull, 0xdf6b4c6134656905ull, 0x607a5b7e2a33765eull }, { // base nonce 1000000 - 0x04a41389bf3dfd3dull, 0xc509164def9207dfull, 0x4a8ffdbdf46e429dull, 0xff13bf0dc1b39aebull, 0xb852acc8e24133d7ull, 0x4bdd991ae56252acull, 0xa7739e74b3a054e9ull, 0xb4e36218d4b45fdcull, - 0x8bbd323155f5edc5ull, 0xb7b56a90659e7fd2ull, 0xdff7c495b7027480ull, 0xffa8adb5c0302b06ull, 0xe97d7967d89a5672ull, 0x0d0c2d4e6493926eull, 0xe9a5cda333cf2043ull, 0xdc95256d0986e5d8ull, - 0xd0dc211b811d6843ull, 0x68dfa3d0fb9a569bull, 0xa9e0028dfd9178c0ull, 0x4a36ca1fc40b20a9ull, 0xe7c765c5a735294bull, 0xf08954b015cb2628ull, 0xc69ee66ecf2740c5ull, 0xe3d01e899e46b089ull, - 0xc3558c74159c8603ull, 0x4c7aeb196bd01b04ull, 0x13c17119385f1910ull, 0xda7fca98e0989b8aull, 0x6f95baf340817945ull, 0x1af52756fd3afcabull, 0xb8eefc370bbe7e4bull, 0x94a55055b48bd4dbull + 0x16b8e21167f3437cull, 0xfd28ad2d0f75a03cull, 0xf8e70bdb604cfff7ull, 0xeba037043c5ece4bull, 0xa2cb7d31f4d25317ull, 0xc3e8b85a50bdab1dull, 0xd7bd65a4353eab2eull, 0x23540281fae8cce3ull, + 0x37f4deac1a67cc5full, 0xd482f81bec2535a5ull, 0xc18f3f46f812b870ull, 0x582514aab0cf566dull, 0xb3b1a7424a758ac6ull, 0x83bbf70ed4151fa4ull, 0x72e2fed205f44a00ull, 0x2f81d15c1a8e17feull, + 0xbe7875e7927ca851ull, 0x1a72a20d292cc17bull, 0x859dd2c75675a04bull, 0xe12711e4d81b1e04ull, 0xfafefd6afdda6b35ull, 0x30ebb4d12f5cf4e3ull, 0x6d4aae24eee724a2ull, 0x317d83e64bf5e9c6ull, + 0x2beba8ecb8b0b26eull, 0x5cf2eacb7a58bd99ull, 0x2d56441aac88a037ull, 0x01708cc58adfeb95ull, 0xb3bc095b95418a2full, 0xf804c273322638e1ull, 0x9d89d48818056f24ull, 0xe07cbfd9aa54ccf1ull } }; diff --git a/proto-cuda/packs-readwidth/scr4/vectors.json b/proto-cuda/packs-readwidth/scr4k128/vectors.json similarity index 65% rename from proto-cuda/packs-readwidth/scr4/vectors.json rename to proto-cuda/packs-readwidth/scr4k128/vectors.json index a0f951221..c394f4c58 100644 --- a/proto-cuda/packs-readwidth/scr4/vectors.json +++ b/proto-cuda/packs-readwidth/scr4k128/vectors.json @@ -8,22 +8,22 @@ "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", "warps": [ {"base_nonce": 0, "expected": [ - "0x66cffcc97c46e625", "0xbc9019f8df50fbfd", "0x65629c90dde6016e", "0xd647a41effa03d3b", "0x86da3b6bbd751b99", "0x6ccf4240a0fb2d19", "0xeb39a1e06f17378c", "0x2ea6b349b289fb10", - "0x6067211e6c220500", "0x6e6095dfedd1360f", "0xbd1190d8b50e1b48", "0x216dc72a0c08d5b5", "0x5be1f8c080836b0c", "0x2a32932a5953ed73", "0xcc2a3d68be83c802", "0xe46daca15338278f", - "0xb2de43b96761e459", "0x9004acd06588cbea", "0x6a9a3543cf93004f", "0xff956d859cb6e408", "0x4397ec6e3c5fb045", "0x521dea569cd481d5", "0x89832b34108759f0", "0xf66e393836ffe4ea", - "0xb4e39af6c40ea2f4", "0x3adc22085dd8d648", "0x27efe270958bbfbb", "0x6c80be0e8dca60d8", "0xa0afbc6a60260d59", "0x5d9a257fb9189537", "0xeb837aeef55dc3ed", "0xd381174dc14f8951" + "0xd48ade5043a1a440", "0x0236be04051c86ab", "0x4a368a9d6ce0d6ea", "0x58518f225df7cd40", "0xa82cf7b3417df902", "0x4f2722d25e5aa401", "0xf94296c41a24bdb4", "0x4fca15537288398b", + "0x9ba276d89178f795", "0x7205ccbf3772e4e5", "0x4149e0fbf1c35bb9", "0x91d484090f0b09e0", "0xb95010c64c2d27db", "0x2b9fc3c52f733771", "0x1f2f6dd44045e10c", "0xb78170533d4b16a6", + "0x679056d6c824110f", "0x0f81ba5b1476c314", "0xeb3d83ce6ccd9cd2", "0x370575fe0b0e7119", "0x7b4ae5a0f119315f", "0x12ceac820c840cfc", "0xd190a33169dd6d61", "0xf7f90f768fc7ec83", + "0x90f11a170fce1e73", "0xca168b60b8a41d61", "0xab9da4f155dc3d4c", "0x884d93725ea8fc2f", "0x202bf861848ba637", "0x7d508d345e67589e", "0x8dfcbd9365f96dda", "0x5e8315d79af30be3" ]}, {"base_nonce": 4096, "expected": [ - "0xfb1f61aaeaeaef94", "0x184c2160963d8b57", "0x42a053c628625778", "0xeaad0e41c4770812", "0x1d6d389ceb462ce1", "0x4a639827672bdbd4", "0x857e42aa5a42dd6f", "0xb4ef399e5339979c", - "0x497a29225b099233", "0x71d8b42862d81954", "0x0af995663313bf04", "0xf436fd126619d7a1", "0x199e4e3333cff269", "0x64077952f3775768", "0x51af1d126c5e8388", "0xffbaf44fe6b15cfd", - "0xfc8fed86ecae34a7", "0x4cb548616f7a7d6b", "0xc21d938c8b5bef35", "0x34789cbdd7088f71", "0xacb099a2c207d891", "0xfe1902d162374413", "0x26f7831c28f4020b", "0xdf5192952b4af6b0", - "0xecab61fe88dbaff4", "0x941c491f7fdb86e5", "0x2b1900c53f746e77", "0x8c40507b1caffeb2", "0x7532a1ec2b9169ef", "0x1cf399b0c8bfb520", "0xdf003d2bb8a2cc0c", "0x4da853307fc977a9" + "0x0925cd0a405af0f8", "0x9eeae6619738a6a0", "0x69d83343d36fa441", "0xef6e0dda67f22db7", "0xf5a5baddc99fc6e8", "0x048243c3a33d6313", "0xa2ed984433185d72", "0x5ce6444f5132231e", + "0xf9c92e489ee1479b", "0x13df97d418a1bb1d", "0x54d8a4aa14eb5bf2", "0xc93ebe91c3aa1860", "0x12e1ca6f27af870b", "0xa37cd8c938ec675b", "0x0085b0d9144040d0", "0x389ce17c45d36ec5", + "0xc82d6694437e2f54", "0xf5a7b7357bc34eb6", "0x5e8e9c4cdaedb41d", "0x18bff888b957603a", "0x661b918790cf28f7", "0xf1411538bea6c80f", "0x0b3f28dd15dd2a0f", "0x076cd4e3230c6857", + "0x2cf5028d8fe8a18f", "0x270327a0333a8520", "0x53461f279f163486", "0x834a13f157378136", "0x5f8609aa7fbda1ae", "0xd57be373dec55a73", "0xdf6b4c6134656905", "0x607a5b7e2a33765e" ]}, {"base_nonce": 1000000, "expected": [ - "0x3d094bd04694b96f", "0xaecbd76cecd1a20a", "0xbcb86febe56b17fe", "0x98082b557ba97517", "0xbb5f94108888564b", "0xea3284877a30fc87", "0xc608fa4d5a8bb2ad", "0x946c721c511e0729", - "0x46c6eba292083aed", "0x936cb97231eb6795", "0xb1413c434c712cbb", "0xedfd554d3948c1bd", "0xa8a20cbef2faccd5", "0x5fe39d756cadbcad", "0x208b2627380791fe", "0xf52f9374ce480218", - "0xb9db7cd8814eb29e", "0xf32ed2192b5a8719", "0x4f1b06a054940aef", "0x406df498e4365eb5", "0x1982075caad345ef", "0x590f725623dbbbbd", "0xa26d9192dedfefa5", "0x36219ec00da18980", - "0x5361d0dcb0f8b1a3", "0x35bdefa2fbb5ffc3", "0xba4c2a4e473a9c80", "0x107d9d3030f8b9d3", "0xa8bb094266d6b987", "0x86164fdfbb1426e8", "0xa6e8cb895021cbbd", "0xfe12809e9d99a243" + "0x16b8e21167f3437c", "0xfd28ad2d0f75a03c", "0xf8e70bdb604cfff7", "0xeba037043c5ece4b", "0xa2cb7d31f4d25317", "0xc3e8b85a50bdab1d", "0xd7bd65a4353eab2e", "0x23540281fae8cce3", + "0x37f4deac1a67cc5f", "0xd482f81bec2535a5", "0xc18f3f46f812b870", "0x582514aab0cf566d", "0xb3b1a7424a758ac6", "0x83bbf70ed4151fa4", "0x72e2fed205f44a00", "0x2f81d15c1a8e17fe", + "0xbe7875e7927ca851", "0x1a72a20d292cc17b", "0x859dd2c75675a04b", "0xe12711e4d81b1e04", "0xfafefd6afdda6b35", "0x30ebb4d12f5cf4e3", "0x6d4aae24eee724a2", "0x317d83e64bf5e9c6", + "0x2beba8ecb8b0b26e", "0x5cf2eacb7a58bd99", "0x2d56441aac88a037", "0x01708cc58adfeb95", "0xb3bc095b95418a2f", "0xf804c273322638e1", "0x9d89d48818056f24", "0xe07cbfd9aa54ccf1" ]} ], "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], diff --git a/proto-cuda/packs-readwidth/scr4k32/kernel.cl b/proto-cuda/packs-readwidth/scr4k32/kernel.cl new file mode 100644 index 000000000..00f785968 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k32/kernel.cl @@ -0,0 +1,291 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/scr4k32/kernel.cu b/proto-cuda/packs-readwidth/scr4k32/kernel.cu new file mode 100644 index 000000000..d5a4717ed --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k32/kernel.cu @@ -0,0 +1,177 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t s_ = r6 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint32_t s_ = r0 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t s_ = r2 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint32_t s_ = r4 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr4k32/kernel_bound.cl b/proto-cuda/packs-readwidth/scr4k32/kernel_bound.cl new file mode 100644 index 000000000..46034e31b --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k32/kernel_bound.cl @@ -0,0 +1,393 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr4k32/kernel_bound.cu b/proto-cuda/packs-readwidth/scr4k32/kernel_bound.cu new file mode 100644 index 000000000..dc5ec90df --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k32/kernel_bound.cu @@ -0,0 +1,136 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + r7 = r7 ^ ds[r2 & mask]; // 4 load + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t s_ = r6 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + r3 = r3 ^ ds[r7 & mask]; // 23 load + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint32_t s_ = r0 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t s_ = r2 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + r4 = r4 ^ ds[r0 & mask]; // 37 load + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + r3 = r3 ^ ds[r5 & mask]; // 44 load + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint32_t s_ = r4 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr4k32/memhard.h b/proto-cuda/packs-readwidth/scr4k32/memhard.h new file mode 100644 index 000000000..4803d8e40 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k32/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/scr4k32/memhard.metal b/proto-cuda/packs-readwidth/scr4k32/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k32/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/scr4/program.h b/proto-cuda/packs-readwidth/scr4k32/program.h similarity index 91% rename from proto-cuda/packs-readwidth/scr4/program.h rename to proto-cuda/packs-readwidth/scr4k32/program.h index f8f7a6823..869816f8c 100644 --- a/proto-cuda/packs-readwidth/scr4/program.h +++ b/proto-cuda/packs-readwidth/scr4k32/program.h @@ -15,7 +15,7 @@ #define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" #define IGNEUM_GENERATOR 2 #define IGNEUM_PROGRAM_ATTEMPT 0 -#define IGNEUM_PROGRAM_ID 0x2f098ee568f386f5ull +#define IGNEUM_PROGRAM_ID 0xe0c4a4d155ce1befull #define IGNEUM_DAY_STRING "2026-10-03" #define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" #define IGNEUM_DAY0 0x3067619fu @@ -30,20 +30,20 @@ #define IGNEUM_OP_MIX "load=12 add=9 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 scratch=4 shfl=4 rotr=2 rotl=1" // Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads // the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. -#define IGNEUM_LOAD_CLASS "scr4" +#define IGNEUM_LOAD_CLASS "scr4k32" #define IGNEUM_LOAD_SLOTS 16 #define IGNEUM_LOAD_MIX { 100, 0, 0 } #define IGNEUM_LOAD_WIDTH_COUNTS { 12, 0, 0 } // loads of 4, 16, 64 bytes per program #define IGNEUM_BYTES_PER_HASH 384 #define IGNEUM_FOLD_ROT 11 #define IGNEUM_FOLD_MUL 0x9e3779b1u -// Variant 5: persistent warps, a 1 MiB scratch per launched warp (the host launches N warps and passes scratch, +// Variant 5: persistent warps, a 32 KiB scratch per launched warp (the host launches N warps and passes scratch, // groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). #define IGNEUM_PERSISTENT_WARPS 1 #define IGNEUM_SCRATCH_OPS 4 // scratch read-modify-writes per program (32 per hash) -#define IGNEUM_SCRATCH_SLOTS 2048u -#define IGNEUM_SCRATCH_WORDS_PER_LANE 8192u -#define IGNEUM_SCRATCH_BYTES_PER_WARP 1048576u +#define IGNEUM_SCRATCH_SLOTS 64u +#define IGNEUM_SCRATCH_WORDS_PER_LANE 256u +#define IGNEUM_SCRATCH_BYTES_PER_WARP 32768u // 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) #define IGNEUM_DATASET_MODE 1 diff --git a/proto-cuda/packs-readwidth/scr4k32/program.json b/proto-cuda/packs-readwidth/scr4k32/program.json new file mode 100644 index 000000000..9bd6cf0f5 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k32/program.json @@ -0,0 +1,130 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0xe0c4a4d155ce1bef", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-genesis", + "seed_bytes": "69676e65756d2d67656e65736973", + "seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "scr4k32", + "load_slots": 16, + "load_mix_percent_4_16_64": [100, 0, 0], + "load_width_counts_4_16_64": [12, 0, 0], + "bytes_per_hash": 384, + "scratch_ops_per_hash": 32, + "scratch_kib_per_warp": 32, + "scratch": "variant 5 (measurement only): persistent warps; a 32 KiB scratch per warp of 64 16-byte slots per lane (lane-major); slot = src & 0x3f; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"load": 12, "add": 9, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "scratch": 4, "shfl": 4, "rotr": 2, "rotl": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2, "width": 1}, + {"i": 1, "op": "add", "dst": 1, "src": 7, "src2": 2, "imm": "0x42da7657", "imm2": "0xc3bd2355", "rot": 25, "bit": 4, "mask": 16, "width": 1}, + {"i": 2, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x61f0b51c", "imm2": "0x2735a174", "rot": 4, "bit": 26, "mask": 2, "width": 1}, + {"i": 3, "op": "mad", "dst": 4, "src": 0, "src2": 6, "imm": "0x679648a8", "imm2": "0x3044ba32", "rot": 31, "bit": 31, "mask": 4, "width": 1}, + {"i": 4, "op": "load", "dst": 7, "src": 2, "src2": 6, "imm": "0x5d1ca2a2", "imm2": "0xe2481807", "rot": 24, "bit": 3, "mask": 1, "width": 1}, + {"i": 5, "op": "load", "dst": 4, "src": 1, "src2": 2, "imm": "0x987c017a", "imm2": "0xf4d60559", "rot": 2, "bit": 0, "mask": 4, "width": 1}, + {"i": 6, "op": "shfl", "dst": 6, "src": 3, "src2": 7, "imm": "0x6ea7b2df", "imm2": "0x9fce5071", "rot": 7, "bit": 15, "mask": 4, "width": 1}, + {"i": 7, "op": "shfl", "dst": 1, "src": 5, "src2": 1, "imm": "0x26a2ecde", "imm2": "0xfec6ad22", "rot": 15, "bit": 11, "mask": 8, "width": 1}, + {"i": 8, "op": "xor", "dst": 7, "src": 5, "src2": 2, "imm": "0xbe4b445c", "imm2": "0x17a5a9c7", "rot": 8, "bit": 8, "mask": 1, "width": 1}, + {"i": 9, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4, "width": 1}, + {"i": 10, "op": "or", "dst": 1, "src": 2, "src2": 3, "imm": "0x4e7dc10d", "imm2": "0x196d165c", "rot": 14, "bit": 27, "mask": 16, "width": 1}, + {"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 1, "imm": "0xc5c3b55d", "imm2": "0xec061424", "rot": 26, "bit": 27, "mask": 8, "width": 1}, + {"i": 12, "op": "or", "dst": 6, "src": 2, "src2": 3, "imm": "0x306542fe", "imm2": "0x1bb1b429", "rot": 31, "bit": 0, "mask": 2, "width": 1}, + {"i": 13, "op": "mul", "dst": 2, "src": 5, "src2": 6, "imm": "0xa672cdd3", "imm2": "0x59a4829c", "rot": 22, "bit": 13, "mask": 16, "width": 1}, + {"i": 14, "op": "load", "dst": 1, "src": 2, "src2": 5, "imm": "0x028b4d37", "imm2": "0x7bbd78ea", "rot": 15, "bit": 2, "mask": 8, "width": 1}, + {"i": 15, "op": "rotl", "dst": 7, "src": 6, "src2": 6, "imm": "0x5c88a1a7", "imm2": "0x5c628769", "rot": 1, "bit": 3, "mask": 8, "width": 1}, + {"i": 16, "op": "scratch", "dst": 3, "src": 6, "src2": 7, "imm": "0xbac2ae81", "imm2": "0xcbbc7bdb", "rot": 18, "bit": 8, "mask": 8, "width": 1}, + {"i": 17, "op": "load", "dst": 7, "src": 4, "src2": 2, "imm": "0xe8ab93e9", "imm2": "0xa00de107", "rot": 2, "bit": 1, "mask": 16, "width": 1}, + {"i": 18, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2, "width": 1}, + {"i": 19, "op": "mad", "dst": 4, "src": 0, "src2": 2, "imm": "0x5fba7bc2", "imm2": "0xdf099cfb", "rot": 4, "bit": 15, "mask": 16, "width": 1}, + {"i": 20, "op": "shfl", "dst": 0, "src": 6, "src2": 3, "imm": "0x0a3056de", "imm2": "0x7f0c25c3", "rot": 27, "bit": 13, "mask": 8, "width": 1}, + {"i": 21, "op": "xor", "dst": 5, "src": 7, "src2": 4, "imm": "0xbd066e1d", "imm2": "0x6d3ddc5a", "rot": 2, "bit": 29, "mask": 1, "width": 1}, + {"i": 22, "op": "mulhi", "dst": 2, "src": 5, "src2": 0, "imm": "0xc7e9887a", "imm2": "0x19ec898f", "rot": 14, "bit": 9, "mask": 1, "width": 1}, + {"i": 23, "op": "load", "dst": 3, "src": 7, "src2": 2, "imm": "0xc7fcfc8f", "imm2": "0x8528b94f", "rot": 17, "bit": 13, "mask": 4, "width": 1}, + {"i": 24, "op": "mulhi", "dst": 7, "src": 3, "src2": 5, "imm": "0xd91641e8", "imm2": "0xaf77faf2", "rot": 22, "bit": 21, "mask": 1, "width": 1}, + {"i": 25, "op": "or", "dst": 5, "src": 4, "src2": 0, "imm": "0x84c03868", "imm2": "0xf6c691b7", "rot": 29, "bit": 14, "mask": 8, "width": 1}, + {"i": 26, "op": "mad", "dst": 4, "src": 5, "src2": 2, "imm": "0x3bb2b6ba", "imm2": "0x49d95fd5", "rot": 1, "bit": 5, "mask": 8, "width": 1}, + {"i": 27, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4, "width": 1}, + {"i": 28, "op": "mulhi", "dst": 6, "src": 7, "src2": 6, "imm": "0xd69c4715", "imm2": "0xe0ebc4ce", "rot": 29, "bit": 2, "mask": 8, "width": 1}, + {"i": 29, "op": "add", "dst": 6, "src": 1, "src2": 2, "imm": "0x3b2d2124", "imm2": "0x187a9128", "rot": 1, "bit": 9, "mask": 16, "width": 1}, + {"i": 30, "op": "rotr", "dst": 6, "src": 7, "src2": 0, "imm": "0x5c64a589", "imm2": "0x61c9a38d", "rot": 17, "bit": 21, "mask": 16, "width": 1}, + {"i": 31, "op": "load", "dst": 3, "src": 1, "src2": 7, "imm": "0xc37723fa", "imm2": "0xf3b024da", "rot": 16, "bit": 27, "mask": 16, "width": 1}, + {"i": 32, "op": "scratch", "dst": 1, "src": 0, "src2": 7, "imm": "0xcc7972c4", "imm2": "0xad098d15", "rot": 30, "bit": 21, "mask": 8, "width": 1}, + {"i": 33, "op": "add", "dst": 0, "src": 4, "src2": 4, "imm": "0x2c35699f", "imm2": "0x351dde38", "rot": 21, "bit": 18, "mask": 4, "width": 1}, + {"i": 34, "op": "scratch", "dst": 0, "src": 2, "src2": 3, "imm": "0xfae8902b", "imm2": "0x5cd8306f", "rot": 5, "bit": 28, "mask": 16, "width": 1}, + {"i": 35, "op": "mul", "dst": 0, "src": 3, "src2": 1, "imm": "0x4fa3f3db", "imm2": "0xdbf37e75", "rot": 7, "bit": 18, "mask": 4, "width": 1}, + {"i": 36, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2, "width": 1}, + {"i": 37, "op": "load", "dst": 4, "src": 0, "src2": 0, "imm": "0x04cc1d55", "imm2": "0x35c52d04", "rot": 11, "bit": 14, "mask": 2, "width": 1}, + {"i": 38, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0x3958f280", "imm2": "0x8713c7e1", "rot": 5, "bit": 23, "mask": 16, "width": 1}, + {"i": 39, "op": "add", "dst": 0, "src": 3, "src2": 3, "imm": "0xa907b90b", "imm2": "0x1b053acf", "rot": 30, "bit": 25, "mask": 16, "width": 1}, + {"i": 40, "op": "rotr", "dst": 2, "src": 5, "src2": 4, "imm": "0xf8662282", "imm2": "0x10bb9e30", "rot": 8, "bit": 6, "mask": 2, "width": 1}, + {"i": 41, "op": "mul", "dst": 3, "src": 2, "src2": 4, "imm": "0x49087d74", "imm2": "0x6348b489", "rot": 17, "bit": 9, "mask": 16, "width": 1}, + {"i": 42, "op": "add", "dst": 1, "src": 5, "src2": 1, "imm": "0xa32e000c", "imm2": "0x6058c2e3", "rot": 25, "bit": 20, "mask": 8, "width": 1}, + {"i": 43, "op": "xor", "dst": 3, "src": 4, "src2": 2, "imm": "0x3dad0eb6", "imm2": "0xb97578cb", "rot": 3, "bit": 27, "mask": 1, "width": 1}, + {"i": 44, "op": "load", "dst": 3, "src": 5, "src2": 7, "imm": "0x374aec92", "imm2": "0x626f11df", "rot": 20, "bit": 18, "mask": 8, "width": 1}, + {"i": 45, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4, "width": 1}, + {"i": 46, "op": "xor", "dst": 7, "src": 1, "src2": 0, "imm": "0xef6ac348", "imm2": "0x963bb7e6", "rot": 26, "bit": 3, "mask": 8, "width": 1}, + {"i": 47, "op": "add", "dst": 0, "src": 3, "src2": 0, "imm": "0x838b5065", "imm2": "0x36360066", "rot": 3, "bit": 31, "mask": 4, "width": 1}, + {"i": 48, "op": "mulhi", "dst": 7, "src": 5, "src2": 0, "imm": "0x8458f7ac", "imm2": "0xc1c15026", "rot": 27, "bit": 15, "mask": 8, "width": 1}, + {"i": 49, "op": "load", "dst": 0, "src": 2, "src2": 4, "imm": "0x636a9dc4", "imm2": "0xac023d9b", "rot": 22, "bit": 29, "mask": 1, "width": 1}, + {"i": 50, "op": "sub", "dst": 2, "src": 6, "src2": 0, "imm": "0x2baec8c9", "imm2": "0x4390f156", "rot": 3, "bit": 12, "mask": 8, "width": 1}, + {"i": 51, "op": "sub", "dst": 7, "src": 5, "src2": 7, "imm": "0x19234061", "imm2": "0xe84dfade", "rot": 4, "bit": 19, "mask": 1, "width": 1}, + {"i": 52, "op": "xor", "dst": 2, "src": 3, "src2": 5, "imm": "0xdc2cd71e", "imm2": "0x1b5d334b", "rot": 9, "bit": 8, "mask": 8, "width": 1}, + {"i": 53, "op": "sub", "dst": 7, "src": 0, "src2": 4, "imm": "0x605c31ec", "imm2": "0x9923ff88", "rot": 28, "bit": 25, "mask": 4, "width": 1}, + {"i": 54, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4, "width": 1}, + {"i": 55, "op": "xor", "dst": 7, "src": 5, "src2": 5, "imm": "0xad7493e7", "imm2": "0x3e400372", "rot": 13, "bit": 8, "mask": 1, "width": 1}, + {"i": 56, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0x87e933c9", "imm2": "0x8c854c1b", "rot": 17, "bit": 3, "mask": 8, "width": 1}, + {"i": 57, "op": "sub", "dst": 5, "src": 6, "src2": 5, "imm": "0x11be3bc9", "imm2": "0xbbaa8e24", "rot": 6, "bit": 5, "mask": 16, "width": 1}, + {"i": 58, "op": "load", "dst": 1, "src": 3, "src2": 2, "imm": "0xa732351a", "imm2": "0xc01349cd", "rot": 14, "bit": 17, "mask": 16, "width": 1}, + {"i": 59, "op": "scratch", "dst": 1, "src": 4, "src2": 0, "imm": "0xb20547b2", "imm2": "0xc94655de", "rot": 27, "bit": 30, "mask": 1, "width": 1}, + {"i": 60, "op": "sub", "dst": 4, "src": 6, "src2": 7, "imm": "0x67cf904c", "imm2": "0x6873b216", "rot": 27, "bit": 7, "mask": 16, "width": 1}, + {"i": 61, "op": "mul", "dst": 1, "src": 2, "src2": 7, "imm": "0x93ab0bf4", "imm2": "0x96158375", "rot": 14, "bit": 0, "mask": 16, "width": 1}, + {"i": 62, "op": "mad", "dst": 3, "src": 6, "src2": 0, "imm": "0x41a443a3", "imm2": "0xe69d7919", "rot": 9, "bit": 0, "mask": 16, "width": 1}, + {"i": 63, "op": "add", "dst": 0, "src": 1, "src2": 3, "imm": "0x2fe0e98b", "imm2": "0xc88e2942", "rot": 5, "bit": 16, "mask": 16, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/scr4k32/program.metal b/proto-cuda/packs-readwidth/scr4k32/program.metal new file mode 100644 index 000000000..ecfc2be8d --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k32/program.metal @@ -0,0 +1,126 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + device uint* scratch [[buffer(3)]], + constant uint& groups [[buffer(4)]], + constant uint& salt [[buffer(5)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + r7 = r7 ^ dataset[r2 & MASK]; // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint s_ = r6 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + r3 = r3 ^ dataset[r7 & MASK]; // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + { uint s_ = r0 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint s_ = r2 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + r4 = r4 ^ dataset[r0 & MASK]; // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + r3 = r3 ^ dataset[r5 & MASK]; // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + { uint s_ = r4 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr4k32/program_bound.metal b/proto-cuda/packs-readwidth/scr4k32/program_bound.metal new file mode 100644 index 000000000..cf840249d --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k32/program_bound.metal @@ -0,0 +1,128 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + device uint* scratch [[buffer(4)]], + constant uint& groups [[buffer(5)]], + constant uint& salt [[buffer(6)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + r7 = r7 ^ dataset[r2 & MASK]; // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint s_ = r6 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + r3 = r3 ^ dataset[r7 & MASK]; // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + { uint s_ = r0 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint s_ = r2 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + r4 = r4 ^ dataset[r0 & MASK]; // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + r3 = r3 ^ dataset[r5 & MASK]; // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + { uint s_ = r4 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr4k32/vectors.h b/proto-cuda/packs-readwidth/scr4k32/vectors.h new file mode 100644 index 000000000..e15e27475 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k32/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x62cab4be0ed880e1ull, 0x90842c849268cf52ull, 0x53d4d591f4a12749ull, 0x020429c3d1279eddull, 0x7086876f3a9183fbull, 0xc1f954e8065b6d02ull, 0x7550abd24b6ed9edull, 0x7dcc57f17b255fc9ull, + 0x01f16667b6ea326dull, 0x05907ad2b28423a5ull, 0x27e8aad8889ea702ull, 0xe93e61d2ff565fabull, 0x21aa77db95a8790full, 0xdc361b8011f7a395ull, 0xbabf590a4cf6dd58ull, 0x7a9b7eb46cf2b88dull, + 0xeb43f11d39490205ull, 0x3c2e3b2ed2b0c6fbull, 0x2bc2293577914bc8ull, 0x66cc059287de6cc0ull, 0x9a2d3a0f30169b20ull, 0xa5756e5027ec459full, 0x734a7fb8546a08b5ull, 0xbc59502ef67511d7ull, + 0x869d40616fa13209ull, 0x7e2713e23d7c3e06ull, 0x640f81856492a8d3ull, 0x1f7ce7b42962bc41ull, 0xc474fdd339993867ull, 0xb08d199c21d08397ull, 0x6b021b07d5dd9fdfull, 0x574adb548a4f3be8ull + }, + { // base nonce 4096 + 0x2956e7703c1553fcull, 0xe06e9c5dd64f0cffull, 0x41d967b788797c3eull, 0xc9beec7571ab7808ull, 0x0d0d99d51ac72942ull, 0xac60f79ae1bb46d6ull, 0xd9b9a509bc33d145ull, 0x512d444977d257f5ull, + 0x0e9759a10d773b28ull, 0xb270e841265b2b3dull, 0x90b97771870e53dcull, 0xf8b0bacb1ead0c1bull, 0x32165ca85108736cull, 0x5a907f0cb371d6d1ull, 0x89d4ccf9b6323847ull, 0x495e23339db371dbull, + 0xa50d559aa7911894ull, 0xbe561e2c2e64f0ffull, 0x862b141ae3b898eaull, 0x69b52c3068f0544aull, 0x2769f3b4051f9e80ull, 0xb679a28f140a5ccaull, 0x037d194732dce935ull, 0xec9f1e85406dbee9ull, + 0x65a0d3f10857795dull, 0x5da0b4908b5cda66ull, 0x1cbdf4f47dad39e4ull, 0x5472317d40d55545ull, 0x24ec2fb5eeff7691ull, 0x4c56104e2454b9beull, 0x8b896d9e85dbf491ull, 0xe80882d5975e09ecull + }, + { // base nonce 1000000 + 0xa417c0e0494f5f0dull, 0x44cfa8bbf55cb40bull, 0xee534b970664a105ull, 0x2814b1857db92d67ull, 0xd358f7e47b35ff5full, 0xa08faa47e58221c3ull, 0xbb76559b9a4a447bull, 0xd438a3fd1fa2976eull, + 0xa08d0e2c88abe900ull, 0x58b3c3ab097c416dull, 0x705de177cf28ccdcull, 0x263f35e27d8cf3aaull, 0xa2c304ccfeb9ae9bull, 0x470ba6ea4e8f661aull, 0xa59e5f33cd8613d9ull, 0xdb887848353dc91cull, + 0xd948d1c36b6a98e1ull, 0xd78806062c54882aull, 0x0744e029194938aaull, 0x42b613ec3d9074c6ull, 0x88c75044753d1496ull, 0x940239d09cbd80e1ull, 0x4534dd536f53adf5ull, 0x6f44b9bd4d6564deull, + 0xb6d8142c422857d5ull, 0x3b61f0c20b8fb04bull, 0x17eaf3a88b49c9fdull, 0xec536015d760eb1eull, 0xf37e9ad57045cc05ull, 0x588b808bc29ab6caull, 0x7cfa7ae3c6e483c2ull, 0x71e3cc45c07b530cull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/scr4k32/vectors.json b/proto-cuda/packs-readwidth/scr4k32/vectors.json new file mode 100644 index 000000000..4b970a994 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr4k32/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-genesis", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x62cab4be0ed880e1", "0x90842c849268cf52", "0x53d4d591f4a12749", "0x020429c3d1279edd", "0x7086876f3a9183fb", "0xc1f954e8065b6d02", "0x7550abd24b6ed9ed", "0x7dcc57f17b255fc9", + "0x01f16667b6ea326d", "0x05907ad2b28423a5", "0x27e8aad8889ea702", "0xe93e61d2ff565fab", "0x21aa77db95a8790f", "0xdc361b8011f7a395", "0xbabf590a4cf6dd58", "0x7a9b7eb46cf2b88d", + "0xeb43f11d39490205", "0x3c2e3b2ed2b0c6fb", "0x2bc2293577914bc8", "0x66cc059287de6cc0", "0x9a2d3a0f30169b20", "0xa5756e5027ec459f", "0x734a7fb8546a08b5", "0xbc59502ef67511d7", + "0x869d40616fa13209", "0x7e2713e23d7c3e06", "0x640f81856492a8d3", "0x1f7ce7b42962bc41", "0xc474fdd339993867", "0xb08d199c21d08397", "0x6b021b07d5dd9fdf", "0x574adb548a4f3be8" + ]}, + {"base_nonce": 4096, "expected": [ + "0x2956e7703c1553fc", "0xe06e9c5dd64f0cff", "0x41d967b788797c3e", "0xc9beec7571ab7808", "0x0d0d99d51ac72942", "0xac60f79ae1bb46d6", "0xd9b9a509bc33d145", "0x512d444977d257f5", + "0x0e9759a10d773b28", "0xb270e841265b2b3d", "0x90b97771870e53dc", "0xf8b0bacb1ead0c1b", "0x32165ca85108736c", "0x5a907f0cb371d6d1", "0x89d4ccf9b6323847", "0x495e23339db371db", + "0xa50d559aa7911894", "0xbe561e2c2e64f0ff", "0x862b141ae3b898ea", "0x69b52c3068f0544a", "0x2769f3b4051f9e80", "0xb679a28f140a5cca", "0x037d194732dce935", "0xec9f1e85406dbee9", + "0x65a0d3f10857795d", "0x5da0b4908b5cda66", "0x1cbdf4f47dad39e4", "0x5472317d40d55545", "0x24ec2fb5eeff7691", "0x4c56104e2454b9be", "0x8b896d9e85dbf491", "0xe80882d5975e09ec" + ]}, + {"base_nonce": 1000000, "expected": [ + "0xa417c0e0494f5f0d", "0x44cfa8bbf55cb40b", "0xee534b970664a105", "0x2814b1857db92d67", "0xd358f7e47b35ff5f", "0xa08faa47e58221c3", "0xbb76559b9a4a447b", "0xd438a3fd1fa2976e", + "0xa08d0e2c88abe900", "0x58b3c3ab097c416d", "0x705de177cf28ccdc", "0x263f35e27d8cf3aa", "0xa2c304ccfeb9ae9b", "0x470ba6ea4e8f661a", "0xa59e5f33cd8613d9", "0xdb887848353dc91c", + "0xd948d1c36b6a98e1", "0xd78806062c54882a", "0x0744e029194938aa", "0x42b613ec3d9074c6", "0x88c75044753d1496", "0x940239d09cbd80e1", "0x4534dd536f53adf5", "0x6f44b9bd4d6564de", + "0xb6d8142c422857d5", "0x3b61f0c20b8fb04b", "0x17eaf3a88b49c9fd", "0xec536015d760eb1e", "0xf37e9ad57045cc05", "0x588b808bc29ab6ca", "0x7cfa7ae3c6e483c2", "0x71e3cc45c07b530c" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/scr8k128/kernel.cl b/proto-cuda/packs-readwidth/scr8k128/kernel.cl new file mode 100644 index 000000000..03840512b --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k128/kernel.cl @@ -0,0 +1,291 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint s_ = r2 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint s_ = r7 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 23 scratch + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint s_ = r0 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint s_ = r5 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/scr8k128/kernel.cu b/proto-cuda/packs-readwidth/scr8k128/kernel.cu new file mode 100644 index 000000000..9db090564 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k128/kernel.cu @@ -0,0 +1,177 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint32_t s_ = r2 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t s_ = r6 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + { uint32_t s_ = r7 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 scratch + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint32_t s_ = r0 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t s_ = r2 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint32_t s_ = r0 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint32_t s_ = r5 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint32_t s_ = r4 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr8k128/kernel_bound.cl b/proto-cuda/packs-readwidth/scr8k128/kernel_bound.cl new file mode 100644 index 000000000..9cb72a38f --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k128/kernel_bound.cl @@ -0,0 +1,393 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint s_ = r2 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint s_ = r7 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 23 scratch + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint s_ = r0 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint s_ = r5 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint s_ = r2 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint s_ = r7 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 23 scratch + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint s_ = r0 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint s_ = r5 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 255u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr8k128/kernel_bound.cu b/proto-cuda/packs-readwidth/scr8k128/kernel_bound.cu new file mode 100644 index 000000000..4224a414c --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k128/kernel_bound.cu @@ -0,0 +1,136 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint32_t s_ = r2 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t s_ = r6 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + { uint32_t s_ = r7 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 scratch + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint32_t s_ = r0 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t s_ = r2 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint32_t s_ = r0 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint32_t s_ = r5 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint32_t s_ = r4 & 255u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr8k128/memhard.h b/proto-cuda/packs-readwidth/scr8k128/memhard.h new file mode 100644 index 000000000..4803d8e40 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k128/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/scr8k128/memhard.metal b/proto-cuda/packs-readwidth/scr8k128/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k128/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/scr8k128/program.h b/proto-cuda/packs-readwidth/scr8k128/program.h new file mode 100644 index 000000000..2da41c27e --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k128/program.h @@ -0,0 +1,67 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Program metadata for host.cu plus the launch wrappers defined in kernel.cu. +// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#ifndef IGNEUM_NO_CUDA +#include +#endif + +#define IGNEUM_SEED_STRING "igneum-genesis" +#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" +#define IGNEUM_GENERATOR 2 +#define IGNEUM_PROGRAM_ATTEMPT 0 +#define IGNEUM_PROGRAM_ID 0xe0d1dcd155d90573ull +#define IGNEUM_DAY_STRING "2026-10-03" +#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" +#define IGNEUM_DAY0 0x3067619fu +#define IGNEUM_DAY1 0x3c269176u +#define IGNEUM_DATASET_LOG2 28 +#define IGNEUM_MASK 0x0fffffffu +#define IGNEUM_LANES 32 +#define IGNEUM_ITERATIONS 8 +#define IGNEUM_INSTR_COUNT 64 +#define IGNEUM_LOADS_PER_HASH 128 +#define IGNEUM_WIDE_LOADS_PER_HASH 0 +#define IGNEUM_OP_MIX "add=9 load=8 scratch=8 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 shfl=4 rotr=2 rotl=1" +// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads +// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. +#define IGNEUM_LOAD_CLASS "scr8k128" +#define IGNEUM_LOAD_SLOTS 16 +#define IGNEUM_LOAD_MIX { 100, 0, 0 } +#define IGNEUM_LOAD_WIDTH_COUNTS { 8, 0, 0 } // loads of 4, 16, 64 bytes per program +#define IGNEUM_BYTES_PER_HASH 256 +#define IGNEUM_FOLD_ROT 11 +#define IGNEUM_FOLD_MUL 0x9e3779b1u +// Variant 5: persistent warps, a 128 KiB scratch per launched warp (the host launches N warps and passes scratch, +// groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). +#define IGNEUM_PERSISTENT_WARPS 1 +#define IGNEUM_SCRATCH_OPS 8 // scratch read-modify-writes per program (64 per hash) +#define IGNEUM_SCRATCH_SLOTS 256u +#define IGNEUM_SCRATCH_WORDS_PER_LANE 1024u +#define IGNEUM_SCRATCH_BYTES_PER_WARP 131072u +// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) +#define IGNEUM_DATASET_MODE 1 + +#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u } +#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u } +#define IGNEUM_CACHE_LOG2_WORDS 26 +#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6 +#define IGNEUM_CACHE_SEGMENTS 65536u +#define IGNEUM_ITEM_ROUNDS 8 +#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u } +#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u } +#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u } + +#ifndef IGNEUM_NO_CUDA +// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError(). +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments); +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems); +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt); +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#endif diff --git a/proto-cuda/packs-readwidth/scr8/program.json b/proto-cuda/packs-readwidth/scr8k128/program.json similarity index 98% rename from proto-cuda/packs-readwidth/scr8/program.json rename to proto-cuda/packs-readwidth/scr8k128/program.json index 25348b09f..8cbee7a2b 100644 --- a/proto-cuda/packs-readwidth/scr8/program.json +++ b/proto-cuda/packs-readwidth/scr8k128/program.json @@ -2,7 +2,7 @@ "format": "igneum-program-pack-3", "generator": 2, "attempt": 0, - "program_id": "0x2f0992e568f38dc1", + "program_id": "0xe0d1dcd155d90573", "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", "dataset_mode": "memory-hard", "seed": "igneum-genesis", @@ -15,13 +15,14 @@ "iterations": 8, "instruction_count": 64, "loads_per_hash": 128, - "load_class": "scr8", + "load_class": "scr8k128", "load_slots": 16, "load_mix_percent_4_16_64": [100, 0, 0], "load_width_counts_4_16_64": [8, 0, 0], "bytes_per_hash": 256, "scratch_ops_per_hash": 64, - "scratch": "variant 5 (measurement only): persistent warps; a 1 MiB scratch per warp of 2048 16-byte slots per lane (lane-major); slot = src & 0x7ff; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", + "scratch_kib_per_warp": 128, + "scratch": "variant 5 (measurement only): persistent warps; a 128 KiB scratch per warp of 256 16-byte slots per lane (lane-major); slot = src & 0xff; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", "op_mix": {"add": 9, "load": 8, "scratch": 8, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "shfl": 4, "rotr": 2, "rotl": 1}, "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", diff --git a/proto-cuda/packs-readwidth/scr8/program.metal b/proto-cuda/packs-readwidth/scr8k128/program.metal similarity index 56% rename from proto-cuda/packs-readwidth/scr8/program.metal rename to proto-cuda/packs-readwidth/scr8k128/program.metal index 73532c454..c30f95277 100644 --- a/proto-cuda/packs-readwidth/scr8/program.metal +++ b/proto-cuda/packs-readwidth/scr8k128/program.metal @@ -21,7 +21,7 @@ inline uint ds_elem(uint i, uint d0, uint d1) { return x; } -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -36,7 +36,7 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], uint lane = tid & 31u; uint warp_ = tid >> 5; uint nwarps_ = nthreads >> 5; - device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -58,7 +58,7 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 r4 = r0 * r6 + r4; // 3 - { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 + { uint s_ = r2 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 r4 = r4 ^ dataset[r1 & MASK]; // 5 r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 @@ -70,14 +70,14 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], r2 = r2 * r5; // 13 r1 = r1 ^ dataset[r2 & MASK]; // 14 r7 = rotl_imm(r7, 1u); // 15 - { uint s_ = r6 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + { uint s_ = r6 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 r7 = r7 ^ dataset[r4 & MASK]; // 17 r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 r4 = r0 * r2 + r4; // 19 r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 r5 = r5 ^ r7; // 21 r2 = mulhi(r2, r5); // 22 - { uint s_ = r7 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 + { uint s_ = r7 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 r7 = mulhi(r7, r3); // 24 r5 = r5 | r4; // 25 r4 = r5 * r2 + r4; // 26 @@ -86,19 +86,19 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 r6 = rotr_var(r6, r7); // 30 r3 = r3 ^ dataset[r1 & MASK]; // 31 - { uint s_ = r0 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + { uint s_ = r0 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 - { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + { uint s_ = r2 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 r0 = r0 * r3; // 35 r2 = r2 ^ r5; // 36 - { uint s_ = r0 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 + { uint s_ = r0 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 r1 = r3 * r5 + r1; // 38 r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 r2 = rotr_var(r2, r5); // 40 r3 = r3 * r2; // 41 r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 r3 = r3 ^ r4; // 43 - { uint s_ = r5 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 + { uint s_ = r5 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 r7 = r7 ^ r1; // 46 r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 @@ -113,7 +113,7 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]], r2 = r2 ^ dataset[r7 & MASK]; // 56 r5 = r5 - r6; // 57 r1 = r1 ^ dataset[r3 & MASK]; // 58 - { uint s_ = r4 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + { uint s_ = r4 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 r4 = r4 - r6; // 60 r1 = r1 * r2; // 61 r3 = r6 * r0 + r3; // 62 diff --git a/proto-cuda/packs-readwidth/scr8/program_bound.metal b/proto-cuda/packs-readwidth/scr8k128/program_bound.metal similarity index 57% rename from proto-cuda/packs-readwidth/scr8/program_bound.metal rename to proto-cuda/packs-readwidth/scr8k128/program_bound.metal index d260c34fe..665f2721e 100644 --- a/proto-cuda/packs-readwidth/scr8/program_bound.metal +++ b/proto-cuda/packs-readwidth/scr8k128/program_bound.metal @@ -21,7 +21,7 @@ inline uint ds_elem(uint i, uint d0, uint d1) { return x; } -// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 1 MiB scratch per warp, 2048 slots of +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 128 KiB scratch per warp, 256 slots of // 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not // this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } @@ -38,7 +38,7 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], uint lane = tid & 31u; uint warp_ = tid >> 5; uint nwarps_ = nthreads >> 5; - device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 8192u; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { uint gid = g_ * 32u + lane; uint gbase = baseNonce + g_ * 32u; @@ -60,7 +60,7 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 r4 = r0 * r6 + r4; // 3 - { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 + { uint s_ = r2 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 r4 = r4 ^ dataset[r1 & MASK]; // 5 r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 @@ -72,14 +72,14 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], r2 = r2 * r5; // 13 r1 = r1 ^ dataset[r2 & MASK]; // 14 r7 = rotl_imm(r7, 1u); // 15 - { uint s_ = r6 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + { uint s_ = r6 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 r7 = r7 ^ dataset[r4 & MASK]; // 17 r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 r4 = r0 * r2 + r4; // 19 r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 r5 = r5 ^ r7; // 21 r2 = mulhi(r2, r5); // 22 - { uint s_ = r7 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 + { uint s_ = r7 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 r7 = mulhi(r7, r3); // 24 r5 = r5 | r4; // 25 r4 = r5 * r2 + r4; // 26 @@ -88,19 +88,19 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 r6 = rotr_var(r6, r7); // 30 r3 = r3 ^ dataset[r1 & MASK]; // 31 - { uint s_ = r0 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + { uint s_ = r0 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 - { uint s_ = r2 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + { uint s_ = r2 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 r0 = r0 * r3; // 35 r2 = r2 ^ r5; // 36 - { uint s_ = r0 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 + { uint s_ = r0 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 r1 = r3 * r5 + r1; // 38 r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 r2 = rotr_var(r2, r5); // 40 r3 = r3 * r2; // 41 r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 r3 = r3 ^ r4; // 43 - { uint s_ = r5 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 + { uint s_ = r5 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 r7 = r7 ^ r1; // 46 r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 @@ -115,7 +115,7 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], r2 = r2 ^ dataset[r7 & MASK]; // 56 r5 = r5 - r6; // 57 r1 = r1 ^ dataset[r3 & MASK]; // 58 - { uint s_ = r4 & 2047u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + { uint s_ = r4 & 255u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 r4 = r4 - r6; // 60 r1 = r1 * r2; // 61 r3 = r6 * r0 + r3; // 62 diff --git a/proto-cuda/packs-readwidth/scr8k128/vectors.h b/proto-cuda/packs-readwidth/scr8k128/vectors.h new file mode 100644 index 000000000..4555a7d44 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k128/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0xe814771d17db365cull, 0xc4d4a4ae8b6033caull, 0xa4ea884e5e961196ull, 0x5a4f2e353f13502bull, 0xb5bfae78c5048e4cull, 0x75ed7ff2f8ccb9c1ull, 0xc9d37d26079b1916ull, 0xef3c29ccb9d46163ull, + 0xb4cae00d3e73ae8eull, 0x88bc44f26a90e913ull, 0xeb91c51e87da69f8ull, 0x2ee04eeac0ba97b2ull, 0x43e33706056bb735ull, 0x88acef8db41e6bbfull, 0x87295bf633750804ull, 0x4a8310fa3c5393f3ull, + 0x9c33366c7aa6a5a6ull, 0x79ba6d5674f78e5cull, 0x03168e07fb7ae416ull, 0xdd45a1b5f54270aeull, 0x0d0084aa74f95b51ull, 0xa04060d3f711930bull, 0xa080e7988516297full, 0xcefa3e2e8de2806full, + 0x95dc7a55e10c4010ull, 0x62e809b8cef37be6ull, 0x0366273048e795cfull, 0xc2b04c7deacb1dffull, 0xe83446f7db3a4686ull, 0xe6c7e6575c34651full, 0xee15067b419606d6ull, 0xfe67f3ae51faade1ull + }, + { // base nonce 4096 + 0xdde6034f4b5824c9ull, 0x07b807430ab9effbull, 0xa581d141cb4bc74aull, 0x0a1b4129b3690618ull, 0xb4d10af1daaec58bull, 0xcf0b63a9aa6b8a96ull, 0x07c20bd30e3eb88cull, 0x32ebffcaafa5df9eull, + 0xe1b528deb263ddc3ull, 0xaa6d2e1e7f45c995ull, 0x6017aaa938e837cfull, 0x23445a9b8c8e5addull, 0x024ebd232a344f41ull, 0x67aebe3e79435f84ull, 0xa7d0b7522e88814aull, 0x1d4d9633b57ac637ull, + 0x77f500325912fcbdull, 0x9f4bc5d04fbb13b9ull, 0xd7e081a23934d582ull, 0x10992aa1c8a93afeull, 0x1596ba0b47520be7ull, 0x344ed3c63b5a75bcull, 0xdb75c50a7a39c7beull, 0x0e1ccc942ec1fad7ull, + 0x81c9d97c4d9605c6ull, 0x1e60918f98df7dc9ull, 0x3e60e5b90d94fe34ull, 0xec7163b65fc01cfdull, 0xb0786922940f66e3ull, 0xb1d049c3e24a38e2ull, 0x5a6a7ac9a0c8ed50ull, 0x7bc50eab43e83a01ull + }, + { // base nonce 1000000 + 0x1ebe406e6227f5e9ull, 0xc9d07c89dd188990ull, 0x3346fbacbe00f719ull, 0x437f4d678259e06dull, 0xd664758bc7508b7cull, 0xa3428dfd2b480593ull, 0xdfbb1aca3c18aeb6ull, 0xe362f4d90b64ab9full, + 0xfaeac4630c51e291ull, 0x2b81c4de8eaa689aull, 0x661b54d4d8763782ull, 0x48839cd831ea402bull, 0xa98dbad17e5f3a49ull, 0x35161bdc5dc87db7ull, 0xcc503293dddde770ull, 0x8960f9fbf15a0e47ull, + 0x3d391f32613d80d5ull, 0xd3fb60a843c31905ull, 0x7f70cbe3a4be1f1eull, 0xb67d653a57d143c7ull, 0x08a9210687821d3dull, 0x54adfc465596fd4cull, 0x12cf1cdd36c931e4ull, 0x62fd59b6a002a406ull, + 0xfef506618454af46ull, 0x8ab9f4c86cfefbe0ull, 0xb54d7b600ee5cbdbull, 0x9400685a172c31ccull, 0x0cdc0ca4f87996dcull, 0x050be1f1631ab65full, 0x79884be8f2e9a1f2ull, 0x824d03dfe16bcb91ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/scr8k128/vectors.json b/proto-cuda/packs-readwidth/scr8k128/vectors.json new file mode 100644 index 000000000..6f5c09848 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k128/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-genesis", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0xe814771d17db365c", "0xc4d4a4ae8b6033ca", "0xa4ea884e5e961196", "0x5a4f2e353f13502b", "0xb5bfae78c5048e4c", "0x75ed7ff2f8ccb9c1", "0xc9d37d26079b1916", "0xef3c29ccb9d46163", + "0xb4cae00d3e73ae8e", "0x88bc44f26a90e913", "0xeb91c51e87da69f8", "0x2ee04eeac0ba97b2", "0x43e33706056bb735", "0x88acef8db41e6bbf", "0x87295bf633750804", "0x4a8310fa3c5393f3", + "0x9c33366c7aa6a5a6", "0x79ba6d5674f78e5c", "0x03168e07fb7ae416", "0xdd45a1b5f54270ae", "0x0d0084aa74f95b51", "0xa04060d3f711930b", "0xa080e7988516297f", "0xcefa3e2e8de2806f", + "0x95dc7a55e10c4010", "0x62e809b8cef37be6", "0x0366273048e795cf", "0xc2b04c7deacb1dff", "0xe83446f7db3a4686", "0xe6c7e6575c34651f", "0xee15067b419606d6", "0xfe67f3ae51faade1" + ]}, + {"base_nonce": 4096, "expected": [ + "0xdde6034f4b5824c9", "0x07b807430ab9effb", "0xa581d141cb4bc74a", "0x0a1b4129b3690618", "0xb4d10af1daaec58b", "0xcf0b63a9aa6b8a96", "0x07c20bd30e3eb88c", "0x32ebffcaafa5df9e", + "0xe1b528deb263ddc3", "0xaa6d2e1e7f45c995", "0x6017aaa938e837cf", "0x23445a9b8c8e5add", "0x024ebd232a344f41", "0x67aebe3e79435f84", "0xa7d0b7522e88814a", "0x1d4d9633b57ac637", + "0x77f500325912fcbd", "0x9f4bc5d04fbb13b9", "0xd7e081a23934d582", "0x10992aa1c8a93afe", "0x1596ba0b47520be7", "0x344ed3c63b5a75bc", "0xdb75c50a7a39c7be", "0x0e1ccc942ec1fad7", + "0x81c9d97c4d9605c6", "0x1e60918f98df7dc9", "0x3e60e5b90d94fe34", "0xec7163b65fc01cfd", "0xb0786922940f66e3", "0xb1d049c3e24a38e2", "0x5a6a7ac9a0c8ed50", "0x7bc50eab43e83a01" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x1ebe406e6227f5e9", "0xc9d07c89dd188990", "0x3346fbacbe00f719", "0x437f4d678259e06d", "0xd664758bc7508b7c", "0xa3428dfd2b480593", "0xdfbb1aca3c18aeb6", "0xe362f4d90b64ab9f", + "0xfaeac4630c51e291", "0x2b81c4de8eaa689a", "0x661b54d4d8763782", "0x48839cd831ea402b", "0xa98dbad17e5f3a49", "0x35161bdc5dc87db7", "0xcc503293dddde770", "0x8960f9fbf15a0e47", + "0x3d391f32613d80d5", "0xd3fb60a843c31905", "0x7f70cbe3a4be1f1e", "0xb67d653a57d143c7", "0x08a9210687821d3d", "0x54adfc465596fd4c", "0x12cf1cdd36c931e4", "0x62fd59b6a002a406", + "0xfef506618454af46", "0x8ab9f4c86cfefbe0", "0xb54d7b600ee5cbdb", "0x9400685a172c31cc", "0x0cdc0ca4f87996dc", "0x050be1f1631ab65f", "0x79884be8f2e9a1f2", "0x824d03dfe16bcb91" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-cuda/packs-readwidth/scr8k32/kernel.cl b/proto-cuda/packs-readwidth/scr8k32/kernel.cl new file mode 100644 index 000000000..54168b39f --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k32/kernel.cl @@ -0,0 +1,291 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint s_ = r2 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint s_ = r7 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 23 scratch + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint s_ = r0 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint s_ = r5 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif diff --git a/proto-cuda/packs-readwidth/scr8k32/kernel.cu b/proto-cuda/packs-readwidth/scr8k32/kernel.cu new file mode 100644 index 000000000..b43200fdd --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k32/kernel.cu @@ -0,0 +1,177 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal). +// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC. +#include +#include +#include "program.h" +#include "memhard.h" + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31. +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x. +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) { + uint32_t x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item. +// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host. +__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) { + uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x; + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + uint32_t t = blockIdx.x * blockDim.x + threadIdx.x; + if (t < nItems) { + uint32_t s[16]; + mh_item(cache, t, s); + uint32_t* d = ds + (size_t)t * 16u; + for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every +// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a +// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid. +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint32_t s_ = r2 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t s_ = r6 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + { uint32_t s_ = r7 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 scratch + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint32_t s_ = r0 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t s_ = r2 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint32_t s_ = r0 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint32_t s_ = r5 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint32_t s_ = r4 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Host-side launch wrappers. Declared in program.h, called from host.cu. +cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) { + if (nSegments == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nSegments + block - 1u) / block; + igneum_cache_fill<<>>(cache, nSegments); + return cudaGetLastError(); +} + +cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) { + if (nItems == 0u) return cudaErrorInvalidValue; + uint32_t block = 256u; + uint32_t grid = (nItems + block - 1u) / block; + igneum_build<<>>(ds, cache, nItems); + return cudaGetLastError(); +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash<<>>(ds, out, baseNonce, mask, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr8k32/kernel_bound.cl b/proto-cuda/packs-readwidth/scr8k32/kernel_bound.cl new file mode 100644 index 000000000..922739e4a --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k32/kernel_bound.cl @@ -0,0 +1,393 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal). +// Built from source at runtime by proto-opencl/host.c, which passes these defines: +// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit) +// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default) +// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32 +// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition +// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units; +// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32. +#ifndef IGNEUM_GROUP +#define IGNEUM_GROUP 32 +#endif +#ifndef IGNEUM_EXCHANGE +#define IGNEUM_EXCHANGE 0 +#endif +#ifdef __OPENCL_VERSION__ +#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1))) +#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n] +#if IGNEUM_EXCHANGE == 1 +#ifdef cl_khr_subgroups +#pragma OPENCL EXTENSION cl_khr_subgroups : enable +#endif +#ifdef cl_khr_subgroup_shuffle +#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable +#endif +#elif IGNEUM_EXCHANGE == 2 +#pragma OPENCL EXTENSION cl_intel_subgroups : enable +#endif +#define IGNEUM_U4(a, b, c, d) ((uint4)((a), (b), (c), (d))) +#else +// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros. +#include "emu_opencl.h" +#endif + +#if IGNEUM_EXCHANGE == 1 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#elif IGNEUM_EXCHANGE == 2 +#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m)) +#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u) +#else +// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per +// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane +// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's +// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier. +#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; } +#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; } +#endif + +static inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32. +static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); } +// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x. +static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); } +static inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +static inline void mh_chacha_block(const uint* x, uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +static inline void mh_cache_segment(__global uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +static inline void mh_mixer(uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +static inline void mh_item(__global const uint* cache, uint t, uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item. +// The same constants as memhard.h in this pack (one emitter, three dialects). +__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) { + uint seg = (uint)get_global_id(0); + if (seg < nSegments) mh_cache_segment(cache, seg); +} +__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) { + uint t = (uint)get_global_id(0); + if (t < nItems) { + uint s[16]; + mh_item(cache, t, s); + __global uint* d = ds + ((ulong)t * 16u); + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; + } +} + +// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the +// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and +// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all). +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +static inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] + { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] + { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] + { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4] + { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5] + { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6] + { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7] + { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0] + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint s_ = r2 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint s_ = r7 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 23 scratch + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint s_ = r0 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint s_ = r5 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} + +#if IGNEUM_EXCHANGE != 0 +// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the +// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a +// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md. +IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) { + if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); } +} +#endif + +// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash. +IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw, __global uint* scratch, uint groups, uint salt) { + uint lane = (uint)get_global_id(0) & 31u; + uint warp_ = (uint)get_global_id(0) >> 5; + uint nwarps_ = (uint)get_global_size(0) >> 5; + __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint lid = (uint)get_local_id(0); + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; +#if IGNEUM_EXCHANGE == 0 + IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); + uint xk = 0u; +#else + (void)lid; +#endif + { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } + { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } + { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } + { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; } + { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; } + { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; } + { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; } + { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint s_ = r2 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 6 shfl + { uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r1 = r1 ^ t_; } // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint s_ = r6 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + { uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r0 = r0 ^ t_; } // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = mul_hi(r2, r5); // 22 mulhi + { uint s_ = r7 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 23 scratch + r7 = mul_hi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = mul_hi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint s_ = r0 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint s_ = r2 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint s_ = r0 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint s_ = r5 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = mul_hi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint s_ = r4 & 63u; uint4 v_ = vload4(s_, arena); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr8k32/kernel_bound.cu b/proto-cuda/packs-readwidth/scr8k32/kernel_bound.cu new file mode 100644 index 000000000..ae117b313 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k32/kernel_bound.cu @@ -0,0 +1,136 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW. +// Host declarations (also in program_bound.h if present): +// struct IgneumInitWords { uint32_t w[8]; }; +// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, +// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps); +// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps); +#include +#include +#include "program.h" + +struct IgneumInitWords { uint32_t w[8]; }; + +__device__ __forceinline__ uint32_t splitmix32(uint32_t x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } +__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +__device__ __forceinline__ uint32_t scr_fill(uint32_t gbase, uint32_t lane, uint32_t slot, uint32_t j) { uint32_t sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw, uint32_t* scratch, uint32_t groups, uint32_t salt) { + uint32_t lane = (blockIdx.x * blockDim.x + threadIdx.x) & 31u; + uint32_t warp_ = (blockIdx.x * blockDim.x + threadIdx.x) >> 5; + uint32_t nwarps_ = (gridDim.x * blockDim.x) >> 5; + uint32_t* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint32_t g_ = warp_; g_ < groups; g_ += nwarps_) { + uint32_t gid = g_ * 32u + lane; + uint32_t gbase = baseNonce + g_ * 32u; + uint32_t tag = salt + g_; + uint32_t nonce = baseNonce + gid; + uint32_t r0, r1, r2, r3, r4, r5, r6, r7; + { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; } + { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; } + { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; } + { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; } + { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; } + { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; } + { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; } + { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; } + + for (uint32_t it = 0u; it < 8u; ++it) { + uint32_t sel = r0; + r2 = r3 * r4 + r2; // 0 mad + r1 = r1 + r7 + ((((sel >> 4u) & 1u) != 0u) ? 0xc3bd2355u : 0x42da7657u); // 1 add + r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 2 add + r4 = r0 * r6 + r4; // 3 mad + { uint32_t s_ = r2 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 scratch + r4 = r4 ^ ds[r1 & mask]; // 5 load + r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 6 shfl + r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 7 shfl + r7 = r7 ^ r5; // 8 xor + r3 = r3 | r4; // 9 or + r1 = r1 | r2; // 10 or + r4 = r4 ^ ds[r3 & mask]; // 11 load + r6 = r6 | r2; // 12 or + r2 = r2 * r5; // 13 mul + r1 = r1 ^ ds[r2 & mask]; // 14 load + r7 = rotl_imm(r7, 1u); // 15 rotl + { uint32_t s_ = r6 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 scratch + r7 = r7 ^ ds[r4 & mask]; // 17 load + r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 18 shfl + r4 = r0 * r2 + r4; // 19 mad + r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 20 shfl + r5 = r5 ^ r7; // 21 xor + r2 = __umulhi(r2, r5); // 22 mulhi + { uint32_t s_ = r7 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 scratch + r7 = __umulhi(r7, r3); // 24 mulhi + r5 = r5 | r4; // 25 or + r4 = r5 * r2 + r4; // 26 mad + r5 = r5 * r1; // 27 mul + r6 = __umulhi(r6, r7); // 28 mulhi + r6 = r6 + r1 + ((((sel >> 9u) & 1u) != 0u) ? 0x187a9128u : 0x3b2d2124u); // 29 add + r6 = rotr_var(r6, r7); // 30 rotr + r3 = r3 ^ ds[r1 & mask]; // 31 load + { uint32_t s_ = r0 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 scratch + r0 = r0 + r4 + ((((sel >> 18u) & 1u) != 0u) ? 0x351dde38u : 0x2c35699fu); // 33 add + { uint32_t s_ = r2 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 scratch + r0 = r0 * r3; // 35 mul + r2 = r2 ^ r5; // 36 xor + { uint32_t s_ = r0 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 scratch + r1 = r3 * r5 + r1; // 38 mad + r0 = r0 + r3 + ((((sel >> 25u) & 1u) != 0u) ? 0x1b053acfu : 0xa907b90bu); // 39 add + r2 = rotr_var(r2, r5); // 40 rotr + r3 = r3 * r2; // 41 mul + r1 = r1 + r5 + ((((sel >> 20u) & 1u) != 0u) ? 0x6058c2e3u : 0xa32e000cu); // 42 add + r3 = r3 ^ r4; // 43 xor + { uint32_t s_ = r5 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 scratch + r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 45 add + r7 = r7 ^ r1; // 46 xor + r0 = r0 + r3 + ((((sel >> 31u) & 1u) != 0u) ? 0x36360066u : 0x838b5065u); // 47 add + r7 = __umulhi(r7, r5); // 48 mulhi + r0 = r0 ^ ds[r2 & mask]; // 49 load + r2 = r2 - r6; // 50 sub + r7 = r7 - r5; // 51 sub + r2 = r2 ^ r3; // 52 xor + r7 = r7 - r0; // 53 sub + r3 = r5 * r0 + r3; // 54 mad + r7 = r7 ^ r5; // 55 xor + r2 = r2 ^ ds[r7 & mask]; // 56 load + r5 = r5 - r6; // 57 sub + r1 = r1 ^ ds[r3 & mask]; // 58 load + { uint32_t s_ = r4 & 63u; uint4 v_ = *(const uint4*)(arena + s_ * 4u); uint32_t m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint32_t w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint32_t w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint32_t w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint32_t x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 scratch + r4 = r4 - r6; // 60 sub + r1 = r1 * r2; // 61 mul + r3 = r6 * r0 + r3; // 62 mad + r0 = r0 + r1 + ((((sel >> 16u) & 1u) != 0u) ? 0xc88e2942u : 0x2fe0e98bu); // 63 add + } + uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo; + } +} + +// Variant 5: the wrapper launches `warps` persistent warps over `nonces / 32` units (host.cu does not use it). +cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, + IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps, uint32_t* scratch, uint32_t warps, uint32_t salt) { + if (blockWarps == 0u || blockWarps > 32u || warps == 0u || (warps % blockWarps) != 0u) return cudaErrorInvalidValue; + uint32_t block = 32u * blockWarps; + if (nonces == 0u || (nonces % (32u * warps)) != 0u) return cudaErrorInvalidValue; + igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw, scratch, nonces / 32u, salt); + return cudaGetLastError(); +} + +cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) { + cudaFuncAttributes attr; + cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound); + if (e != cudaSuccess) return e; + *numRegs = attr.numRegs; + return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0); +} diff --git a/proto-cuda/packs-readwidth/scr8k32/memhard.h b/proto-cuda/packs-readwidth/scr8k32/memhard.h new file mode 100644 index 000000000..4803d8e40 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k32/memhard.h @@ -0,0 +1,108 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against. +// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference). +// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C. +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif +#if defined(__CUDACC__) +#define IGNEUM_HD __host__ __device__ __forceinline__ +#elif defined(_MSC_VER) && !defined(__cplusplus) +#define IGNEUM_HD static __inline +#else +#define IGNEUM_HD static inline +#endif +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) { + for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint32_t r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) { + uint32_t prev[16]; uint32_t x[16]; uint32_t y[16]; + for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint32_t r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } diff --git a/proto-cuda/packs-readwidth/scr8k32/memhard.metal b/proto-cuda/packs-readwidth/scr8k32/memhard.metal new file mode 100644 index 000000000..241866369 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k32/memhard.metal @@ -0,0 +1,106 @@ +#include +using namespace metal; +// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines. +// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals. +#define MH_CACHE_LINE_MASK 0x003fffffu +#define MH_SEGMENT_LINES 64u +#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); } +inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site + +// y = ChaCha12 core(x) + x +inline void mh_chacha_block(const thread uint* x, thread uint* y) { + for (uint i = 0u; i < 16u; ++i) y[i] = x[i]; + for (uint r = 0u; r < 6u; ++r) { + MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u) + MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u) + MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u) + } + for (uint i = 0u; i < 16u; ++i) y[i] += x[i]; +} + +// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0. +inline void mh_cache_segment(device uint* cache, uint seg) { + uint prev[16]; uint x[16]; uint y[16]; + for (uint i = 0u; i < 16u; ++i) prev[i] = 0u; + for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) { + x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3]; + x[4] = 0x3067619fu ^ prev[4]; + x[5] = 0x3c269176u ^ prev[5]; + x[6] = 0x84a03b03u ^ prev[6]; + x[7] = 0xf8c63294u ^ prev[7]; + x[8] = 0xff977c5bu ^ prev[8]; + x[9] = 0xe60def3eu ^ prev[9]; + x[10] = 0x63630141u ^ prev[10]; + x[11] = 0xb8fbcb58u ^ prev[11]; + x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15]; + mh_chacha_block(x, y); + device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u); + for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; } + } +} + +// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations. +inline void mh_mixer(thread uint* s, uint rk) { + s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u; + s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu; + s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u; + s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu; + s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u; + s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u; + s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u; + s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u; + s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du; + s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u; + s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du; + s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu; + s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du; + s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu; + s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u; + s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u; + MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u) + MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u) + MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u) + MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u) +} + +// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer. +inline void mh_item(device const uint* cache, uint t, thread uint* s) { + s[0] = 0x3067619fu; + s[1] = 0x3c269176u; + s[2] = 0x84a03b03u; + s[3] = 0xf8c63294u; + s[4] = 0xff977c5bu; + s[5] = 0xe60def3eu; + s[6] = 0x63630141u; + s[7] = 0xb8fbcb58u; + s[8] = t * 0x42146205u + 0xbab68293u; + s[9] = t * 0x52cbe0fbu + 0xcc162340u; + s[10] = t * 0x7ecf4a03u + 0x6ce151ccu; + s[11] = t * 0x6728907fu + 0xe62b8997u; + s[12] = t * 0xd81d9751u + 0xc9c80297u; + s[13] = t * 0x132952c3u + 0xf74a1654u; + s[14] = t * 0xf60de277u + 0x3d704af5u; + s[15] = t * 0x05358035u + 0x3cf522b7u; + for (uint r = 0u; r < 8u; ++r) { + mh_mixer(s, 0x9E3779B9u * (r + 1u)); + device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); + for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i]; + } + mh_mixer(s, 0x9E3779B9u * 9u); +} +// dataset[w] without the dataset: derive item w >> 4 and take word w & 15. +inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; } + +// One thread per segment (2^16 threads). +kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) { + mh_cache_segment(cache, gid); +} +// One thread per 64-byte item (dataset words / 16 threads). +kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]], + uint gid [[thread_position_in_grid]]) { + uint s[16]; + mh_item(cache, gid, s); + device uint* d = dataset + gid * 16u; + for (uint i = 0u; i < 16u; ++i) d[i] = s[i]; +} diff --git a/proto-cuda/packs-readwidth/scr8/program.h b/proto-cuda/packs-readwidth/scr8k32/program.h similarity index 91% rename from proto-cuda/packs-readwidth/scr8/program.h rename to proto-cuda/packs-readwidth/scr8k32/program.h index 769c814b6..d201dd8b0 100644 --- a/proto-cuda/packs-readwidth/scr8/program.h +++ b/proto-cuda/packs-readwidth/scr8k32/program.h @@ -15,7 +15,7 @@ #define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973" #define IGNEUM_GENERATOR 2 #define IGNEUM_PROGRAM_ATTEMPT 0 -#define IGNEUM_PROGRAM_ID 0x2f0992e568f38dc1ull +#define IGNEUM_PROGRAM_ID 0xe0d27cd155da1553ull #define IGNEUM_DAY_STRING "2026-10-03" #define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033" #define IGNEUM_DAY0 0x3067619fu @@ -30,20 +30,20 @@ #define IGNEUM_OP_MIX "add=9 load=8 scratch=8 mad=7 xor=7 mul=5 sub=5 mulhi=4 or=4 shfl=4 rotr=2 rotl=1" // Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads // the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x. -#define IGNEUM_LOAD_CLASS "scr8" +#define IGNEUM_LOAD_CLASS "scr8k32" #define IGNEUM_LOAD_SLOTS 16 #define IGNEUM_LOAD_MIX { 100, 0, 0 } #define IGNEUM_LOAD_WIDTH_COUNTS { 8, 0, 0 } // loads of 4, 16, 64 bytes per program #define IGNEUM_BYTES_PER_HASH 256 #define IGNEUM_FOLD_ROT 11 #define IGNEUM_FOLD_MUL 0x9e3779b1u -// Variant 5: persistent warps, a 1 MiB scratch per launched warp (the host launches N warps and passes scratch, +// Variant 5: persistent warps, a 32 KiB scratch per launched warp (the host launches N warps and passes scratch, // groups and salt as the last three kernel arguments; groups must be a multiple of N; the tag of a unit is salt + unit). #define IGNEUM_PERSISTENT_WARPS 1 #define IGNEUM_SCRATCH_OPS 8 // scratch read-modify-writes per program (64 per hash) -#define IGNEUM_SCRATCH_SLOTS 2048u -#define IGNEUM_SCRATCH_WORDS_PER_LANE 8192u -#define IGNEUM_SCRATCH_BYTES_PER_WARP 1048576u +#define IGNEUM_SCRATCH_SLOTS 64u +#define IGNEUM_SCRATCH_WORDS_PER_LANE 256u +#define IGNEUM_SCRATCH_BYTES_PER_WARP 32768u // 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h) #define IGNEUM_DATASET_MODE 1 diff --git a/proto-cuda/packs-readwidth/scr8k32/program.json b/proto-cuda/packs-readwidth/scr8k32/program.json new file mode 100644 index 000000000..3b0163f57 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k32/program.json @@ -0,0 +1,130 @@ +{ + "format": "igneum-program-pack-3", + "generator": 2, + "attempt": 0, + "program_id": "0xe0d27cd155da1553", + "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32", + "dataset_mode": "memory-hard", + "seed": "igneum-genesis", + "seed_bytes": "69676e65756d2d67656e65736973", + "seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"], + "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32", + "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried", + "lanes": 32, + "registers": 8, + "iterations": 8, + "instruction_count": 64, + "loads_per_hash": 128, + "load_class": "scr8k32", + "load_slots": 16, + "load_mix_percent_4_16_64": [100, 0, 0], + "load_width_counts_4_16_64": [8, 0, 0], + "bytes_per_hash": 256, + "scratch_ops_per_hash": 64, + "scratch_kib_per_warp": 32, + "scratch": "variant 5 (measurement only): persistent warps; a 32 KiB scratch per warp of 64 16-byte slots per lane (lane-major); slot = src & 0x3f; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)", + "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots", + "op_mix": {"add": 9, "load": 8, "scratch": 8, "mad": 7, "xor": 7, "mul": 5, "sub": 5, "mulhi": 4, "or": 4, "shfl": 4, "rotr": 2, "rotl": 1}, + "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]", + "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16", + "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order", + "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo", + "op_semantics": { + "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)", + "sub": "dst = dst - src", + "mul": "dst = dst * src (low 32)", + "mulhi": "dst = high 32 bits of dst * src", + "xor": "dst = dst ^ src", + "or": "dst = dst | src", + "rotl": "dst = rotl(dst, rot), rot in 1..31", + "rotr": "dst = rotr(dst, src & 31)", + "mad": "dst = src * src2 + dst", + "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp", + "load": "dst = dst ^ dataset[src & dataset.mask]", + "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)" + }, + "dataset": { + "log2_words": 28, + "bytes": 1073741824, + "mask": "0x0fffffff", + "day": "2026-10-03", + "day_bytes": "6461792f323032362d31302d3033", + "day_words_from": "seed_words_from_bytes(day_bytes)", + "d0": "0x3067619f", + "d1": "0x3c269176", + "mode": "memory-hard", + "spec": "proto-metal/MEMHARD.md", + "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"], + "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]", + "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"}, + "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"}, + "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s", + "word": "dataset[w] = item(w >> 4)[w & 15]" + }, + "instructions": [ + {"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2, "width": 1}, + {"i": 1, "op": "add", "dst": 1, "src": 7, "src2": 2, "imm": "0x42da7657", "imm2": "0xc3bd2355", "rot": 25, "bit": 4, "mask": 16, "width": 1}, + {"i": 2, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x61f0b51c", "imm2": "0x2735a174", "rot": 4, "bit": 26, "mask": 2, "width": 1}, + {"i": 3, "op": "mad", "dst": 4, "src": 0, "src2": 6, "imm": "0x679648a8", "imm2": "0x3044ba32", "rot": 31, "bit": 31, "mask": 4, "width": 1}, + {"i": 4, "op": "scratch", "dst": 7, "src": 2, "src2": 6, "imm": "0x5d1ca2a2", "imm2": "0xe2481807", "rot": 24, "bit": 3, "mask": 1, "width": 1}, + {"i": 5, "op": "load", "dst": 4, "src": 1, "src2": 2, "imm": "0x987c017a", "imm2": "0xf4d60559", "rot": 2, "bit": 0, "mask": 4, "width": 1}, + {"i": 6, "op": "shfl", "dst": 6, "src": 3, "src2": 7, "imm": "0x6ea7b2df", "imm2": "0x9fce5071", "rot": 7, "bit": 15, "mask": 4, "width": 1}, + {"i": 7, "op": "shfl", "dst": 1, "src": 5, "src2": 1, "imm": "0x26a2ecde", "imm2": "0xfec6ad22", "rot": 15, "bit": 11, "mask": 8, "width": 1}, + {"i": 8, "op": "xor", "dst": 7, "src": 5, "src2": 2, "imm": "0xbe4b445c", "imm2": "0x17a5a9c7", "rot": 8, "bit": 8, "mask": 1, "width": 1}, + {"i": 9, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4, "width": 1}, + {"i": 10, "op": "or", "dst": 1, "src": 2, "src2": 3, "imm": "0x4e7dc10d", "imm2": "0x196d165c", "rot": 14, "bit": 27, "mask": 16, "width": 1}, + {"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 1, "imm": "0xc5c3b55d", "imm2": "0xec061424", "rot": 26, "bit": 27, "mask": 8, "width": 1}, + {"i": 12, "op": "or", "dst": 6, "src": 2, "src2": 3, "imm": "0x306542fe", "imm2": "0x1bb1b429", "rot": 31, "bit": 0, "mask": 2, "width": 1}, + {"i": 13, "op": "mul", "dst": 2, "src": 5, "src2": 6, "imm": "0xa672cdd3", "imm2": "0x59a4829c", "rot": 22, "bit": 13, "mask": 16, "width": 1}, + {"i": 14, "op": "load", "dst": 1, "src": 2, "src2": 5, "imm": "0x028b4d37", "imm2": "0x7bbd78ea", "rot": 15, "bit": 2, "mask": 8, "width": 1}, + {"i": 15, "op": "rotl", "dst": 7, "src": 6, "src2": 6, "imm": "0x5c88a1a7", "imm2": "0x5c628769", "rot": 1, "bit": 3, "mask": 8, "width": 1}, + {"i": 16, "op": "scratch", "dst": 3, "src": 6, "src2": 7, "imm": "0xbac2ae81", "imm2": "0xcbbc7bdb", "rot": 18, "bit": 8, "mask": 8, "width": 1}, + {"i": 17, "op": "load", "dst": 7, "src": 4, "src2": 2, "imm": "0xe8ab93e9", "imm2": "0xa00de107", "rot": 2, "bit": 1, "mask": 16, "width": 1}, + {"i": 18, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2, "width": 1}, + {"i": 19, "op": "mad", "dst": 4, "src": 0, "src2": 2, "imm": "0x5fba7bc2", "imm2": "0xdf099cfb", "rot": 4, "bit": 15, "mask": 16, "width": 1}, + {"i": 20, "op": "shfl", "dst": 0, "src": 6, "src2": 3, "imm": "0x0a3056de", "imm2": "0x7f0c25c3", "rot": 27, "bit": 13, "mask": 8, "width": 1}, + {"i": 21, "op": "xor", "dst": 5, "src": 7, "src2": 4, "imm": "0xbd066e1d", "imm2": "0x6d3ddc5a", "rot": 2, "bit": 29, "mask": 1, "width": 1}, + {"i": 22, "op": "mulhi", "dst": 2, "src": 5, "src2": 0, "imm": "0xc7e9887a", "imm2": "0x19ec898f", "rot": 14, "bit": 9, "mask": 1, "width": 1}, + {"i": 23, "op": "scratch", "dst": 3, "src": 7, "src2": 2, "imm": "0xc7fcfc8f", "imm2": "0x8528b94f", "rot": 17, "bit": 13, "mask": 4, "width": 1}, + {"i": 24, "op": "mulhi", "dst": 7, "src": 3, "src2": 5, "imm": "0xd91641e8", "imm2": "0xaf77faf2", "rot": 22, "bit": 21, "mask": 1, "width": 1}, + {"i": 25, "op": "or", "dst": 5, "src": 4, "src2": 0, "imm": "0x84c03868", "imm2": "0xf6c691b7", "rot": 29, "bit": 14, "mask": 8, "width": 1}, + {"i": 26, "op": "mad", "dst": 4, "src": 5, "src2": 2, "imm": "0x3bb2b6ba", "imm2": "0x49d95fd5", "rot": 1, "bit": 5, "mask": 8, "width": 1}, + {"i": 27, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4, "width": 1}, + {"i": 28, "op": "mulhi", "dst": 6, "src": 7, "src2": 6, "imm": "0xd69c4715", "imm2": "0xe0ebc4ce", "rot": 29, "bit": 2, "mask": 8, "width": 1}, + {"i": 29, "op": "add", "dst": 6, "src": 1, "src2": 2, "imm": "0x3b2d2124", "imm2": "0x187a9128", "rot": 1, "bit": 9, "mask": 16, "width": 1}, + {"i": 30, "op": "rotr", "dst": 6, "src": 7, "src2": 0, "imm": "0x5c64a589", "imm2": "0x61c9a38d", "rot": 17, "bit": 21, "mask": 16, "width": 1}, + {"i": 31, "op": "load", "dst": 3, "src": 1, "src2": 7, "imm": "0xc37723fa", "imm2": "0xf3b024da", "rot": 16, "bit": 27, "mask": 16, "width": 1}, + {"i": 32, "op": "scratch", "dst": 1, "src": 0, "src2": 7, "imm": "0xcc7972c4", "imm2": "0xad098d15", "rot": 30, "bit": 21, "mask": 8, "width": 1}, + {"i": 33, "op": "add", "dst": 0, "src": 4, "src2": 4, "imm": "0x2c35699f", "imm2": "0x351dde38", "rot": 21, "bit": 18, "mask": 4, "width": 1}, + {"i": 34, "op": "scratch", "dst": 0, "src": 2, "src2": 3, "imm": "0xfae8902b", "imm2": "0x5cd8306f", "rot": 5, "bit": 28, "mask": 16, "width": 1}, + {"i": 35, "op": "mul", "dst": 0, "src": 3, "src2": 1, "imm": "0x4fa3f3db", "imm2": "0xdbf37e75", "rot": 7, "bit": 18, "mask": 4, "width": 1}, + {"i": 36, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2, "width": 1}, + {"i": 37, "op": "scratch", "dst": 4, "src": 0, "src2": 0, "imm": "0x04cc1d55", "imm2": "0x35c52d04", "rot": 11, "bit": 14, "mask": 2, "width": 1}, + {"i": 38, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0x3958f280", "imm2": "0x8713c7e1", "rot": 5, "bit": 23, "mask": 16, "width": 1}, + {"i": 39, "op": "add", "dst": 0, "src": 3, "src2": 3, "imm": "0xa907b90b", "imm2": "0x1b053acf", "rot": 30, "bit": 25, "mask": 16, "width": 1}, + {"i": 40, "op": "rotr", "dst": 2, "src": 5, "src2": 4, "imm": "0xf8662282", "imm2": "0x10bb9e30", "rot": 8, "bit": 6, "mask": 2, "width": 1}, + {"i": 41, "op": "mul", "dst": 3, "src": 2, "src2": 4, "imm": "0x49087d74", "imm2": "0x6348b489", "rot": 17, "bit": 9, "mask": 16, "width": 1}, + {"i": 42, "op": "add", "dst": 1, "src": 5, "src2": 1, "imm": "0xa32e000c", "imm2": "0x6058c2e3", "rot": 25, "bit": 20, "mask": 8, "width": 1}, + {"i": 43, "op": "xor", "dst": 3, "src": 4, "src2": 2, "imm": "0x3dad0eb6", "imm2": "0xb97578cb", "rot": 3, "bit": 27, "mask": 1, "width": 1}, + {"i": 44, "op": "scratch", "dst": 3, "src": 5, "src2": 7, "imm": "0x374aec92", "imm2": "0x626f11df", "rot": 20, "bit": 18, "mask": 8, "width": 1}, + {"i": 45, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4, "width": 1}, + {"i": 46, "op": "xor", "dst": 7, "src": 1, "src2": 0, "imm": "0xef6ac348", "imm2": "0x963bb7e6", "rot": 26, "bit": 3, "mask": 8, "width": 1}, + {"i": 47, "op": "add", "dst": 0, "src": 3, "src2": 0, "imm": "0x838b5065", "imm2": "0x36360066", "rot": 3, "bit": 31, "mask": 4, "width": 1}, + {"i": 48, "op": "mulhi", "dst": 7, "src": 5, "src2": 0, "imm": "0x8458f7ac", "imm2": "0xc1c15026", "rot": 27, "bit": 15, "mask": 8, "width": 1}, + {"i": 49, "op": "load", "dst": 0, "src": 2, "src2": 4, "imm": "0x636a9dc4", "imm2": "0xac023d9b", "rot": 22, "bit": 29, "mask": 1, "width": 1}, + {"i": 50, "op": "sub", "dst": 2, "src": 6, "src2": 0, "imm": "0x2baec8c9", "imm2": "0x4390f156", "rot": 3, "bit": 12, "mask": 8, "width": 1}, + {"i": 51, "op": "sub", "dst": 7, "src": 5, "src2": 7, "imm": "0x19234061", "imm2": "0xe84dfade", "rot": 4, "bit": 19, "mask": 1, "width": 1}, + {"i": 52, "op": "xor", "dst": 2, "src": 3, "src2": 5, "imm": "0xdc2cd71e", "imm2": "0x1b5d334b", "rot": 9, "bit": 8, "mask": 8, "width": 1}, + {"i": 53, "op": "sub", "dst": 7, "src": 0, "src2": 4, "imm": "0x605c31ec", "imm2": "0x9923ff88", "rot": 28, "bit": 25, "mask": 4, "width": 1}, + {"i": 54, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4, "width": 1}, + {"i": 55, "op": "xor", "dst": 7, "src": 5, "src2": 5, "imm": "0xad7493e7", "imm2": "0x3e400372", "rot": 13, "bit": 8, "mask": 1, "width": 1}, + {"i": 56, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0x87e933c9", "imm2": "0x8c854c1b", "rot": 17, "bit": 3, "mask": 8, "width": 1}, + {"i": 57, "op": "sub", "dst": 5, "src": 6, "src2": 5, "imm": "0x11be3bc9", "imm2": "0xbbaa8e24", "rot": 6, "bit": 5, "mask": 16, "width": 1}, + {"i": 58, "op": "load", "dst": 1, "src": 3, "src2": 2, "imm": "0xa732351a", "imm2": "0xc01349cd", "rot": 14, "bit": 17, "mask": 16, "width": 1}, + {"i": 59, "op": "scratch", "dst": 1, "src": 4, "src2": 0, "imm": "0xb20547b2", "imm2": "0xc94655de", "rot": 27, "bit": 30, "mask": 1, "width": 1}, + {"i": 60, "op": "sub", "dst": 4, "src": 6, "src2": 7, "imm": "0x67cf904c", "imm2": "0x6873b216", "rot": 27, "bit": 7, "mask": 16, "width": 1}, + {"i": 61, "op": "mul", "dst": 1, "src": 2, "src2": 7, "imm": "0x93ab0bf4", "imm2": "0x96158375", "rot": 14, "bit": 0, "mask": 16, "width": 1}, + {"i": 62, "op": "mad", "dst": 3, "src": 6, "src2": 0, "imm": "0x41a443a3", "imm2": "0xe69d7919", "rot": 9, "bit": 0, "mask": 16, "width": 1}, + {"i": 63, "op": "add", "dst": 0, "src": 1, "src2": 3, "imm": "0x2fe0e98b", "imm2": "0xc88e2942", "rot": 5, "bit": 16, "mask": 16, "width": 1} + ] +} diff --git a/proto-cuda/packs-readwidth/scr8k32/program.metal b/proto-cuda/packs-readwidth/scr8k32/program.metal new file mode 100644 index 000000000..573e63513 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k32/program.metal @@ -0,0 +1,126 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +kernel void igneum_hash(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + device uint* scratch [[buffer(3)]], + constant uint& groups [[buffer(4)]], + constant uint& salt [[buffer(5)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; } + { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; } + { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; } + { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; } + { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; } + { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; } + { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; } + { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + { uint s_ = r2 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint s_ = r6 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + { uint s_ = r7 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + { uint s_ = r0 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint s_ = r2 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + { uint s_ = r0 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + { uint s_ = r5 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + { uint s_ = r4 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr8k32/program_bound.metal b/proto-cuda/packs-readwidth/scr8k32/program_bound.metal new file mode 100644 index 000000000..e3ec2c039 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k32/program_bound.metal @@ -0,0 +1,128 @@ +#include +using namespace metal; + +#define MASK 0x0fffffffu +constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }; + +inline uint splitmix32(uint x) { + x ^= x >> 16; x *= 0x7feb352du; + x ^= x >> 15; x *= 0x846ca68bu; + x ^= x >> 16; + return x; +} +inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 +inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); } +inline uint ds_elem(uint i, uint d0, uint d1) { + uint x = i ^ d0; + x *= 0x9E3779B1u; x ^= x >> 15; + x += d1; + x *= 0x85EBCA77u; x ^= x >> 13; + x *= 0xC2B2AE3Du; x ^= x >> 16; + return x; +} + +// Variant 5 (read-width experiment, 5 October 2026, NOT the lottery hash): a 32 KiB scratch per warp, 64 slots of +// 16 bytes per lane, lane-major. A slot starts the unit as the fill words below (tagged lazily: a slot whose tag is not +// this unit's reads as its fill) and holds what the unit wrote afterwards. scr_fill mirrors verify::scratch_fill. +inline uint scr_fill(uint gbase, uint lane, uint slot, uint j) { uint sw = (j == 0u) ? 0x67a9a7beu : ((j == 1u) ? 0x1a155b25u : 0xfddfb732u); return splitmix32(((gbase + lane) ^ sw) + slot * 0x9e3779b1u + (j + 1u) * 0x85ebca77u); } +// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW. +kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]], + device ulong* out [[buffer(1)]], + constant uint& baseNonce [[buffer(2)]], + constant uint* initw [[buffer(3)]], + device uint* scratch [[buffer(4)]], + constant uint& groups [[buffer(5)]], + constant uint& salt [[buffer(6)]], + uint tid [[thread_position_in_grid]], + uint nthreads [[threads_per_grid]]) { + uint lane = tid & 31u; + uint warp_ = tid >> 5; + uint nwarps_ = nthreads >> 5; + device uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; } + { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; } + { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; } + { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; } + { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; } + { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; } + { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; } + { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; } + + for (uint it = 0u; it < 8u; ++it) { + uint sel = r0; + r2 = r3 * r4 + r2; // 0 + r1 = r1 + r7 + select(0x42da7657u, 0xc3bd2355u, ((sel >> 4u) & 1u) != 0u); // 1 + r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 2 + r4 = r0 * r6 + r4; // 3 + { uint s_ = r2 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r7 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r7 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 4 + r4 = r4 ^ dataset[r1 & MASK]; // 5 + r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 6 + r1 = r1 ^ simd_shuffle_xor(r5, (ushort)8); // 7 + r7 = r7 ^ r5; // 8 + r3 = r3 | r4; // 9 + r1 = r1 | r2; // 10 + r4 = r4 ^ dataset[r3 & MASK]; // 11 + r6 = r6 | r2; // 12 + r2 = r2 * r5; // 13 + r1 = r1 ^ dataset[r2 & MASK]; // 14 + r7 = rotl_imm(r7, 1u); // 15 + { uint s_ = r6 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 16 + r7 = r7 ^ dataset[r4 & MASK]; // 17 + r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 18 + r4 = r0 * r2 + r4; // 19 + r0 = r0 ^ simd_shuffle_xor(r6, (ushort)8); // 20 + r5 = r5 ^ r7; // 21 + r2 = mulhi(r2, r5); // 22 + { uint s_ = r7 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 23 + r7 = mulhi(r7, r3); // 24 + r5 = r5 | r4; // 25 + r4 = r5 * r2 + r4; // 26 + r5 = r5 * r1; // 27 + r6 = mulhi(r6, r7); // 28 + r6 = r6 + r1 + select(0x3b2d2124u, 0x187a9128u, ((sel >> 9u) & 1u) != 0u); // 29 + r6 = rotr_var(r6, r7); // 30 + r3 = r3 ^ dataset[r1 & MASK]; // 31 + { uint s_ = r0 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 32 + r0 = r0 + r4 + select(0x2c35699fu, 0x351dde38u, ((sel >> 18u) & 1u) != 0u); // 33 + { uint s_ = r2 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r0 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r0 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 34 + r0 = r0 * r3; // 35 + r2 = r2 ^ r5; // 36 + { uint s_ = r0 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r4 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r4 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 37 + r1 = r3 * r5 + r1; // 38 + r0 = r0 + r3 + select(0xa907b90bu, 0x1b053acfu, ((sel >> 25u) & 1u) != 0u); // 39 + r2 = rotr_var(r2, r5); // 40 + r3 = r3 * r2; // 41 + r1 = r1 + r5 + select(0xa32e000cu, 0x6058c2e3u, ((sel >> 20u) & 1u) != 0u); // 42 + r3 = r3 ^ r4; // 43 + { uint s_ = r5 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r3 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r3 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 44 + r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 45 + r7 = r7 ^ r1; // 46 + r0 = r0 + r3 + select(0x838b5065u, 0x36360066u, ((sel >> 31u) & 1u) != 0u); // 47 + r7 = mulhi(r7, r5); // 48 + r0 = r0 ^ dataset[r2 & MASK]; // 49 + r2 = r2 - r6; // 50 + r7 = r7 - r5; // 51 + r2 = r2 ^ r3; // 52 + r7 = r7 - r0; // 53 + r3 = r5 * r0 + r3; // 54 + r7 = r7 ^ r5; // 55 + r2 = r2 ^ dataset[r7 & MASK]; // 56 + r5 = r5 - r6; // 57 + r1 = r1 ^ dataset[r3 & MASK]; // 58 + { uint s_ = r4 & 63u; uint4 v_ = *(device const uint4*)(arena + s_ * 4u); uint m_ = (v_.x == tag) ? 0xffffffffu : 0u; uint w0_ = (v_.y & m_) | (scr_fill(gbase, lane, s_, 0u) & ~m_); uint w1_ = (v_.z & m_) | (scr_fill(gbase, lane, s_, 1u) & ~m_); uint w2_ = (v_.w & m_) | (scr_fill(gbase, lane, s_, 2u) & ~m_); uint x_ = r1 ^ w0_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w1_; x_ = (rotl_imm(x_, 11u) * 0x9e3779b1u) ^ w2_; r1 = x_; *(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); } // 59 + r4 = r4 - r6; // 60 + r1 = r1 * r2; // 61 + r3 = r6 * r0 + r3; // 62 + r0 = r0 + r1 + select(0x2fe0e98bu, 0xc88e2942u, ((sel >> 16u) & 1u) != 0u); // 63 + } + uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u); + uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u); + out[gid] = ((ulong)hi << 32) | (ulong)lo; + } +} diff --git a/proto-cuda/packs-readwidth/scr8k32/vectors.h b/proto-cuda/packs-readwidth/scr8k32/vectors.h new file mode 100644 index 000000000..3d095f577 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k32/vectors.h @@ -0,0 +1,57 @@ +// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand. +// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset +#pragma once +#ifdef __cplusplus +#include +#else +#include +#endif + +#define IGNEUM_VEC_WARPS 3 +static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u }; +static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = { + { // base nonce 0 + 0x82097370c10ee4eaull, 0xa0f16965a422e758ull, 0x2b7513e2593b170cull, 0x8fbb69f8974c4533ull, 0x500c384970d19f32ull, 0x72fb96e1820fa2b8ull, 0xca1521df6a934a65ull, 0x7407397e10548cbcull, + 0xa619f127d1aa896bull, 0x36ce1104c1517d86ull, 0x6078ddd46d6d26b2ull, 0x2d22ded89010fbd3ull, 0xb030b743355bf6f9ull, 0x8280c8c1f4946ef7ull, 0xbfdae5712c7b9ad4ull, 0x50e9136a89e2e046ull, + 0xfe0230d874a94f3aull, 0x8b6dbb8ff0546c32ull, 0x41438f3ab2d35cbcull, 0x6faf2dd5c5997f9cull, 0x52c3aae40019a0abull, 0xc455bf9926661472ull, 0x4eaf9e34af6c7c22ull, 0x9dcd05c3a3991115ull, + 0x3a4f0ed406a66796ull, 0x7615703c08eabd8aull, 0x6007bcf34b8c6341ull, 0x1fb11900437441aeull, 0x540cff895991c6acull, 0x367a10121393f054ull, 0xf0dd7aa43fbdc2f4ull, 0x53d26f9c120f8a64ull + }, + { // base nonce 4096 + 0xc04adc4c89d1593bull, 0xc938ee393ee383dcull, 0xaef535444b6968fbull, 0x83581d97cebb4b30ull, 0x365315d0762bf586ull, 0xc26dfc2c40131920ull, 0x2f560c6fa351ead5ull, 0xc235996f6596d536ull, + 0x1adf3384c9f23712ull, 0x0c87a8ad0dc872c3ull, 0x82471bfd2a3182d4ull, 0xe8828b0fb7560877ull, 0x3fd9044dd153dd7full, 0x3a91a9618ce2c525ull, 0xf8c503abac8481f9ull, 0x56b0989f6014f8dbull, + 0xa1107a2732c588baull, 0xfc47a39e530d7efaull, 0x9237f7fc01727703ull, 0x01b5cbb25e6629c1ull, 0xae8e66fcec949f3eull, 0xb31655fbd75d3dbdull, 0x5a7750b585b3d0f1ull, 0x72f0484703cf30a0ull, + 0x4cea6de3f76e970bull, 0x64fd914fb1fb5096ull, 0xbdcca26f07e7182cull, 0xc86ba79905164a5bull, 0x521b7d3ac36f06cbull, 0x6b8618a80cef2e71ull, 0x13df43bef1f7852full, 0x202e92ce7b597a33ull + }, + { // base nonce 1000000 + 0x4db5bf37f24811f2ull, 0x5e93450594a45e5full, 0xebe57bb921f163aaull, 0x0457e6f5ac702b1eull, 0x442d8926fceec0d4ull, 0x417256d19cce43d3ull, 0x3d61563a8303daceull, 0x4a43c5efac6f6d47ull, + 0x9ef8a8d0a98f5c7eull, 0xbee88175926bb251ull, 0xb3b8e73f1e427be1ull, 0x515405b57446beecull, 0xba4f1765e616bf9aull, 0x148e5d9895c48299ull, 0x303fed05bdcdd7d1ull, 0x44d07bf31dba804full, + 0x8616fd225f3851beull, 0x426a7a80ac1b462full, 0x6e0163361c5ec30full, 0x065b3666feb8d0e5ull, 0xf9bc697886c9983eull, 0x90bc61f358b511f4ull, 0xfa47afe197811c66ull, 0x12a39e4e67aa2e97ull, + 0xde01f49ccab26a42ull, 0x2c6e837e74897413ull, 0x62d22c8acdeb1d09ull, 0x7fa8da035f65bb0bull, 0xdc4ce47ce0d48b6bull, 0x6199717653754041ull, 0x3a5e113c0d160d86ull, 0x718b3357f2391b60ull + } +}; + +// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455). +static const uint32_t IGNEUM_DS_HEAD[16] = { + 0xffc3cd94u, 0x5920ccd8u, 0x392f44bbu, 0x5e57f67au, 0x2f2bc2a9u, 0x620b0e36u, 0xbdc09014u, 0x436654bfu, + 0x311e0b48u, 0x1abd93adu, 0x59cc7ce8u, 0xee5247b2u, 0x86171fe8u, 0x6d874751u, 0xc9f7728fu, 0x7c2a435du +}; +static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u; +static const uint32_t IGNEUM_DS_LAST = 0xa33ada72u; +// 64 sampled dataset words (index, value) computed on the Mac. +#define IGNEUM_DS_SAMPLES 64 +static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = { + 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u +}; +static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = { + 0xe8b73d94u, 0x337028b5u, 0xafe148c9u, 0xab99f7aeu, 0x434ea619u, 0xd85cb880u, 0x54764c7fu, 0x82c7e420u, 0xedf4cb9eu, 0x9884c959u, 0x223ee793u, 0x3a9ccf69u, 0x81da4fd2u, 0xd6ce8cb9u, 0xe3922dcau, 0x3e7e6bdeu, 0x382a3acau, 0x567e7f7fu, 0x25a0f084u, 0xbfeef128u, 0xe338abfbu, 0x7c3b5280u, 0x909bc5f1u, 0xd8b74b9cu, 0x8e31a22eu, 0x26b5f1d8u, 0x79122c00u, 0xcafc3340u, 0xd5e02ea3u, 0x1aee1afdu, 0xdb090d9au, 0xb049f435u, 0x4954d8bau, 0x03797ba0u, 0x196eefbdu, 0xd153412au, 0xbe5d2c4bu, 0xdaa14f0eu, 0x8e61ed07u, 0x9e9a64c6u, 0x2e29ff36u, 0x392a8589u, 0xb56a5912u, 0xfa6e8b57u, 0xd1a737cbu, 0xb0fa841au, 0xbe1c341fu, 0xe25be0f1u, 0xe937f543u, 0xebab2248u, 0x8e1b607au, 0x202a2fedu, 0x95e2819cu, 0x9c9652d4u, 0x32fedef0u, 0xdecfff82u, 0xcb5d43e5u, 0xb735806au, 0x8905939cu, 0xfbf8472du, 0xada74e5du, 0x7ebdeeeau, 0x0119f2b3u, 0xa9a376b8u +}; +// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words. +static const uint32_t IGNEUM_CACHE_HEAD[16] = { + 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u, + 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u +}; +static const uint32_t IGNEUM_CACHE_LAST[16] = { + 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du, + 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu +}; +static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull; diff --git a/proto-cuda/packs-readwidth/scr8k32/vectors.json b/proto-cuda/packs-readwidth/scr8k32/vectors.json new file mode 100644 index 000000000..ce548d538 --- /dev/null +++ b/proto-cuda/packs-readwidth/scr8k32/vectors.json @@ -0,0 +1,36 @@ +{ + "seed": "igneum-genesis", + "day": "2026-10-03", + "dataset_mode": "memory-hard", + "dataset_log2_words": 28, + "mask": "0x0fffffff", + "lanes": 32, + "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset", + "warps": [ + {"base_nonce": 0, "expected": [ + "0x82097370c10ee4ea", "0xa0f16965a422e758", "0x2b7513e2593b170c", "0x8fbb69f8974c4533", "0x500c384970d19f32", "0x72fb96e1820fa2b8", "0xca1521df6a934a65", "0x7407397e10548cbc", + "0xa619f127d1aa896b", "0x36ce1104c1517d86", "0x6078ddd46d6d26b2", "0x2d22ded89010fbd3", "0xb030b743355bf6f9", "0x8280c8c1f4946ef7", "0xbfdae5712c7b9ad4", "0x50e9136a89e2e046", + "0xfe0230d874a94f3a", "0x8b6dbb8ff0546c32", "0x41438f3ab2d35cbc", "0x6faf2dd5c5997f9c", "0x52c3aae40019a0ab", "0xc455bf9926661472", "0x4eaf9e34af6c7c22", "0x9dcd05c3a3991115", + "0x3a4f0ed406a66796", "0x7615703c08eabd8a", "0x6007bcf34b8c6341", "0x1fb11900437441ae", "0x540cff895991c6ac", "0x367a10121393f054", "0xf0dd7aa43fbdc2f4", "0x53d26f9c120f8a64" + ]}, + {"base_nonce": 4096, "expected": [ + "0xc04adc4c89d1593b", "0xc938ee393ee383dc", "0xaef535444b6968fb", "0x83581d97cebb4b30", "0x365315d0762bf586", "0xc26dfc2c40131920", "0x2f560c6fa351ead5", "0xc235996f6596d536", + "0x1adf3384c9f23712", "0x0c87a8ad0dc872c3", "0x82471bfd2a3182d4", "0xe8828b0fb7560877", "0x3fd9044dd153dd7f", "0x3a91a9618ce2c525", "0xf8c503abac8481f9", "0x56b0989f6014f8db", + "0xa1107a2732c588ba", "0xfc47a39e530d7efa", "0x9237f7fc01727703", "0x01b5cbb25e6629c1", "0xae8e66fcec949f3e", "0xb31655fbd75d3dbd", "0x5a7750b585b3d0f1", "0x72f0484703cf30a0", + "0x4cea6de3f76e970b", "0x64fd914fb1fb5096", "0xbdcca26f07e7182c", "0xc86ba79905164a5b", "0x521b7d3ac36f06cb", "0x6b8618a80cef2e71", "0x13df43bef1f7852f", "0x202e92ce7b597a33" + ]}, + {"base_nonce": 1000000, "expected": [ + "0x4db5bf37f24811f2", "0x5e93450594a45e5f", "0xebe57bb921f163aa", "0x0457e6f5ac702b1e", "0x442d8926fceec0d4", "0x417256d19cce43d3", "0x3d61563a8303dace", "0x4a43c5efac6f6d47", + "0x9ef8a8d0a98f5c7e", "0xbee88175926bb251", "0xb3b8e73f1e427be1", "0x515405b57446beec", "0xba4f1765e616bf9a", "0x148e5d9895c48299", "0x303fed05bdcdd7d1", "0x44d07bf31dba804f", + "0x8616fd225f3851be", "0x426a7a80ac1b462f", "0x6e0163361c5ec30f", "0x065b3666feb8d0e5", "0xf9bc697886c9983e", "0x90bc61f358b511f4", "0xfa47afe197811c66", "0x12a39e4e67aa2e97", + "0xde01f49ccab26a42", "0x2c6e837e74897413", "0x62d22c8acdeb1d09", "0x7fa8da035f65bb0b", "0xdc4ce47ce0d48b6b", "0x6199717653754041", "0x3a5e113c0d160d86", "0x718b3357f2391b60" + ]} + ], + "dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"], + "dataset_last_index": 268435455, + "dataset_last": "0xa33ada72", + "dataset_samples": [{"index": 59471966, "value": "0xe8b73d94"}, {"index": 217795994, "value": "0x337028b5"}, {"index": 208353206, "value": "0xafe148c9"}, {"index": 42483309, "value": "0xab99f7ae"}, {"index": 172547758, "value": "0x434ea619"}, {"index": 148076330, "value": "0xd85cb880"}, {"index": 183853158, "value": "0x54764c7f"}, {"index": 214389424, "value": "0x82c7e420"}, {"index": 267488061, "value": "0xedf4cb9e"}, {"index": 169781097, "value": "0x9884c959"}, {"index": 184093494, "value": "0x223ee793"}, {"index": 153880993, "value": "0x3a9ccf69"}, {"index": 84977930, "value": "0x81da4fd2"}, {"index": 46426879, "value": "0xd6ce8cb9"}, {"index": 3093825, "value": "0xe3922dca"}, {"index": 225364072, "value": "0x3e7e6bde"}, {"index": 44593546, "value": "0x382a3aca"}, {"index": 260713159, "value": "0x567e7f7f"}, {"index": 168250303, "value": "0x25a0f084"}, {"index": 52384140, "value": "0xbfeef128"}, {"index": 223401610, "value": "0xe338abfb"}, {"index": 45554030, "value": "0x7c3b5280"}, {"index": 95410555, "value": "0x909bc5f1"}, {"index": 175039924, "value": "0xd8b74b9c"}, {"index": 79171087, "value": "0x8e31a22e"}, {"index": 267580473, "value": "0x26b5f1d8"}, {"index": 24168642, "value": "0x79122c00"}, {"index": 37981670, "value": "0xcafc3340"}, {"index": 171551130, "value": "0xd5e02ea3"}, {"index": 195559979, "value": "0x1aee1afd"}, {"index": 204611762, "value": "0xdb090d9a"}, {"index": 140997658, "value": "0xb049f435"}, {"index": 138925853, "value": "0x4954d8ba"}, {"index": 86637313, "value": "0x03797ba0"}, {"index": 20736778, "value": "0x196eefbd"}, {"index": 219665210, "value": "0xd153412a"}, {"index": 160430336, "value": "0xbe5d2c4b"}, {"index": 264654675, "value": "0xdaa14f0e"}, {"index": 8013395, "value": "0x8e61ed07"}, {"index": 228945585, "value": "0x9e9a64c6"}, {"index": 213884386, "value": "0x2e29ff36"}, {"index": 104419827, "value": "0x392a8589"}, {"index": 44185464, "value": "0xb56a5912"}, {"index": 142737231, "value": "0xfa6e8b57"}, {"index": 99284897, "value": "0xd1a737cb"}, {"index": 132475900, "value": "0xb0fa841a"}, {"index": 61861762, "value": "0xbe1c341f"}, {"index": 132056166, "value": "0xe25be0f1"}, {"index": 262388043, "value": "0xe937f543"}, {"index": 91878046, "value": "0xebab2248"}, {"index": 117353561, "value": "0x8e1b607a"}, {"index": 124768597, "value": "0x202a2fed"}, {"index": 71352993, "value": "0x95e2819c"}, {"index": 190698941, "value": "0x9c9652d4"}, {"index": 46055428, "value": "0x32fedef0"}, {"index": 55281366, "value": "0xdecfff82"}, {"index": 165145231, "value": "0xcb5d43e5"}, {"index": 106810753, "value": "0xb735806a"}, {"index": 171985651, "value": "0x8905939c"}, {"index": 232085256, "value": "0xfbf8472d"}, {"index": 159510492, "value": "0xada74e5d"}, {"index": 40072060, "value": "0x7ebdeeea"}, {"index": 209107596, "value": "0x0119f2b3"}, {"index": 39023794, "value": "0xa9a376b8"}], + "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"], + "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"], + "cache_fnv1a64": "0x48c4f5bf24166b2e" +} diff --git a/proto-metal/packbench.swift b/proto-metal/packbench.swift new file mode 100644 index 000000000..75b6fc4f3 --- /dev/null +++ b/proto-metal/packbench.swift @@ -0,0 +1,209 @@ +// packbench: runs a program pack (igneum-pow export) on the Mac's Metal GPU from its files alone: memhard.metal (cache +// fill, dataset build), program.metal (igneum_hash), program.h (constants), vectors.json (the CPU reference's vectors). +// Read-width experiment, 5 October 2026 (docs/plans/read-width.md): the Swift bench generates its own programs and does +// not know the experiment's load classes; this harness runs whatever text the Rust emitter wrote, so Metal is checked +// against the Rust CPU reference and timed without a Swift mirror of the generator. One file, no packages. +// +// swiftc -O -target arm64-apple-macos11 -o packbench packbench.swift -framework Metal +// ./packbench --pack [--batches 5] [--batch-log2 24] [--group 256] [--warps 2048] +// +// Prints one RESULT line per run: vectors, cache and dataset checks, the batch fingerprint (FNV-1a 64 over the 2^B +// outputs at base nonce 0) and MH/s by wall and by GPU time. Variant 5 packs (IGNEUM_PERSISTENT_WARPS) are launched +// as --warps persistent warps with a 1 MiB scratch each; the batch is rounded to a multiple of 32 x warps. +import Foundation +import Metal + +func nowMs() -> Double { return Double(DispatchTime.now().uptimeNanoseconds) / 1e6 } +func fail(_ m: String) -> Never { print("FAIL: \(m)"); exit(1) } + +struct Opts { var pack = ""; var batches = 5; var batchLog2 = 24; var group = 256; var warps = 2048 } +var opts = Opts() +var args = Array(CommandLine.arguments.dropFirst()) +while !args.isEmpty { + let a = args.removeFirst() + func next() -> String { if args.isEmpty { fail("missing value for \(a)") }; return args.removeFirst() } + switch a { + case "--pack": opts.pack = next() + case "--batches": opts.batches = Int(next())! + case "--batch-log2": opts.batchLog2 = Int(next())! + case "--group": opts.group = Int(next())! + case "--warps": opts.warps = Int(next())! + default: fail("unknown argument \(a)") + } +} +if opts.pack.isEmpty { fail("--pack is required") } + +func readText(_ name: String) -> String { + guard let s = try? String(contentsOfFile: opts.pack + "/" + name, encoding: .utf8) else { fail("cannot read \(opts.pack)/\(name)") } + return s +} +let programH = readText("program.h") +func defineU32(_ name: String) -> UInt32? { + let pat = "#define \(name) ([0-9a-fA-Fx]+)" + guard let re = try? NSRegularExpression(pattern: pat), let m = re.firstMatch(in: programH, range: NSRange(programH.startIndex..., in: programH)) else { return nil } + let v = String(programH[Range(m.range(at: 1), in: programH)!]).replacingOccurrences(of: "u", with: "") + if v.hasPrefix("0x") { return UInt32(v.dropFirst(2), radix: 16) } + return UInt32(v) +} +func defineStr(_ name: String) -> String? { + let pat = "#define \(name) \"([^\"]*)\"" + guard let re = try? NSRegularExpression(pattern: pat), let m = re.firstMatch(in: programH, range: NSRange(programH.startIndex..., in: programH)) else { return nil } + return String(programH[Range(m.range(at: 1), in: programH)!]) +} +let datasetLog2 = Int(defineU32("IGNEUM_DATASET_LOG2") ?? 28) +let datasetMode = defineU32("IGNEUM_DATASET_MODE") ?? 1 +if datasetMode != 1 { fail("packbench runs memory-hard packs only") } +let cacheLog2 = Int(defineU32("IGNEUM_CACHE_LOG2_WORDS") ?? 26) +let cacheSegments = Int(defineU32("IGNEUM_CACHE_SEGMENTS") ?? 65536) +let loadsPerHash = Int(defineU32("IGNEUM_LOADS_PER_HASH") ?? 128) +let bytesPerHash = Int(defineU32("IGNEUM_BYTES_PER_HASH") ?? UInt32(loadsPerHash * 4)) +let scratchOps = Int(defineU32("IGNEUM_SCRATCH_OPS") ?? 0) +let persistent = (defineU32("IGNEUM_PERSISTENT_WARPS") ?? 0) == 1 +let scratchWordsPerLane = Int(defineU32("IGNEUM_SCRATCH_WORDS_PER_LANE") ?? 8192) +let className = defineStr("IGNEUM_LOAD_CLASS") ?? "v2" +let seedString = defineStr("IGNEUM_SEED_STRING") ?? "?" +let programId = defineStr("IGNEUM_PROGRAM_ID") ?? "" + +// vectors.json: bases, expected outputs, dataset head / last, cache fingerprint +let vj = try! JSONSerialization.jsonObject(with: Data(contentsOf: URL(fileURLWithPath: opts.pack + "/vectors.json"))) as! [String: Any] +func hex64(_ s: String) -> UInt64 { return UInt64(s.dropFirst(2), radix: 16)! } +func hex32(_ s: String) -> UInt32 { return UInt32(s.dropFirst(2), radix: 16)! } +let warpsJ = vj["warps"] as! [[String: Any]] +let vecBases = warpsJ.map { UInt32(($0["base_nonce"] as! NSNumber).uint64Value) } +let vecOuts = warpsJ.map { ($0["expected"] as! [String]).map(hex64) } +let dsHead = (vj["dataset_head"] as! [String]).map(hex32) +let dsLastIndex = UInt32((vj["dataset_last_index"] as! NSNumber).uint64Value) +let dsLast = hex32(vj["dataset_last"] as! String) +let cacheFnvWant = hex64(vj["cache_fnv1a64"] as! String) + +guard let device = MTLCreateSystemDefaultDevice(), let queue = device.makeCommandQueue() else { fail("no Metal device") } +let words = 1 << datasetLog2 +let mask = UInt32(words - 1) +let cacheWords = 1 << cacheLog2 + +func compile(_ file: String) -> MTLLibrary { + do { return try device.makeLibrary(source: readText(file), options: MTLCompileOptions()) } catch { fail("Metal compile of \(file): \(error)") } +} +let t0 = nowMs() +let mhLib = compile("memhard.metal") +let progLib = compile("program.metal") +guard let fillFn = mhLib.makeFunction(name: "igneum_cache_fill"), let buildFn = mhLib.makeFunction(name: "igneum_build"), let hashFn = progLib.makeFunction(name: "igneum_hash") else { fail("kernel functions missing") } +let fillPipe = try! device.makeComputePipelineState(function: fillFn) +let buildPipe = try! device.makeComputePipelineState(function: buildFn) +let hashPipe = try! device.makeComputePipelineState(function: hashFn) +let compileMs = nowMs() - t0 +if hashPipe.threadExecutionWidth != 32 { print("WARNING: threadExecutionWidth \(hashPipe.threadExecutionWidth), not 32") } + +guard let cache = device.makeBuffer(length: cacheWords * 4, options: .storageModePrivate) else { fail("cache alloc") } +guard let dataset = device.makeBuffer(length: words * 4, options: .storageModePrivate) else { fail("dataset alloc") } + +func run(_ body: (MTLComputeCommandEncoder) -> Void) -> (Double, Double) { + let cb = queue.makeCommandBuffer()! + let enc = cb.makeComputeCommandEncoder()! + body(enc) + enc.endEncoding() + let w0 = nowMs() + cb.commit(); cb.waitUntilCompleted() + if let e = cb.error { fail("command buffer: \(e)") } + return (nowMs() - w0, (cb.gpuEndTime - cb.gpuStartTime) * 1000) +} +let (cacheWall, cacheGpu) = run { enc in + enc.setComputePipelineState(fillPipe); enc.setBuffer(cache, offset: 0, index: 0) + enc.dispatchThreadgroups(MTLSize(width: cacheSegments / 256, height: 1, depth: 1), threadsPerThreadgroup: MTLSize(width: 256, height: 1, depth: 1)) +} +let items = words / 16 +let (buildWall, buildGpu) = run { enc in + enc.setComputePipelineState(buildPipe); enc.setBuffer(cache, offset: 0, index: 0); enc.setBuffer(dataset, offset: 0, index: 1) + enc.dispatchThreadgroups(MTLSize(width: items / 256, height: 1, depth: 1), threadsPerThreadgroup: MTLSize(width: 256, height: 1, depth: 1)) +} +// cache fingerprint and dataset head/last through a blit to shared memory +func blit(_ src: MTLBuffer, _ offset: Int, _ n: Int) -> MTLBuffer { + let dst = device.makeBuffer(length: n, options: .storageModeShared)! + let cb = queue.makeCommandBuffer()!; let b = cb.makeBlitCommandEncoder()! + b.copy(from: src, sourceOffset: offset, to: dst, destinationOffset: 0, size: n); b.endEncoding(); cb.commit(); cb.waitUntilCompleted() + return dst +} +func fnv1a64(_ p: UnsafeRawPointer, _ n: Int) -> UInt64 { + var h: UInt64 = 0xcbf29ce484222325 + let b = p.bindMemory(to: UInt8.self, capacity: n) + for i in 0..> 20) MiB") } + scratch = s +} +// One hash launch: `nonces` outputs from `base`. Persistent: warpsN warps loop over nonces / 32 units. +func encodeHash(_ enc: MTLComputeCommandEncoder, out: MTLBuffer, base: UInt32, nonces: Int, group: Int) { + enc.setComputePipelineState(hashPipe) + enc.setBuffer(dataset, offset: 0, index: 0) + enc.setBuffer(out, offset: 0, index: 1) + var b = base; enc.setBytes(&b, length: 4, index: 2) + if persistent { + let units = nonces / 32 + let nw = min(warpsN, units) + if units % nw != 0 { fail("nonces \(nonces) is not a multiple of 32 x \(nw) warps") } + enc.setBuffer(scratch!, offset: 0, index: 3) + var g = UInt32(units); enc.setBytes(&g, length: 4, index: 4) + var s = salt; enc.setBytes(&s, length: 4, index: 5) + salt = salt &+ UInt32(units) + let threads = nw * 32 + let tg = min(group, threads) + enc.dispatchThreadgroups(MTLSize(width: threads / tg, height: 1, depth: 1), threadsPerThreadgroup: MTLSize(width: tg, height: 1, depth: 1)) + } else { + let tg = min(group, nonces) + enc.dispatchThreadgroups(MTLSize(width: nonces / tg, height: 1, depth: 1), threadsPerThreadgroup: MTLSize(width: tg, height: 1, depth: 1)) + } +} +// Vectors, standalone (one unit per launch) +var vecPass = 0 +let vecOut = device.makeBuffer(length: 32 * 8, options: .storageModeShared)! +for (i, base) in vecBases.enumerated() { + _ = run { enc in encodeHash(enc, out: vecOut, base: base, nonces: 32, group: 32) } + let p = vecOut.contents().bindMemory(to: UInt64.self, capacity: 32) + var ok = true + for l in 0..<32 where p[l] != vecOuts[i][l] { ok = false; print("vector warp base \(base) lane \(l): GPU \(String(format: "%016llx", p[l])) expected \(String(format: "%016llx", vecOuts[i][l]))"); break } + if ok { vecPass += 1 } +} +// Batch at base 0: fingerprint and the vectors inside the batch +var nonces = 1 << opts.batchLog2 +if persistent { let unit = 32 * min(warpsN, nonces / 32); nonces = (nonces / unit) * unit } +let out = device.makeBuffer(length: nonces * 8, options: .storageModeShared)! +let (warmWall, warmGpu) = run { enc in encodeHash(enc, out: out, base: 0, nonces: nonces, group: opts.group) } +let outPtr = out.contents().bindMemory(to: UInt64.self, capacity: nonces) +var batchVecPass = 0, batchVecN = 0 +for (i, base) in vecBases.enumerated() where Int(base) + 32 <= nonces { + batchVecN += 1 + if (0..<32).allSatisfy({ outPtr[Int(base) + $0] == vecOuts[i][$0] }) { batchVecPass += 1 } +} +let fingerprint = fnv1a64(out.contents(), nonces * 8) +// Timed batches +var wallSum = 0.0, gpuSum = 0.0 +for b in 0.. out(nonces); +#ifdef IGNEUM_PERSISTENT_WARPS + // Variant 5: EMU_WARPS persistent warps (one per work-group of 32), each with IGNEUM_SCRATCH_WORDS_PER_LANE x 32 words. + const unsigned emuWarps = 8; + std::vector scratchArena((size_t)emuWarps * 32u * IGNEUM_SCRATCH_WORDS_PER_LANE, 0u); + uint salt = 1u; + if (IGNEUM_GROUP != 32) { std::printf("variant 5 needs IGNEUM_GROUP 32 (one warp per work-group)\n"); return 2; } + std::printf("variant 5: %u persistent warps, scratch arena %u MiB, lazy tagged fill\n", emuWarps, (unsigned)((scratchArena.size() * 4u) >> 20)); + auto launchHash = [&](size_t units, uint base) { + size_t nw = units < emuWarps ? units : emuWarps; + emu_launch(igneum_hash, nw * 32u, 32u, (const uint*)ds.data(), out.data(), base, mask, scratchArena.data(), (uint)units, salt); + salt += (uint)units; + }; +#else + auto launchHash = [&](size_t nonceCount, uint base) { + emu_launch(igneum_hash, nonceCount, (unsigned)IGNEUM_GROUP, (const uint*)ds.data(), out.data(), base, mask); + }; +#endif for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) { - emu_launch(igneum_hash, (size_t)IGNEUM_GROUP, (unsigned)IGNEUM_GROUP, (const uint*)ds.data(), out.data(), (uint)IGNEUM_VEC_BASE[w], mask); +#ifdef IGNEUM_PERSISTENT_WARPS + launchHash(1, (uint)IGNEUM_VEC_BASE[w]); +#else + launchHash((size_t)IGNEUM_GROUP, (uint)IGNEUM_VEC_BASE[w]); +#endif char how[96]; std::snprintf(how, sizeof(how), "standalone, work-group %d, sub-group width %u", IGNEUM_GROUP, gSubGroupWidth); overall = compareWarp((const uint64_t*)out.data(), IGNEUM_VEC_OUT[w], IGNEUM_VEC_BASE[w], how) && overall; } // In batch: every vector warp that fits in 2^batchLog2 nonces. - emu_launch(igneum_hash, (size_t)nonces, (unsigned)IGNEUM_GROUP, (const uint*)ds.data(), out.data(), 0u, mask); +#ifdef IGNEUM_PERSISTENT_WARPS + launchHash(nonces / 32u, 0u); +#else + launchHash((size_t)nonces, 0u); +#endif for (int w = 0; w < IGNEUM_VEC_WARPS; ++w) { if ((uint64_t)IGNEUM_VEC_BASE[w] + 32ull > nonces) { std::printf("verify warp base %u in batch: skipped (batch has %u nonces)\n", IGNEUM_VEC_BASE[w], nonces); continue; } char how[96]; diff --git a/proto-opencl/host.c b/proto-opencl/host.c index 31be99b18..bd04cfe56 100644 --- a/proto-opencl/host.c +++ b/proto-opencl/host.c @@ -207,6 +207,9 @@ typedef struct { int memprobe; // --memprobe: dependent-load latency and throughput, independent-load throughput and an ALU // chain on the chosen device, no pack needed (5 October 2026, the 9070 XT on the eGPU) int probeMib; // --probe-mib N: --memprobe at that one buffer size only (default 0 = 4, 64 and 1024 MiB) + int benchPack; // --bench-pack: with --pack D, build and self-test the pack at run time (as --serve does) and time + // igneum_hash_bound with the pack's seed words as init words; read-width experiment, 5 October 2026 + int warps; // --warps N: persistent warps for a variant-5 pack (IGNEUM_PERSISTENT_WARPS); 0 = 2048 } Options; static int packMib(void) { return (int)(((1ull << IGNEUM_DATASET_LOG2) * 4ull) >> 20); } @@ -236,6 +239,8 @@ static void usage(void) { " pack this exe was built against; it is self-tested against its vectors.h first (the one-click worker)\n" " --readback M with --serve: select (default) reads back only the hits and 34 sentinel words of each dispatch through a\n" " GPU-side pass; full reads back every output (8 bytes per nonce). IGNEUM_READBACK=full does the same.\n" + " --bench-pack with --pack D: read the pack at run time, build and self-test it, time its bound kernel (one exe, any pack)\n" + " --warps N persistent warps for a variant-5 pack (a 1 MiB scratch each; default 2048; the batch rounds to 32 x N)\n" " --memprobe no pack: dependent random loads (latency and throughput against lanes in flight), independent random\n" " loads and an ALU chain on the chosen device, at 4, 64 and 1024 MiB (--probe-mib N for one size)\n", packMib(), IGNEUM_KERNEL_PATH); } @@ -248,7 +253,7 @@ static Options parseArgs(int argc, char** argv) { int i; o.datasetMib = 1024; o.batchLog2 = 24; o.batches = 5; o.groupWarps = 1; o.sweep = 0; o.device = -1; o.exchange = 0; o.list = 0; o.timeWall = -1; o.kernelPath = IGNEUM_KERNEL_PATH; o.extraOpts = ""; o.serve = 0; o.noPrepare = 0; o.kernelGiven = 0; o.vendor = NULL; o.packDir = NULL; - o.readback = (getenv("IGNEUM_READBACK") && strcmp(getenv("IGNEUM_READBACK"), "full") == 0) ? 1 : 0; o.memprobe = 0; o.probeMib = 0; + o.readback = (getenv("IGNEUM_READBACK") && strcmp(getenv("IGNEUM_READBACK"), "full") == 0) ? 1 : 0; o.memprobe = 0; o.probeMib = 0; o.benchPack = 0; o.warps = 0; for (i = 1; i < argc; ++i) { const char* a = argv[i]; int needs = (strcmp(a, "--dataset-mib") == 0 || strcmp(a, "--batch-log2") == 0 || strcmp(a, "--batches") == 0 || @@ -274,6 +279,8 @@ static Options parseArgs(int argc, char** argv) { else { printf("--readback must be select or full\n"); exit(2); } } else if (strcmp(a, "--memprobe") == 0) o.memprobe = 1; + else if (strcmp(a, "--bench-pack") == 0) o.benchPack = 1; + else if (strcmp(a, "--warps") == 0) { if (i + 1 >= argc) { usage(); exit(2); } o.warps = atoi(argv[++i]); } else if (strcmp(a, "--probe-mib") == 0) { if (i + 1 >= argc) { usage(); exit(2); } o.probeMib = atoi(argv[++i]); } else if (strcmp(a, "--build-opts") == 0) o.extraOpts = argv[++i]; else if (strcmp(a, "--time") == 0) { @@ -1013,6 +1020,19 @@ static int unhexBuf(const char* s, uint8_t* out, size_t cap, size_t* len) { * from the compiled-in program.h, so one prebuilt exe serves every pack. The compiled-in values are the defaults. */ static PfPack gPack; static int gGeneric = 0; +/* Variant 5 of the read-width experiment (5 October 2026): the scratch arena of a persistent-warp pack, its warp count + * and the running tag salt; set by --bench-pack before the self-test. Serve mode does not support these packs. */ +static cl_mem gScratch = NULL; +static cl_uint gScratchWarps = 0; +static cl_uint gSalt = 1; +/* Sets the three extra arguments of a variant-5 kernel (after the five of igneum_hash_bound) for `units` units. */ +static cl_int setScratchArgs(cl_kernel k, cl_uint firstArg, cl_uint units) { + cl_int e = clSetKernelArg(k, firstArg, sizeof(cl_mem), &gScratch); + if (e == CL_SUCCESS) e = clSetKernelArg(k, firstArg + 1, sizeof(cl_uint), &units); + if (e == CL_SUCCESS) e = clSetKernelArg(k, firstArg + 2, sizeof(cl_uint), &gSalt); + gSalt += units; + return e; +} static uint32_t gServeWords = 1u << IGNEUM_DATASET_LOG2; #if IGNEUM_DATASET_MODE == 1 static uint32_t gServeCacheWords = 1u << IGNEUM_CACHE_LOG2_WORDS; @@ -1102,6 +1122,10 @@ static int pairSelfTest(Device* dv, const DeviceInfo* di, cl_command_queue q, Se out = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, g * sizeof(uint64_t), NULL, &e); if (e == CL_SUCCESS) { ++gMemCreated; init = clCreateBuffer(dv->ctx, CL_MEM_READ_ONLY | CL_MEM_COPY_HOST_PTR, 32, pk.seedw, &e); } if (e == CL_SUCCESS) ++gMemCreated; + if (pk.persistent) { + if (!gScratch) { snprintf(err, errCap, "self-test: a variant-5 pack (persistent warps) needs --bench-pack (serve mode does not carry a scratch)"); return 0; } + g = 32; local = 32; /* one persistent warp runs the one unit */ + } for (w = 0; w < pk.vecWarps && e == CL_SUCCESS; ++w) { cl_uint base = pk.vecBase[w], mask = words - 1u; e = clSetKernelArg(p->kHashBound, 0, sizeof(cl_mem), &p->ds); @@ -1109,6 +1133,7 @@ static int pairSelfTest(Device* dv, const DeviceInfo* di, cl_command_queue q, Se if (e == CL_SUCCESS) e = clSetKernelArg(p->kHashBound, 2, sizeof(cl_uint), &base); if (e == CL_SUCCESS) e = clSetKernelArg(p->kHashBound, 3, sizeof(cl_uint), &mask); if (e == CL_SUCCESS) e = clSetKernelArg(p->kHashBound, 4, sizeof(cl_mem), &init); + if (e == CL_SUCCESS && pk.persistent) e = setScratchArgs(p->kHashBound, 5, 1u); if (e == CL_SUCCESS) e = clEnqueueNDRangeKernel(q, p->kHashBound, 1, NULL, &g, &local, 0, NULL, NULL); if (e == CL_SUCCESS) e = clEnqueueReadBuffer(q, out, CL_TRUE, 0, 32 * sizeof(uint64_t), &vec[w * 32], 0, NULL, NULL); } @@ -1209,6 +1234,92 @@ static int startPrepareThread(PrepareTask* t) { pthread_t th; if (pthread_create #endif #endif +/* --bench-pack (read-width experiment, 5 October 2026): the pack in --pack is built and self-tested exactly as the + * first pair of --serve (pairBuffers: cache, dataset, cache FNV, dataset words, the vector warps through + * igneum_hash_bound with the pack's seed words), then the bound kernel is timed over --batches dispatches of + * 2^--batch-log2 nonces with device event time, and the 2^B outputs at base nonce 0 are fingerprinted (FNV-1a 64) so + * the same pack can be compared bit for bit across vendors. A variant-5 pack is launched as --warps persistent warps + * with a 1 MiB scratch each. One line per run starts with RESULT. */ +static int runBenchPack(Device* dv, const DeviceInfo* di, const Options* o) { +#if IGNEUM_DATASET_MODE != 1 + (void)dv; (void)di; (void)o; + printf("FAIL: --bench-pack needs a memory-hard placeholder pack\n"); + return 2; +#else + const uint32_t words = gServeWords, mask = words - 1u; + uint32_t nonces = 1u << o->batchLog2; + size_t groupSize = 32 * (size_t)o->groupWarps, g; + cl_int err = 0; + cl_mem dOut, dInit; + uint64_t* hOut; + ServePair* cur; + char perr[512], devName[256]; + double t0 = wallMs(), sum = 0, warmMs; + uint64_t fp; + int b, k; + cl_uint warps = (cl_uint)(o->warps > 0 ? o->warps : 2048), units = nonces / 32u; + if (!dv->kHashBound) { printf("FAIL: the kernel source has no igneum_hash_bound\n"); return 2; } + if (gPack.persistent) { + size_t arena; + if (o->groupWarps != 1) { printf("FAIL: a variant-5 pack needs --group-warps 1 (one warp per work-group: the loop trip count must be uniform)\n"); return 2; } + while (warps > 1 && units % warps != 0) warps >>= 1; + arena = (size_t)warps * 32u * (size_t)gPack.scratchWordsPerLane * 4u; + if ((uint64_t)arena > di->maxAlloc) { printf("FAIL: scratch arena %llu MiB exceeds the device's max alloc %llu MiB; lower --warps\n", (unsigned long long)(arena >> 20), (unsigned long long)(di->maxAlloc >> 20)); return 2; } + gScratch = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, arena, NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer scratch"); + gScratchWarps = warps; + printf("variant 5: %u persistent warps (%u x 32 work-items, work-group 32), scratch arena %llu MiB, %u units per dispatch, lazy tagged fill\n", warps, warps, (unsigned long long)(arena >> 20), units); + } + cur = (ServePair*)calloc(1, sizeof(ServePair)); + cur->kHashBound = dv->kHashBound; cur->kCacheFill = dv->kCacheFill; cur->kBuild = dv->kBuild; cur->prog = dv->prog; + dv->kHashBound = dv->kCacheFill = dv->kBuild = NULL; dv->prog = NULL; + memcpy(cur->sw, gPack.seedw, 32); memcpy(cur->kw, gPack.keyw, 32); + if (!pairBuffers(dv, di, dv->q, cur, words, gServeCacheWords, gServeSegments, o->packDir, perr, sizeof(perr))) { printf("FAIL: pack %s: %s\n", o->packDir, perr); return 1; } + printf("pack %s: cache %.0f dataset %.0f check %.0f ms (%.0f ms in all); %s\n", o->packDir, cur->cacheMs, cur->datasetMs, cur->checkMs, wallMs() - t0, cur->check); + printKernelInfo(di, cur->kHashBound, "igneum_hash_bound", (int)groupSize, ""); + dOut = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, (size_t)nonces * sizeof(uint64_t), NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer out"); + dInit = clCreateBuffer(dv->ctx, CL_MEM_READ_ONLY | CL_MEM_COPY_HOST_PTR, 32, cur->sw, &err); CL_CHECK_ERR(err, "clCreateBuffer init words"); + hOut = (uint64_t*)malloc((size_t)nonces * sizeof(uint64_t)); + strncpy(devName, di->name, 255); devName[255] = 0; + for (k = 0; devName[k]; ++k) if (devName[k] == ' ') devName[k] = '_'; + g = gPack.persistent ? (size_t)warps * 32u : (size_t)nonces; + for (b = -1; b < o->batches; ++b) { + cl_uint base = (cl_uint)((uint32_t)(b + 1) * nonces); + cl_event ev = NULL; + double ms; + CL_CHECK(clSetKernelArg(cur->kHashBound, 0, sizeof(cl_mem), &cur->ds)); + CL_CHECK(clSetKernelArg(cur->kHashBound, 1, sizeof(cl_mem), &dOut)); + CL_CHECK(clSetKernelArg(cur->kHashBound, 2, sizeof(cl_uint), &base)); + CL_CHECK(clSetKernelArg(cur->kHashBound, 3, sizeof(cl_uint), &mask)); + CL_CHECK(clSetKernelArg(cur->kHashBound, 4, sizeof(cl_mem), &dInit)); + if (gPack.persistent) CL_CHECK(setScratchArgs(cur->kHashBound, 5, units)); + { + double w0 = wallMs(); + CL_CHECK(clEnqueueNDRangeKernel(dv->q, cur->kHashBound, 1, NULL, &g, &groupSize, 0, NULL, &ev)); + CL_CHECK(clWaitForEvents(1, &ev)); + ms = o->timeWall ? wallMs() - w0 : eventMs(ev); + if (ms < 0) ms = wallMs() - w0; + clReleaseEvent(ev); + } + if (b < 0) { + warmMs = ms; + CL_CHECK(clEnqueueReadBuffer(dv->q, dOut, CL_TRUE, 0, (size_t)nonces * sizeof(uint64_t), hOut, 0, NULL, NULL)); + fp = pf_fnv1a64((const uint32_t*)hOut, (size_t)nonces * 8u); + } else sum += ms; + } + printf("warm-up dispatch (base 0): %.2f ms; %d timed dispatches of 2^%d nonces: mean %.2f ms\n", warmMs, o->batches, o->batchLog2, sum / o->batches); + printf("RESULT pack=%s class=%s device=%s platform=%s group=%d warps=%u arena_mib=%llu nonces=%u batches=%d check=%s fingerprint=%016llx mhs=%.3f loads=%u bytes=%u scratch_ops=%u time=%s\n", + o->packDir, gPack.loadClass, devName, strcmp(di->platformName, "Apple") == 0 ? "Apple" : "other", (int)groupSize, gPack.persistent ? warps : 0u, + gPack.persistent ? (unsigned long long)(((size_t)warps * 32u * gPack.scratchWordsPerLane * 4u) >> 20) : 0ull, nonces, o->batches, + cur->checked ? "PASS" : "skipped", (unsigned long long)fp, (double)nonces * (double)o->batches / (sum / 1000.0) / 1e6, + gPack.loadsPerHash, gPack.bytesPerHash, gPack.scratchOps * 8u, o->timeWall ? "wall" : "event"); + free(hOut); + clReleaseMemObject(dOut); clReleaseMemObject(dInit); + if (gScratch) clReleaseMemObject(gScratch); + releasePair(cur); + return 0; +#endif +} + static int runServe(Device* dv, const DeviceInfo* di, const Options* o) { #if IGNEUM_DATASET_MODE != 1 (void)dv; (void)di; (void)o; @@ -1612,6 +1723,11 @@ static const char* PROBE_SRC = " }\n" " out[g] = x0 ^ x1 ^ x2 ^ x3 ^ x4 ^ x5 ^ x6 ^ x7;\n" "}\n" + "__kernel void probe_line16(__global const uint4* ds, uint vecMask, uint steps, uint seed, __global uint* out) {\n" + " uint x = pm_mix((uint)get_global_id(0) ^ seed);\n" + " for (uint s = 0u; s < steps; ++s) { uint4 a = ds[x & vecMask]; x = (a.x ^ a.y ^ a.z ^ a.w) ^ (x * 0x9E3779B1u + s); }\n" + " out[get_global_id(0)] = x;\n" + "}\n" "__kernel void probe_line(__global const uint4* ds, uint lineMask, uint steps, uint seed, __global uint* out) {\n" " uint x = pm_mix((uint)get_global_id(0) ^ seed);\n" " for (uint s = 0u; s < steps; ++s) {\n" @@ -1656,7 +1772,7 @@ static double probeLaunch(Device* dv, const Options* o, cl_kernel k, size_t glob static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) { cl_int err = 0; cl_program prog; - cl_kernel kFill, kChase, kIndep, kAlu, kLine, kStream; + cl_kernel kFill, kChase, kIndep, kAlu, kLine, kStream, kLine16; size_t srcLen = strlen(PROBE_SRC); int sizes[3] = { 4, 64, 1024 }, nSizes = 3, si; size_t lanesList[8] = { 256, 1024, 1u << 12, 1u << 14, 1u << 16, 1u << 18, 1u << 20, 1u << 22 }; @@ -1682,6 +1798,7 @@ static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) { kIndep = clCreateKernel(prog, "probe_indep", &err); CL_CHECK_ERR(err, "probe_indep"); kAlu = clCreateKernel(prog, "probe_alu", &err); CL_CHECK_ERR(err, "probe_alu"); kLine = clCreateKernel(prog, "probe_line", &err); CL_CHECK_ERR(err, "probe_line"); + kLine16 = clCreateKernel(prog, "probe_line16", &err); CL_CHECK_ERR(err, "probe_line16"); kStream = clCreateKernel(prog, "probe_stream", &err); CL_CHECK_ERR(err, "probe_stream"); dOut = clCreateBuffer(dv->ctx, CL_MEM_READ_WRITE, maxLanes * 4u, NULL, &err); CL_CHECK_ERR(err, "clCreateBuffer probe out"); printf("memprobe on [%s] %s, driver %s, %u compute units, %u MHz, %s time\n", di->platformName, di->name, di->driver, di->computeUnits, di->clockMHz, o->timeWall ? "wall" : "device event"); @@ -1737,6 +1854,25 @@ static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) { fflush(stdout); } } + { + /* Random 16-byte reads (one uint4) in a dependent chain: the W = 16 width of the read-width experiment. */ + size_t local = di->maxWorkGroup < 256 ? di->maxWorkGroup : 256; + size_t lanes; + cl_uint vecMask = (words / 4u) - 1u; + for (lanes = 1u << 14; lanes <= maxLanes; lanes <<= 2) { + cl_uint seed = 0x2718281u; + double ms; + CL_CHECK(clSetKernelArg(kLine16, 0, sizeof(cl_mem), &dDs)); + CL_CHECK(clSetKernelArg(kLine16, 1, sizeof(cl_uint), &vecMask)); + CL_CHECK(clSetKernelArg(kLine16, 2, sizeof(cl_uint), &STEPS)); + CL_CHECK(clSetKernelArg(kLine16, 3, sizeof(cl_uint), &seed)); + CL_CHECK(clSetKernelArg(kLine16, 4, sizeof(cl_mem), &dOut)); + ms = probeLaunch(dv, o, kLine16, lanes, local, 3, 3, seed); + printf("| line 16 B | %d | %llu | %llu | %u | %.3f | %.3f G reads/s | %.1f GB/s in 16 B reads |\n", mib, (unsigned long long)local, (unsigned long long)lanes, STEPS, ms, + (double)lanes * (double)STEPS / (ms / 1000.0) / 1e9, (double)lanes * (double)STEPS * 16.0 / (ms / 1000.0) / 1e9); + fflush(stdout); + } + } { /* Random 64-byte lines (16 words, four uint4 loads) in a dependent chain: lines per second against the * 4-byte chase above says what one random 4-byte read costs the memory system. If the two rates are @@ -1792,7 +1928,7 @@ static int runMemprobe(Device* dv, const DeviceInfo* di, const Options* o) { (double)lanes * (double)ALU_STEPS / (ms / 1000.0) / 1e9 / (double)(di->computeUnits ? di->computeUnits : 1)); } clReleaseMemObject(dOut); - clReleaseKernel(kFill); clReleaseKernel(kChase); clReleaseKernel(kIndep); clReleaseKernel(kAlu); clReleaseKernel(kLine); clReleaseKernel(kStream); + clReleaseKernel(kFill); clReleaseKernel(kChase); clReleaseKernel(kIndep); clReleaseKernel(kAlu); clReleaseKernel(kLine); clReleaseKernel(kStream); clReleaseKernel(kLine16); clReleaseProgram(prog); printf("memprobe: done\n"); return 0; @@ -1821,7 +1957,7 @@ int main(int argc, char** argv) { static char boundPath[1200]; char perr[512]; size_t n = strlen(o.packDir); - if (!o.serve) { printf("FAIL: --pack goes with --serve (the bench runs the compiled-in pack)\n"); return 2; } + if (!o.serve && !o.benchPack) { printf("FAIL: --pack goes with --serve or --bench-pack (the plain bench runs the compiled-in pack)\n"); return 2; } if (n > 1 && (o.packDir[n - 1] == '/' || o.packDir[n - 1] == '\\')) ((char*)o.packDir)[n - 1] = 0; if (!pf_load(o.packDir, &gPack, perr, sizeof(perr))) { printf("error 0 pack %s: %s\n", o.packDir, perr); fflush(stdout); return 2; } gGeneric = 1; @@ -1877,6 +2013,7 @@ int main(int argc, char** argv) { clReleaseContext(dv.ctx); return rc; } + if (o.benchPack && !o.packDir) { printf("FAIL: --bench-pack needs --pack \n"); return 2; } if (o.serve && !o.kernelGiven) { /* The bound kernel lives next to the compiled-in kernel.cl as kernel_bound.cl (packs from igneum-pow or igneum-miner export-pack). */ static char boundPath[1024]; @@ -1894,6 +2031,7 @@ int main(int argc, char** argv) { printf("build options: %s\n", dv.buildOptions); printf("exchange: %s\n", dv.exchangeNote); if (o.serve) return runServe(&dv, di, &o); + if (o.benchPack) return runBenchPack(&dv, di, &o); printKernelInfo(di, dv.kHash, "igneum_hash", dv.groupSize, ""); printf("program: %d instructions x %d iterations, loads/hash %d, op mix %s\n", IGNEUM_INSTR_COUNT, IGNEUM_ITERATIONS, IGNEUM_LOADS_PER_HASH, IGNEUM_OP_MIX); From 4badcee1ff5b5e79627a7d566fd14e80a154ed04 Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:05:57 +0100 Subject: [PATCH 003/311] read-width: CUDA worker --bench and --memprobe, scratch arena from the occupancy capacity; PC playbooks (card under test off in the app, restored after); harness arena label Co-Authored-By: Claude Fable 5.1 --- proto-cuda/nvrtc/worker.cpp | 268 ++++++++++++++++++++++++++++- proto-metal/packbench.swift | 2 +- relay/playbooks/readwidth-5090.ps1 | 67 ++++++++ relay/playbooks/readwidth-9070.ps1 | 69 ++++++++ 4 files changed, 400 insertions(+), 6 deletions(-) create mode 100644 relay/playbooks/readwidth-5090.ps1 create mode 100644 relay/playbooks/readwidth-9070.ps1 diff --git a/proto-cuda/nvrtc/worker.cpp b/proto-cuda/nvrtc/worker.cpp index 11597fd2a..94705562d 100644 --- a/proto-cuda/nvrtc/worker.cpp +++ b/proto-cuda/nvrtc/worker.cpp @@ -241,6 +241,10 @@ static bool loadNvrtc(Rtc& r, std::string& err, std::string& libName) { struct Ctx { Drv drv; Rtc rtc; + // read-width experiment, variant 5 (5 October 2026): persistent warps and their scratch; --warps caps the launch + int warps = 0; // 0 = the resident capacity from the occupancy query, rounded down to a power of two + uint32_t salt = 1; // the running per-unit tag salt (+= units per launch) + int batches = 5; // --bench: timed dispatches CUdevice dev = 0; CUcontext ctx = nullptr; std::string name; @@ -409,6 +413,14 @@ struct Pair { std::string variant = "base"; // the bound kernel in service: a variant name (see allVariants) std::string raceLine; // the race's one-line report, emitted by the main thread with "prepared" double raceMs = 0; + // read-width experiment (5 October 2026): the pack's load class and, for variant 5, the persistent-warp scratch + std::string loadClass = "v2"; + uint32_t loadsPerHash = 128, bytesPerHash = 512, scratchOps = 0; + bool persistent = false; + CUdeviceptr scratch = 0; + int warps = 0; // persistent warps launched (the arena holds this many) + int residentWarps = 0; // the occupancy query's capacity: blocks/SM x warps/block x SMs + size_t scratchBytes = 0; }; // The job loop and a race take turns on the card: a variant is timed with no job running (exclusive numbers), and @@ -419,6 +431,7 @@ static void releasePair(Ctx& c, Pair* p) { if (!p) return; if (p->ds) c.drv.memFree(p->ds); if (p->cache) c.drv.memFree(p->cache); + if (p->scratch) c.drv.memFree(p->scratch); if (p->modBound) c.drv.moduleUnload(p->modBound); if (p->modKernel) c.drv.moduleUnload(p->modKernel); delete p; @@ -430,6 +443,17 @@ struct IgneumInitWordsArg { uint32_t w[8]; }; static bool launchHash(Ctx& c, Pair* p, CUdeviceptr out, uint32_t baseNonce, const uint32_t iw[8], uint32_t nonces, uint32_t block, CUstream s, std::string& err) { uint32_t mask = p->words - 1u; IgneumInitWordsArg a; std::memcpy(a.w, iw, 32); + if (p->persistent) { + // Variant 5: N persistent warps over nonces / 32 units; the arena was sized for p->warps warps in buildPair. + uint32_t units = nonces / 32u, warps = (uint32_t)p->warps; + if (warps > units) warps = units; + while (warps > 1u && units % warps != 0u) warps >>= 1; + if (block != 32u) { err = "a variant-5 pack runs one warp per block (--block-warps 1)"; return false; } + uint32_t salt = c.salt; c.salt += units; + void* args[8] = { &p->ds, &out, &baseNonce, &mask, &a, &p->scratch, &units, &salt }; + DRV_CHECK(c, c.drv.launchKernel(p->fHashBound, warps, 1, 1, 32, 1, 1, 0, s, args, nullptr), "cuLaunchKernel igneum_hash_bound (persistent)"); + return true; + } void* args[5] = { &p->ds, &out, &baseNonce, &mask, &a }; DRV_CHECK(c, c.drv.launchKernel(p->fHashBound, nonces / block, 1, 1, block, 1, 1, 0, s, args, nullptr), "cuLaunchKernel igneum_hash_bound"); return true; @@ -768,16 +792,39 @@ static Pair* buildPair(Ctx& c, const std::string& dir, CUstream s, std::string& c.drv.funcGetAttribute(&p->regs, CU_FUNC_ATTRIBUTE_NUM_REGS, p->fHashBound); c.drv.occupancy(&p->blocksPerSM, p->fHashBound, 32 * c.blockWarps, 0); } + p->loadClass = pk.loadClass; p->loadsPerHash = pk.loadsPerHash; p->bytesPerHash = pk.bytesPerHash; p->scratchOps = pk.scratchOps; + p->persistent = pk.persistent != 0; + p->residentWarps = p->blocksPerSM * c.blockWarps * c.sms; + size_t scratchBytes = 0; + if (p->persistent) { + // Variant 5: one arena per launched warp. The launch is the resident capacity (the occupancy query), rounded + // down to a power of two so it divides every batch, or --warps; the allocation cannot change the occupancy + // (registers and shared memory decide it), and the number is re-queried after the allocation below to show it. + if (c.blockWarps != 1) { err = "a variant-5 pack runs one warp per block: use --block-warps 1"; releasePair(c, p); return nullptr; } + int w = c.warps > 0 ? c.warps : p->residentWarps; + int pw = 1; while (pw * 2 <= w) pw *= 2; + p->warps = pw; + scratchBytes = (size_t)p->warps * 32u * (size_t)pk.scratchWordsPerLane * 4u; + } // Cache double t0 = wallMs(); size_t cacheBytes = (size_t)p->cacheWords * 4u, dsBytes = (size_t)p->words * 4u; { size_t freeB = 0, totalB = 0; - if (c.drv.memGetInfo(&freeB, &totalB) == CUDA_SUCCESS && freeB < cacheBytes + dsBytes + (64u << 20)) { - err = fmt("%llu MiB free on the device, this pack needs %llu MiB (cache %llu + dataset %llu)", (unsigned long long)(freeB >> 20), (unsigned long long)((cacheBytes + dsBytes) >> 20), (unsigned long long)(cacheBytes >> 20), (unsigned long long)(dsBytes >> 20)); + if (c.drv.memGetInfo(&freeB, &totalB) == CUDA_SUCCESS && freeB < cacheBytes + dsBytes + scratchBytes + (64u << 20)) { + err = fmt("%llu MiB free on the device, this pack needs %llu MiB (cache %llu + dataset %llu + scratch %llu)", (unsigned long long)(freeB >> 20), (unsigned long long)((cacheBytes + dsBytes + scratchBytes) >> 20), (unsigned long long)(cacheBytes >> 20), (unsigned long long)(dsBytes >> 20), (unsigned long long)(scratchBytes >> 20)); releasePair(c, p); return nullptr; } } + if (p->persistent) { + CUresult r = c.drv.memAlloc(&p->scratch, scratchBytes); + if (r != CUDA_SUCCESS) { err = "cuMemAlloc scratch: " + c.err(r); p->scratch = 0; releasePair(c, p); return nullptr; } + p->scratchBytes = scratchBytes; + int after = 0; + c.drv.occupancy(&after, p->fHashBound, 32 * c.blockWarps, 0); + info(fmt("variant 5: %d persistent warps (resident capacity %d = %d blocks/SM x %d warps/block x %d SMs; occupancy query after the allocation %d blocks/SM), scratch %llu MiB (%u KiB per warp)", + p->warps, p->residentWarps, p->blocksPerSM, c.blockWarps, c.sms, after, (unsigned long long)(scratchBytes >> 20), pk.scratchWordsPerLane * 4u * 32u / 1024u)); + } { CUresult r = c.drv.memAlloc(&p->cache, cacheBytes); if (r != CUDA_SUCCESS) { err = "cuMemAlloc cache: " + c.err(r); p->cache = 0; releasePair(c, p); return nullptr; } @@ -876,6 +923,8 @@ static void prepareRun(Ctx* c, PrepareTask* t) { struct Options { bool serve = false, check = false, raceOnly = false; + bool bench = false, memprobe = false; // read-width experiment (5 October 2026) + int batches = 5, warps = 0, probeMib = 0; int device = 0, batchLog2 = 22, blockWarps = 1; std::string pack, arch = "auto"; std::string race = "on", pinned, tuningPath; @@ -890,6 +939,12 @@ static void usage() { " --batch-log2 B nonces per dispatch = 2^B (default 22)\n" " --block-warps W warps per thread block (default 1)\n" " --arch sm_XY|compute_XY|auto NVRTC target (default auto: the device's architecture)\n" + " --bench --pack read-width experiment: build and self-test the pack, time --batches dispatches of 2^B nonces,\n" + " print the 2^B fingerprint at base nonce 0 (one RESULT line); a variant-5 pack runs --warps persistent warps\n" + " --memprobe [--probe-mib N] no pack: dependent random 4, 16 and 64-byte reads, independent reads, a coalesced stream and an\n" + " integer chain at 4, 64 and 1024 MiB (the same table as igneum-worker-opencl --memprobe)\n" + " --batches N --bench: timed dispatches (default 5)\n" + " --warps N --bench on a variant-5 pack: persistent warps (default: the occupancy capacity, rounded down to a power of two)\n" " --race --pack the variant race alone (3 rounds): one line per variant, the race line, exit 0 or 1\n" " --race on|off|a,b,c in --serve: race every variant (default), none, or these names\n" " --race-bench-ms N timed window per variant (default 2000)\n" @@ -906,6 +961,11 @@ static Options parseArgs(int argc, char** argv) { auto next = [&]() -> std::string { if (i + 1 >= argc) { usage(); std::exit(2); } return argv[++i]; }; if (a == "--serve") o.serve = true; else if (a == "--check") o.check = true; + else if (a == "--bench") o.bench = true; + else if (a == "--memprobe") o.memprobe = true; + else if (a == "--batches") o.batches = std::atoi(next().c_str()); + else if (a == "--warps") o.warps = std::atoi(next().c_str()); + else if (a == "--probe-mib") o.probeMib = std::atoi(next().c_str()); else if (a == "--race" && (i + 1 >= argc || std::string(argv[i + 1]).rfind("--", 0) == 0)) o.raceOnly = true; else if (a == "--race") o.race = next(); else if (a == "--race-bench-ms") o.raceBenchMs = std::atoi(next().c_str()); @@ -924,12 +984,12 @@ static Options parseArgs(int argc, char** argv) { } if (o.batchLog2 < 10 || o.batchLog2 > 28) { std::printf("--batch-log2 must be between 10 and 28\n"); std::exit(2); } if (o.blockWarps < 1 || o.blockWarps > 32) { std::printf("--block-warps must be between 1 and 32\n"); std::exit(2); } - if (!o.serve && !o.check && !o.raceOnly) { usage(); std::exit(2); } + if (!o.serve && !o.check && !o.raceOnly && !o.bench && !o.memprobe) { usage(); std::exit(2); } if (o.raceBenchMs < 200 || o.raceBenchMs > 20000) { std::printf("--race-bench-ms must be between 200 and 20000\n"); std::exit(2); } if (o.raceBudgetS < 5 || o.raceBudgetS > 540) { std::printf("--race-budget-s must be between 5 and 540 (the prepare lead is 600 DAA)\n"); std::exit(2); } if (o.raceRounds == 0) o.raceRounds = o.raceOnly ? 3 : 1; if (o.tuningPath.empty()) if (const char* t = std::getenv("IGNEUM_TUNING_FILE")) o.tuningPath = t; - if (o.pack.empty()) { std::printf("--pack is required (igneum-miner export-pack writes one)\n"); std::exit(2); } + if (o.pack.empty() && !o.memprobe) { std::printf("--pack is required (igneum-miner export-pack writes one)\n"); std::exit(2); } while (o.pack.size() > 1 && (o.pack.back() == '/' || o.pack.back() == '\\')) o.pack.pop_back(); return o; } @@ -1127,11 +1187,201 @@ static int runServe(Ctx& c, const Options& o, Pair* cur) { // --------------------------------------------------------------------------------------------- // Main +// --------------------------------------------------------------------------------------------- +// Read-width experiment (5 October 2026, docs/plans/read-width.md): --bench and --memprobe + +static uint64_t fnv1a64Bytes(const void* p, size_t n) { + const uint8_t* b = (const uint8_t*)p; + uint64_t h = 0xcbf29ce484222325ull; + for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; } + return h; +} + +// --bench: the pair is built and self-tested (vectors through the bound kernel with the seed words); then a warm-up +// dispatch at base nonce 0 (fingerprinted) and --batches timed dispatches of 2^B nonces, wall time around +// cuStreamSynchronize (the driver API path loads no event symbols; a 2^24 dispatch is 60 to 900 ms on the cards here, +// so the launch overhead is under 1 percent). +static int runBench(Ctx& c, const Options& o, Pair* p) { + uint32_t nonces = 1u << o.batchLog2, block = 32u * (uint32_t)c.blockWarps; + if (p->persistent) { uint32_t unit = 32u * (uint32_t)p->warps; nonces = (nonces / unit) * unit; if (nonces == 0) nonces = unit; } + CUdeviceptr dOut = 0; + std::string err; + if (c.drv.memAlloc(&dOut, (size_t)nonces * 8u) != CUDA_SUCCESS) { std::printf("FAIL: cuMemAlloc out\n"); return 2; } + std::vector hOut(nonces); + double sum = 0, warm = 0; + uint64_t fp = 0; + for (int b = -1; b < o.batches; ++b) { + double t0 = wallMs(); + if (!launchHash(c, p, dOut, (uint32_t)(b + 1) * nonces, p->sw, nonces, block, nullptr, err)) { std::printf("FAIL: %s\n", err.c_str()); return 2; } + CUresult r = c.drv.streamSynchronize(nullptr); + if (r != CUDA_SUCCESS) { std::printf("FAIL: dispatch %d: %s\n", b, c.err(r).c_str()); return 2; } + double ms = wallMs() - t0; + if (b < 0) { + warm = ms; + if (c.drv.memcpyDtoH(hOut.data(), dOut, (size_t)nonces * 8u) != CUDA_SUCCESS) { std::printf("FAIL: read-back\n"); return 2; } + fp = fnv1a64Bytes(hOut.data(), (size_t)nonces * 8u); + } else sum += ms; + } + c.drv.memFree(dOut); + std::string dev = c.name; for (char& ch : dev) if (ch == ' ') ch = '_'; + std::printf("warm-up dispatch (base 0): %.2f ms; %d timed dispatches of %u nonces: mean %.2f ms\n", warm, o.batches, nonces, sum / o.batches); + std::printf("RESULT pack=%s class=%s device=%s arch=%s regs=%d blocks_per_sm=%d warps=%d resident=%d arena_mib=%llu nonces=%u batches=%d check=%s fingerprint=%016llx mhs=%.3f loads=%u bytes=%u scratch_ops=%u time=wall\n", + p->dir.c_str(), p->loadClass.c_str(), dev.c_str(), c.archOpt.c_str(), p->regs, p->blocksPerSM, p->warps, p->residentWarps, (unsigned long long)(p->scratchBytes >> 20), nonces, o.batches, + p->checked ? (p->checkPass ? "PASS" : "FAIL") : "skipped", (unsigned long long)fp, (double)nonces * (double)o.batches / (sum / 1000.0) / 1e6, + p->loadsPerHash, p->bytesPerHash, p->scratchOps * 8u); + return 0; +} + +// --memprobe: the OpenCL worker's table (proto-opencl/host.c, 5 October 2026) in CUDA C through NVRTC, so the two +// vendors are probed with the same access patterns: a dependent chain of random 4-byte reads (the hash's pattern), +// eight independent chains per lane, dependent random 16-byte and 64-byte reads, a coalesced stream and an integer +// chain, at 4, 64 and 1024 MiB. Wall time around cuStreamSynchronize, best of 3, a fresh seed per repetition. +static const char* PROBE_CUDA = + "#include \n" + "__device__ __forceinline__ uint32_t pm_mix(uint32_t x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }\n" + "extern \"C\" __global__ void probe_fill(uint32_t* ds, uint32_t n) { uint32_t i = blockIdx.x * blockDim.x + threadIdx.x; if (i < n) ds[i] = pm_mix(i ^ 0x9E3779B9u); }\n" + "extern \"C\" __global__ void probe_chase(const uint32_t* ds, uint32_t mask, uint32_t steps, uint32_t seed, uint32_t* out) {\n" + " uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t x = pm_mix(g ^ seed);\n" + " for (uint32_t s = 0u; s < steps; ++s) x = ds[x & mask] ^ (x * 0x9E3779B1u + s);\n" + " out[g] = x;\n" + "}\n" + "extern \"C\" __global__ void probe_indep(const uint32_t* ds, uint32_t mask, uint32_t steps, uint32_t seed, uint32_t* out) {\n" + " uint32_t g = blockIdx.x * blockDim.x + threadIdx.x;\n" + " uint32_t x0 = pm_mix(g * 8u ^ seed), x1 = pm_mix((g * 8u + 1u) ^ seed), x2 = pm_mix((g * 8u + 2u) ^ seed), x3 = pm_mix((g * 8u + 3u) ^ seed);\n" + " uint32_t x4 = pm_mix((g * 8u + 4u) ^ seed), x5 = pm_mix((g * 8u + 5u) ^ seed), x6 = pm_mix((g * 8u + 6u) ^ seed), x7 = pm_mix((g * 8u + 7u) ^ seed);\n" + " for (uint32_t s = 0u; s < steps; ++s) {\n" + " x0 = ds[x0 & mask] ^ (x0 * 0x9E3779B1u + s); x1 = ds[x1 & mask] ^ (x1 * 0x9E3779B1u + s);\n" + " x2 = ds[x2 & mask] ^ (x2 * 0x9E3779B1u + s); x3 = ds[x3 & mask] ^ (x3 * 0x9E3779B1u + s);\n" + " x4 = ds[x4 & mask] ^ (x4 * 0x9E3779B1u + s); x5 = ds[x5 & mask] ^ (x5 * 0x9E3779B1u + s);\n" + " x6 = ds[x6 & mask] ^ (x6 * 0x9E3779B1u + s); x7 = ds[x7 & mask] ^ (x7 * 0x9E3779B1u + s);\n" + " }\n" + " out[g] = x0 ^ x1 ^ x2 ^ x3 ^ x4 ^ x5 ^ x6 ^ x7;\n" + "}\n" + "extern \"C\" __global__ void probe_line16(const uint4* ds, uint32_t vecMask, uint32_t steps, uint32_t seed, uint32_t* out) {\n" + " uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t x = pm_mix(g ^ seed);\n" + " for (uint32_t s = 0u; s < steps; ++s) { uint4 a = ds[x & vecMask]; x = (a.x ^ a.y ^ a.z ^ a.w) ^ (x * 0x9E3779B1u + s); }\n" + " out[g] = x;\n" + "}\n" + "extern \"C\" __global__ void probe_line(const uint4* ds, uint32_t lineMask, uint32_t steps, uint32_t seed, uint32_t* out) {\n" + " uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t x = pm_mix(g ^ seed);\n" + " for (uint32_t s = 0u; s < steps; ++s) { uint32_t l = (x & lineMask) * 4u; uint4 a = ds[l], b = ds[l + 1u], c = ds[l + 2u], d = ds[l + 3u]; x = (a.x ^ b.y ^ c.z ^ d.w) ^ (x * 0x9E3779B1u + s); }\n" + " out[g] = x;\n" + "}\n" + "extern \"C\" __global__ void probe_stream(const uint4* ds, uint32_t perLane, uint32_t* out) {\n" + " uint32_t g = blockIdx.x * blockDim.x + threadIdx.x, n = gridDim.x * blockDim.x; uint4 acc = make_uint4(0u, 0u, 0u, 0u);\n" + " for (uint32_t s = 0u; s < perLane; ++s) { uint4 v = ds[s * n + g]; acc.x ^= v.x; acc.y ^= v.y; acc.z ^= v.z; acc.w ^= v.w; }\n" + " out[g] = acc.x ^ acc.y ^ acc.z ^ acc.w;\n" + "}\n" + "extern \"C\" __global__ void probe_alu(uint32_t steps, uint32_t seed, uint32_t* out) {\n" + " uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u;\n" + " for (uint32_t s = 0u; s < steps; ++s) { x = x * 0x9E3779B1u + ((y << 7u) | (y >> 25u)); y = (y ^ x) + s; }\n" + " out[g] = x ^ y;\n" + "}\n"; + +static double probeLaunch(Ctx& c, CUfunction f, size_t lanes, size_t local, int reps, int seedArg, uint32_t seed, void** args) { + double best = -1; + for (int r = 0; r < reps; ++r) { + uint32_t s = seed + (uint32_t)r * 0x9E3779B9u; + if (seedArg >= 0) args[seedArg] = &s; + double t0 = wallMs(); + if (c.drv.launchKernel(f, (unsigned)(lanes / local), 1, 1, (unsigned)local, 1, 1, 0, nullptr, args, nullptr) != CUDA_SUCCESS) return -1; + if (c.drv.streamSynchronize(nullptr) != CUDA_SUCCESS) return -1; + double ms = wallMs() - t0; + if (best < 0 || ms < best) best = ms; + } + return best; +} + +static int runMemprobe(Ctx& c, const Options& o) { + Compiled cp; + std::string err; + if (!rtcCompile(c, PROBE_CUDA, "probe.cu", "", "", {}, cp, err)) { std::printf("memprobe: build FAILED: %s\n", err.c_str()); return 2; } + CUmodule mod = nullptr; + if (c.drv.moduleLoadData(&mod, cp.image.data()) != CUDA_SUCCESS) { std::printf("memprobe: cuModuleLoadData failed\n"); return 2; } + CUfunction kFill, kChase, kIndep, kLine16, kLine, kStream, kAlu; + const char* names[7] = { "probe_fill", "probe_chase", "probe_indep", "probe_line16", "probe_line", "probe_stream", "probe_alu" }; + CUfunction* fns[7] = { &kFill, &kChase, &kIndep, &kLine16, &kLine, &kStream, &kAlu }; + for (int i = 0; i < 7; ++i) if (c.drv.moduleGetFunction(fns[i], mod, names[i]) != CUDA_SUCCESS) { std::printf("memprobe: %s not in the module\n", names[i]); return 2; } + int sizes[3] = { 4, 64, 1024 }, nSizes = 3; + if (o.probeMib > 0) { sizes[0] = o.probeMib; nSizes = 1; } + const size_t lanesList[8] = { 256, 1024, 1u << 12, 1u << 14, 1u << 16, 1u << 18, 1u << 20, 1u << 22 }; + const size_t groups[2] = { 32, 256 }; + const uint32_t STEPS = 256u, ALU_STEPS = 4096u; + const size_t maxLanes = 1u << 22; + CUdeviceptr dOut = 0; + if (c.drv.memAlloc(&dOut, maxLanes * 4u) != CUDA_SUCCESS) { std::printf("memprobe: cuMemAlloc out\n"); return 2; } + std::printf("memprobe on %s (sm_%d%d, %d SMs, driver %d.%d, NVRTC %d.%d), wall time around cuStreamSynchronize\n", c.name.c_str(), c.major, c.minor, c.sms, c.driverVersion / 1000, (c.driverVersion % 100) / 10, c.rtcMajor, c.rtcMinor); + std::printf("| probe | MiB | block | lanes in flight | steps per lane | best ms | G loads/s | ns per dependent load |\n|---|---|---|---|---|---|---|---|\n"); + for (int si = 0; si < nSizes; ++si) { + int mib = sizes[si]; + uint64_t bytes = (uint64_t)mib << 20; + uint32_t words = (uint32_t)(bytes / 4ull), mask = words - 1u, n = words; + CUdeviceptr dDs = 0; + if (c.drv.memAlloc(&dDs, (size_t)bytes) != CUDA_SUCCESS) { std::printf("| chase | %d | skipped: cuMemAlloc failed | | | | | |\n", mib); continue; } + { void* a[2] = { &dDs, &n }; probeLaunch(c, kFill, ((size_t)words + 255) / 256 * 256, 256, 1, -1, 0, a); } + for (int gi = 0; gi < 2; ++gi) { + size_t local = groups[gi]; + for (int li = 0; li < 8; ++li) { + size_t lanes = lanesList[li]; + if (lanes < local) continue; + uint32_t seed = 0x1234567u + (uint32_t)li * 977u, steps = STEPS; + void* a[5] = { &dDs, &mask, &steps, &seed, &dOut }; + double ms = probeLaunch(c, kChase, lanes, local, 3, 3, seed, a); + std::printf("| chase | %d | %zu | %zu | %u | %.3f | %.3f | %.0f |\n", mib, local, lanes, STEPS, ms, (double)lanes * STEPS / (ms / 1000.0) / 1e9, ms * 1e6 / STEPS); + std::fflush(stdout); + } + } + for (size_t lanes = 1u << 16; lanes <= maxLanes; lanes <<= 2) { + uint32_t seed = 0x7654321u, steps = STEPS; + void* a[5] = { &dDs, &mask, &steps, &seed, &dOut }; + double ms = probeLaunch(c, kIndep, lanes, 256, 3, 3, seed, a); + std::printf("| indep x8 | %d | 256 | %zu | %u | %.3f | %.3f | (8 loads in flight per lane) |\n", mib, lanes, STEPS, ms, (double)lanes * 8.0 * STEPS / (ms / 1000.0) / 1e9); + } + for (size_t lanes = 1u << 14; lanes <= maxLanes; lanes <<= 2) { + uint32_t vecMask = (words / 4u) - 1u, seed = 0x2718281u, steps = STEPS; + void* a[5] = { &dDs, &vecMask, &steps, &seed, &dOut }; + double ms = probeLaunch(c, kLine16, lanes, 256, 3, 3, seed, a); + std::printf("| line 16 B | %d | 256 | %zu | %u | %.3f | %.3f G reads/s | %.1f GB/s in 16 B reads |\n", mib, lanes, STEPS, ms, (double)lanes * STEPS / (ms / 1000.0) / 1e9, (double)lanes * STEPS * 16.0 / (ms / 1000.0) / 1e9); + } + for (size_t lanes = 1u << 14; lanes <= maxLanes; lanes <<= 2) { + uint32_t lineMask = (words / 16u) - 1u, seed = 0x3141592u, steps = STEPS; + void* a[5] = { &dDs, &lineMask, &steps, &seed, &dOut }; + double ms = probeLaunch(c, kLine, lanes, 256, 3, 3, seed, a); + std::printf("| line 64 B | %d | 256 | %zu | %u | %.3f | %.3f G lines/s | %.1f GB/s in lines |\n", mib, lanes, STEPS, ms, (double)lanes * STEPS / (ms / 1000.0) / 1e9, (double)lanes * STEPS * 64.0 / (ms / 1000.0) / 1e9); + } + { + size_t lanes = 1u << 20; + uint32_t perLane = (uint32_t)((uint64_t)words / 4ull / (uint64_t)lanes); + if (perLane == 0) { perLane = 1; lanes = (size_t)words / 4u; } + double bytesRead = (double)perLane * (double)lanes * 16.0; + void* a[3] = { &dDs, &perLane, &dOut }; + double ms = probeLaunch(c, kStream, lanes, 256, 3, -1, 0, a); + std::printf("| stream | %d | 256 | %zu | %u | %.3f | %.1f GB/s coalesced | (%.0f MiB read once) |\n", mib, lanes, perLane, ms, bytesRead / (ms / 1000.0) / 1e9, bytesRead / 1048576.0); + } + c.drv.memFree(dDs); + std::fflush(stdout); + } + { + size_t lanes = 1u << 20; + uint32_t seed = 0x2468aceu, steps = ALU_STEPS; + void* a[3] = { &steps, &seed, &dOut }; + double ms = probeLaunch(c, kAlu, lanes, 256, 3, 1, seed, a); + double ops = (double)lanes * ALU_STEPS * 5.0; + std::printf("| alu | 0 | 256 | %zu | %u | %.3f | %.1f G int ops/s | %.3f G steps/s per SM (approximate: 5 ops per step counted) |\n", lanes, ALU_STEPS, ms, ops / (ms / 1000.0) / 1e9, (double)lanes * ALU_STEPS / (ms / 1000.0) / 1e9 / (c.sms ? c.sms : 1)); + } + c.drv.memFree(dOut); + c.drv.moduleUnload(mod); + std::printf("memprobe: done\n"); + return 0; +} + int main(int argc, char** argv) { Options o = parseArgs(argc, argv); Ctx c; c.blockWarps = o.blockWarps; c.race = o.race; c.raceBenchMs = o.raceBenchMs; c.raceBudgetS = o.raceBudgetS; c.raceRounds = o.raceRounds; c.batchLog2 = o.batchLog2; c.pinned = o.pinned; + c.warps = o.warps; c.batches = o.batches; + if (o.bench || o.memprobe) c.race = "off"; if (!o.tuningPath.empty()) { bool ok = false; c.tuning = readText(o.tuningPath, ok); if (!ok) c.tuning.clear(); } std::string err, drvLib, rtcLib; if (!loadDriver(c.drv, err, drvLib)) { emit("error 0 " + err); return 2; } @@ -1140,9 +1390,17 @@ int main(int argc, char** argv) { info(fmt("igneum-worker-cuda %s: device %d %s (sm_%d%d, %d SMs), driver %d.%d from %s, NVRTC %d.%d from %s, target %s (%s)", WORKER_VERSION, o.device, c.name.c_str(), c.major, c.minor, c.sms, c.driverVersion / 1000, (c.driverVersion % 100) / 10, drvLib.c_str(), c.rtcMajor, c.rtcMinor, rtcLib.c_str(), c.archOpt.c_str(), c.why.c_str())); if (!c.tuning.empty()) info(fmt("tuning file %s (%zu bytes): %s", o.tuningPath.c_str(), c.tuning.size(), readTuning(c.tuning, c.name).found ? "has an entry for this card" : "no entry for this card")); + if (o.memprobe) { int rc = runMemprobe(c, o); c.drv.primaryCtxRelease(c.dev); return rc; } double t0 = wallMs(); - Pair* cur = buildPair(c, o.pack, nullptr, err, !o.check); + Pair* cur = buildPair(c, o.pack, nullptr, err, !o.check && !o.bench); if (!cur) { emit("error 0 " + err); return 1; } + if (o.bench) { + std::printf("pack %s on %s: %s\n", o.pack.c_str(), c.name.c_str(), pairSummary(cur).c_str()); + int rc = runBench(c, o, cur); + releasePair(c, cur); + c.drv.primaryCtxRelease(c.dev); + return rc; + } if (o.raceOnly) { std::printf("race %s on %s (%s, %d SMs, driver %d.%d, NVRTC %d.%d, %s): %s\n", o.pack.c_str(), c.name.c_str(), c.archOpt.c_str(), c.sms, c.driverVersion / 1000, (c.driverVersion % 100) / 10, c.rtcMajor, c.rtcMinor, c.why.c_str(), pairSummary(cur).c_str()); std::printf("%s\n", cur->raceLine.c_str()); diff --git a/proto-metal/packbench.swift b/proto-metal/packbench.swift index 75b6fc4f3..3110bbe12 100644 --- a/proto-metal/packbench.swift +++ b/proto-metal/packbench.swift @@ -205,5 +205,5 @@ print("device \(device.name); compile \(String(format: "%.0f", compileMs)) ms; c print("cache FNV-1a 64 \(String(format: "%016llx", cacheFnv)) \(cacheOk ? "PASS" : "FAIL"); dataset head and last \(dsOk ? "PASS" : "FAIL"); vectors standalone \(vecPass)/\(vecBases.count), in batch \(batchVecPass)/\(batchVecN)") print("warm-up batch \(nonces) hashes: \(String(format: "%.1f", warmGpu)) ms GPU, \(String(format: "%.1f", warmWall)) ms wall") let overall = cacheOk && dsOk && vecPass == vecBases.count && batchVecPass == batchVecN -print("RESULT pack=\(packName) class=\(className) device=\(device.name.replacingOccurrences(of: " ", with: "_")) group=\(opts.group) warps=\(warpsN) arena_mib=\(persistent ? warpsN : 0) nonces=\(nonces) batches=\(opts.batches) vectors=\(vecPass)/\(vecBases.count) batch_vectors=\(batchVecPass)/\(batchVecN) cache=\(cacheOk ? "PASS" : "FAIL") dataset=\(dsOk ? "PASS" : "FAIL") fingerprint=\(String(format: "%016llx", fingerprint)) mhs_gpu=\(String(format: "%.3f", mhsGpu)) mhs_wall=\(String(format: "%.3f", mhsWall)) loads=\(loadsPerHash) bytes=\(bytesPerHash) scratch_ops=\(scratchOps * 8) overall=\(overall ? "PASS" : "FAIL")") +print("RESULT pack=\(packName) class=\(className) device=\(device.name.replacingOccurrences(of: " ", with: "_")) group=\(opts.group) warps=\(warpsN) arena_mib=\(persistent ? warpsN * 32 * scratchWordsPerLane * 4 / 1048576 : 0) nonces=\(nonces) batches=\(opts.batches) vectors=\(vecPass)/\(vecBases.count) batch_vectors=\(batchVecPass)/\(batchVecN) cache=\(cacheOk ? "PASS" : "FAIL") dataset=\(dsOk ? "PASS" : "FAIL") fingerprint=\(String(format: "%016llx", fingerprint)) mhs_gpu=\(String(format: "%.3f", mhsGpu)) mhs_wall=\(String(format: "%.3f", mhsWall)) loads=\(loadsPerHash) bytes=\(bytesPerHash) scratch_ops=\(scratchOps * 8) overall=\(overall ? "PASS" : "FAIL")") exit(overall ? 0 : 1) diff --git a/relay/playbooks/readwidth-5090.ps1 b/relay/playbooks/readwidth-5090.ps1 new file mode 100644 index 000000000..c9384f214 --- /dev/null +++ b/relay/playbooks/readwidth-5090.ps1 @@ -0,0 +1,67 @@ +# Igneum run job: the read-width experiment on PC 2's RTX 5090 (machine 1ccfe586), 5 October 2026 (docs/plans/read-width.md). +# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining; this script switches +# off ONLY the NVIDIA card in the app through POST api/cards, waits for its worker to stop, runs the fetched +# igneum-worker-cuda.exe (--memprobe, then --bench on every pack of the fetched packs folder), and switches the card back +# on with the settings it had. Every result line starts with RESULT so `node tools/jobs.mjs ` shows them. +$ErrorActionPreference = 'Continue' +function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) } +$jobs = Split-Path $env:IGNEUM_JOB_DIR +$fetched = Join-Path $jobs 'fetch-readwidth-20261005' +$exe = Join-Path $fetched 'igneum-worker-cuda.exe' +$packs = Join-Path $fetched 'packs-readwidth' +if (-not (Test-Path $exe)) { Write-Output "RESULT error worker missing at $exe (the fetch job runs first)"; exit 2 } +if (-not (Test-Path $packs)) { Write-Output "RESULT error packs missing at $packs"; exit 2 } +$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1 +if (-not $inst) { Write-Output 'RESULT error no installed igneum-worker-cuda.exe (the NVRTC DLLs come from there)'; exit 2 } +Get-ChildItem $inst -Filter 'nvrtc*.dll' | Copy-Item -Destination $fetched -Force +Write-Output "RESULT worker $exe sha256 $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower()) with $((Get-ChildItem $fetched -Filter 'nvrtc*.dll').Count) NVRTC DLL(s) from $inst" + +# the app: switch off the NVIDIA card only, remember its settings +$appDir = $env:IGNEUM_APP_DIR +if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' } +$urlFile = Join-Path $appDir 'app.url' +$url = $null +if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() } +$card = $null +if ($url) { + try { + $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10 + $card = $st.mining.cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1 + if (-not $card) { $card = $st.cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1 } + } catch { Say ("api/state: " + $_.Exception.Message) } +} +if ($card) { + Write-Output ("RESULT card " + $card.key + " enabled=" + $card.enabled + " identities=" + $card.identities + " power_pct=" + $card.power_pct + " state=" + $card.state) + $body = @{ cards = @(@{ key = $card.key; enabled = $false; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "card off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) } + $t = 0 + while ($t -lt 90) { + Start-Sleep -Seconds 5; $t += 5 + try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $c2 = $st.mining.cards | Where-Object { $_.key -eq $card.key }; if (-not $c2) { $c2 = $st.cards | Where-Object { $_.key -eq $card.key } }; if ($c2 -and $c2.state -eq 'off' -and $c2.pid -eq 0) { break } } catch { } + } + Write-Output ("RESULT card-off after " + $t + " s") + Start-Sleep -Seconds 5 +} else { Write-Output 'RESULT card none-found (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' } + +& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" } +Write-Output "RESULT memprobe start $(Get-Date -Format HH:mm:ss)" +& $exe --memprobe 2>&1 | ForEach-Object { "RESULT $_" } +foreach ($pk in @('w4', 'w16', 'w64', 'w64x4', 'mixA-0', 'mixA-1', 'mixA-2', 'mixA-3', 'mixA-4', 'mixA-5', 'mixB-0', 'mixB-1', 'mixB-2', 'mixB-3', 'mixB-4', 'mixB-5')) { + $d = Join-Path $packs $pk + Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)" + & $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT $_" } + & $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL' } | ForEach-Object { "RESULT $_" } +} +foreach ($pk in @('scr0k32', 'scr2k32', 'scr4k32', 'scr8k32', 'scr2k128', 'scr4k128', 'scr8k128')) { + $d = Join-Path $packs $pk + Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)" + & $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT $_" } + & $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 --warps 4096 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|variant 5' } | ForEach-Object { "RESULT $_" } +} +& nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" } + +if ($card) { + $body = @{ cards = @(@{ key = $card.key; enabled = [bool]$card.enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored enabled=" + $card.enabled) } catch { Write-Output ("RESULT error card restore: " + $_.Exception.Message) } +} +exit 0 diff --git a/relay/playbooks/readwidth-9070.ps1 b/relay/playbooks/readwidth-9070.ps1 new file mode 100644 index 000000000..4e44dea09 --- /dev/null +++ b/relay/playbooks/readwidth-9070.ps1 @@ -0,0 +1,69 @@ +# Igneum run job: the read-width experiment on PC 1's RX 9070 XT on the eGPU (machine ae432dc7), 5 October 2026 (docs/plans/read-width.md). +# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining; this script switches +# off ONLY the NVIDIA card in the app through POST api/cards, waits for its worker to stop, runs the fetched +# igneum-worker-opencl.exe (--memprobe, then --bench on every pack of the fetched packs folder), and switches the card back +# on with the settings it had. Every result line starts with RESULT so `node tools/jobs.mjs ` shows them. +$ErrorActionPreference = 'Continue' +function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) } +$jobs = Split-Path $env:IGNEUM_JOB_DIR +$fetched = Join-Path $jobs 'fetch-readwidth-20261005' +$exe = Join-Path $fetched 'igneum-worker-opencl.exe' +$packs = Join-Path $fetched 'packs-readwidth' +if (-not (Test-Path $exe)) { Write-Output "RESULT error worker missing at $exe (the fetch job runs first)"; exit 2 } +if (-not (Test-Path $packs)) { Write-Output "RESULT error packs missing at $packs"; exit 2 } +Write-Output "RESULT worker $exe sha256 $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower())" +# the card's OpenCL device index on the current (3683.0) platform, from the worker's own list (the older platform's duplicate is marked dup) +$list = & $exe --list 2>&1 +$list | ForEach-Object { "RESULT list $_" } +$dev = $null +foreach ($l in $list) { if ($l -match '^\s*\[(\d+)\].*gfx1201' -and $l -notmatch 'dup') { $dev = [int]$Matches[1]; break } } +if ($null -eq $dev) { Write-Output 'RESULT error no gfx1201 device in --list'; exit 2 } +Write-Output "RESULT device $dev" + +# the app: switch off the NVIDIA card only, remember its settings +$appDir = $env:IGNEUM_APP_DIR +if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' } +$urlFile = Join-Path $appDir 'app.url' +$url = $null +if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() } +$card = $null +if ($url) { + try { + $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10 + $card = $st.mining.cards | Where-Object { ($_.vendor -eq 'amd' -and $_.key -match 'gfx1201') } | Select-Object -First 1 + if (-not $card) { $card = $st.cards | Where-Object { ($_.vendor -eq 'amd' -and $_.key -match 'gfx1201') } | Select-Object -First 1 } + } catch { Say ("api/state: " + $_.Exception.Message) } +} +if ($card) { + Write-Output ("RESULT card " + $card.key + " enabled=" + $card.enabled + " identities=" + $card.identities + " power_pct=" + $card.power_pct + " state=" + $card.state) + $body = @{ cards = @(@{ key = $card.key; enabled = $false; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "card off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) } + $t = 0 + while ($t -lt 90) { + Start-Sleep -Seconds 5; $t += 5 + try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $c2 = $st.mining.cards | Where-Object { $_.key -eq $card.key }; if (-not $c2) { $c2 = $st.cards | Where-Object { $_.key -eq $card.key } }; if ($c2 -and $c2.state -eq 'off' -and $c2.pid -eq 0) { break } } catch { } + } + Write-Output ("RESULT card-off after " + $t + " s") + Start-Sleep -Seconds 5 +} else { Write-Output 'RESULT card none-found (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' } + +Write-Output "RESULT memprobe start $(Get-Date -Format HH:mm:ss)" +& $exe --device $dev --memprobe 2>&1 | ForEach-Object { "RESULT $_" } +foreach ($pk in @('w4', 'w16', 'w64', 'w64x4', 'mixA-0', 'mixA-1', 'mixA-2', 'mixA-3', 'mixA-4', 'mixA-5', 'mixB-0', 'mixB-1', 'mixB-2', 'mixB-3', 'mixB-4', 'mixB-5')) { + $d = Join-Path $packs $pk + Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)" + & $exe --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev 2>&1 | ForEach-Object { "RESULT $_" } + & $exe --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev --group-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL' } | ForEach-Object { "RESULT $_" } +} +foreach ($pk in @('scr0k32', 'scr2k32', 'scr4k32', 'scr8k32', 'scr2k128', 'scr4k128', 'scr8k128')) { + $d = Join-Path $packs $pk + Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)" + & $exe --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev 2>&1 | ForEach-Object { "RESULT $_" } + & $exe --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev --warps 4096 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|variant 5' } | ForEach-Object { "RESULT $_" } +} + +if ($card) { + $body = @{ cards = @(@{ key = $card.key; enabled = [bool]$card.enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored enabled=" + $card.enabled) } catch { Write-Output ("RESULT error card restore: " + $_.Exception.Message) } +} +exit 0 From 5d5ba15ed738ccf71bb11230a15e853f089e2afa Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:09:10 +0000 Subject: [PATCH 004/311] Counter ASIC 2.0 layer 6: SRAM mirror analysis (cited bit cells N7 to N2 and 18A, area and cost per node, no cache growth rule in the spec, options A to E for Josh); layer 7 dp4a probes for Metal, CUDA and OpenCL (standalone, no lottery kernel) Co-Authored-By: Claude Fable 5.1 --- docs/analysis/sram-mirror.md | 227 +++++++++++++++++++++++++++++++++++ proto-cuda/dot4-probe.cu | 115 ++++++++++++++++++ proto-metal/dot4-probe.swift | 165 +++++++++++++++++++++++++ proto-opencl/dot4-probe.c | 187 +++++++++++++++++++++++++++++ 4 files changed, 694 insertions(+) create mode 100644 docs/analysis/sram-mirror.md create mode 100644 proto-cuda/dot4-probe.cu create mode 100644 proto-metal/dot4-probe.swift create mode 100644 proto-opencl/dot4-probe.c diff --git a/docs/analysis/sram-mirror.md b/docs/analysis/sram-mirror.md new file mode 100644 index 000000000..970239776 --- /dev/null +++ b/docs/analysis/sram-mirror.md @@ -0,0 +1,227 @@ +# Layer 6: the SRAM mirror of the cache against published SRAM density, year 0 to 10 + +5 October 2026 (night), Counter ASIC 2.0 (`docs/plans/counter-asic-2.md`, layer 6), branch `ca2-analysis`. Every figure +below is either cited (paper, vendor document, URL, date) or labelled approximate. Nothing here is a measurement of a +chip. Numbers in this file were computed with the arithmetic shown; the script is in section 9. + +## 1. The question + +The lottery hash derives every dataset item from a 256 MiB cache (spec 01 sections 1.5 and 1.8). A chip that holds the +cache in on-die SRAM can recompute items instead of reading the dataset (ledger M16, the recompute attacker). Layer 6 +asks whether the cache size, as the specification schedules it, keeps that SRAM mirror unaffordable for ten years of +the genesis schedule, and if not what growth rule would. + +Two things also sit in a chip's SRAM budget if it mirrors the full read-only working set: the layer 5 hot table (32, +64 or 96 MB, a class parameter on `readwidth` b970dda, coordinator's note of 5 October) beside the 256 MiB cache. The +per-warp scratch of layer 3 (32 or 128 KB per warp, written, not read-only) is not mirrorable and is left out of the +mirror; it is counted in the 6 GB working-set budget in section 7. + +## 2. What the specification schedules for the cache + +| Quantity | Rule | Where | +|---|---|---| +| Dataset | 2 GiB at genesis plus 0.5 GiB per year (`N_d` grows about 23 KiB per day) | spec 01 section 1.13.3, Designed | +| Cache | 256 MiB, "prototype value, to be fixed at gate 1"; the rule that fixes it: "the cache must exceed the largest on-chip cache of any card that mines, and 96 MiB of L2 on the 5090 is the figure to beat" | spec 01 sections 1.5 and 1.16 | +| Cache growth | None. No section of `docs/spec/` grows the cache (grep of `docs/spec` for cache growth, schedule, doubling: only the dataset rule of 1.13.3 and the README's "growth" word, which refers to it) | this analysis, 5 October 2026 | + +So the plan's layer 6 row ("already in the design; confirm the schedule") is half right: dataset growth is in the +design, cache growth is not. The cache is flat at 256 MiB for every year of the schedule as the spec stands. M16's +closing line names the rule the cache should get ("exceeds what one die can hold, and grows") as a gate 1 decision +that has not been taken. + +## 3. SRAM bit cell per node, cited + +| Node (vendor) | HD 6T bit cell, um^2 | Raw density, Mbit/mm^2 (1/cell) | Year of volume (approximate) | Source | +|---|---|---|---|---| +| N7 (TSMC) | 0.027 | 37.0 | 2018 | WikiChip, "TSMC Details 5 nm" (ISSCC/IEDM disclosures), https://fuse.wikichip.org/news/3398/tsmc-details-5-nm/ | +| N5 (TSMC) | 0.021 | 47.6 | 2020 | same (two N5 cells: HD 0.021, HP 0.025) | +| N3B (TSMC) | 0.0199 | 50.3 | 2022 to 2023 | WikiChip, "IEDM 2022: Did We Just Witness The Death Of SRAM?", https://fuse.wikichip.org/news/7343/iedm-2022-did-we-just-witness-the-death-of-sram/ (TSMC's IEDM 2022 N3 paper) | +| N3E (TSMC) | 0.021 | 47.6 | 2023 | same; Tom's Hardware, "TSMC's 3nm Node: No SRAM Scaling", https://www.tomshardware.com/news/no-sram-scaling-implies-on-more-expensive-cpus-and-gpus | +| N2 (TSMC) | 0.0175 | 57.1 | 2025 to 2026 | TSMC at IEDM 2024, reported by Tom's Hardware, https://www.tomshardware.com/tech-industry/tsmc-shares-deep-dive-details-about-its-cutting-edge-2nm-process-node-at-iedm-2024-35-percent-less-power-or-15-percent-more-performance ; ISSCC 2025 paper "A 38.1Mb/mm2 SRAM in a 2nm-CMOS-Nanosheet Technology", https://research.tsmc.com/page/memory/4.html | +| Intel 18A | 0.021 | 47.6 | 2025 to 2026 | ISSCC 2025 paper 29.2, "A 0.021 um^2 High-Density SRAM in Intel 18A RibbonFET Technology with PowerVia", https://www.researchgate.net/publication/389644177 ; IEEE Spectrum 26 Feb 2025, https://spectrum.ieee.org/sram-intel-tsmc | +| Samsung SF3 / SF2 | not disclosed as a bit cell area in anything found tonight (Samsung's ISSCC papers give assist circuits and macro figures, not the HD cell) | | | search of ISSCC 2021 to 2025 coverage, 5 October 2026; left out of the tables | + +The stall. N3B's cell is 5% smaller than N5's and N3E's is the same size as N5's (0.021 um^2 both): zero SRAM +scaling from N5 to N3E (WikiChip IEDM 2022 article above; Tom's Hardware above; SemiAnalysis "TSMC's 3nm Conundrum", +https://newsletter.semianalysis.com/p/tsmcs-3nm-conundrum-does-it-even). N2's nanosheet cell recovers 17% (0.021 to +0.0175 um^2). So across 2020 to 2026 the HD bit cell shrank once, by 17%. + +Array efficiency (bit cell to macro). The usable density of a macro is below 1/cell because of word-line and +bit-line drivers, sense amplifiers, decoders and redundancy. The factor used here is 0.70, WikiChip's convention +(their 31.8 Mib/mm^2 for the 0.021 um^2 N3E cell is 1/0.021 x 0.70 in Mib). The two ISSCC 2025 macros bracket it: +TSMC N2 38.1 Mb/mm^2 at a 0.0175 um^2 cell is 67%; Intel 18A 38.1 Mb/mm^2 array density and 34.3 Mb/mm^2 for the +volume macro at a 0.021 um^2 cell are 80% and 72% (ISSCC 2025 29.2, above). Both lie within 10% of 0.70. + +GPU on-die SRAM for scale: the RTX 5090 carries 96 MB of L2 (98,304 KB) on a 750 mm^2 TSMC 4N die with 92.2 billion +transistors; the full GB202 has 128 MB; the RTX 4090 had 72 MB and the RTX 3090 6 MB (NVIDIA, "RTX Blackwell GPU +Architecture" whitepaper v1.1, appendix table "L2 Cache Size", https://images.nvidia.com/aem-dam/Solutions/geforce/blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf). +At the N5-class cell and 0.70 that L2 is about 24 mm^2 of the 750 (3%), approximate. The RX 9070 XT carries 64 MB +of Infinity Cache plus 8 MB of L2 (vendor figures, approximate, bench-log "the 9070 XT on the eGPU"). + +Reticle: the EUV field is 26 x 33 mm = 858 mm^2, about 830 mm^2 usable after scribe lanes (SemiAnalysis, "Die Size +And Reticle Conundrum", https://newsletter.semianalysis.com/p/die-size-and-reticle-conundrum-cost ; WikiChip "Mask", +https://en.wikichip.org/wiki/mask). The 5090's 750 mm^2 is 90% of it. + +Wafer prices (approximate; TSMC publishes none, every figure is supply-chain reporting): N7 about $9,500, N5 and N3 +about $20,000 (Silicon Analysts, "Wafer Pricing by Node", September 2026, https://siliconanalysts.com/data/wafer-pricing); +N2 about $30,000 (Tom's Hardware, https://www.tomshardware.com/tech-industry/semiconductors/tsmc-could-charge-up-to-usd45-000-for-1-6nm-wafers-rumors-allege-a-50-percent-increase-in-pricing-over-prior-gen-wafers). + +## 4. Die area to mirror the cache, per node + +Area = bits / (raw density x 0.70). The columns are the 256 MiB cache alone, the cache plus each hot-table size of +layer 5 (32, 64, 96 MB taken as MiB), and the larger caches of the options in section 6. + +| Node | Macro Mbit/mm^2 at 0.70 | 256 MiB | 256 + 32 | 256 + 64 | 256 + 96 | 512 MiB | 1 GiB | 2 GiB | 4 GiB | +|---|---|---|---|---|---|---|---|---|---| +| N7 | 25.9 | 83 mm^2 | 93 | 104 | 114 | 166 | 331 | 663 | 1,325 (2 dies) | +| N5 | 33.3 | 64 | 72 | 81 | 89 | 129 | 258 | 515 | 1,031 (2 dies) | +| N3B | 35.2 | 61 | 69 | 76 | 84 | 122 | 244 | 488 | 977 (2 dies) | +| N3E, Intel 18A | 33.3 | 64 | 72 | 81 | 89 | 129 | 258 | 515 | 1,031 (2 dies) | +| N2 | 40.0 | 54 | 60 | 67 | 74 | 107 | 215 | 429 | 859 (2 dies) | + +One reticle (830 mm^2) holds 2.5 GiB of SRAM at N7, 3.2 GiB at N5, N3E and 18A, 3.9 GiB at N2 (same arithmetic). + +Against the figures the ledger carries: M16's "100 to 300 mm^2" (low end from a 0.02 um^2 cell with overhead, high +end from wafer-scale parts at about 1 MB/mm^2) and the plan's "about 45 mm^2 at a leading node" both bracket the +cited 54 to 64 mm^2; the wafer-scale high end is a different efficiency (Cerebras-class arrays sit beside logic) and +is not the right number for a pure SRAM die. The right figure for the ledger is 54 to 83 mm^2 depending on node, +cited above. + +## 5. Cost per good die + +Dies per 300 mm wafer by the usual approximation pi x 150^2 / A minus the edge term pi x 300 / sqrt(2A); yield by +Poisson exp(-A x D0) with D0 = 0.1 defects per cm^2 (an assumption, approximate; SRAM arrays carry redundancy so +real yield is higher, which lowers these costs). Cost per good die = wafer price / (dies x yield). Packaging, test, +the logic beside the SRAM and the design (masks at N5 and below run into the tens of millions of dollars, +approximate) are not in these numbers; they are per-die silicon only. + +| Node, wafer price | 256 MiB | 256 + 96 MiB | 1 GiB | 4 GiB (2 dies) | +|---|---|---|---|---| +| N7, $9,500 | 83 mm^2, 780 dies, yield 0.92, $13 | $19 | 331 mm^2, 177 dies, 0.72, $75 | $456 | +| N5, $20,000 | 64 mm^2, 1,014 dies, 0.94, $21 | $30 | 258 mm^2, 233 dies, 0.77, $111 | $621 | +| N3E, $20,000 | $21 | $30 | $111 | $621 | +| N2, $30,000 | 54 mm^2, 1,226 dies, 0.95, $26 | $37 | 215 mm^2, 284 dies, 0.81, $131 | $696 | + +Reading. The silicon for a 256 MiB mirror is tens of dollars per die on any node from N7 up. With the hot table it +is still under $40. It was never the SRAM that priced the recompute attacker out; the plan's premise for layer 6 +("the SRAM mirror stays unaffordable") does not hold for the cache as a mirror and did not hold at genesis either. + +## 6. What the mirror buys the attacker, year by year + +From M16 (`docs/analysis/m16-recompute-attacker-2026-10-05.md`): with the cache on die the attacker recomputes 128 +items per hash at about 1,170 integer operations and 8 dependent 64-byte cache reads each, about 150,000 operations +and 1,024 dependent SRAM reads per hash. At a 5090-class integer budget (about 50 T op/s, approximate) that is +0.33 Ghash/s against the honest 141 Mhash/s projected for version 2 programs: 2.4x at equal silicon before any +fixed-function factor, 3x to 6x with one (approximate). The SRAM is 54 to 83 mm^2 of that chip (7 to 11% of a +750 mm^2 die), so the mirror is cheap and the recompute route is bound by integer throughput, not by SRAM. + +The layer 5 hot table changes nothing in that arithmetic: the hot table is read-only and derived from the day key +like the cache, so a chip mirrors it in the same SRAM (another 32 to 96 MB, 7 to 24 mm^2 at N5) and reads it at +SRAM latency, which is exactly what a GPU's L2 does with it. Layer 5 taxes the DRAM-only chip (the one without SRAM); +it does not tax the SRAM chip. + +Dataset growth does not touch the recompute attacker: the attacker never holds the dataset. It taxes the +partial-store attacker (O-1.6, the time-memory curve, not drawn) and the honest card. + +Year by year under the schedule as it stands (flat 256 MiB), the mirror's area at the best node available that +year. Node years are approximate; the density trend from 2018 to 2025 is 37.0 to 57.1 Mbit/mm^2 raw, 1.54x in +7 years, about 6% per year, and it came in one step (N2); the extrapolation past 2026 assumes that average holds +(approximate, and optimistic for the attacker: A16 and A14 have no disclosed SRAM cell yet). + +| Year | Calendar (approximate) | Dataset, GiB | Cache (spec) | Best node, raw Mbit/mm^2 | Mirror of the cache, mm^2 | With a 96 MiB hot table, mm^2 | Mirror as a share of a 750 mm^2 die | +|---|---|---|---|---|---|---|---| +| 0 | 2027 | 2.0 | 256 MiB | N2, 57.1 (cited) | 54 | 74 | 7% | +| 1 | 2028 | 2.5 | 256 MiB | N2 or A16, 57 to 61 | 50 to 54 | 69 to 74 | 7% | +| 2 | 2029 | 3.0 | 256 MiB | about 64 (trend) | 48 | 66 | 6% | +| 3 | 2030 | 3.5 | 256 MiB | about 68 | 45 | 62 | 6% | +| 4 | 2031 | 4.0 | 256 MiB | about 72 | 43 | 59 | 6% | +| 5 | 2032 | 4.5 | 256 MiB | about 76 | 40 | 55 | 5% | +| 6 | 2033 | 5.0 | 256 MiB | about 81 | 38 | 52 | 5% | +| 7 | 2034 | 5.5 | 256 MiB | about 86 | 36 | 49 | 5% | +| 8 | 2035 | 6.0 | 256 MiB | about 91 | 34 | 46 | 5% | +| 9 | 2036 | 6.5 | 256 MiB | about 97 | 32 | 44 | 4% | +| 10 | 2037 | 7.0 | 256 MiB | about 102 | 30 | 41 | 4% | + +Reading. A flat cache's mirror shrinks from 7% to 4% of a large die over the decade, and a 5090-class consumer GPU +already carries 96 MB of L2 on one die with the full GB202 at 128 MB; at the 2020 to 2025 pace of GPU L2 growth +(6 MB, 72 MB, 96 MB on the three NVIDIA flagships in the whitepaper table) a consumer GPU could hold 256 MiB on die +within the decade. The spec's own rule for the cache ("must exceed the largest on-chip cache of any card that +mines") would then be broken by a flat cache. That is the real reason to grow it: not to price a chip out (section +5 shows the SRAM cannot do that) but to keep the cache out of every GPU's own cache, so the honest hash stays +DRAM-latency-bound and the recompute route stays a route only a custom chip can take. + +## 7. Answer to the layer 6 question, and the options + +Does the flat 256 MiB cache keep the SRAM mirror unaffordable through year 10? No. It is affordable at year 0 +(tens of dollars of silicon per die, section 5) and gets cheaper. What keeps the recompute attacker near 1x is +M16's integer arithmetic and the mixer-cost lever (4x the mixer cost puts the equal-silicon gain at 0.36x, bounded +by the CPU verify gate), not the cache size. The cache size does one other job, keeping the cache larger than any +GPU's L2, and that job needs growth. + +Options for the cache rule, with the honest costs each implies. Verifier fill time is 0.2 s per 256 MiB on one core +(spec 1.12: "a 0.2 s CPU cache fill", from the measured 175 to 190 ms of section 1.8.3), scaled linearly; the +verifier holds the whole cache (section 1.11), so its memory is the cache size plus the program and the interpreter. +GPU fill: 0.67 ms per 256 MiB on the 5090 (section 1.8.3), linear. The GPU dataset build (13.4 ms per 1 GiB on the +5090, section 1.8.3) depends on the dataset size, not the cache size; a larger cache spreads the build's 8 dependent +reads per item over more memory, which on a GPU means more of them miss L2 and the build slows by some factor +between 1x and the L2-to-DRAM latency ratio, which is a measurement to take (approximate; owed). Mirror area is at N2 +(cited density), the node of the first years; at the trend's year-10 density divide by about 1.8. + +| Option | Rule | Cache at year 0 / 4 / 10 | Mirror at N2, year 0 / 4 / 10 (mm^2) | Dies at year 10 (830 mm^2 reticle) | Verifier fill, one core, year 0 / 10 | Verifier memory, year 10 | GPU cache fill (5090), year 10 | Keeps the cache above a 96 MB L2 at year 10 | Keeps it above a 256 MB L2 | +|---|---|---|---|---|---|---|---|---|---| +| A, as specified | flat 256 MiB | 256 / 256 / 256 MiB | 54 / 54 / 54 | 1 | 0.2 / 0.2 s | 256 MiB | 0.7 ms | yes, 2.7x | no | +| B | cache = dataset / 8 (today's ratio) | 256 / 512 / 896 MiB | 54 / 107 / 188 | 1 | 0.2 / 0.7 s | 896 MiB | 2.3 ms | yes, 9.3x | yes, 3.5x | +| C | cache doubles when the dataset doubles (the dataset's own clock: year 4, then year 12) | 256 / 512 / 512 MiB | 54 / 107 / 107 | 1 | 0.2 / 0.4 s | 512 MiB | 1.3 ms | yes, 5.3x | yes, 2x | +| D | cache = dataset / 4 | 512 / 1,024 / 1,792 MiB | 107 / 215 / 376 | 1 | 0.4 / 1.4 s | 1.75 GiB | 4.7 ms | yes | yes, 7x | +| E, one reticle | cache sized so the mirror exceeds one reticle at the node of the day: 4 GiB at N2 (section 4), growing with density | 4 GiB / about 4.5 / about 7 GiB | 859 / 860 / 860 (by construction) | 2 | 3.2 / 5.6 s | 7 GiB | 11 / 19 ms | yes | yes | + +Where the working set enters (coordinator's budget: 1 GiB table + hot table + scratch for every resident warp + +buffers under 6 GB on an 8 GB card): the cache is not in the miner's working set at hash time (the dataset is built +from it once a day and the cache can be dropped or kept), so options A to D do not move that budget; the dataset's own +growth does (2 GiB at genesis, 4 GiB at year 4, 7 GiB at year 10, which is past an 8 GB card at about year 8 on its +own). Option E's 4 GiB cache would have to be built on the card and dropped, which is fine for a 16 GB card and tight +on an 8 GB one at build time (4 GiB cache + 2 GiB dataset + hot table). The per-warp scratch at 170 SMs x 64 warps +(approximate, readwidth) is 340 MB at 32 KB and 1.36 GB at 128 KB per warp; with the 1 GiB table, a 96 MB hot table +and buffers that is 1.5 to 2.5 GB at the prototype dataset size, 2.5 to 3.5 GB at the 2 GiB genesis size, inside +6 GB either way. + +Recommendation. Option C (the cache doubles when the dataset doubles) is the one that keeps the spec's own rule true +with the smallest verifier cost: it ties the cache to a clock the spec already has, keeps `AND MASK` (a power of two +every step, which is the 1.13.3 option (b) argument again), costs the verifier 0.4 s and 512 MiB at year 4 and nothing +more until year 12, and keeps the cache 2x above a 256 MB GPU L2 if one appears. It does not price a chip out; nothing +about cache size does (section 5). The lever that does is the mixer cost multiplier of M16, which is the gate 1 +decision to take beside this one. Option B is the same idea in a smooth form and costs the verifier 0.7 s at year 10. +Option E is the only one that makes the mirror a multi-die part and it costs every verifier 3.2 s and 4 GiB at +genesis, which fails the spirit of the 10 ms verify gate (the fill is once a day, but a light node joining pays it on +every day it syncs across). + +Decision for Josh, at gate 1: A, B, C, D or E above, together with M16's mixer multiplier. Nothing here changes a +vector today: the cache size is a prototype value of spec 1.16 and the growth rule would be a new sentence in 1.13.3. + +## 8. What is cited, what is approximate, what is owed + +| Item | Status | +|---|---| +| Bit cells for N7, N5, N3B, N3E, N2, Intel 18A | cited (section 3) | +| Samsung SF2 or SF3 bit cell | not found; left out | +| Array efficiency 0.70 | WikiChip's convention, bracketed by two ISSCC 2025 macros (67 to 80%) | +| Wafer prices | approximate, supply-chain reporting, cited | +| D0 = 0.1 per cm^2, Poisson yield | assumption, stated | +| Node years and the 6% per year density trend past 2026 | approximate, extrapolated from cited 2018 to 2025 points | +| GPU L2 sizes | cited (NVIDIA whitepaper); AMD Infinity Cache approximate | +| Recompute attacker arithmetic | M16, which is itself arithmetic on measured rates, not a chip measurement | +| Dataset-build slowdown at a larger cache on a GPU | owed, a measurement (5090 at a 512 MiB and 1 GiB cache) | +| The on-die emulation of M16 (inline kernel with a 64 MiB cache inside the 5090's L2) | still a PC job (M16) | + +## 9. The arithmetic + +``` +MiB = 2^20; bits = cache_MiB * MiB * 8 +raw_Mbit_per_mm2 = 1 / cell_um2 (1e6 cells per mm2 per um2 of cell) +area_mm2 = bits / (raw * 0.70 * 1e6) +dies_per_wafer = pi * 150^2 / area - pi * 300 / sqrt(2 * area) +yield = exp(-area_mm2 * 0.001) (D0 = 0.1 per cm2) +cost_per_good_die = wafer_price / (dies * yield) +reticle_GiB = 830 * raw * 0.70 * 1e6 / 8 / 2^30 +``` +Run on 5 October 2026 with Python 3 on the M5 Max; the printed tables are the ones above, rounded. diff --git a/proto-cuda/dot4-probe.cu b/proto-cuda/dot4-probe.cu new file mode 100644 index 000000000..05b66fb68 --- /dev/null +++ b/proto-cuda/dot4-probe.cu @@ -0,0 +1,115 @@ +// dot4-probe (CUDA): dp4a-class throughput on NVIDIA, standalone (no pack, no lottery kernel). +// Counter ASIC 2.0 layer 7 (docs/analysis/int8-matrix-family.md), 5 October 2026. PC job; not run on the Mac. +// +// Three dependent chains, same shape as the OpenCL --memprobe ALU chain (proto-opencl/host.c, probe_alu) and the +// Metal probe (proto-metal/dot4-probe.swift): 1,048,576 lanes x 4,096 steps, best of 3, event time. +// alu x = x * K + rotl(y, 7); y = (y ^ x) + s the card's integer baseline, 5 ops per step counted +// dot4i acc = __dp4a(x, y, acc) (PTX dp4a.s32.s32, sm_61+) one hardware dot4 per step per lane +// dot4e the scalar emulation of the same (4 sign-extended byte products summed, wrapping int32) +// The emulation and the intrinsic must agree bit for bit with the CPU reference (checked on two lanes per run). +// +// Build (Windows, CUDA Toolkit): nvcc -O2 -arch=sm_120 -o dot4-probe-cuda.exe dot4-probe.cu +// Build (Linux): nvcc -O2 -arch=sm_120 -o dot4-probe-cuda dot4-probe.cu +// Run: dot4-probe-cuda [--lanes N] [--steps N] [--reps N] [--device N] +#include +#include +#include +#include +#include + +__host__ __device__ inline uint32_t pm_mix(uint32_t x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; } +__host__ __device__ inline uint32_t rotl32(uint32_t v, uint32_t n) { return (v << n) | (v >> (32u - n)); } +__host__ __device__ inline int32_t dot4_emul(uint32_t a, uint32_t b, int32_t acc) { + int32_t r = acc; + for (int i = 0; i < 4; ++i) { + int32_t ba = (int32_t)(int8_t)((a >> (8 * i)) & 0xffu); + int32_t bb = (int32_t)(int8_t)((b >> (8 * i)) & 0xffu); + r = (int32_t)((uint32_t)r + (uint32_t)(ba * bb)); + } + return r; +} + +__global__ void probe_alu(uint32_t steps, uint32_t seed, uint32_t* out) { + uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u; + for (uint32_t s = 0; s < steps; ++s) { x = x * 0x9E3779B1u + rotl32(y, 7u); y = (y ^ x) + s; } + out[g] = x ^ y; +} +__global__ void probe_dot4i(uint32_t steps, uint32_t seed, uint32_t* out) { + uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u; + int32_t acc = (int32_t)pm_mix(x); + for (uint32_t s = 0; s < steps; ++s) { + acc = __dp4a((int)x, (int)y, acc); + x = x * 0x9E3779B1u + (uint32_t)acc; + y = rotl32(y, 7u) ^ ((uint32_t)acc + s); + } + out[g] = (uint32_t)acc ^ x ^ y; +} +__global__ void probe_dot4e(uint32_t steps, uint32_t seed, uint32_t* out) { + uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; + uint32_t x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u; + int32_t acc = (int32_t)pm_mix(x); + for (uint32_t s = 0; s < steps; ++s) { + acc = dot4_emul(x, y, acc); + x = x * 0x9E3779B1u + (uint32_t)acc; + y = rotl32(y, 7u) ^ ((uint32_t)acc + s); + } + out[g] = (uint32_t)acc ^ x ^ y; +} + +static uint32_t lane_ref(const char* name, uint32_t g, uint32_t seed, uint32_t steps) { + uint32_t x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u; + if (strcmp(name, "alu") == 0) { + for (uint32_t s = 0; s < steps; ++s) { x = x * 0x9E3779B1u + rotl32(y, 7u); y = (y ^ x) + s; } + return x ^ y; + } + int32_t acc = (int32_t)pm_mix(x); + for (uint32_t s = 0; s < steps; ++s) { + acc = dot4_emul(x, y, acc); + x = x * 0x9E3779B1u + (uint32_t)acc; + y = rotl32(y, 7u) ^ ((uint32_t)acc + s); + } + return (uint32_t)acc ^ x ^ y; +} + +#define CK(x) do { cudaError_t e = (x); if (e != cudaSuccess) { printf("CUDA error %s at %s:%d\n", cudaGetErrorString(e), __FILE__, __LINE__); return 1; } } while (0) + +int main(int argc, char** argv) { + uint32_t lanes = 1u << 20, steps = 4096u; int reps = 3, device = 0; + for (int i = 1; i < argc; ++i) { + if (!strcmp(argv[i], "--lanes") && i + 1 < argc) lanes = (uint32_t)strtoul(argv[++i], 0, 10); + else if (!strcmp(argv[i], "--steps") && i + 1 < argc) steps = (uint32_t)strtoul(argv[++i], 0, 10); + else if (!strcmp(argv[i], "--reps") && i + 1 < argc) reps = atoi(argv[++i]); + else if (!strcmp(argv[i], "--device") && i + 1 < argc) device = atoi(argv[++i]); + else { printf("unknown argument %s\n", argv[i]); return 2; } + } + CK(cudaSetDevice(device)); + cudaDeviceProp p; CK(cudaGetDeviceProperties(&p, device)); + printf("dot4-probe (CUDA) on %s, sm_%d%d, %d SMs, %d MHz, lanes %u, steps %u, best of %d, event time\n", p.name, p.major, p.minor, p.multiProcessorCount, p.clockRate / 1000, lanes, steps, reps); + uint32_t* d_out; CK(cudaMalloc(&d_out, (size_t)lanes * 4)); + uint32_t* h_out = (uint32_t*)malloc((size_t)lanes * 4); + cudaEvent_t e0, e1; CK(cudaEventCreate(&e0)); CK(cudaEventCreate(&e1)); + printf("| kernel | lanes | steps | best ms | G steps/s (= G dot4/s for dot4 rows) | ns per dependent step | lanes 0 and last ok |\n|---|---|---|---|---|---|---|\n"); + const char* names[3] = { "alu", "dot4i", "dot4e" }; + for (int k = 0; k < 3; ++k) { + float best = 1e30f; int ok = 1; + for (int r = 0; r < reps; ++r) { + uint32_t seed = 0x2468aceu + (uint32_t)r * 0x9E3779B9u; + CK(cudaEventRecord(e0)); + if (k == 0) probe_alu<<>>(steps, seed, d_out); + else if (k == 1) probe_dot4i<<>>(steps, seed, d_out); + else probe_dot4e<<>>(steps, seed, d_out); + CK(cudaEventRecord(e1)); CK(cudaEventSynchronize(e1)); CK(cudaGetLastError()); + float ms = 0; CK(cudaEventElapsedTime(&ms, e0, e1)); if (ms < best) best = ms; + CK(cudaMemcpy(h_out, d_out, (size_t)lanes * 4, cudaMemcpyDeviceToHost)); + uint32_t gs[2] = { 0u, lanes - 1u }; + for (int j = 0; j < 2; ++j) { uint32_t want = lane_ref(names[k], gs[j], seed, steps); if (h_out[gs[j]] != want) { ok = 0; printf("MISMATCH %s lane %u: gpu %08x cpu %08x\n", names[k], gs[j], h_out[gs[j]], want); } } + } + double sps = (double)lanes * (double)steps / (best / 1000.0); + printf("| %s | %u | %u | %.3f | %.2f | %.3f | %s |\n", names[k], lanes, steps, best, sps / 1e9, best * 1e6 / (double)steps, ok ? "yes" : "NO"); + printf("RESULT DOT4 vendor=nvidia device=\"%s\" kernel=%s lanes=%u steps=%u best_ms=%.3f gsteps_per_s=%.2f ok=%d\n", p.name, names[k], lanes, steps, best, sps / 1e9, ok); + } + printf("dot4-probe: done\n"); + return 0; +} diff --git a/proto-metal/dot4-probe.swift b/proto-metal/dot4-probe.swift new file mode 100644 index 000000000..2fa1b4024 --- /dev/null +++ b/proto-metal/dot4-probe.swift @@ -0,0 +1,165 @@ +// dot4-probe: dp4a-class throughput on Apple silicon, standalone (no pack, no lottery kernel). +// Counter ASIC 2.0 layer 7 (docs/analysis/int8-matrix-family.md), 5 October 2026. +// +// Metal has no dp4a intrinsic and no integer simdgroup_matrix (MSL 4.1 section 2.4 lists half, bfloat and float only; +// the Metal 4 tensor op matmul2d does carry char x char -> int, MSL 4.1 table 7.3, measured separately when it is). +// So the per-lane dot4 here is the scalar emulation a conforming Apple miner would run: four sign-extended bytes of +// each operand multiplied and summed into a wrapping int32 accumulator, exactly the PTX dp4a semantics +// (PTX ISA 9.4 section 9.7.1.24: d = c; d += Va[i] * Vb[i] for i in 0..3, bytes sign- or zero-extended). +// +// Two kernels, same shape as the OpenCL --memprobe ALU chain (proto-opencl/host.c, probe_alu: 1,048,576 lanes x 4,096 +// steps, best of 3): +// alu x = x * K + rotate(y, 7); y = (y ^ x) + s the card's integer baseline, 5 ops per step counted +// dot4 acc = dot4(x, y, acc); x = x * K + acc; y = rotate(y, 7) ^ acc one dependent dot4 per step per lane +// Rates: G steps/s per lane-step, so G dot4/s for the second kernel. Timing is the command buffer's GPU start to end. +// +// Build: swiftc -O -o dot4-probe dot4-probe.swift -framework Metal +// Run: ./dot4-probe [--lanes N] [--steps N] [--reps N] [--signed|--unsigned] +import Foundation +import Metal + +let source = """ +#include +using namespace metal; + +inline uint pm_mix(uint x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; } + +// dp4a, signed bytes, wrapping int32 accumulate: the exact PTX dp4a.s32.s32 semantics. +inline int dot4_s(uint a, uint b, int acc) { + int4 va = int4(as_type(a)); + int4 vb = int4(as_type(b)); + return acc + va.x * vb.x + va.y * vb.y + va.z * vb.z + va.w * vb.w; +} +// dp4a, unsigned bytes, wrapping uint32 accumulate: dp4a.u32.u32. +inline uint dot4_u(uint a, uint b, uint acc) { + uint4 va = uint4(as_type(a)); + uint4 vb = uint4(as_type(b)); + return acc + va.x * vb.x + va.y * vb.y + va.z * vb.z + va.w * vb.w; +} + +kernel void probe_alu(constant uint& steps [[buffer(0)]], constant uint& seed [[buffer(1)]], + device uint* out [[buffer(2)]], uint g [[thread_position_in_grid]]) { + uint x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u; + for (uint s = 0u; s < steps; ++s) { x = x * 0x9E3779B1u + rotate(y, 7u); y = (y ^ x) + s; } + out[g] = x ^ y; +} + +kernel void probe_dot4s(constant uint& steps [[buffer(0)]], constant uint& seed [[buffer(1)]], + device uint* out [[buffer(2)]], uint g [[thread_position_in_grid]]) { + uint x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u; + int acc = int(pm_mix(x)); + for (uint s = 0u; s < steps; ++s) { + acc = dot4_s(x, y, acc); + x = x * 0x9E3779B1u + uint(acc); + y = rotate(y, 7u) ^ (uint(acc) + s); + } + out[g] = uint(acc) ^ x ^ y; +} + +kernel void probe_dot4u(constant uint& steps [[buffer(0)]], constant uint& seed [[buffer(1)]], + device uint* out [[buffer(2)]], uint g [[thread_position_in_grid]]) { + uint x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u; + uint acc = pm_mix(x); + for (uint s = 0u; s < steps; ++s) { + acc = dot4_u(x, y, acc); + x = x * 0x9E3779B1u + acc; + y = rotate(y, 7u) ^ (acc + s); + } + out[g] = acc ^ x ^ y; +} +""" + +// CPU reference of the dot4 chain for one lane, to check the kernel is the arithmetic it claims (bit-exact). +func pmMix(_ v: UInt32) -> UInt32 { + var x = v + x ^= x >> 16; x = x &* 0x7feb352d; x ^= x >> 15; x = x &* 0x846ca68b; x ^= x >> 16 + return x +} +func dot4sRef(_ a: UInt32, _ b: UInt32, _ acc: Int32) -> Int32 { + var r = acc + for i in 0..<4 { + let ba = Int32(Int8(truncatingIfNeeded: a >> (8 * UInt32(i)))) + let bb = Int32(Int8(truncatingIfNeeded: b >> (8 * UInt32(i)))) + r = r &+ ba &* bb + } + return r +} +func dot4uRef(_ a: UInt32, _ b: UInt32, _ acc: UInt32) -> UInt32 { + var r = acc + for i in 0..<4 { + let ba = UInt32(UInt8(truncatingIfNeeded: a >> (8 * UInt32(i)))) + let bb = UInt32(UInt8(truncatingIfNeeded: b >> (8 * UInt32(i)))) + r = r &+ ba &* bb + } + return r +} +func rotl(_ v: UInt32, _ n: UInt32) -> UInt32 { (v << n) | (v >> (32 - n)) } +func laneRef(kernel: String, g: UInt32, seed: UInt32, steps: UInt32) -> UInt32 { + var x = pmMix(g ^ seed), y = x ^ 0x5bd1e995 + switch kernel { + case "probe_alu": + for s in 0../include -o dot4-probe-cl.exe dot4-probe.c (OpenCL.dll loaded at run time) + * Run: dot4-probe-cl [--list] [--device N] [--lanes N] [--steps N] [--reps N] + */ +#define CL_TARGET_OPENCL_VERSION 120 +#ifdef __APPLE__ +#include +#else +#include +#endif +#ifdef IGNEUM_CL_DYNAMIC +#include "cl_dynamic.h" +#endif +#include +#include +#include +#include + +static const char* COMMON = + "static inline uint pm_mix(uint x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }\n" + "#define CHAIN_HEAD uint g = (uint)get_global_id(0); uint x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u; int acc = (int)pm_mix(x);\n" + "#define CHAIN_TAIL x = x * 0x9E3779B1u + (uint)acc; y = rotate(y, 7u) ^ ((uint)acc + s);\n" + "#define CHAIN_OUT out[g] = (uint)acc ^ x ^ y;\n"; + +static const char* K_ALU = + "__kernel void probe(uint steps, uint seed, __global uint* out) {\n" + " uint g = (uint)get_global_id(0); uint x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u;\n" + " for (uint s = 0u; s < steps; ++s) { x = x * 0x9E3779B1u + rotate(y, 7u); y = (y ^ x) + s; }\n" + " out[g] = x ^ y;\n" + "}\n"; +static const char* K_EMUL = + "static inline int dot4e(uint a, uint b, int acc) {\n" + " int4 va = convert_int4(as_char4(a)); int4 vb = convert_int4(as_char4(b));\n" + " return acc + va.x * vb.x + va.y * vb.y + va.z * vb.z + va.w * vb.w;\n" + "}\n" + "__kernel void probe(uint steps, uint seed, __global uint* out) {\n" + " CHAIN_HEAD\n" + " for (uint s = 0u; s < steps; ++s) { acc = dot4e(x, y, acc); CHAIN_TAIL }\n" + " CHAIN_OUT\n" + "}\n"; +static const char* K_AMD = + "__kernel void probe(uint steps, uint seed, __global uint* out) {\n" + " CHAIN_HEAD\n" + " for (uint s = 0u; s < steps; ++s) { acc = __builtin_amdgcn_sudot4(true, (int)x, true, (int)y, acc, false); CHAIN_TAIL }\n" + " CHAIN_OUT\n" + "}\n"; +static const char* K_KHR = + "#pragma OPENCL EXTENSION cl_khr_integer_dot_product : enable\n" + "__kernel void probe(uint steps, uint seed, __global uint* out) {\n" + " CHAIN_HEAD\n" + " for (uint s = 0u; s < steps; ++s) { acc = acc + dot(as_char4(x), as_char4(y)); CHAIN_TAIL }\n" + " CHAIN_OUT\n" + "}\n"; +static const char* K_NV = + "static inline int dot4nv(uint a, uint b, int acc) { int d; asm(\"dp4a.s32.s32 %0, %1, %2, %3;\" : \"=r\"(d) : \"r\"(a), \"r\"(b), \"r\"(acc)); return d; }\n" + "__kernel void probe(uint steps, uint seed, __global uint* out) {\n" + " CHAIN_HEAD\n" + " for (uint s = 0u; s < steps; ++s) { acc = dot4nv(x, y, acc); CHAIN_TAIL }\n" + " CHAIN_OUT\n" + "}\n"; + +static uint32_t pm_mix(uint32_t x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; } +static uint32_t rotl32(uint32_t v, uint32_t n) { return (v << n) | (v >> (32u - n)); } +static int32_t dot4_ref(uint32_t a, uint32_t b, int32_t acc) { + int32_t r = acc; int i; + for (i = 0; i < 4; ++i) { int32_t ba = (int8_t)((a >> (8 * i)) & 0xffu), bb = (int8_t)((b >> (8 * i)) & 0xffu); r = (int32_t)((uint32_t)r + (uint32_t)(ba * bb)); } + return r; +} +static uint32_t lane_ref(int alu, uint32_t g, uint32_t seed, uint32_t steps) { + uint32_t x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u, s; int32_t acc; + if (alu) { for (s = 0; s < steps; ++s) { x = x * 0x9E3779B1u + rotl32(y, 7u); y = (y ^ x) + s; } return x ^ y; } + acc = (int32_t)pm_mix(x); + for (s = 0; s < steps; ++s) { acc = dot4_ref(x, y, acc); x = x * 0x9E3779B1u + (uint32_t)acc; y = rotl32(y, 7u) ^ ((uint32_t)acc + s); } + return (uint32_t)acc ^ x ^ y; +} + +typedef struct { cl_platform_id p; cl_device_id d; char pname[128], dname[128], driver[64], ver[64]; } Dev; +static Dev devs[32]; static int ndevs = 0; +static void enumerate(void) { + cl_platform_id ps[8]; cl_uint np = 0, i; + if (clGetPlatformIDs(8, ps, &np) != CL_SUCCESS) return; + for (i = 0; i < np; ++i) { + cl_device_id ds[8]; cl_uint nd = 0, j; + if (clGetDeviceIDs(ps[i], CL_DEVICE_TYPE_GPU, 8, ds, &nd) != CL_SUCCESS) continue; + for (j = 0; j < nd && ndevs < 32; ++j) { + Dev* v = &devs[ndevs++]; v->p = ps[i]; v->d = ds[j]; + clGetPlatformInfo(ps[i], CL_PLATFORM_NAME, sizeof v->pname, v->pname, NULL); + clGetDeviceInfo(ds[j], CL_DEVICE_NAME, sizeof v->dname, v->dname, NULL); + clGetDeviceInfo(ds[j], CL_DRIVER_VERSION, sizeof v->driver, v->driver, NULL); + clGetDeviceInfo(ds[j], CL_DEVICE_VERSION, sizeof v->ver, v->ver, NULL); + } + } +} + +static int run_variant(cl_context ctx, cl_command_queue q, cl_device_id dev, const char* name, const char* body, int alu, cl_uint lanes, cl_uint steps, int reps, cl_mem out, uint32_t* host, const char* vendor, const char* dname) { + const char* srcs[2] = { COMMON, body }; cl_int err; cl_program prog; cl_kernel k; int r, ok = 1; double best = -1; + prog = clCreateProgramWithSource(ctx, 2, srcs, NULL, &err); + if (err != CL_SUCCESS) { printf("| %s | build failed | clCreateProgramWithSource %d | | | | |\n", name, (int)err); return 0; } + err = clBuildProgram(prog, 1, &dev, "-cl-std=CL1.2", NULL, NULL); + if (err != CL_SUCCESS) { + size_t n = 0; char* log; char* nl; + clGetProgramBuildInfo(prog, dev, CL_PROGRAM_BUILD_LOG, 0, NULL, &n); log = (char*)calloc(n + 1, 1); + if (n) clGetProgramBuildInfo(prog, dev, CL_PROGRAM_BUILD_LOG, n, log, NULL); + while (*log == '\n' || *log == '\r') ++log; + nl = strpbrk(log, "\r\n"); if (nl) *nl = 0; + printf("| %s | build failed | %.160s | | | | |\n", name, log); + printf("RESULT DOT4 vendor=%s device=\"%s\" kernel=%s build=failed\n", vendor, dname, name); + clReleaseProgram(prog); return 0; + } + k = clCreateKernel(prog, "probe", &err); + if (err != CL_SUCCESS) { printf("| %s | build failed | clCreateKernel %d | | | | |\n", name, (int)err); clReleaseProgram(prog); return 0; } + for (r = 0; r < reps; ++r) { + cl_uint seed = 0x2468aceu + (cl_uint)r * 0x9E3779B9u; size_t global = lanes, local = 256; cl_event ev; cl_ulong t0, t1; double ms; uint32_t gs[2]; int j; + clSetKernelArg(k, 0, sizeof(cl_uint), &steps); clSetKernelArg(k, 1, sizeof(cl_uint), &seed); clSetKernelArg(k, 2, sizeof(cl_mem), &out); + err = clEnqueueNDRangeKernel(q, k, 1, NULL, &global, &local, 0, NULL, &ev); + if (err != CL_SUCCESS) { printf("| %s | launch failed | %d | | | | |\n", name, (int)err); clReleaseKernel(k); clReleaseProgram(prog); return 0; } + clWaitForEvents(1, &ev); + clGetEventProfilingInfo(ev, CL_PROFILING_COMMAND_START, sizeof t0, &t0, NULL); clGetEventProfilingInfo(ev, CL_PROFILING_COMMAND_END, sizeof t1, &t1, NULL); + ms = (double)(t1 - t0) / 1e6; clReleaseEvent(ev); + if (best < 0 || ms < best) best = ms; + clEnqueueReadBuffer(q, out, CL_TRUE, 0, (size_t)lanes * 4, host, 0, NULL, NULL); + gs[0] = 0; gs[1] = lanes - 1; + for (j = 0; j < 2; ++j) { uint32_t want = lane_ref(alu, gs[j], seed, steps); if (host[gs[j]] != want) { ok = 0; printf("MISMATCH %s lane %u: gpu %08x cpu %08x\n", name, gs[j], host[gs[j]], want); } } + } + { + double sps = (double)lanes * (double)steps / (best / 1000.0); + printf("| %s | %u | %u | %.3f | %.2f | %.3f | %s |\n", name, lanes, steps, best, sps / 1e9, best * 1e6 / (double)steps, ok ? "yes" : "NO"); + printf("RESULT DOT4 vendor=%s device=\"%s\" kernel=%s lanes=%u steps=%u best_ms=%.3f gsteps_per_s=%.2f ok=%d\n", vendor, dname, name, lanes, steps, best, sps / 1e9, ok); + } + clReleaseKernel(k); clReleaseProgram(prog); return 1; +} + +int main(int argc, char** argv) { + cl_uint lanes = 1u << 20, steps = 4096u; int reps = 3, device = 0, list = 0, i; Dev* v; cl_int err; cl_context ctx; cl_command_queue q; cl_mem out; uint32_t* host; const char* vendor; + for (i = 1; i < argc; ++i) { + if (!strcmp(argv[i], "--list")) list = 1; + else if (!strcmp(argv[i], "--device") && i + 1 < argc) device = atoi(argv[++i]); + else if (!strcmp(argv[i], "--lanes") && i + 1 < argc) lanes = (cl_uint)strtoul(argv[++i], 0, 10); + else if (!strcmp(argv[i], "--steps") && i + 1 < argc) steps = (cl_uint)strtoul(argv[++i], 0, 10); + else if (!strcmp(argv[i], "--reps") && i + 1 < argc) reps = atoi(argv[++i]); + else { printf("unknown argument %s\n", argv[i]); return 2; } + } +#ifdef IGNEUM_CL_DYNAMIC + if (!ig_cl_load()) { printf("%s\n", ig_cl_error); return 1; } +#endif + enumerate(); + if (list || ndevs == 0) { for (i = 0; i < ndevs; ++i) printf("[%d] %s | %s | driver %s | %s\n", i, devs[i].dname, devs[i].pname, devs[i].driver, devs[i].ver); if (ndevs == 0) printf("no OpenCL GPU devices\n"); return ndevs ? 0 : 1; } + if (device < 0 || device >= ndevs) { printf("no device %d (have %d)\n", device, ndevs); return 2; } + v = &devs[device]; + vendor = strstr(v->pname, "NVIDIA") ? "nvidia" : (strstr(v->pname, "AMD") ? "amd" : (strstr(v->pname, "Apple") ? "apple" : "other")); + ctx = clCreateContext(NULL, 1, &v->d, NULL, NULL, &err); if (err != CL_SUCCESS) { printf("clCreateContext %d\n", (int)err); return 1; } + q = clCreateCommandQueue(ctx, v->d, CL_QUEUE_PROFILING_ENABLE, &err); if (err != CL_SUCCESS) { printf("clCreateCommandQueue %d\n", (int)err); return 1; } + out = clCreateBuffer(ctx, CL_MEM_READ_WRITE, (size_t)lanes * 4, NULL, &err); if (err != CL_SUCCESS) { printf("clCreateBuffer %d\n", (int)err); return 1; } + host = (uint32_t*)malloc((size_t)lanes * 4); + printf("dot4-probe (OpenCL) on [%d] %s | %s | driver %s | %s, lanes %u, steps %u, best of %d, device event time\n", device, v->dname, v->pname, v->driver, v->ver, lanes, steps, reps); + { + char ext[8192]; ext[0] = 0; clGetDeviceInfo(v->d, CL_DEVICE_EXTENSIONS, sizeof ext, ext, NULL); + printf("cl_khr_integer_dot_product listed: %s\n", strstr(ext, "cl_khr_integer_dot_product") ? "yes" : "no"); + } + printf("| kernel | lanes | steps | best ms | G steps/s (= G dot4/s for dot4 rows) | ns per dependent step | lanes 0 and last ok |\n|---|---|---|---|---|---|---|\n"); + run_variant(ctx, q, v->d, "alu", K_ALU, 1, lanes, steps, reps, out, host, vendor, v->dname); + run_variant(ctx, q, v->d, "dot4e", K_EMUL, 0, lanes, steps, reps, out, host, vendor, v->dname); + run_variant(ctx, q, v->d, "dot4_khr", K_KHR, 0, lanes, steps, reps, out, host, vendor, v->dname); + if (!strcmp(vendor, "amd")) run_variant(ctx, q, v->d, "dot4_amd", K_AMD, 0, lanes, steps, reps, out, host, vendor, v->dname); + if (!strcmp(vendor, "nvidia")) run_variant(ctx, q, v->d, "dot4_nv", K_NV, 0, lanes, steps, reps, out, host, vendor, v->dname); + clReleaseMemObject(out); clReleaseCommandQueue(q); clReleaseContext(ctx); free(host); + printf("dot4-probe: done\n"); + return 0; +} From a9e002c47d646e7d9bdc2a6131970155336e8e67 Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:09:58 +0100 Subject: [PATCH 005/311] read-width: packfile accepts a string-seed pack (any even-length hex epoch seed; the seed-word re-derivation still checks it); bench-only playbooks for the second PC round Round 1 (run-readwidth-{5090,9070}-20261005) delivered the probes and refused every pack: pf_load demanded the chain's 32-byte epoch seed and the experiment packs carry igneum-pow --seed strings. Co-Authored-By: Claude Fable 5.1 --- proto-cuda/nvrtc/packfile.h | 10 ++-- relay/playbooks/readwidth-5090-bench.ps1 | 66 +++++++++++++++++++++++ relay/playbooks/readwidth-9070-bench.ps1 | 68 ++++++++++++++++++++++++ 3 files changed, 141 insertions(+), 3 deletions(-) create mode 100644 relay/playbooks/readwidth-5090-bench.ps1 create mode 100644 relay/playbooks/readwidth-9070-bench.ps1 diff --git a/proto-cuda/nvrtc/packfile.h b/proto-cuda/nvrtc/packfile.h index ba1719e8e..4f2315f23 100644 --- a/proto-cuda/nvrtc/packfile.h +++ b/proto-cuda/nvrtc/packfile.h @@ -283,13 +283,17 @@ static int pf_load(const char* dir, PfPack* pk, char* err, size_t cap) { strcpy(ehex, e2); strcpy(dhex, d2); } if (!ehex[0] || !dhex[0]) return pf_fail(err, cap, "no seeds: neither seeds.txt nor IGNEUM_SEED_BYTES_HEX / IGNEUM_DAY_BYTES_HEX in program.h (a pack from igneum-pow export --seed has no byte seeds)"); - if (strlen(ehex) != 64) return pf_fail(err, cap, "epoch seed is not 64 hex characters"); + // The chain's epoch seed is 32 bytes (64 hex characters). A pack exported from a seed STRING (igneum-pow export + // --seed , the read-width experiment's packs of 5 October 2026) carries the string's bytes instead; the + // seed-word re-derivation below checks either form, so any even-length hex seed is accepted here. The serve + // protocol still carries 64-hex seeds; a string-seed pack can only be benched (--bench, --bench-pack, --check). + if (strlen(ehex) < 2 || strlen(ehex) % 2 != 0) return pf_fail(err, cap, "epoch seed is not an even-length hex string"); strcpy(pk->epochHex, ehex); strcpy(pk->dayHex, dhex); // The seed words derived from the bytes must be the pack's own words: otherwise the pack and the seeds disagree { uint32_t w[8]; - if (!pf_unhex(ehex, bytes, 32, &blen) || blen != 32) return pf_fail(err, cap, "epoch seed hex is malformed"); - pf_seed_words_from_bytes(bytes, 32, w); + if (!pf_unhex(ehex, bytes, sizeof(bytes), &blen) || blen == 0) return pf_fail(err, cap, "epoch seed hex is malformed"); + pf_seed_words_from_bytes(bytes, blen, w); if (memcmp(w, pk->seedw, 32) != 0) return pf_fail(err, cap, "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT (wrong seeds.txt for this pack?)"); if (!pf_unhex(dhex, bytes, sizeof(bytes), &blen)) return pf_fail(err, cap, "day seed hex is malformed"); pf_seed_words_from_bytes(bytes, blen, w); diff --git a/relay/playbooks/readwidth-5090-bench.ps1 b/relay/playbooks/readwidth-5090-bench.ps1 new file mode 100644 index 000000000..2a2e5190e --- /dev/null +++ b/relay/playbooks/readwidth-5090-bench.ps1 @@ -0,0 +1,66 @@ +# Igneum run job (second round, bench only; the first round's packs were refused for a string seed): the read-width experiment on PC 2's RTX 5090 (machine 1ccfe586), 5 October 2026 (docs/plans/read-width.md). +# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining; this script switches +# off ONLY the NVIDIA card in the app through POST api/cards, waits for its worker to stop, runs the fetched +# igneum-worker-cuda.exe (--memprobe, then --bench on every pack of the fetched packs folder), and switches the card back +# on with the settings it had. Every result line starts with RESULT so `node tools/jobs.mjs ` shows them. +$ErrorActionPreference = 'Continue' +function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) } +$jobs = Split-Path $env:IGNEUM_JOB_DIR +$fetched = Join-Path $jobs 'fetch-readwidth-20261005c' +$exe = Join-Path $fetched 'igneum-worker-cuda.exe' +$packs = Join-Path $fetched 'packs-readwidth' +if (-not (Test-Path $exe)) { Write-Output "RESULT error worker missing at $exe (the fetch job runs first)"; exit 2 } +if (-not (Test-Path $packs)) { Write-Output "RESULT error packs missing at $packs"; exit 2 } +$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1 +if (-not $inst) { Write-Output 'RESULT error no installed igneum-worker-cuda.exe (the NVRTC DLLs come from there)'; exit 2 } +Get-ChildItem $inst -Filter 'nvrtc*.dll' | Copy-Item -Destination $fetched -Force +Write-Output "RESULT worker $exe sha256 $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower()) with $((Get-ChildItem $fetched -Filter 'nvrtc*.dll').Count) NVRTC DLL(s) from $inst" + +# the app: switch off the NVIDIA card only, remember its settings +$appDir = $env:IGNEUM_APP_DIR +if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' } +$urlFile = Join-Path $appDir 'app.url' +$url = $null +if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() } +$card = $null +if ($url) { + try { + $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10 + $card = $st.mining.cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1 + if (-not $card) { $card = $st.cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1 } + } catch { Say ("api/state: " + $_.Exception.Message) } +} +if ($card) { + Write-Output ("RESULT card " + $card.key + " enabled=" + $card.enabled + " identities=" + $card.identities + " power_pct=" + $card.power_pct + " state=" + $card.state) + $body = @{ cards = @(@{ key = $card.key; enabled = $false; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "card off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) } + $t = 0 + while ($t -lt 90) { + Start-Sleep -Seconds 5; $t += 5 + try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $c2 = $st.mining.cards | Where-Object { $_.key -eq $card.key }; if (-not $c2) { $c2 = $st.cards | Where-Object { $_.key -eq $card.key } }; if ($c2 -and $c2.state -eq 'off' -and $c2.pid -eq 0) { break } } catch { } + } + Write-Output ("RESULT card-off after " + $t + " s") + Start-Sleep -Seconds 5 +} else { Write-Output 'RESULT card none-found (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' } + +& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" } +# the probe ran in run-readwidth-5090-20261005 (bench-log, 5 October 2026); this second run is the bench only +foreach ($pk in @('w4', 'w16', 'w64', 'w64x4', 'mixA-0', 'mixA-1', 'mixA-2', 'mixA-3', 'mixA-4', 'mixA-5', 'mixB-0', 'mixB-1', 'mixB-2', 'mixB-3', 'mixB-4', 'mixB-5')) { + $d = Join-Path $packs $pk + Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)" + & $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT $_" } + & $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL' } | ForEach-Object { "RESULT $_" } +} +foreach ($pk in @('scr0k32', 'scr2k32', 'scr4k32', 'scr8k32', 'scr2k128', 'scr4k128', 'scr8k128')) { + $d = Join-Path $packs $pk + Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)" + & $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT $_" } + & $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 --warps 4096 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|variant 5' } | ForEach-Object { "RESULT $_" } +} +& nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" } + +if ($card) { + $body = @{ cards = @(@{ key = $card.key; enabled = [bool]$card.enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored enabled=" + $card.enabled) } catch { Write-Output ("RESULT error card restore: " + $_.Exception.Message) } +} +exit 0 diff --git a/relay/playbooks/readwidth-9070-bench.ps1 b/relay/playbooks/readwidth-9070-bench.ps1 new file mode 100644 index 000000000..7ce46fe37 --- /dev/null +++ b/relay/playbooks/readwidth-9070-bench.ps1 @@ -0,0 +1,68 @@ +# Igneum run job (second round, bench only; the first round's packs were refused for a string seed): the read-width experiment on PC 1's RX 9070 XT on the eGPU (machine ae432dc7), 5 October 2026 (docs/plans/read-width.md). +# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining; this script switches +# off ONLY the NVIDIA card in the app through POST api/cards, waits for its worker to stop, runs the fetched +# igneum-worker-opencl.exe (--memprobe, then --bench on every pack of the fetched packs folder), and switches the card back +# on with the settings it had. Every result line starts with RESULT so `node tools/jobs.mjs ` shows them. +$ErrorActionPreference = 'Continue' +function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) } +$jobs = Split-Path $env:IGNEUM_JOB_DIR +$fetched = Join-Path $jobs 'fetch-readwidth-20261005c' +$exe = Join-Path $fetched 'igneum-worker-opencl.exe' +$packs = Join-Path $fetched 'packs-readwidth' +if (-not (Test-Path $exe)) { Write-Output "RESULT error worker missing at $exe (the fetch job runs first)"; exit 2 } +if (-not (Test-Path $packs)) { Write-Output "RESULT error packs missing at $packs"; exit 2 } +Write-Output "RESULT worker $exe sha256 $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower())" +# the card's OpenCL device index on the current (3683.0) platform, from the worker's own list (the older platform's duplicate is marked dup) +$list = & $exe --list 2>&1 +$list | ForEach-Object { "RESULT list $_" } +$dev = $null +foreach ($l in $list) { if ($l -match '^\s*\[(\d+)\].*gfx1201' -and $l -notmatch 'dup') { $dev = [int]$Matches[1]; break } } +if ($null -eq $dev) { Write-Output 'RESULT error no gfx1201 device in --list'; exit 2 } +Write-Output "RESULT device $dev" + +# the app: switch off the NVIDIA card only, remember its settings +$appDir = $env:IGNEUM_APP_DIR +if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' } +$urlFile = Join-Path $appDir 'app.url' +$url = $null +if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() } +$card = $null +if ($url) { + try { + $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10 + $card = $st.mining.cards | Where-Object { ($_.vendor -eq 'amd' -and $_.key -match 'gfx1201') } | Select-Object -First 1 + if (-not $card) { $card = $st.cards | Where-Object { ($_.vendor -eq 'amd' -and $_.key -match 'gfx1201') } | Select-Object -First 1 } + } catch { Say ("api/state: " + $_.Exception.Message) } +} +if ($card) { + Write-Output ("RESULT card " + $card.key + " enabled=" + $card.enabled + " identities=" + $card.identities + " power_pct=" + $card.power_pct + " state=" + $card.state) + $body = @{ cards = @(@{ key = $card.key; enabled = $false; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "card off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) } + $t = 0 + while ($t -lt 90) { + Start-Sleep -Seconds 5; $t += 5 + try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $c2 = $st.mining.cards | Where-Object { $_.key -eq $card.key }; if (-not $c2) { $c2 = $st.cards | Where-Object { $_.key -eq $card.key } }; if ($c2 -and $c2.state -eq 'off' -and $c2.pid -eq 0) { break } } catch { } + } + Write-Output ("RESULT card-off after " + $t + " s") + Start-Sleep -Seconds 5 +} else { Write-Output 'RESULT card none-found (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' } + +# the probe ran in run-readwidth-9070-20261005 (bench-log, 5 October 2026); this second run is the bench only +foreach ($pk in @('w4', 'w16', 'w64', 'w64x4', 'mixA-0', 'mixA-1', 'mixA-2', 'mixA-3', 'mixA-4', 'mixA-5', 'mixB-0', 'mixB-1', 'mixB-2', 'mixB-3', 'mixB-4', 'mixB-5')) { + $d = Join-Path $packs $pk + Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)" + & $exe --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev 2>&1 | ForEach-Object { "RESULT $_" } + & $exe --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev --group-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL' } | ForEach-Object { "RESULT $_" } +} +foreach ($pk in @('scr0k32', 'scr2k32', 'scr4k32', 'scr8k32', 'scr2k128', 'scr4k128', 'scr8k128')) { + $d = Join-Path $packs $pk + Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)" + & $exe --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev 2>&1 | ForEach-Object { "RESULT $_" } + & $exe --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev --warps 4096 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|variant 5' } | ForEach-Object { "RESULT $_" } +} + +if ($card) { + $body = @{ cards = @(@{ key = $card.key; enabled = [bool]$card.enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored enabled=" + $card.enabled) } catch { Write-Output ("RESULT error card restore: " + $_.Exception.Message) } +} +exit 0 From f59708dff7ff76e15e3a35454d549649d4f246e2 Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:12:53 +0000 Subject: [PATCH 006/311] Counter ASIC 2.0 layer 7: integer matrix family design (vendor primitives cited: PTX dp4a and mma .u8/.s8, AMD v_dot4_i32_iu8 and WMMA iu8 on RDNA 3 and 4 via LLVM and GPUOpen, CDNA 3 MFMA i8, Metal 4 matmul2d char x char -> int found in MSL 4.1 table 7.3; dot4 and mm8 semantics, reserve entry R1, emulation rule); M5 Max dot4 probe numbers in the bench log; PC playbook dot4-probe.ps1 prepared, not published Co-Authored-By: Claude Fable 5.1 --- docs/analysis/int8-matrix-family.md | 159 ++++++++++++++++++++++++++++ docs/bench-log.md | 15 +++ relay/playbooks/dot4-probe.ps1 | 67 ++++++++++++ 3 files changed, 241 insertions(+) create mode 100644 docs/analysis/int8-matrix-family.md create mode 100644 relay/playbooks/dot4-probe.ps1 diff --git a/docs/analysis/int8-matrix-family.md b/docs/analysis/int8-matrix-family.md new file mode 100644 index 000000000..3a4dd03ed --- /dev/null +++ b/docs/analysis/int8-matrix-family.md @@ -0,0 +1,159 @@ +# Layer 7: the integer matrix family (INT8 x INT8 into INT32) as a reserved instruction family, design + +5 October 2026 (night), Counter ASIC 2.0 (`docs/plans/counter-asic-2.md`, layer 7), branch `ca2-analysis`. Design +only: nothing here touches the generator, a vector or a node. Every figure is cited (vendor document, URL, section) or +measured (machine, date, command) or labelled approximate. + +## 1. The primitive per vendor, from the vendor documents + +| Vendor, hardware | Per-lane dot4 (4 bytes x 4 bytes into a 32-bit integer) | Warp or wave matrix (int8 tiles, int32 accumulate) | Source | +|---|---|---|---| +| NVIDIA, sm_61 and later (Pascal on) | PTX `dp4a.atype.btype d, a, b, c` with `.atype = .btype = {.u32, .s32}`: "Four-way byte dot product which is accumulated in 32-bit result"; semantics `d = c; for i in 0..3: d += Va[i] * Vb[i]` with the bytes sign- or zero-extended by type; introduced in PTX ISA 5.0, "Requires sm_61 or higher". CUDA: `__device__ int __dp4a(int srcA, int srcB, int c)` ("Four-way signed int8 dot product with int32 accumulate") and the unsigned form, plus `char4`/`uchar4` overloads | `mma.sync` with `.u8`/`.s8` A and B and `.s32` C and D: shape `.m8n8k16` "requires sm_75 or higher" (Turing on, PTX 6.5); shapes `.m16n8k16` and `.m16n8k32` require sm_80 (Ampere on, PTX 7.0); sparse `.m16n8k32` and `.m16n8k64` with `.u8`/`.s8` also exist | PTX ISA 9.4, section 9.7.1.24 (dp4a) and 9.7.16.5 (mma), https://docs.nvidia.com/cuda/parallel-thread-execution/index.html ; CUDA Math API, integer intrinsics, https://docs.nvidia.com/cuda/cuda-math-api/cuda_math_api/group__CUDA__MATH__INTRINSIC__INT.html ; read 5 October 2026 | +| AMD RDNA 3 (gfx11) | `v_dot4_i32_iu8` (VOP3P; each operand signed or unsigned by a per-operand bit, optional clamp) reached from clang/HIP/OpenCL C as `__builtin_amdgcn_sudot4(bool a_signed, int a, bool b_signed, int b, int acc, bool clamp)` (LLVM feature `dot8-insts`: "Has v_dot4_i32_iu8, v_dot8_i32_iu4 instructions"); `v_dot4_u32_u8` as `__builtin_amdgcn_udot4` (`dot7-insts`: "Has v_dot4_u32_u8, v_dot8_u32_u4"); `v_dot4_i32_i8` as `__builtin_amdgcn_sdot4` (`dot1-insts`: "Has v_dot4_i32_i8 and v_dot8_i32_i4"). gfx11's common feature set carries dot7, dot8, dot9, dot10 and dot12 | `V_WMMA_I32_16X16X16_IU8`: `__builtin_amdgcn_wmma_i32_16x16x16_iu8_w32` and `_w64` (feature `wmma-256b-insts`), a 16x16x16 tile per wave | LLVM `clang/include/clang/Basic/BuiltinsAMDGPU.td` (main, read 5 October 2026), lines defining `__builtin_amdgcn_sdot4`, `udot4`, `sudot4`, `wmma_i32_16x16x16_iu8_w32`; AMD GPUOpen, "How to accelerate AI applications on RDNA 3 using WMMA", https://gpuopen.com/learn/wmma_on_rdna3/ ; the RDNA 3 ISA PDF itself did not download tonight (AMD's CDN refused curl and the fetcher timed out), so the instruction names are from the compiler and GPUOpen, not quoted from the ISA guide | +| AMD RDNA 4 (gfx12, the 9070 XT) | the same `sudot4` and `udot4` builtins: LLVM's `FeatureISAVersion12_Generic` carries `FeatureDot7Insts` and `FeatureDot8Insts` and not `FeatureDot1Insts`, so `__builtin_amdgcn_sdot4` is NOT exposed on gfx12 and `sudot4` with both operands signed is the signed form to use | `__builtin_amdgcn_wmma_i32_16x16x16_iu8_w32_gfx12` and `_w64_gfx12` (feature `wmma-128b-insts`): the int8 WMMA exists on RDNA 4 with a narrower per-lane operand (2 ints per lane for A and B against 4 on RDNA 3); AMD's RDNA 4 WMMA guide names the same builtin | LLVM `llvm/lib/Target/AMDGPU/AMDGPU.td` (`FeatureISAVersion12_Generic`) and `BuiltinsAMDGPU.td` (main, 5 October 2026); AMD GPUOpen, "WMMA guide for AMD RDNA 4 architecture GPUs, part 2", https://gpuopen.com/learn/wmma-guide-amd-rdna-4-gpus-part-2/ ; the RDNA 4 ISA guide (AMD document 70651, April 2025) was not readable tonight (docs.amd.com returned 401 to a direct fetch) | +| AMD CDNA 3 (MI300) | the same VOP3P dot instructions (approximate: not checked in the CDNA 3 guide tonight) | `V_MFMA_I32_16X16X32_I8` and `V_MFMA_I32_32X32X16_I8` (opcodes 87 and 86 in the VOP3P-MFMA table) | AMD Instinct MI300 CDNA 3 ISA Reference Guide, 5 August 2025, https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/instruction-set-architectures/amd-instinct-mi300-cdna3-instruction-set-architecture.pdf (downloaded and grepped 5 October 2026) | +| AMD, OpenCL on Adrenalin (Windows) | `cl_khr_integer_dot_product` is NOT in the 24 extensions Adrenalin lists for gfx1201 (the list: fp64, the int32 and int64 atomics, 3d image writes, byte addressable store, fp16, gl sharing, amd device attribute query, amd media ops and media ops2, d3d10, d3d11 and dx9 sharing, image2d from buffer, subgroups, gl event, depth images, mipmap image and writes, amd copy buffer p2p); the platform is OpenCL 2.1 so the OpenCL C 3.0 feature macro `__opencl_c_integer_dot_product_input_4x8bit` is not expected. What IS reachable: the Adrenalin OpenCL compiler is clang (driver string `PAL,LC`), and `__builtin_amdgcn_sudot4` from OpenCL C has been shown to emit `V_DOT4_I32_IU8` on a Radeon 780M (gfx1103, RDNA 3, driver 32.0.31041, Windows 11) at 2.7x the scalar fallback (1.27 to 3.42 TMAC/s) | not from OpenCL C | Adrenalin 26.9.2 extension list, https://geeks3d.com/20260904/amd-radeon-adrenalin-26-9-x-graphics-driver/ ; the OpenCL C route: https://github.com/1640675651/CPPminer/pull/1 (third party, one machine; the 9070 XT run of this document's probe is the check) ; `cl_khr_integer_dot_product` itself: OpenCL C 3.0 specification section 6.2.2.16, `int dot(char4, char4)` and `int dot_acc_sat(char4, char4, int)`, https://registry.khronos.org/OpenCL/specs/3.0-unified/html/OpenCL_Ext.html | +| Apple, Metal (MSL 4.1, 4 June 2026) | none. MSL has no dp4a or packed byte dot product: the built-in `dot(T x, T y)` is a geometric function on floating-point vectors (section 6.9); the integer functions of section 6.4 have no dot form. A per-lane dot4 is scalar emulation (section 4 below measures it) | `simdgroup_matrix` exists for T = half, bfloat (Metal 3.1 and later) and float only (section 2.4: "T is half, bfloat ... or float"); no integer SIMD-group matrix. BUT Metal 4's tensor operation `mpp::tensor_ops::matmul2d` (section 7.2.1, table 7.3, "MatMul2D data type supported") lists A `char` x B `char` into C `int` (Metal 4) and `uchar` x `uchar` into `int` (Metal 4 and OS 26.4), plus `char` x `int4b_format` into `int`. So Apple has an exact int8 x int8 into int32 matrix path, on tensors (device or threadgroup memory, or a `cooperative_tensor` per SIMD-group or threadgroup), not on registers, and only through Metal 4's tensor API. Which GPU families run it in hardware (the M5's neural accelerators) against emulation is in the Metal Feature Set Tables, which the spec defers to and which were not read tonight | Metal Shading Language Specification version 4.1, https://developer.apple.com/metal/Metal-Shading-Language-Specification.pdf , sections 2.4, 6.9, 7.2.1 table 7.3 (PDF downloaded and text-extracted 5 October 2026) | + +The Apple finding, stated plainly: the brief's expectation ("Apple has no int8 matrix or dot path") is half right. There +is no per-lane dot4 and no integer `simdgroup_matrix`. There is an exact `char x char -> int` matmul2d in Metal 4 +(table 7.3). Two things about it are unverified tonight and matter for conformance: whether the int accumulate wraps or +saturates (the spec text I extracted says nothing either way; a vector at the int32 edge on the M5 settles it), and the +feature-set table (which Apple GPUs run it natively). What is settled: a generator op that is a per-lane dot4 has no +Apple intrinsic and costs scalar emulation; a generator op that is a whole-unit 8x8x16 or 16x16x16 int8 tile has a +native path on all three vendors (mma.sync on sm_75+, WMMA on RDNA 3 and 4, matmul2d on Metal 4), with Apple's path +living in a different API shape (tensors, not register fragments). + +## 2. The family's semantics as a generator op (integer only, bit-exact) + +Two forms are proposed; the reserve can hold both as separate families or one. + +### 2.1 `dot4`: per-lane + +``` +dot4 dst = dst + dot4_u8(src, src2) + where dot4_u8(a, b) = sum over i in 0..3 of byte_i(a) * byte_i(b), bytes zero-extended, sum modulo 2^32 +``` + +- Bytes are UNSIGNED. Reason, measured below: on Apple the unsigned emulation costs 1.6 ALU-chain steps per dot4 + and the signed one 4.7 (section 4), while NVIDIA (`dp4a.u32.u32`) and AMD (`V_DOT4_U32_U8`, `udot4`, `dot7-insts`) + carry the unsigned form natively as they carry the signed one. Signed bytes buy nothing for the hash (the input is a + pseudo-random register) and cost the vendor without the intrinsic 3x more. +- Accumulation wraps modulo 2^32 like every other op in section 1 (spec 1.14 item 5). The maximum dot of four unsigned + bytes is 4 x 255 x 255 = 260,100, so no single dot4 overflows; the wrap is in the running sum, which is why the + AMD `clamp` bit and the OpenCL `dot_acc_sat` form are NOT the primitive (saturation would change results). +- Operands: `dst`, `src`, `src2` with `src != dst` as for `mad`; `src2` may equal either. +- Verifier: one closed-form integer expression per lane; the register-major interpreter of 1.11 adds four byte + multiplies and adds per lane. The CPU reference in the probes (`dot4_ref` in `proto-opencl/dot4-probe.c`) is this + expression. + +### 2.2 `mm8`: the 32-lane unit as one int8 tile + +The 32 lanes of a unit (spec 1.9) hold, in `src`, a 4-byte row fragment of an 8 x 16 int8 matrix A and, in `src2`, +a 4-byte column fragment of a 16 x 8 int8 matrix B, in exactly the layout of PTX `mma.m8n8k16` with `.u8` operands +(PTX ISA 9.4 section 9.7.16.5, "Matrix Fragments for mma.m8n8k16", the integer-type layout): + +``` +lane l (0..31): A[row = l >> 2][k = 4 * (l & 3) .. 4 * (l & 3) + 3] = the 4 bytes of src (byte 0 = lowest k) + B[k = 4 * (l & 3) .. +3][col = l >> 2] = the 4 bytes of src2 +result C[r][c] = sum over k in 0..15 of A[r][k] * B[k][c] (uint8 x uint8, 16 products, exact, at most 1,040,400) +mm8 dst = dst + C[l >> 2][2 * (l & 3) + bit] bit = an immediate 0 or 1 drawn by the generator +``` + +Every lane receives one of the two C elements its lane position owns in the PTX fragment (`c0` for bit 0, `c1` for +bit 1), added into `dst` modulo 2^32. The whole op is a function of the unit's `src` and `src2` across all 32 lanes, +like `shfl`, so it needs the unit to be exactly 32 logical lanes (the wave64 rule of 1.9 applies: a wave64 device +holds two units and the local-memory path is used). + +How each vendor runs it: + +| Vendor | Native form | Cost per `mm8` (approximate until measured) | +|---|---|---| +| NVIDIA sm_75+ | one `mma.sync.aligned.m8n8k16.row.col.s32.u8.u8.s32` per warp, A and B fragments straight from `src` and `src2`, C = 0 in, `c0`/`c1` out, one add | one tensor instruction plus one add | +| AMD RDNA 3 and 4 | one `V_WMMA_I32_16X16X16_IU8` per wave32 with the 8x16 and 16x8 tiles zero-padded into 16x16 (the WMMA fragment layout differs from PTX's: a fixed permutation of bytes between lanes, which is a few `ds_bpermute` or `v_perm` operations, bit-exact) | one WMMA plus the permutation and the pad | +| AMD CDNA | `V_MFMA_I32_16X16X32_I8` with padding | as above | +| Apple, Metal 4 | `matmul2d` on `uchar` A and B into an `int` cooperative tensor (table 7.3 row "uchar, uchar, int", OS 26.4), the fragments written from registers into a threadgroup tensor first (32 lanes x 8 bytes = 256 bytes), the C element read back per lane | one tensor op plus two threadgroup round trips; on Apple GPUs without the neural accelerators the runtime's emulation, unmeasured | +| Any vendor, fallback | 16 scalar byte products per lane after gathering the 16 bytes of B's column from the 4 lanes that hold them (4 shuffles or one 64-byte threadgroup exchange) | 4 shuffles plus 4 `dot4` emulations: on Apple about 4 x 1.6 = 6.4 ALU steps plus the shuffles (approximate, from the probe) | + +Verifier: the unit evaluates C as 8 x 8 x 16 = 1,024 unsigned byte products once per `mm8` instruction and hands +each lane its element. That is 1,024 multiply-adds per instruction per unit, against 64 x 8 = 512 instructions per +hash: at W_new = 4 (section 3) a program carries about 2.6 `mm8` per iteration, 21 per hash, 21,500 multiply-adds per +unit per hash, under 10 microseconds on one core (approximate), far inside the 0.63 ms the verifier already spends per +unit (spec 1.11). The simulation stays exact because every product and sum is an integer with a defined wrap. + +### 2.3 Which form to reserve + +`mm8` is the one that takes matrix hardware at GPU scale from a chip (the plan's layer 7 row): a chip without tensor +units pays 1,024 products per unit per instruction where a GPU pays one tensor instruction. `dot4` is a per-lane ALU op +that a chip matches with four 8-bit multipliers, which is cheap silicon; it adds little chip resistance and costs Apple +emulation. Recommendation: reserve `mm8`; keep `dot4` out, or in only as `dot4_u8` behind `mm8`. + +## 3. The genesis reserve entry (spec text for 1.13.2) + +Proposed wording, to go under 1.13.2 as the first named reserve family once the conformance runs of section 5 pass: + +> Reserve family R1, `mm8` (integer matrix). Semantics: section 2.2 of `docs/analysis/int8-matrix-family.md`, +> uint8 operands from `src` and `src2` in the m8n8k16 fragment layout, one int32 element of C per lane selected by +> the immediate `bit`, added into `dst` modulo 2^32. Weight at unlock `W_new = 4` points, taken proportionally from the +> ten live non-load families (the load weight and count are untouched, 1.13.1). Edge vectors, each a hand-built unit +> run on every vendor: all bytes 0xFF in A and B (C = 16 x 65,025 = 1,040,400 everywhere); all bytes 0x80 (C = 16 x +> 16,384 = 262,144); A all zero (C = 0); `dst` = 0xFFFFFFFF with a nonzero C (the wrap); alternating 0x00 and 0xFF by +> lane (the fragment mapping: C[r][c] nonzero only where the row and column bytes meet); `bit` = 0 and 1 on the same +> fragments. Unlock: at the start of era n = 4 (two years after genesis, DAA 62,208,000), or earlier by the 90% +> signalling path of section 5.7; never by a release. + +The era-4 choice is deliberate: two years is long enough for the three vendors' tensor paths (and Apple's Metal 4 +feature-set coverage) to be in every miner's driver, and short enough to land before any chip built against the +launch instruction set has paid back (approximate; a chip programme is 12 to 24 months, approximate, from memory). + +Reserve rule for a vendor that can only emulate. Spec 1.13.2 as written requires conformance on every vendor; it says +nothing about cost. Proposed addition: + +> A family enters the reserve when it is bit-exact on every vendor of 1.15. A vendor that reaches the result only by +> emulation (no instruction or library path) does not block entry if the measured penalty of the emulation on that +> vendor, on the family's own probe (a dependent chain of the op, G ops/s against the same vendor's integer ALU chain), +> is at most 8x per op, AND the family's weight at unlock keeps the emulating vendor's hash-rate loss under 5% on the +> memory-hard hash (the hash is latency-bound, so a per-op penalty on 4% of the instructions is a small fraction of a +> hash whose time is 128 dependent DRAM reads; the 5% is checked on the vendor's card with the family live, not +> computed). A family whose emulation exceeds either bound stays out of the reserve until the vendor ships a path. + +With tonight's numbers: on Apple the unsigned `dot4` emulation is 1.6x per op (inside the bound); the signed one 4.7x +(inside, but why pay it); `mm8` through Metal 4's matmul2d is a path, not an emulation, and its cost is owed. + +## 4. dp4a-class throughput, measured so far + +Probe: a dependent chain of one dot4 per step per lane (`acc = dot4(x, y, acc); x = x * K + acc; y = rotl(y, 7) ^ +(acc + s)`), 1,048,576 lanes x 4,096 steps, best of 3, device time, bit-exact against a CPU reference on two lanes +per run, beside the ALU chain of the 9070 XT bench-log entry (`x = x * K + rotl(y, 7); y = (y ^ x) + s`, 5 ops per +step counted). Sources: `proto-metal/dot4-probe.swift` (Metal), `proto-opencl/dot4-probe.c` (OpenCL: scalar, the +`cl_khr_integer_dot_product` `dot`, AMD `__builtin_amdgcn_sudot4`, NVIDIA inline PTX `dp4a.s32.s32`), +`proto-cuda/dot4-probe.cu` (CUDA `__dp4a` and the scalar emulation, for a PC with nvcc). Each OpenCL variant is built on +its own and a variant the platform cannot compile prints a "build failed" row. + +| Card, API | Date, command | ALU chain, G steps/s | dot4 signed emulation, G dot4/s | dot4 unsigned emulation, G dot4/s | dot4 intrinsic, G dot4/s | Penalty of the emulation per op (ALU steps per dot4) | ok (bit-exact) | +|---|---|---|---|---|---|---|---| +| Apple M5 Max, Metal | 5 October 2026 20:0x UTC, `with-lock.sh measure ./dot4-probe` (swiftc -O), GPU start-to-end time | 879.8 (4.882 ms) | 188.2 (22.82 ms) | 548.2 (7.834 ms) | none exists | signed 4.7x, unsigned 1.6x | yes, all three kernels | +| Apple M5 Max, Apple OpenCL 1.2 | same, `with-lock.sh measure ./dot4-probe-cl --device 0`, event time | 871.5 (4.928 ms) | 188.4 (22.80 ms) | not in this probe | `cl_khr_integer_dot_product` not listed; the kernel using `dot(char4, char4)` compiled anyway and ran at 846 G/s but MISMATCHED the CPU reference on every lane checked (Apple's `dot` on char4 is not an integer dot; the extension macro must gate it) | signed 4.6x | alu and dot4e yes; dot4_khr NO | +| RTX 5090 (PC 1 or 2), NVIDIA OpenCL inline PTX, and CUDA `__dp4a` | owed: PC job prepared (`relay/playbooks/dot4-probe.ps1`, exe `dot4-probe-cl.exe` cross-compiled, sha256 in the bench-log entry); not published until the coordinator's "go PC" | | | | | | | +| RX 9070 XT (PC 1), AMD OpenCL `__builtin_amdgcn_sudot4` | owed, same job (the job runs every listed device: the 5090 on NVIDIA's OpenCL, the 9070 XT, the gfx1036) | | | | | | | + +Reading of the Mac numbers. The ALU chain's 880 G steps/s on the M5 Max is the integer baseline (5 ops per step +counted, so about 4.4 T int ops/s, approximate). A signed dot4 emulated as `int4(as_type(a))` products costs +4.7 of those steps; the unsigned form 1.6 steps. The 3x gap between the two is the sign extension (Metal lowers the +unsigned byte extraction to masks that fold into the multiplies, approximate reading of the result, not of the +compiled code). Both are far under the 8x bound of section 3, and the hash spends its time on DRAM reads, so a per-lane +`dot4` family would cost Apple a few percent at W_new = 4 (to be measured with the family live, not computed). The +Apple OpenCL `dot(char4, char4)` mismatch is the kind of thing the edge vectors of section 3 exist to catch. + +## 5. What is owed or unverified + +| Item | State | +|---|---| +| dp4a throughput on the RTX 5090 (inline PTX through NVIDIA OpenCL; `__dp4a` through CUDA if the PC has nvcc) | owed, PC job prepared, waiting for "go PC" | +| `sudot4` on the 9070 XT through Adrenalin's OpenCL C, and whether `cl_khr_integer_dot_product` appears on the 3683.0 platform | owed, same job; the probe prints both | +| Metal 4 `matmul2d` uchar x uchar into int on the M5 Max: wrap or saturate at the int32 edge, native or emulated, throughput | owed (a second Metal probe; the API needs a tensor set-up the dot4 probe does not have) | +| Metal Feature Set Tables: which Apple GPU families run int8 matmul2d natively | not read tonight | +| RDNA 3 and RDNA 4 ISA guides: the instruction text itself (names taken from LLVM and GPUOpen) | AMD's CDN refused the downloads tonight | +| `mm8` on AMD: the exact byte permutation between the PTX m8n8k16 fragment layout and the RDNA WMMA 16x16x16 layout | design, to be written with the kernel | +| The hash-rate cost of the family live at W_new = 4 on each vendor (the 5% rule of section 3) | owed, needs the generator change (not tonight) | +| Edge vectors of section 3 as files | owed, with the generator change | diff --git a/docs/bench-log.md b/docs/bench-log.md index 30b299cbc..a060c5c90 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1524,3 +1524,18 @@ What is measured: one BLS12-381 aggregate signature over 16 summed G1 keys plus | on, split 90 s | v3 | 0 / 2 | none / 3 | 278 / 265 | apart | none | 3 on n0 | 2 (n0 reconnected 6 s after the heal, A's chain at about 58 DAA, inside the table) | Reading (the NEW finding, ledger C4). With the module off GHOSTDAG alone converges on the heavier chain and the losing side's records re-determine (F24 works when the chain moves). With the module on the overlay holds during the split (A, with 30% of the frozen table, locks nothing; B locks 7 and 8) and then fails at the heal in the shipped node: B's certificates for blocks off n0's chain are "kept pending until the chain decides (no lock at this index)", n0's chain never decides because GHOSTDAG keeps its heavier tip and nothing turns the certificate into a fork-choice constraint, and once n0's last lock (index 7, DAA 209) is one window old (DAA 329) the frozen table stops applying on A's chain ("no frozen table (no lock on this chain inside the window)"), A's two keys are 100% of A's own window (B's post-cut blocks are red there) and n0 locks 10, 11, 12 alone; B's certificates for 10 and 11 then log CONFLICTING on n0 (n0 log, 17:27:04 to 17:29:54 BST). A finality fork from a 96-s honest partition, no attacker, table intact at the heal; the 150-s run and the v2 control end the same way. The spec's fork choice ("GHOSTDAG among tips through all certified checkpoints", 3.5) is therefore implemented only for certificates over blocks already on the node's chain. Fix named in the ledger entry: verify an off-chain certificate against the table at its own block and let it constrain fork choice (a certificate-driven reorg), then re-determine. Raw: `scratchpad fud-a/c4-results-*.md`, node logs `c4-on90-tmp/`, `c4-v2-control-tmp/`. + +## 5 October 2026 (night), dp4a-class throughput on the M5 Max: the dot4 emulation against the ALU chain (Counter ASIC 2.0 layer 7) + +Apple M5 Max, macOS 26, branch `ca2-analysis` (base `readwidth` 4badcee). The probes are standalone (no pack, no lottery kernel): `proto-metal/dot4-probe.swift` (built `swiftc -O -o dot4-probe dot4-probe.swift -framework Metal` under `with-lock.sh build`), `proto-opencl/dot4-probe.c` (built `cc -std=c99 -O2 -o dot4-probe-cl dot4-probe.c -framework OpenCL`), both run under `with-lock.sh measure` (exclusive; nothing else built or measured on the Mac during the runs). Shape: a dependent chain of one dot4 per step per lane, `acc = dot4(x, y, acc); x = x * 0x9E3779B1 + acc; y = rotl(y, 7) ^ (acc + s)`, 1,048,576 lanes x 4,096 steps, work-group 256, best of 3 with a fresh seed per repetition, device time (Metal: command buffer GPU start to end; OpenCL: event profiling). Beside it the ALU chain of the 9070 XT entry (`x = x * K + rotl(y, 7); y = (y ^ x) + s`, 5 ops per step counted). Every kernel is checked bit for bit against a CPU reference on lanes 0 and 1,048,575 in every repetition ("ok"). Design context: `docs/analysis/int8-matrix-family.md`. + +| API, kernel | What one step is | best ms | G steps/s | ns per dependent step | ok | +|---|---|---|---|---|---| +| Metal, `probe_alu` | mul, add, rotate, xor, add | 4.882 | 879.8 (about 4.4 T int ops/s at 5 per step, approximate) | 1,192 | yes | +| Metal, `probe_dot4s` | signed dot4 emulated: `int4(as_type(a))` x same for b, 4 products summed into a wrapping int, plus the 3-op chain | 22.820 | 188.2 G dot4/s | 5,571 | yes | +| Metal, `probe_dot4u` | unsigned dot4 emulated: `uint4(as_type(a))`, same chain | 7.834 | 548.2 G dot4/s | 1,913 | yes | +| Apple OpenCL 1.2, `alu` | as Metal | 4.928 | 871.5 | 1,203 | yes | +| Apple OpenCL 1.2, `dot4e` | signed dot4 emulated with `convert_int4(as_char4(a))` | 22.797 | 188.4 G dot4/s | 5,566 | yes | +| Apple OpenCL 1.2, `dot4_khr` | `acc + dot(as_char4(x), as_char4(y))` under `#pragma OPENCL EXTENSION cl_khr_integer_dot_product : enable` | 5.076 | 846.2 | 1,239 | NO: mismatched the CPU reference on every lane checked in all 3 repetitions | + +Reading: on this GPU a signed-byte dot4 costs 4.7 ALU-chain steps and an unsigned-byte one 1.6; Metal has no dp4a and no integer simdgroup matrix (MSL 4.1 sections 2.4 and 6.9), so these are the honest Apple costs of a per-lane dot4 family, and an unsigned definition is 3x cheaper for Apple at no cost to NVIDIA or AMD (both carry the unsigned form, PTX `dp4a.u32.u32`, AMD `v_dot4_u32_u8`). Apple's OpenCL does not list `cl_khr_integer_dot_product`; its `dot` on `char4` compiled anyway and returned something other than the integer dot (the mismatch), which is why a family's conformance vectors must gate every vendor path on the feature macro, not on "it compiled". Not run here: NVIDIA and AMD. The PC job is prepared and not published (coordinator's rule): `relay/playbooks/dot4-probe.ps1` with `dot4-probe-cl.exe` (proto-opencl/dot4-probe.c cross-compiled with mingw as `x86_64-w64-mingw32-gcc -std=c99 -O2 -static -DIGNEUM_CL_DYNAMIC -DCL_TARGET_OPENCL_VERSION=120 -I proto-cuda/nvrtc/redist/include`, sha256 `5adaeb1aceb03dc41135baabe0b53f1ed5fac891a5b3c3849645b03efe4416f4`, 161,863 bytes); it runs the scalar, KHR, AMD `__builtin_amdgcn_sudot4` and NVIDIA inline-PTX `dp4a` variants on every OpenCL GPU of the machine with the mining cards switched off through `/api/cards` and restored after. The CUDA form (`proto-cuda/dot4-probe.cu`, `__dp4a`) needs nvcc on the PC and is the cross-check. diff --git a/relay/playbooks/dot4-probe.ps1 b/relay/playbooks/dot4-probe.ps1 new file mode 100644 index 000000000..2d90cda50 --- /dev/null +++ b/relay/playbooks/dot4-probe.ps1 @@ -0,0 +1,67 @@ +# Igneum run job: dp4a-class throughput (Counter ASIC 2.0 layer 7, docs/analysis/int8-matrix-family.md) on every OpenCL +# GPU of the machine: PC 1 (ae432dc7: RTX 5090 on NVIDIA's OpenCL, RX 9070 XT on the eGPU on AMD's, the gfx1036) or +# PC 2 (1ccfe586: RTX 5090). 5 October 2026. NOT published until the coordinator says "go PC ". +# Published as a plain `run` job (NOT --stop-miners), after a `fetch` job with --id fetch-dot4-20261005 that places +# dot4-probe-cl.exe (proto-opencl/dot4-probe.c cross-compiled with mingw, OpenCL.dll loaded at run time) in the jobs +# folder. The script switches off every NVIDIA and gfx1201 card in the app through POST api/cards, waits for +# their workers to stop, runs the probe on every device the exe lists (the whole run is under a minute: three to five +# kernels x 3 repetitions x about 10 ms each), and switches the cards back on with the settings they had. Every result +# line starts with RESULT so `node tools/jobs.mjs ` shows them; the probe's own "RESULT DOT4 ..." lines carry +# vendor, device, kernel, best ms and G steps/s, and ok=1 means bit-exact against the CPU reference on two lanes. +$ErrorActionPreference = 'Continue' +function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) } +$jobs = Split-Path $env:IGNEUM_JOB_DIR +$fetched = Join-Path $jobs 'fetch-dot4-20261005' +$exe = Join-Path $fetched 'dot4-probe-cl.exe' +if (-not (Test-Path $exe)) { Write-Output "RESULT error probe missing at $exe (the fetch job runs first)"; exit 2 } +Write-Output "RESULT probe $exe sha256 $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower())" +$list = & $exe --list 2>&1 +$list | ForEach-Object { "RESULT list $_" } +$devs = @() +foreach ($l in $list) { if ($l -match '^\[(\d+)\]') { $devs += [int]$Matches[1] } } +if ($devs.Count -eq 0) { Write-Output 'RESULT error no OpenCL GPU device listed'; exit 2 } + +# the app: switch off the NVIDIA card(s) and the 9070 XT, remember their settings +$appDir = $env:IGNEUM_APP_DIR +if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' } +$urlFile = Join-Path $appDir 'app.url' +$url = $null +if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() } +$cards = @() +if ($url) { + try { + $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10 + $all = $st.mining.cards; if (-not $all) { $all = $st.cards } + $cards = @($all | Where-Object { $_.vendor -eq 'nvidia' -or ($_.vendor -eq 'amd' -and $_.key -match 'gfx1201') }) + } catch { Say ("api/state: " + $_.Exception.Message) } +} +if ($cards.Count -gt 0) { + foreach ($c in $cards) { Write-Output ("RESULT card " + $c.key + " enabled=" + $c.enabled + " identities=" + $c.identities + " power_pct=" + $c.power_pct + " state=" + $c.state) } + $body = @{ cards = @($cards | ForEach-Object { @{ key = $_.key; enabled = $false; identities = [int]$_.identities; power_pct = [int]$_.power_pct } }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "cards off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) } + $t = 0 + while ($t -lt 90) { + Start-Sleep -Seconds 5; $t += 5 + try { + $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10 + $all = $st.mining.cards; if (-not $all) { $all = $st.cards } + $still = @($all | Where-Object { ($cards.key -contains $_.key) -and -not ($_.state -eq 'off' -and $_.pid -eq 0) }) + if ($still.Count -eq 0) { break } + } catch { } + } + Write-Output ("RESULT cards-off after " + $t + " s") + Start-Sleep -Seconds 5 +} else { Write-Output 'RESULT card none-found (the app is not running or has no NVIDIA or gfx1201 card); measuring with whatever else runs on the GPUs' } + +if (Get-Command nvidia-smi -ErrorAction SilentlyContinue) { & nvidia-smi --query-gpu=name,driver_version,clocks.sm,clocks.mem,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" } } +foreach ($d in $devs) { + Write-Output "RESULT probe device $d start $(Get-Date -Format HH:mm:ss)" + & $exe --device $d 2>&1 | ForEach-Object { if ($_ -match '^RESULT ') { $_ } else { "RESULT $_" } } +} +if (Get-Command nvidia-smi -ErrorAction SilentlyContinue) { & nvidia-smi --query-gpu=clocks.sm,clocks.mem,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" } } + +if ($cards.Count -gt 0) { + $body = @{ cards = @($cards | ForEach-Object { @{ key = $_.key; enabled = [bool]$_.enabled; identities = [int]$_.identities; power_pct = [int]$_.power_pct } }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output "RESULT cards restored" } catch { Write-Output ("RESULT error card restore: " + $_.Exception.Message) } +} +exit 0 From 0d8f7450825626077098c621ec32003d0a6fcd7a Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:13:04 +0000 Subject: [PATCH 007/311] scratch soundness (layer 3 of Counter ASIC 2.0): verify.rs scratch trace hook; tests/scratch.rs: rewrite and fill bijections, written-word bias and re-hit rates per class, hand-built slot edge programs against a hand model, static scratch-mask check over every emitted kernel of every scr pack with six deliberate breaks, 200-program CPU fuzz and pack writer for the Metal runs; packbench --batch-base for the launch-level nonce wrap Co-Authored-By: Claude Fable 5.1 --- igneum-pow/src/verify.rs | 44 ++- igneum-pow/tests/scratch.rs | 765 ++++++++++++++++++++++++++++++++++++ proto-metal/packbench.swift | 14 +- 3 files changed, 815 insertions(+), 8 deletions(-) create mode 100644 igneum-pow/tests/scratch.rs diff --git a/igneum-pow/src/verify.rs b/igneum-pow/src/verify.rs index fedfe1b91..2f32928e9 100644 --- a/igneum-pow/src/verify.rs +++ b/igneum-pow/src/verify.rs @@ -51,6 +51,22 @@ pub struct ScratchModel { data: Vec<[u32; 3]>, pub reads: usize, pub writes: usize, + /// Soundness tests (`tests/scratch.rs`, `docs/analysis/scratch-soundness.md`): when `Some`, every + /// read-modify-write is appended as it happened. `None` on every verification path. + pub trace: Option>, +} + +/// One scratch read-modify-write as the interpreter saw it (variant 5 soundness tests). +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct ScratchEvent { + pub lane: u8, + pub slot: u32, + /// The slot had been written earlier in this unit (a re-hit): the words read were a rewrite, not the fill. + pub hit: bool, + pub read: [u32; 3], + /// The fold result, the new value of `dst`. + pub x: u32, + pub written: [u32; 3], } impl ScratchModel { @@ -61,6 +77,7 @@ impl ScratchModel { data: vec![[0; 3]; LANES * slots_per_lane], reads: 0, writes: 0, + trace: None, } } /// Read slot `slot` of `lane`, then rewrite it from the fold result `x`. Returns the three words read. @@ -77,7 +94,11 @@ impl ScratchModel { ] }; let x = fold_words(dst, &w); - self.data[i] = scratch_rewrite(x, &w); + let out = scratch_rewrite(x, &w); + if let Some(t) = self.trace.as_mut() { + t.push(ScratchEvent { lane: lane as u8, slot, hit: self.written[i], read: w, x, written: out }); + } + self.data[i] = out; self.written[i] = true; self.reads += 1; self.writes += 1; @@ -243,6 +264,19 @@ pub fn interpret_warp(program: &Program, base_nonce: u32, ds: &DatasetSource) -> /// [`interpret_warp`] with explicit init words `I` (section 1.6 of the spec). The packs use `I = program.seed`; /// a block uses `I = bind::block_init_words(H, nonce)`. pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32, ds: &DatasetSource) -> WarpResult { + interpret_warp_scratch(program, seed, base_nonce, ds, false).0 +} + +/// [`interpret_warp_init`] that also returns every scratch read-modify-write of the unit in execution order +/// (lane-minor within an instruction, as the interpreter runs them) when `trace` is set; empty otherwise and for +/// a class without a scratch. For the soundness tests of variant 5 only. +pub fn interpret_warp_scratch( + program: &Program, + seed: &[u32; 8], + base_nonce: u32, + ds: &DatasetSource, + trace: bool, +) -> (WarpResult, Vec) { let mask = ds.mask; let mut r = [[0u32; LANES]; 8]; for lane in 0..LANES { @@ -258,6 +292,11 @@ pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32, let mut idx = [0u32; LANES]; let mut val = [0u32; LANES]; let mut scratch = if program.has_scratch() { Some(ScratchModel::new(program.class.scratch_slots_per_lane())) } else { None }; + if trace { + if let Some(m) = scratch.as_mut() { + m.trace = Some(Vec::new()); + } + } let slot_mask = program.class.scratch_slot_mask(); for _ in 0..ITERATIONS { let sel = r[0]; @@ -279,7 +318,8 @@ pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32, let hi = r[4][lane] ^ r[5][lane].rotate_left(9) ^ r[6][lane].rotate_left(18) ^ r[7][lane].rotate_left(27); hashes[lane] = ((hi as u64) << 32) | lo as u64; } - WarpResult { hashes, items_derived } + let events = scratch.and_then(|m| m.trace).unwrap_or_default(); + (WarpResult { hashes, items_derived }, events) } #[inline(always)] diff --git a/igneum-pow/tests/scratch.rs b/igneum-pow/tests/scratch.rs new file mode 100644 index 000000000..d00bbfe4e --- /dev/null +++ b/igneum-pow/tests/scratch.rs @@ -0,0 +1,765 @@ +//! Soundness tests of layer 3 of `docs/plans/counter-asic-2.md`: the per-warp scratch with read-modify-writes +//! (variant 5 of the read-width experiment, `LoadClass::scratch(k, kb)`). Analysis and results: +//! `docs/analysis/scratch-soundness.md`. Every test is parametric over the class's slot count +//! (`scratch_slots_per_lane()`), so the 32 and 128 KiB geometries and any later one run the same checks. +//! +//! What runs under plain `cargo test`: +//! 1. `rewrite_is_a_bijection_of_the_fold_value`, `fill_is_a_bijection_of_the_nonce`: the written words as +//! functions (question 1). +//! 2. `written_words_unbiased_and_rehit_rates`: bit bias of every written word over 2^11 units x 3 seeds per class +//! (the TESTS.md section 3 shape), and the measured slot re-hit rate against the birthday formula (question 2). +//! 3. `edge_programs_match_the_hand_model`: hand-built programs that drive every read-modify-write of a hash to +//! slot 0, slot MASK, through out-of-range registers, to one slot per lane, alternating two slots, and 16 +//! read-modify-writes per iteration on one slot; the interpreter against an independent hand model, and the +//! hand model shown to have teeth (question 3, CPU half). +//! 4. `scr_packs_regenerate_and_pass_the_static_scratch_check`: every emitted kernel of every scr pack under +//! `proto-cuda/packs-readwidth` regenerates from its program.json and passes the static scratch-mask check; +//! the check is shown to fail on four deliberate breaks (question 4). +//! 5. `fuzz_scr_programs_cpu`: 200 generated scratch programs over the six classes, generator contract on every +//! instruction, 4 units each at base nonces across the 32-bit range including the wrap; with +//! `IGNEUM_SCRATCH_PACKS_OUT=` it also writes the packs (and the edge packs) for the Metal runs of +//! `proto-metal/packbench` (question 3 GPU half, question 4, `TESTS.md` section 9 shape). + +use igneum_pow::emit::{ + cuda_kernel, cuda_kernel_bound, export_pack, metal_program, metal_program_bound, opencl_kernel, + opencl_kernel_bound, vectors_json, LoadSource, +}; +use igneum_pow::generator::{ + generate_class, generate_from_seed_bytes_class, Instr, LoadClass, Op, Program, GENERATOR_VERSION, INSTR_COUNT, + ITERATIONS, LANES, +}; +use igneum_pow::seed::{seed_words_from_bytes, SplitMix64}; +use igneum_pow::verify::{ + fold_words, interpret_warp_scratch, scratch_fill, scratch_rewrite, splitmix32, DatasetMode, DatasetSource, + Epoch, ScratchEvent, FOLD_MUL, FOLD_ROT, +}; +use serde_json::Value; +use std::collections::HashMap; +use std::path::PathBuf; + +/// The classes under study: the two capped geometries (32 and 128 KiB per warp: 64 and 256 slots per lane) at the +/// RMW shares the readwidth branch measures. +const CLASSES: [&str; 6] = ["scr2k32", "scr4k32", "scr8k32", "scr2k128", "scr4k128", "scr8k128"]; + +fn class(name: &str) -> LoadClass { + LoadClass::parse(name).unwrap_or_else(|| panic!("class {name}")) +} + +// --------------------------------------------------------------------------------------------------------------- +// 1. The written words as functions (question 1) +// --------------------------------------------------------------------------------------------------------------- + +/// For a fixed slot content `w`, each of the three rewritten words is a bijection of the fold value `x` +/// (`x ^ w1`, `rotl(x, 7) ^ w2`, `x + w0`), so the rewrite is injective in `x` and a uniform `x` gives a uniform +/// word in every position. Checked over 2^16 consecutive `x` for 16 random `w`. +#[test] +fn rewrite_is_a_bijection_of_the_fold_value() { + let mut rng = SplitMix64::new(0x7363_7261_7463_6801); + for _ in 0..16 { + let w = [rng.next() as u32, rng.next() as u32, rng.next() as u32]; + let x0 = rng.next() as u32; + let mut seen = [vec![false; 1 << 16], vec![false; 1 << 16], vec![false; 1 << 16]]; + for i in 0..(1u32 << 16) { + let x = x0.wrapping_add(i); + let out = scratch_rewrite(x, &w); + for j in 0..3 { + // a bijection of x maps 2^16 consecutive x to 2^16 distinct words; the low 16 bits alone are + // distinct for the xor words (x ^ c) and for the add word (x + c), since both act on the low 16 + // bits as bijections of the low 16 bits of x; the rotl word is checked on its rotated-back bits + let key = if j == 1 { out[j].rotate_right(7) & 0xffff } else { out[j] & 0xffff }; + assert!(!seen[j][key as usize], "word {j} repeats inside 2^16 consecutive x"); + seen[j][key as usize] = true; + } + } + } + // The rewrite inverts: from the old content and any ONE written word the fold value is recovered, so a + // rewritten slot carries exactly 32 bits of new state (the point of question 2's arithmetic). + let w = [0x1234_5678, 0x9abc_def0, 0x0fed_cba9]; + let x = 0xdead_beef; + let out = scratch_rewrite(x, &w); + assert_eq!(out[0] ^ w[1], x); + assert_eq!((out[1] ^ w[2]).rotate_right(7), x); + assert_eq!(out[2].wrapping_sub(w[0]), x); +} + +/// For a fixed (seed, slot, j) the fill is a bijection of the lane nonce: `splitmix32` is a bijection of its +/// 32-bit input and the input `((base + lane) ^ s) + c` is a bijection of `base + lane`. Over 2^16 consecutive +/// nonces no fill word repeats, for 8 slots x 3 words. +#[test] +fn fill_is_a_bijection_of_the_nonce() { + let seed = seed_words_from_bytes(b"igneum-genesis"); + for slot in [0u32, 1, 63, 64, 255, 1023, 2047] { + for j in 0..3u32 { + let mut words: Vec = (0..(1u32 << 16)).map(|n| scratch_fill(&seed, n, 0, slot, j)).collect(); + words.sort_unstable(); + words.dedup(); + assert_eq!(words.len(), 1 << 16, "slot {slot} word {j}: fill words of 2^16 consecutive nonces are distinct"); + } + } + // base + lane is the lane nonce: the fill of lane l at base b is the fill of lane 0 at base b + l + assert_eq!(scratch_fill(&seed, 0x1000, 7, 5, 2), scratch_fill(&seed, 0x1007, 0, 5, 2)); + // and it wraps with the nonce: base 0xffffffe0, lane 31 is nonce 0xffffffff; lane 32 would be nonce 0 + assert_eq!(scratch_fill(&seed, 0xffff_ffe0, 32, 5, 2), scratch_fill(&seed, 0, 0, 5, 2)); + // the three word positions of one slot and nonce are three different permutation outputs + let f: Vec = (0..3).map(|j| scratch_fill(&seed, 12345, 7, 17, j)).collect(); + assert!(f[0] != f[1] && f[1] != f[2] && f[0] != f[2]); +} + +// --------------------------------------------------------------------------------------------------------------- +// 2. Uniformity of the written words and the slot re-hit rate (questions 1 and 2) +// --------------------------------------------------------------------------------------------------------------- + +/// Birthday arithmetic: the expected number of distinct slots after `n` uniform draws from `s` slots. +fn expected_distinct(s: usize, n: usize) -> f64 { + let s = s as f64; + s * (1.0 - (1.0 - 1.0 / s).powi(n as i32)) +} + +struct ClassStats { + units: usize, + events: usize, + hits: usize, + /// ones count per bit of the written words, 3 x 32 + ones: [[u64; 32]; 3], + /// ones count per bit of written XOR read (the change the rewrite makes to the slot) + delta_ones: [[u64; 32]; 3], + /// re-hit depth histogram: how many earlier RMWs the slot had seen in this unit (0 = first touch) + depth: Vec, + max_depth: usize, + /// how often each slot index was addressed (the slot comes from a register's low bits) + slot_hist: Vec, +} + +fn class_stats(name: &str, seeds: &[&str], units_per_seed: usize) -> ClassStats { + let c = class(name); + let mut st = ClassStats { + units: 0, + events: 0, + hits: 0, + ones: [[0; 32]; 3], + delta_ones: [[0; 32]; 3], + depth: vec![0; 256], + max_depth: 0, + slot_hist: vec![0; c.scratch_slots_per_lane()], + }; + let ds = DatasetSource::new("2026-10-03", DatasetMode::ClosedForm, 28); + for seed in seeds { + let p = generate_class(seed, c); + assert_eq!(p.scratch_ops_per_hash(), c.scratch_slots() * ITERATIONS); + for u in 0..units_per_seed { + let base = (u as u32).wrapping_mul(32).wrapping_add(0x4000_0000); + let (_, ev) = interpret_warp_scratch(&p, &p.seed, base, &ds, true); + assert_eq!(ev.len(), p.scratch_ops_per_hash() * LANES); + let mut count: HashMap<(u8, u32), usize> = HashMap::new(); + for e in &ev { + assert!(e.slot < c.scratch_slots_per_lane() as u32, "slot inside the lane's scratch"); + let d = count.entry((e.lane, e.slot)).or_insert(0); + assert_eq!(e.hit, *d > 0, "hit flag agrees with the unit's own history"); + assert_eq!(e.written, scratch_rewrite(e.x, &e.read)); + if !e.hit { + let fill = [ + scratch_fill(&p.seed, base, e.lane as u32, e.slot, 0), + scratch_fill(&p.seed, base, e.lane as u32, e.slot, 1), + scratch_fill(&p.seed, base, e.lane as u32, e.slot, 2), + ]; + assert_eq!(e.read, fill, "a first touch reads the fill"); + } + st.depth[(*d).min(255)] += 1; + st.max_depth = st.max_depth.max(*d); + st.slot_hist[e.slot as usize] += 1; + *d += 1; + st.events += 1; + st.hits += e.hit as usize; + for j in 0..3 { + for b in 0..32 { + st.ones[j][b] += ((e.written[j] >> b) & 1) as u64; + st.delta_ones[j][b] += (((e.written[j] ^ e.read[j]) >> b) & 1) as u64; + } + } + } + st.units += 1; + } + } + st +} + +/// Bit bias of every written word (and of the change each rewrite makes) within 6 sigma of a fair coin, over +/// 3 seeds x 2^11 units per class (131,072 hashes per seed set); the slot re-hit rate against the birthday +/// formula within 3 percent relative. The table printed here is the one in the analysis. +#[test] +fn written_words_unbiased_and_rehit_rates() { + let seeds = ["igneum-genesis", "igneum-genesis/stats1", "igneum-genesis/stats2"]; + let units = 1usize << 11; + println!("class | slots/lane | RMW/hash | events | re-hits | re-hit % | birthday % | slot chi2 z (spread) | max depth | max bias sigma | max delta bias sigma"); + for name in CLASSES { + let c = class(name); + let st = class_stats(name, &seeds, units); + let n = st.events as f64; + let sigma = (n / 4.0).sqrt(); + let mut worst = 0.0f64; + let mut worst_delta = 0.0f64; + for j in 0..3 { + for b in 0..32 { + let z = (st.ones[j][b] as f64 - n / 2.0).abs() / sigma; + let zd = (st.delta_ones[j][b] as f64 - n / 2.0).abs() / sigma; + assert!(z <= 6.0, "{name}: written word {j} bit {b} biased: {z:.2} sigma"); + assert!(zd <= 6.0, "{name}: rewrite delta word {j} bit {b} biased: {zd:.2} sigma"); + worst = worst.max(z); + worst_delta = worst_delta.max(zd); + } + } + let per_lane_hash = c.scratch_slots() * ITERATIONS; + let s = c.scratch_slots_per_lane(); + let exp_hits = per_lane_hash as f64 - expected_distinct(s, per_lane_hash); + let exp_pct = 100.0 * exp_hits / per_lane_hash as f64; + let got_pct = 100.0 * st.hits as f64 / st.events as f64; + // chi-square of the slot histogram against uniform (df = s - 1): the slot is a register's low bits, and + // the measured re-hit rate runs above the uniform birthday rate (the finding of the analysis, question 2) + let expect_per_slot = n / s as f64; + let chi2: f64 = st.slot_hist.iter().map(|&h| (h as f64 - expect_per_slot).powi(2) / expect_per_slot).sum(); + let chi2_z = (chi2 - (s as f64 - 1.0)) / (2.0 * (s as f64 - 1.0)).sqrt(); + let hot = *st.slot_hist.iter().max().unwrap() as f64 / expect_per_slot; + let cold = *st.slot_hist.iter().min().unwrap() as f64 / expect_per_slot; + println!( + "{name} | {s} | {per_lane_hash} | {} | {} | {got_pct:.2} | {exp_pct:.2} | {chi2_z:.1} (hottest slot {hot:.2}x, coldest {cold:.2}x) | {} | {worst:.2} | {worst_delta:.2}", + st.events, st.hits, st.max_depth + ); + // a regression band, not a uniformity claim: the rate sits between the uniform birthday rate and twice it + assert!( + got_pct >= 0.9 * exp_pct && got_pct <= 2.0 * exp_pct, + "{name}: re-hit rate {got_pct:.2}% against birthday {exp_pct:.2}%" + ); + // depth histogram: the number of earlier RMWs a re-hit slot had seen in the unit + let shown: Vec = st.depth.iter().take(st.max_depth + 1).enumerate().map(|(d, n)| format!("{d}:{n}")).collect(); + println!(" depth histogram {}", shown.join(" ")); + } +} + +// --------------------------------------------------------------------------------------------------------------- +// 3. Hand-built edge programs against an independent hand model (question 3, CPU half) +// --------------------------------------------------------------------------------------------------------------- + +fn ins(op: Op, dst: u8, src: u8) -> Instr { + Instr { op, dst, src, src2: 0, imm: 0, imm2: 0, rot: 1, bit: 0, mask: 1, width: 1 } +} +fn add_imm(dst: u8, src: u8, imm: u32) -> Instr { + Instr { op: Op::Add, dst, src, src2: 0, imm, imm2: imm, rot: 1, bit: 0, mask: 1, width: 1 } +} + +/// A hand-built program of class `c` named `name` (its seed is the name, so its fill words and init words are +/// its own). These bypass the generator and the acceptance rule, like `TESTS.md` section 2; `sub r, r` zeroes a +/// register as the Swift edge set does. +fn edge(name: &str, c: LoadClass, instrs: Vec) -> Program { + let seed_string = format!("igneum-scratch-edge/{name}"); + let seed_bytes = seed_string.as_bytes().to_vec(); + let k = instrs.iter().filter(|i| i.op == Op::Scratch).count(); + assert_eq!(k, c.scratch_slots(), "{name}: the class carries the program's scratch count"); + Program { + seed: seed_words_from_bytes(&seed_bytes), + seed_string, + seed_bytes, + generator: GENERATOR_VERSION, + attempt: 0, + class: c, + instrs, + } +} + +/// The edge set for a scratch of `kb` KiB per warp. Each entry: (name, what it drives, program). +fn edge_programs(kb: u8) -> Vec<(String, &'static str, Program)> { + let m = LoadClass::scratch(1, kb).scratch_slot_mask(); + let dsts = [2u8, 3, 4, 5, 6, 7, 0, 2, 3, 4, 5, 6, 7, 0, 2, 3]; + let scr = |n: usize, src: u8| -> Vec { (0..n).map(|i| ins(Op::Scratch, dsts[i], src)).collect() }; + let mut v = Vec::new(); + // every RMW of the hash to slot 0 through a zero register: 64 dependent RMWs on one slot per lane + let mut p = vec![ins(Op::Sub, 1, 1)]; + p.extend(scr(8, 1)); + v.push(("slot0".to_string(), "r1 = 0: every RMW to slot 0", edge(&format!("slot0/k{kb}"), LoadClass::scratch(8, kb), p))); + // slot MASK through the in-range register MASK + let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(1, 2, m)]; + p.extend(scr(8, 1)); + v.push(("slotmask".to_string(), "r1 = MASK: every RMW to the last slot", edge(&format!("slotmask/k{kb}"), LoadClass::scratch(8, kb), p))); + // slot MASK through the out-of-range register 0xffffffff + let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(2, 1, 1), ins(Op::Sub, 1, 2)]; + p.extend(scr(8, 1)); + v.push(("ones".to_string(), "r1 = 0xffffffff: masked to the last slot", edge(&format!("ones/k{kb}"), LoadClass::scratch(8, kb), p))); + // slot 0 through the out-of-range register MASK + 1 + let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(1, 2, m.wrapping_add(1))]; + p.extend(scr(8, 1)); + v.push(("maskplus1".to_string(), "r1 = MASK + 1: masked to slot 0", edge(&format!("maskplus1/k{kb}"), LoadClass::scratch(8, kb), p))); + // 16 RMWs per iteration on slot 0: 128 dependent RMWs on one slot per lane per hash + let mut p = vec![ins(Op::Sub, 1, 1)]; + p.extend(scr(16, 1)); + v.push(("sixteen".to_string(), "16 RMWs per iteration on slot 0", edge(&format!("sixteen/k{kb}"), LoadClass::scratch(16, kb), p))); + // one slot per lane from the init words: lanes with equal slots would show any cross-lane aliasing + // (r5 is the slot register and is never a destination here) + let p: Vec = [0u8, 1, 2, 3, 4, 6, 7, 0].iter().map(|&d| ins(Op::Scratch, d, 5)).collect(); + v.push(("lanevar".to_string(), "r5 never written: one init-dependent slot per lane", edge(&format!("lanevar/k{kb}"), LoadClass::scratch(8, kb), p))); + // alternating slot 0 and slot MASK inside one iteration + let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(2, 1, m)]; + for (i, &d) in [3u8, 4, 5, 6, 7, 0, 3, 4].iter().enumerate() { + // r1 and r2 hold the two slots and are never destinations + p.push(ins(Op::Scratch, d, if i % 2 == 0 { 1 } else { 2 })); + } + v.push(("twoslots".to_string(), "slot 0 and slot MASK alternating", edge(&format!("twoslots/k{kb}"), LoadClass::scratch(8, kb), p))); + v +} + +/// The hand model: a second, minimal interpreter for the ops the edge programs use (sub, add, scratch), with its +/// own slot store keyed by (lane, slot). `mutate` swaps the rewrite's words to show the comparison has teeth. +fn hand_model(p: &Program, base: u32, mutate: bool) -> [u64; 32] { + let seed = &p.seed; + let m = p.class.scratch_slot_mask(); + let mut r = [[0u32; LANES]; 8]; + for lane in 0..LANES { + let nonce = base.wrapping_add(lane as u32); + for i in 0..8 { + let mut x = nonce ^ seed[i]; + x = x.wrapping_add(0x9e3779b9u32.wrapping_mul(i as u32 + 1)); + x = splitmix32(x); + r[i][lane] = x ^ seed[(i + 1) & 7]; + } + } + let mut store: HashMap<(usize, u32), [u32; 3]> = HashMap::new(); + for _ in 0..ITERATIONS { + let sel = r[0]; + for ins in &p.instrs { + let (d, a) = (ins.dst as usize, ins.src as usize); + match ins.op { + Op::Sub => { + for lane in 0..LANES { + r[d][lane] = r[d][lane].wrapping_sub(r[a][lane]); + } + } + Op::Add => { + for lane in 0..LANES { + let c = if (sel[lane] >> ins.bit) & 1 != 0 { ins.imm2 } else { ins.imm }; + r[d][lane] = r[d][lane].wrapping_add(r[a][lane]).wrapping_add(c); + } + } + Op::Scratch => { + for lane in 0..LANES { + let slot = r[a][lane] & m; + let w = *store.entry((lane, slot)).or_insert_with(|| { + let mut f = [0u32; 3]; + for j in 0..3u32 { + // the fill, written out in full rather than through verify::scratch_fill + let n = base.wrapping_add(lane as u32); + f[j as usize] = splitmix32( + (n ^ seed[j as usize]) + .wrapping_add(slot.wrapping_mul(0x9E37_79B1)) + .wrapping_add((j + 1).wrapping_mul(0x85EB_CA77)), + ); + } + f + }); + let mut x = r[d][lane] ^ w[0]; + x = x.rotate_left(FOLD_ROT).wrapping_mul(FOLD_MUL) ^ w[1]; + x = x.rotate_left(FOLD_ROT).wrapping_mul(FOLD_MUL) ^ w[2]; + r[d][lane] = x; + let out = if mutate { + [x.rotate_left(7) ^ w[2], x ^ w[1], x.wrapping_add(w[0])] + } else { + [x ^ w[1], x.rotate_left(7) ^ w[2], x.wrapping_add(w[0])] + }; + store.insert((lane, slot), out); + } + } + other => panic!("the hand model does not implement {other:?}"), + } + } + } + let mut out = [0u64; 32]; + for lane in 0..LANES { + let lo = r[0][lane] ^ r[1][lane].rotate_left(7) ^ r[2][lane].rotate_left(14) ^ r[3][lane].rotate_left(21); + let hi = r[4][lane] ^ r[5][lane].rotate_left(9) ^ r[6][lane].rotate_left(18) ^ r[7][lane].rotate_left(27); + out[lane] = ((hi as u64) << 32) | lo as u64; + } + out +} + +/// The four unit bases of every edge vector: 0 and 32 (two consecutive units, the pair a one-warp persistent +/// launch runs on one arena), a unit straddling 2^31, and the unit that wraps past 2^32. +const EDGE_BASES: [u32; 4] = [0, 32, 0x7fff_fff0, 0xffff_ffe0]; + +#[test] +fn edge_programs_match_the_hand_model() { + let ds = DatasetSource::new("2026-10-03", DatasetMode::ClosedForm, 24); + let mut cases = 0; + for kb in [32u8, 128] { + for (name, what, p) in edge_programs(kb) { + let slots = p.class.scratch_slots_per_lane(); + for base in EDGE_BASES { + let (res, ev) = interpret_warp_scratch(&p, &p.seed, base, &ds, true); + let hand = hand_model(&p, base, false); + assert_eq!(res.hashes, hand, "{name} k{kb} base {base:#x}: interpreter against the hand model ({what})"); + assert_ne!(res.hashes, hand_model(&p, base, true), "{name} k{kb}: the comparison has teeth"); + // the slots the trace saw are the ones the program was built to drive + let slot_set: std::collections::BTreeSet = ev.iter().map(|e| e.slot).collect(); + let m = (slots - 1) as u32; + match name.as_str() { + "slot0" | "maskplus1" | "sixteen" => assert_eq!(slot_set.into_iter().collect::>(), vec![0]), + "slotmask" | "ones" => assert_eq!(slot_set.into_iter().collect::>(), vec![m]), + "twoslots" => assert_eq!(slot_set.into_iter().collect::>(), vec![0, m]), + "lanevar" => { + for e in &ev { + assert!(e.slot <= m); + } + } + _ => unreachable!(), + } + // the chain depth on the driven slot: every RMW after the first per lane is a re-hit + let per_lane = p.scratch_ops_per_hash(); + let hits = ev.iter().filter(|e| e.hit).count(); + let expected_hits = match name.as_str() { + "twoslots" => (per_lane - 2) * LANES, + _ => (per_lane - 1) * LANES, + }; + assert_eq!(hits, expected_hits, "{name} k{kb}: re-hits"); + cases += 1; + } + } + } + assert_eq!(cases, 2 * 7 * 4); +} + +// --------------------------------------------------------------------------------------------------------------- +// 4. The static scratch check over every emitted kernel of every scr pack (question 4) +// --------------------------------------------------------------------------------------------------------------- + +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum Dialect { + Metal, + Cuda, + OpenCl, +} + +/// The static scratch check: every scratch read-modify-write in an emitted kernel has the one masked form the +/// emitter writes, the arena is the lane's own `slots x 4` words, the tag is `salt + unit`, and nothing else +/// touches the scratch. Like the dataset mask check of `TESTS.md` section 5 and `tests/packs.rs`, a text check: +/// the guarantee is that the emitter has one template and it masks. +pub fn scratch_text_check(text: &str, dialect: Dialect, k: usize, slots: usize, kernels: usize) -> Result<(), String> { + assert!(kernels >= 1); + // every count below is per hash kernel; an OpenCL bound file carries igneum_hash and igneum_hash_bound + let k = k * kernels; + assert!(slots.is_power_of_two() && slots >= 1); + let mask = (slots - 1) as u32; + let wpl = slots * 4; + let (u, load, store, ptr) = match dialect { + Dialect::Metal => ("uint", "uint4 v_ = *(device const uint4*)(arena + s_ * 4u);", "*(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); }", "device uint* arena"), + Dialect::Cuda => ("uint32_t", "uint4 v_ = *(const uint4*)(arena + s_ * 4u);", "*(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); }", "uint32_t* arena"), + Dialect::OpenCl => ("uint", "uint4 v_ = vload4(s_, arena);", "vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); }", "__global uint* arena"), + }; + let count = |needle: &str| text.matches(needle).count(); + let mut errs = Vec::new(); + let mut expect = |what: &str, got: usize, want: usize| { + if got != want { + errs.push(format!("{what}: {got}, expected {want}")); + } + }; + // k slot computations, each masked with exactly the class's mask and immediately followed by the one load form + expect("slot definitions `{ u s_ = r`", count(&format!("{{ {u} s_ = r")), k); + expect("masked slot followed by the load", count(&format!(" & {mask}u; {load}")), k); + expect("stores of the tagged slot", count(store), k); + expect("tag compares", count("(v_.x == tag)"), k); + expect("fill calls (three per RMW)", count("scr_fill(gbase, lane, s_, "), 3 * k); + // the arena: one definition with the class's words per lane, and 2k uses (one load, one store per RMW) + expect("arena definition", count(&format!("{ptr} = scratch + ((size_t)warp_ * 32u + lane) * {wpl}u;")), kernels); + expect("arena mentions (definition + load + store per RMW)", count("arena"), kernels + 2 * k); + expect("tag definition `tag = salt + g_`", count(&format!("{u} tag = salt + g_;")), kernels); + expect("direct scratch indexing", count("scratch["), 0); + expect("scratch pointer arithmetic outside the arena definition", count("scratch +"), kernels); + // no other mask value on a slot: every `s_ = r` line carries the class mask and nothing else carries ` & Nu; uint4 v_` + let any_mask_load = count(&format!("u; {load}")); + expect("loads preceded by some mask (must all be the class mask)", any_mask_load, k); + if errs.is_empty() { + Ok(()) + } else { + Err(errs.join("; ")) + } +} + +fn packs_rw_dir() -> PathBuf { + PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../proto-cuda/packs-readwidth") +} + +fn scr_packs() -> Vec { + let mut v: Vec = std::fs::read_dir(packs_rw_dir()) + .unwrap() + .map(|d| d.unwrap().file_name().to_string_lossy().to_string()) + .filter(|n| n.starts_with("scr")) + .collect(); + v.sort(); + v +} + +fn read_pack(pack: &str, file: &str) -> String { + let p = packs_rw_dir().join(pack).join(file); + std::fs::read_to_string(&p).unwrap_or_else(|e| panic!("read {}: {e}", p.display())) +} + +/// Every scr pack regenerates from its program.json (seed bytes, class, day bytes, size) to the same six kernel +/// texts, byte for byte, and every one of those texts passes the static scratch check for the class's k and slot +/// count; the check fails on four deliberate breaks of a copy of the Metal text (mask dropped, mask changed, arena +/// stride changed, a stray scratch access) and on the OpenCL and CUDA twins of the first. +#[test] +fn scr_packs_regenerate_and_pass_the_static_scratch_check() { + let packs = scr_packs(); + assert!(packs.len() >= 6, "the scr packs: {packs:?}"); + let mut checked = 0; + let mut sample_metal = String::new(); + let mut sample_cl = String::new(); + let mut sample_cu = String::new(); + let mut sample_k = 0; + let mut sample_slots = 0; + for pack in &packs { + let j: Value = serde_json::from_str(&read_pack(pack, "program.json")).unwrap(); + let name = j["load_class"].as_str().unwrap(); + let c = class(name); + assert_eq!(&format!("{name}"), pack, "pack directory named after its class"); + let seed = j["seed"].as_str().unwrap(); + let seed_bytes = igneum_pow::bind::unhex(j["seed_bytes"].as_str().unwrap()).unwrap(); + let day_bytes = igneum_pow::bind::unhex(j["dataset"]["day_bytes"].as_str().unwrap()).unwrap(); + let log2 = j["dataset"]["log2_words"].as_u64().unwrap() as u32; + assert_eq!(j["dataset_mode"].as_str().unwrap(), "memory-hard"); + let program = generate_from_seed_bytes_class(seed, &seed_bytes, c); + assert_eq!(program.class, c); + assert_eq!(program.program_id(), u64::from_str_radix(j["program_id"].as_str().unwrap().trim_start_matches("0x"), 16).unwrap()); + let mut dataset = DatasetSource::from_key(seed_words_from_bytes(&day_bytes), DatasetMode::MemoryHard, log2); + dataset.key_bytes = day_bytes; + let e = Epoch { program, dataset }; + let p = &e.program; + let mp = e.dataset.memhard().map(|m| &m.params); + let k = c.scratch_slots(); + let slots = c.scratch_slots_per_lane(); + assert_eq!(p.scratch_ops_per_hash(), k * ITERATIONS); + for (file, text, dialect, kernels) in [ + ("program.metal", metal_program(p, log2, LoadSource::Stored), Dialect::Metal, 1), + ("program_bound.metal", metal_program_bound(p, log2), Dialect::Metal, 1), + ("kernel.cu", cuda_kernel(p, mp), Dialect::Cuda, 1), + ("kernel_bound.cu", cuda_kernel_bound(p, mp), Dialect::Cuda, 1), + ("kernel.cl", opencl_kernel(p, mp), Dialect::OpenCl, 1), + // the OpenCL bound file carries igneum_hash and igneum_hash_bound + ("kernel_bound.cl", opencl_kernel_bound(p, mp), Dialect::OpenCl, 2), + ] { + let on_disk = read_pack(pack, file); + assert_eq!(on_disk, text, "{pack}/{file}: the pack is the emitter's text"); + // scr0 is the persistent control: an arena and a tag, no read-modify-write; the check holds with k = 0 + scratch_text_check(&on_disk, dialect, k, slots, kernels).unwrap_or_else(|e| panic!("{pack}/{file}: {e}")); + checked += 1; + } + // the vectors of the pack are the CPU's + let v: Value = serde_json::from_str(&read_pack(pack, "vectors.json")).unwrap(); + for w in v["warps"].as_array().unwrap() { + let base = w["base_nonce"].as_u64().unwrap() as u32; + let got = e.hash_warp(base); + for (lane, x) in w["expected"].as_array().unwrap().iter().enumerate() { + let want = u64::from_str_radix(x.as_str().unwrap().trim_start_matches("0x"), 16).unwrap(); + assert_eq!(got[lane], want, "{pack}: base {base} lane {lane}"); + } + } + if k == 4 && slots == 64 { + sample_metal = read_pack(pack, "program.metal"); + sample_cl = read_pack(pack, "kernel.cl"); + sample_cu = read_pack(pack, "kernel.cu"); + sample_k = k; + sample_slots = slots; + } + } + assert_eq!(checked, packs.len() * 6); + println!("static scratch check: {checked} kernels over {} scr packs", packs.len()); + + // The deliberate breaks (the watcher rule of CLAUDE.md: a check is trusted once it fails on a known-broken + // case). Each must be caught; the message names what. + assert!(sample_k == 4 && sample_slots == 64, "scr4k32 is in the pack set"); + let mask = format!(" & {}u; uint4 v_", sample_slots - 1); + let broken_mask = sample_metal.replacen(&mask, "; uint4 v_", 1); + assert_ne!(broken_mask, sample_metal); + let e = scratch_text_check(&broken_mask, Dialect::Metal, 4, 64, 1).unwrap_err(); + assert!(e.contains("masked slot followed by the load: 3, expected 4"), "{e}"); + println!("break 1 (one mask dropped, Metal): {e}"); + let wrong_mask = sample_metal.replace(" & 63u;", " & 127u;"); + let e = scratch_text_check(&wrong_mask, Dialect::Metal, 4, 64, 1).unwrap_err(); + assert!(e.contains("masked slot followed by the load: 0, expected 4"), "{e}"); + println!("break 2 (mask 63 -> 127 on every RMW, Metal): {e}"); + let wrong_stride = sample_metal.replace("* 256u;", "* 128u;"); + let e = scratch_text_check(&wrong_stride, Dialect::Metal, 4, 64, 1).unwrap_err(); + assert!(e.contains("arena definition: 0, expected 1"), "{e}"); + println!("break 3 (arena stride 256 -> 128 words, Metal): {e}"); + let stray = format!("{sample_metal}\n// stray\n// arena[0] = 0u; scratch[1] = 1u;\n"); + let e = scratch_text_check(&stray, Dialect::Metal, 4, 64, 1).unwrap_err(); + assert!(e.contains("arena mentions") && e.contains("direct scratch indexing: 1, expected 0"), "{e}"); + println!("break 4 (a stray arena and scratch access, Metal): {e}"); + let e = scratch_text_check(&sample_cl.replacen(" & 63u; uint4 v_ = vload4", "; uint4 v_ = vload4", 1), Dialect::OpenCl, 4, 64, 1).unwrap_err(); + assert!(e.contains("masked slot followed by the load: 3, expected 4"), "{e}"); + println!("break 5 (one mask dropped, OpenCL): {e}"); + let e = scratch_text_check(&sample_cu.replacen(" & 63u; uint4 v_ = *(const uint4*)", "; uint4 v_ = *(const uint4*)", 1), Dialect::Cuda, 4, 64, 1).unwrap_err(); + assert!(e.contains("masked slot followed by the load: 3, expected 4"), "{e}"); + println!("break 6 (one mask dropped, CUDA): {e}"); + // and the unbroken texts pass under the same calls + scratch_text_check(&sample_metal, Dialect::Metal, 4, 64, 1).unwrap(); + scratch_text_check(&sample_cl, Dialect::OpenCl, 4, 64, 1).unwrap(); + scratch_text_check(&sample_cu, Dialect::Cuda, 4, 64, 1).unwrap(); + // a wrong slot count, RMW count or kernel count against a right text fails too (the check is tied to the class) + assert!(scratch_text_check(&sample_metal, Dialect::Metal, 4, 256, 1).is_err()); + assert!(scratch_text_check(&sample_metal, Dialect::Metal, 3, 64, 1).is_err()); + assert!(scratch_text_check(&sample_metal, Dialect::Metal, 4, 64, 2).is_err()); +} + +// --------------------------------------------------------------------------------------------------------------- +// 5. The fuzz: 200 generated scratch programs, contract on every instruction, 4 units each across the 32-bit +// range including the wrap; with IGNEUM_SCRATCH_PACKS_OUT the packs for the Metal runs (question 3, 4) +// --------------------------------------------------------------------------------------------------------------- + +/// Write a pack whose vectors.json carries `bases` (any number of units) instead of the three standard bases. +fn write_pack_with_bases(dir: &PathBuf, e: &Epoch, day: &str, bases: &[u32], source: &str) -> Vec<[u64; 32]> { + let mut pack = export_pack(e, day, source); + let outs: Vec<[u64; 32]> = bases.iter().map(|&b| e.hash_warp(b)).collect(); + let vj = vectors_json(&e.program, day, e.dataset.log2_words, bases, &outs, &pack.vectors, e.dataset.mask, source, true); + for f in pack.files.iter_mut() { + if f.0 == "vectors.json" { + f.1 = vj.clone(); + } + } + pack.write_to(dir).unwrap(); + outs +} + +fn contract(p: &Program) { + assert_eq!(p.instrs.len(), INSTR_COUNT); + assert_eq!(p.instrs.iter().filter(|i| i.op == Op::Load).count() + p.instrs.iter().filter(|i| i.op == Op::Scratch).count(), 16); + assert_eq!(p.instrs.iter().filter(|i| i.op == Op::Scratch).count(), p.class.scratch_slots()); + assert!(p.instrs[0].op != Op::Load && p.instrs[0].op != Op::Scratch, "instruction 0 is never a memory op"); + for (k, i) in p.instrs.iter().enumerate() { + assert!(i.src != i.dst, "#{k}: src == dst"); + assert!((1..=31).contains(&i.rot), "#{k}: rot {}", i.rot); + assert!([1u8, 2, 4, 8, 16].contains(&i.mask), "#{k}: mask {}", i.mask); + assert!(i.dst < 8 && i.src < 8 && i.src2 < 8); + assert_eq!(i.width, 1, "#{k}: a scratch class reads one-word loads"); + } + assert!(igneum_pow::accept::check(p).is_ok(), "an accepted program"); +} + +#[test] +fn fuzz_scr_programs_cpu() { + let n: usize = std::env::var("IGNEUM_SCRATCH_FUZZ").ok().and_then(|s| s.parse().ok()).unwrap_or(200); + let out = std::env::var("IGNEUM_SCRATCH_PACKS_OUT").ok().map(PathBuf::from); + let mut rng = SplitMix64::new(0x6967_6e65_756d_2d73); // "igneum-s" + let day = "2026-10-03"; + let closed = DatasetSource::new(day, DatasetMode::ClosedForm, 28); + // memory-hard sources per size, built once each (the cache fill is 0.2 s); only when packs are written + let mut mh: HashMap = HashMap::new(); + let mut manifest = String::from("pack\tclass\tlog2\tprogram_id\tscratch_ops_per_hash\tbases\n"); + let mut per_class: HashMap = HashMap::new(); + let mut units = 0usize; + let mut wraps = 0usize; + if let Some(dir) = &out { + std::fs::create_dir_all(dir).unwrap(); + // the edge packs first: 64 MiB datasets (no dataset load in them), the four edge bases + for kb in [32u8, 128] { + for (name, _what, p) in edge_programs(kb) { + let log2 = 24; + let ds = mh.remove(&log2).unwrap_or_else(|| DatasetSource::new(day, DatasetMode::MemoryHard, log2)); + let e = Epoch { program: p, dataset: ds }; + let pack_name = format!("edge-{name}-k{kb}"); + write_pack_with_bases(&dir.join(&pack_name), &e, day, &EDGE_BASES, "igneum-pow tests/scratch.rs edge"); + manifest.push_str(&format!( + "{pack_name}\t{}\t{log2}\t{:016x}\t{}\t{}\n", + e.program.class.name(), + e.program.program_id(), + e.program.scratch_ops_per_hash(), + EDGE_BASES.iter().map(|b| format!("{b}")).collect::>().join(",") + )); + mh.insert(log2, e.dataset); + } + } + } + for i in 0..n { + let name = CLASSES[rng.below(CLASSES.len() as u64) as usize]; + let c = class(name); + let seed = format!("igneum-scratch-fuzz/{i}"); + let p = generate_class(&seed, c); + contract(&p); + *per_class.entry(name.to_string()).or_insert(0) += 1; + // four bases: one inside a 256-nonce batch (in-batch check on the GPU), one straddling 2^31, one in + // the last 256 nonces (the unit wraps past 2^32 or ends on it), one uniform + let b0 = (rng.below(8) as u32) * 32; + let b1 = 0x8000_0000u32.wrapping_sub(256).wrapping_add((rng.below(16) as u32) * 32); + let b2 = 0xffff_ff00u32.wrapping_add((rng.below(8) as u32) * 32); + let b3 = (rng.next() as u32) & !31; + let bases = [b0, b1, b2, b3]; + // an aligned unit never straddles 2^32 (spec 1.9); the top unit ends on 0xffffffff and the persistent + // kernel's unit sequence wraps inside a launch, which the Metal run checks with packbench --batch-base + wraps += bases.iter().filter(|&&b| b >= 0xffff_ff00).count(); + // the CPU: the interpreter is deterministic and every scratch event is inside the lane's slots + for &b in &bases { + let (r1, ev) = interpret_warp_scratch(&p, &p.seed, b, &closed, true); + let r2 = interpret_warp_scratch(&p, &p.seed, b, &closed, false).0; + assert_eq!(r1.hashes, r2.hashes); + assert_eq!(ev.len(), p.scratch_ops_per_hash() * LANES); + assert!(ev.iter().all(|e: &ScratchEvent| e.slot < c.scratch_slots_per_lane() as u32)); + units += 1; + } + if let Some(dir) = &out { + let log2 = [24u32, 26, 28][rng.below(3) as usize]; + let ds = mh.remove(&log2).unwrap_or_else(|| DatasetSource::new(day, DatasetMode::MemoryHard, log2)); + let e = Epoch { program: p, dataset: ds }; + let pack_name = format!("fuzz-{i:03}-{name}-l{log2}"); + write_pack_with_bases(&dir.join(&pack_name), &e, day, &bases, "igneum-pow tests/scratch.rs fuzz"); + manifest.push_str(&format!( + "{pack_name}\t{name}\t{log2}\t{:016x}\t{}\t{}\n", + e.program.program_id(), + e.program.scratch_ops_per_hash(), + bases.iter().map(|b| format!("{b}")).collect::>().join(",") + )); + mh.insert(log2, e.dataset); + } else { + let _ = rng.below(3); + } + } + let mut classes: Vec<_> = per_class.iter().collect(); + classes.sort(); + println!("fuzz: {n} programs, {units} units on the CPU, {wraps} units in the top 256 nonces, classes {classes:?}"); + assert_eq!(units, 4 * n); + assert_eq!(wraps, n, "every program has a unit in the top 256 nonces"); + if let Some(dir) = &out { + std::fs::write(dir.join("manifest.tsv"), manifest).unwrap(); + println!("packs written to {}", dir.display()); + } +} + +/// The fold and rewrite, restated: a slot after `d` dependent RMWs holds 96 bits that are a function of the fill +/// (3 words, a pure function of nonce, slot and seed) and the `d` fold values; a chip that keeps the `d` fold +/// values (32 bits each) instead of the 96-bit slot recomputes the slot in `d` rewrites. This test pins the +/// arithmetic the analysis uses (question 2): the replay from the fold values reproduces the slot. +#[test] +fn slot_is_replayable_from_its_fold_values() { + let seed = seed_words_from_bytes(b"igneum-genesis"); + let (base, lane, slot) = (0x1234_5600u32, 5u32, 17u32); + let fill = [scratch_fill(&seed, base, lane, slot, 0), scratch_fill(&seed, base, lane, slot, 1), scratch_fill(&seed, base, lane, slot, 2)]; + let mut rng = SplitMix64::new(99); + let dsts: Vec = (0..64).map(|_| rng.next() as u32).collect(); + // the honest sequence: read, fold, rewrite, 64 times + let mut w = fill; + let mut xs = Vec::new(); + for &d in &dsts { + let x = fold_words(d, &w); + xs.push(x); + w = scratch_rewrite(x, &w); + } + // the replay: from the fill and the stored fold values alone + let mut w2 = fill; + for &x in &xs { + w2 = scratch_rewrite(x, &w2); + } + assert_eq!(w, w2); + // and nothing shorter: the fold value at step d depends on the slot content at step d, which depends on + // every earlier fold value (drop one and the chain diverges) + let mut w3 = fill; + for (i, &x) in xs.iter().enumerate() { + if i != 10 { + w3 = scratch_rewrite(x, &w3); + } + } + assert_ne!(w, w3); +} diff --git a/proto-metal/packbench.swift b/proto-metal/packbench.swift index 75b6fc4f3..dfa07e44c 100644 --- a/proto-metal/packbench.swift +++ b/proto-metal/packbench.swift @@ -5,7 +5,7 @@ // against the Rust CPU reference and timed without a Swift mirror of the generator. One file, no packages. // // swiftc -O -target arm64-apple-macos11 -o packbench packbench.swift -framework Metal -// ./packbench --pack [--batches 5] [--batch-log2 24] [--group 256] [--warps 2048] +// ./packbench --pack [--batches 5] [--batch-log2 24] [--group 256] [--warps 2048] [--batch-base 0] // // Prints one RESULT line per run: vectors, cache and dataset checks, the batch fingerprint (FNV-1a 64 over the 2^B // outputs at base nonce 0) and MH/s by wall and by GPU time. Variant 5 packs (IGNEUM_PERSISTENT_WARPS) are launched @@ -16,7 +16,7 @@ import Metal func nowMs() -> Double { return Double(DispatchTime.now().uptimeNanoseconds) / 1e6 } func fail(_ m: String) -> Never { print("FAIL: \(m)"); exit(1) } -struct Opts { var pack = ""; var batches = 5; var batchLog2 = 24; var group = 256; var warps = 2048 } +struct Opts { var pack = ""; var batches = 5; var batchLog2 = 24; var group = 256; var warps = 2048; var batchBase: UInt32 = 0 } var opts = Opts() var args = Array(CommandLine.arguments.dropFirst()) while !args.isEmpty { @@ -28,6 +28,7 @@ while !args.isEmpty { case "--batch-log2": opts.batchLog2 = Int(next())! case "--group": opts.group = Int(next())! case "--warps": opts.warps = Int(next())! + case "--batch-base": opts.batchBase = UInt32(next())! // base nonce of the fingerprint batch (default 0; a base near 2^32 makes the persistent unit sequence wrap inside the launch) default: fail("unknown argument \(a)") } } @@ -183,12 +184,13 @@ for (i, base) in vecBases.enumerated() { var nonces = 1 << opts.batchLog2 if persistent { let unit = 32 * min(warpsN, nonces / 32); nonces = (nonces / unit) * unit } let out = device.makeBuffer(length: nonces * 8, options: .storageModeShared)! -let (warmWall, warmGpu) = run { enc in encodeHash(enc, out: out, base: 0, nonces: nonces, group: opts.group) } +let (warmWall, warmGpu) = run { enc in encodeHash(enc, out: out, base: opts.batchBase, nonces: nonces, group: opts.group) } let outPtr = out.contents().bindMemory(to: UInt64.self, capacity: nonces) var batchVecPass = 0, batchVecN = 0 -for (i, base) in vecBases.enumerated() where Int(base) + 32 <= nonces { +for (i, base) in vecBases.enumerated() where Int(base &- opts.batchBase) + 32 <= nonces { // the batch window, wrapping past 2^32 + let off = Int(base &- opts.batchBase) batchVecN += 1 - if (0..<32).allSatisfy({ outPtr[Int(base) + $0] == vecOuts[i][$0] }) { batchVecPass += 1 } + if (0..<32).allSatisfy({ outPtr[off + $0] == vecOuts[i][$0] }) { batchVecPass += 1 } } let fingerprint = fnv1a64(out.contents(), nonces * 8) // Timed batches @@ -205,5 +207,5 @@ print("device \(device.name); compile \(String(format: "%.0f", compileMs)) ms; c print("cache FNV-1a 64 \(String(format: "%016llx", cacheFnv)) \(cacheOk ? "PASS" : "FAIL"); dataset head and last \(dsOk ? "PASS" : "FAIL"); vectors standalone \(vecPass)/\(vecBases.count), in batch \(batchVecPass)/\(batchVecN)") print("warm-up batch \(nonces) hashes: \(String(format: "%.1f", warmGpu)) ms GPU, \(String(format: "%.1f", warmWall)) ms wall") let overall = cacheOk && dsOk && vecPass == vecBases.count && batchVecPass == batchVecN -print("RESULT pack=\(packName) class=\(className) device=\(device.name.replacingOccurrences(of: " ", with: "_")) group=\(opts.group) warps=\(warpsN) arena_mib=\(persistent ? warpsN : 0) nonces=\(nonces) batches=\(opts.batches) vectors=\(vecPass)/\(vecBases.count) batch_vectors=\(batchVecPass)/\(batchVecN) cache=\(cacheOk ? "PASS" : "FAIL") dataset=\(dsOk ? "PASS" : "FAIL") fingerprint=\(String(format: "%016llx", fingerprint)) mhs_gpu=\(String(format: "%.3f", mhsGpu)) mhs_wall=\(String(format: "%.3f", mhsWall)) loads=\(loadsPerHash) bytes=\(bytesPerHash) scratch_ops=\(scratchOps * 8) overall=\(overall ? "PASS" : "FAIL")") +print("RESULT pack=\(packName) class=\(className) device=\(device.name.replacingOccurrences(of: " ", with: "_")) group=\(opts.group) warps=\(warpsN) batch_base=\(opts.batchBase) arena_mib=\(persistent ? warpsN : 0) nonces=\(nonces) batches=\(opts.batches) vectors=\(vecPass)/\(vecBases.count) batch_vectors=\(batchVecPass)/\(batchVecN) cache=\(cacheOk ? "PASS" : "FAIL") dataset=\(dsOk ? "PASS" : "FAIL") fingerprint=\(String(format: "%016llx", fingerprint)) mhs_gpu=\(String(format: "%.3f", mhsGpu)) mhs_wall=\(String(format: "%.3f", mhsWall)) loads=\(loadsPerHash) bytes=\(bytesPerHash) scratch_ops=\(scratchOps * 8) overall=\(overall ? "PASS" : "FAIL")") exit(overall ? 0 : 1) From d0018cf18aba373b4ac46f041cb3784229d9c3e8 Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:18:35 +0100 Subject: [PATCH 008/311] read-width: OpenCL scratch kernels declare the __local exchange buffer before the unit loop (AMD's compiler requires the outermost scope; round 3 on the 9070 XT); seven scratch packs re-exported, vectors unchanged Co-Authored-By: Claude Fable 5.1 --- igneum-pow/src/emit.rs | 85 +++++++++++++------ proto-cuda/packs-readwidth/scr0k32/kernel.cl | 12 +-- .../packs-readwidth/scr0k32/kernel_bound.cl | 26 +++--- proto-cuda/packs-readwidth/scr2k128/kernel.cl | 12 +-- .../packs-readwidth/scr2k128/kernel_bound.cl | 26 +++--- proto-cuda/packs-readwidth/scr2k32/kernel.cl | 12 +-- .../packs-readwidth/scr2k32/kernel_bound.cl | 26 +++--- proto-cuda/packs-readwidth/scr4k128/kernel.cl | 12 +-- .../packs-readwidth/scr4k128/kernel_bound.cl | 26 +++--- proto-cuda/packs-readwidth/scr4k32/kernel.cl | 12 +-- .../packs-readwidth/scr4k32/kernel_bound.cl | 26 +++--- proto-cuda/packs-readwidth/scr8k128/kernel.cl | 12 +-- .../packs-readwidth/scr8k128/kernel_bound.cl | 26 +++--- proto-cuda/packs-readwidth/scr8k32/kernel.cl | 12 +-- .../packs-readwidth/scr8k32/kernel_bound.cl | 26 +++--- 15 files changed, 194 insertions(+), 157 deletions(-) diff --git a/igneum-pow/src/emit.rs b/igneum-pow/src/emit.rs index 77029f280..5c8965b77 100644 --- a/igneum-pow/src/emit.rs +++ b/igneum-pow/src/emit.rs @@ -157,6 +157,14 @@ fn scratch_stmt(dialect: CoreDialect, d: &str, a: &str, slot_mask: u32) -> Strin /// lottery hash's text is unchanged: `gid` is the unit's first output index plus the lane. The host MUST launch /// `groups` as a multiple of N (a uniform trip count: the OpenCL local-memory exchange carries a barrier). fn persistent_prologue(dialect: CoreDialect, words_per_lane: usize) -> String { + let (a, b) = persistent_prologue_parts(dialect, words_per_lane); + a + &b +} + +/// The prologue in two parts: the warp's identity and arena, then the unit loop. OpenCL C requires a `__local` +/// variable at the outermost scope of the kernel (AMD's compiler enforces it, 5 October 2026, round 3 on the +/// 9070 XT), so the OpenCL kernels declare the exchange buffer between the two parts. +fn persistent_prologue_parts(dialect: CoreDialect, words_per_lane: usize) -> (String, String) { let (u, tid, nthreads, ptr) = match dialect { CoreDialect::Metal => ("uint", "tid", "nthreads", "device uint*"), CoreDialect::Cuda => ("uint32_t", "(blockIdx.x * blockDim.x + threadIdx.x)", "(gridDim.x * blockDim.x)", "uint32_t*"), @@ -167,11 +175,12 @@ fn persistent_prologue(dialect: CoreDialect, words_per_lane: usize) -> String { s.push_str(&format!(" {u} warp_ = {tid} >> 5;\n")); s.push_str(&format!(" {u} nwarps_ = {nthreads} >> 5;\n")); s.push_str(&format!(" {ptr} arena = scratch + ((size_t)warp_ * 32u + lane) * {words_per_lane}u;\n")); - s.push_str(&format!(" for ({u} g_ = warp_; g_ < groups; g_ += nwarps_) {{\n")); - s.push_str(&format!(" {u} gid = g_ * 32u + lane;\n")); - s.push_str(&format!(" {u} gbase = baseNonce + g_ * 32u;\n")); - s.push_str(&format!(" {u} tag = salt + g_;\n")); - s + let mut l = String::new(); + l.push_str(&format!(" for ({u} g_ = warp_; g_ < groups; g_ += nwarps_) {{\n")); + l.push_str(&format!(" {u} gid = g_ * 32u + lane;\n")); + l.push_str(&format!(" {u} gbase = baseNonce + g_ * 32u;\n")); + l.push_str(&format!(" {u} tag = salt + g_;\n")); + (s, l) } /// The scratch lines of program.h (variant 5). @@ -875,21 +884,36 @@ pub fn opencl_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String { ); let scratch_args = if p.has_scratch() { ", __global uint* scratch, uint groups, uint salt" } else { "" }; s.push_str(&format!("IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw{scratch_args}) {{\n")); + let (setup, unit_loop) = persistent_prologue_parts(CoreDialect::OpenCl, p.class.scratch_words_per_lane()); if p.has_scratch() { - s.push_str(&persistent_prologue(CoreDialect::OpenCl, p.class.scratch_words_per_lane())); + s.push_str(&setup); } else { s.push_str(" uint gid = (uint)get_global_id(0);\n"); } s.push_str(" uint lid = (uint)get_local_id(0);\n"); - s.push_str(" uint nonce = baseNonce + gid;\n"); - s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n"); - s.push_str(" uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];\n"); - s.push_str("#if IGNEUM_EXCHANGE == 0\n"); - s.push_str(" IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);\n"); - s.push_str(" uint xk = 0u;\n"); - s.push_str("#else\n"); - s.push_str(" (void)lid;\n"); - s.push_str("#endif\n"); + if p.has_scratch() { + // the __local exchange buffer must sit at the kernel's outermost scope: declare it, then open the unit loop + s.push_str("#if IGNEUM_EXCHANGE == 0\n"); + s.push_str(" IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);\n"); + s.push_str(" uint xk = 0u;\n"); + s.push_str("#else\n"); + s.push_str(" (void)lid;\n"); + s.push_str("#endif\n"); + s.push_str(&unit_loop); + s.push_str(" uint nonce = baseNonce + gid;\n"); + s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n"); + s.push_str(" uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];\n"); + } else { + s.push_str(" uint nonce = baseNonce + gid;\n"); + s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n"); + s.push_str(" uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];\n"); + s.push_str("#if IGNEUM_EXCHANGE == 0\n"); + s.push_str(" IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);\n"); + s.push_str(" uint xk = 0u;\n"); + s.push_str("#else\n"); + s.push_str(" (void)lid;\n"); + s.push_str("#endif\n"); + } if p.has_wide() { s.push_str(" uint lane = lid & 31u;\n uint wmask = mask & ~31u;\n"); } @@ -1007,20 +1031,33 @@ pub fn opencl_kernel(p: &Program, memhard: Option<&MixParams>) -> String { s.push_str(&scratch_prelude(p, CoreDialect::OpenCl)); let scratch_args = if p.has_scratch() { ", __global uint* scratch, uint groups, uint salt" } else { "" }; s.push_str(&format!("IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask{scratch_args}) {{\n")); + let (setup, unit_loop) = persistent_prologue_parts(CoreDialect::OpenCl, p.class.scratch_words_per_lane()); if p.has_scratch() { - s.push_str(&persistent_prologue(CoreDialect::OpenCl, p.class.scratch_words_per_lane())); + s.push_str(&setup); } else { s.push_str(" uint gid = (uint)get_global_id(0);\n"); } s.push_str(" uint lid = (uint)get_local_id(0);\n"); - s.push_str(" uint nonce = baseNonce + gid;\n"); - s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n"); - s.push_str("#if IGNEUM_EXCHANGE == 0\n"); - s.push_str(" IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);\n"); - s.push_str(" uint xk = 0u;\n"); - s.push_str("#else\n"); - s.push_str(" (void)lid;\n"); - s.push_str("#endif\n"); + if p.has_scratch() { + s.push_str("#if IGNEUM_EXCHANGE == 0\n"); + s.push_str(" IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);\n"); + s.push_str(" uint xk = 0u;\n"); + s.push_str("#else\n"); + s.push_str(" (void)lid;\n"); + s.push_str("#endif\n"); + s.push_str(&unit_loop); + s.push_str(" uint nonce = baseNonce + gid;\n"); + s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n"); + } else { + s.push_str(" uint nonce = baseNonce + gid;\n"); + s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n"); + s.push_str("#if IGNEUM_EXCHANGE == 0\n"); + s.push_str(" IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);\n"); + s.push_str(" uint xk = 0u;\n"); + s.push_str("#else\n"); + s.push_str(" (void)lid;\n"); + s.push_str("#endif\n"); + } if p.has_wide() { s.push_str(" uint lane = lid & 31u;\n uint wmask = mask & ~31u;\n"); } diff --git a/proto-cuda/packs-readwidth/scr0k32/kernel.cl b/proto-cuda/packs-readwidth/scr0k32/kernel.cl index 040d671a2..517a52f3f 100644 --- a/proto-cuda/packs-readwidth/scr0k32/kernel.cl +++ b/proto-cuda/packs-readwidth/scr0k32/kernel.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] diff --git a/proto-cuda/packs-readwidth/scr0k32/kernel_bound.cl b/proto-cuda/packs-readwidth/scr0k32/kernel_bound.cl index 057f228d0..1b033b3cb 100644 --- a/proto-cuda/packs-readwidth/scr0k32/kernel_bound.cl +++ b/proto-cuda/packs-readwidth/scr0k32/kernel_bound.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] @@ -296,20 +296,20 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; - uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } diff --git a/proto-cuda/packs-readwidth/scr2k128/kernel.cl b/proto-cuda/packs-readwidth/scr2k128/kernel.cl index 8fc591da5..706ed7e3e 100644 --- a/proto-cuda/packs-readwidth/scr2k128/kernel.cl +++ b/proto-cuda/packs-readwidth/scr2k128/kernel.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] diff --git a/proto-cuda/packs-readwidth/scr2k128/kernel_bound.cl b/proto-cuda/packs-readwidth/scr2k128/kernel_bound.cl index b08bc77a8..314afb38c 100644 --- a/proto-cuda/packs-readwidth/scr2k128/kernel_bound.cl +++ b/proto-cuda/packs-readwidth/scr2k128/kernel_bound.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] @@ -296,20 +296,20 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; - uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } diff --git a/proto-cuda/packs-readwidth/scr2k32/kernel.cl b/proto-cuda/packs-readwidth/scr2k32/kernel.cl index 6befdde06..59e009d77 100644 --- a/proto-cuda/packs-readwidth/scr2k32/kernel.cl +++ b/proto-cuda/packs-readwidth/scr2k32/kernel.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] diff --git a/proto-cuda/packs-readwidth/scr2k32/kernel_bound.cl b/proto-cuda/packs-readwidth/scr2k32/kernel_bound.cl index 9fb87c712..fa26b1c76 100644 --- a/proto-cuda/packs-readwidth/scr2k32/kernel_bound.cl +++ b/proto-cuda/packs-readwidth/scr2k32/kernel_bound.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] @@ -296,20 +296,20 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; - uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } diff --git a/proto-cuda/packs-readwidth/scr4k128/kernel.cl b/proto-cuda/packs-readwidth/scr4k128/kernel.cl index 3917a927d..babcdd9ed 100644 --- a/proto-cuda/packs-readwidth/scr4k128/kernel.cl +++ b/proto-cuda/packs-readwidth/scr4k128/kernel.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] diff --git a/proto-cuda/packs-readwidth/scr4k128/kernel_bound.cl b/proto-cuda/packs-readwidth/scr4k128/kernel_bound.cl index 669eb1d51..93b2e735f 100644 --- a/proto-cuda/packs-readwidth/scr4k128/kernel_bound.cl +++ b/proto-cuda/packs-readwidth/scr4k128/kernel_bound.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] @@ -296,20 +296,20 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; - uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } diff --git a/proto-cuda/packs-readwidth/scr4k32/kernel.cl b/proto-cuda/packs-readwidth/scr4k32/kernel.cl index 00f785968..b1c967eee 100644 --- a/proto-cuda/packs-readwidth/scr4k32/kernel.cl +++ b/proto-cuda/packs-readwidth/scr4k32/kernel.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] diff --git a/proto-cuda/packs-readwidth/scr4k32/kernel_bound.cl b/proto-cuda/packs-readwidth/scr4k32/kernel_bound.cl index 46034e31b..1090d77b8 100644 --- a/proto-cuda/packs-readwidth/scr4k32/kernel_bound.cl +++ b/proto-cuda/packs-readwidth/scr4k32/kernel_bound.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] @@ -296,20 +296,20 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; - uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } diff --git a/proto-cuda/packs-readwidth/scr8k128/kernel.cl b/proto-cuda/packs-readwidth/scr8k128/kernel.cl index 03840512b..09bfaa3c4 100644 --- a/proto-cuda/packs-readwidth/scr8k128/kernel.cl +++ b/proto-cuda/packs-readwidth/scr8k128/kernel.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] diff --git a/proto-cuda/packs-readwidth/scr8k128/kernel_bound.cl b/proto-cuda/packs-readwidth/scr8k128/kernel_bound.cl index 9cb72a38f..ae27d8663 100644 --- a/proto-cuda/packs-readwidth/scr8k128/kernel_bound.cl +++ b/proto-cuda/packs-readwidth/scr8k128/kernel_bound.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] @@ -296,20 +296,20 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 1024u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; - uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } diff --git a/proto-cuda/packs-readwidth/scr8k32/kernel.cl b/proto-cuda/packs-readwidth/scr8k32/kernel.cl index 54168b39f..39cddf512 100644 --- a/proto-cuda/packs-readwidth/scr8k32/kernel.cl +++ b/proto-cuda/packs-readwidth/scr8k32/kernel.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] diff --git a/proto-cuda/packs-readwidth/scr8k32/kernel_bound.cl b/proto-cuda/packs-readwidth/scr8k32/kernel_bound.cl index 922739e4a..a7356de52 100644 --- a/proto-cuda/packs-readwidth/scr8k32/kernel_bound.cl +++ b/proto-cuda/packs-readwidth/scr8k32/kernel_bound.cl @@ -186,19 +186,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1] { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2] { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3] @@ -296,20 +296,20 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon uint warp_ = (uint)get_global_id(0) >> 5; uint nwarps_ = (uint)get_global_size(0) >> 5; __global uint* arena = scratch + ((size_t)warp_ * 32u + lane) * 256u; - for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { - uint gid = g_ * 32u + lane; - uint gbase = baseNonce + g_ * 32u; - uint tag = salt + g_; uint lid = (uint)get_local_id(0); - uint nonce = baseNonce + gid; - uint r0, r1, r2, r3, r4, r5, r6, r7; - uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; #if IGNEUM_EXCHANGE == 0 IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP); uint xk = 0u; #else (void)lid; #endif + for (uint g_ = warp_; g_ < groups; g_ += nwarps_) { + uint gid = g_ * 32u + lane; + uint gbase = baseNonce + g_ * 32u; + uint tag = salt + g_; + uint nonce = baseNonce + gid; + uint r0, r1, r2, r3, r4, r5, r6, r7; + uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7]; { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; } { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; } { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; } From e6085c66631a0a2f8996dcc3e99fa9fd91651782 Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:20:46 +0000 Subject: [PATCH 009/311] Counter ASIC 2.0 layer 6, revision 2: shipped cache-die density (AMD V-Cache 64 MB on 41 mm^2 at 7 nm, Hot Chips 33) as the headline column beside the bit-cell lower bound; mirror 106 to 164 mm^2 and $30 to $56 per die at 256 MiB; year 0 to 10, hot-table and option tables redone; latency citations section (MEMSYS 2018, Chang 2017, mining-chip memory types) Co-Authored-By: Claude Fable 5.1 --- docs/analysis/sram-mirror.md | 212 ++++++++++++++++++++++------------- 1 file changed, 134 insertions(+), 78 deletions(-) diff --git a/docs/analysis/sram-mirror.md b/docs/analysis/sram-mirror.md index 970239776..e1798fa88 100644 --- a/docs/analysis/sram-mirror.md +++ b/docs/analysis/sram-mirror.md @@ -2,7 +2,13 @@ 5 October 2026 (night), Counter ASIC 2.0 (`docs/plans/counter-asic-2.md`, layer 6), branch `ca2-analysis`. Every figure below is either cited (paper, vendor document, URL, date) or labelled approximate. Nothing here is a measurement of a -chip. Numbers in this file were computed with the arithmetic shown; the script is in section 9. +chip. Numbers in this file were computed with the arithmetic shown; the script is in section 10. + +Revision 2 (same night): the first draft priced the mirror from bit-cell area times a 0.70 array factor. The +coordinator's chip-economics research (sources below) showed that shipped cache-only dies land at about half that +density once assist circuits, redundancy, TSVs, power and test are in. Every table now carries two columns: the +shipped-product density as the headline and the bit-cell figure as the lower bound. The conclusion did not move; the +cost per die rose 2 to 3x. ## 1. The question @@ -29,7 +35,9 @@ design, cache growth is not. The cache is flat at 256 MiB for every year of the closing line names the rule the cache should get ("exceeds what one die can hold, and grows") as a gate 1 decision that has not been taken. -## 3. SRAM bit cell per node, cited +## 3. SRAM density, cited: bit cells per node and shipped cache dies + +### 3.1 Bit cells | Node (vendor) | HD 6T bit cell, um^2 | Raw density, Mbit/mm^2 (1/cell) | Year of volume (approximate) | Source | |---|---|---|---|---| @@ -46,17 +54,35 @@ scaling from N5 to N3E (WikiChip IEDM 2022 article above; Tom's Hardware above; https://newsletter.semianalysis.com/p/tsmcs-3nm-conundrum-does-it-even). N2's nanosheet cell recovers 17% (0.021 to 0.0175 um^2). So across 2020 to 2026 the HD bit cell shrank once, by 17%. -Array efficiency (bit cell to macro). The usable density of a macro is below 1/cell because of word-line and -bit-line drivers, sense amplifiers, decoders and redundancy. The factor used here is 0.70, WikiChip's convention -(their 31.8 Mib/mm^2 for the 0.021 um^2 N3E cell is 1/0.021 x 0.70 in Mib). The two ISSCC 2025 macros bracket it: -TSMC N2 38.1 Mb/mm^2 at a 0.0175 um^2 cell is 67%; Intel 18A 38.1 Mb/mm^2 array density and 34.3 Mb/mm^2 for the -volume macro at a 0.021 um^2 cell are 80% and 72% (ISSCC 2025 29.2, above). Both lie within 10% of 0.70. +Macro density from the bit cell. WikiChip's and SemiAnalysis's convention is bit-cell density times about 0.70 for +the assist and periphery overhead (SemiAnalysis, December 2022: TSMC N5 HD SRAM macro 31.8 Mib/mm^2 after about 30% +assist overhead; WikiChip's 31.8 Mib/mm^2 for the 0.021 um^2 cell is the same arithmetic). The two ISSCC 2025 macros +bracket it: TSMC N2 38.1 Mb/mm^2 at a 0.0175 um^2 cell is 67%; Intel 18A 38.1 Mb/mm^2 array density and 34.3 Mb/mm^2 +for the volume macro at a 0.021 um^2 cell are 80% and 72%. That is a macro on a test chip. It is the LOWER BOUND on +die area, not the die. + +### 3.2 Shipped cache dies (what a whole die of SRAM really holds) + +| Product | SRAM | Die | Node | MB per mm^2 | Source | +|---|---|---|---|---|---| +| AMD 3D V-Cache (Zen 3 SRAM chiplet) | 64 MB | 41 mm^2 | TSMC 7 nm | 1.56 | AMD at Hot Chips 33, reported by Tom's Hardware, August 2021, https://www.tomshardware.com/news/amd-unveils-more-ryzen-3d-packaging-and-v-cache-details-at-hot-chips ("the 3D V-Cache SRAM measures 41 mm^2", "64 MB of 7 nm SRAM"); the densest cache-only die that has shipped | +| Graphcore GC200 (with compute) | 900 MB | 823 mm^2 | 7 nm | 1.09 | coordinator's chip-economics research, 5 October 2026 (vendor figures) | +| Groq TSP | 220 MB | 725 mm^2 | 14 nm | 0.30 | same | + +The V-Cache die is a pure SRAM die with its TSVs, redundancy, test and power: 1.56 MB/mm^2 at N7 against the bit-cell +figure 37.0 Mbit/mm^2 = 4.6 MB/mm^2 and the 0.70-macro figure 3.2 MB/mm^2. The shipped die is 0.48 of the macro +figure. The headline column below scales the V-Cache density to other nodes by the bit-cell ratio (0.027 / cell), an +approximation that assumes the periphery and TSV overheads scale with the cell, which they do not fully (so the +headline column is itself slightly optimistic for the attacker at N5 and below). + +### 3.3 GPU on-die SRAM, the reticle, wafer prices GPU on-die SRAM for scale: the RTX 5090 carries 96 MB of L2 (98,304 KB) on a 750 mm^2 TSMC 4N die with 92.2 billion transistors; the full GB202 has 128 MB; the RTX 4090 had 72 MB and the RTX 3090 6 MB (NVIDIA, "RTX Blackwell GPU Architecture" whitepaper v1.1, appendix table "L2 Cache Size", https://images.nvidia.com/aem-dam/Solutions/geforce/blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf). -At the N5-class cell and 0.70 that L2 is about 24 mm^2 of the 750 (3%), approximate. The RX 9070 XT carries 64 MB -of Infinity Cache plus 8 MB of L2 (vendor figures, approximate, bench-log "the 9070 XT on the eGPU"). +At the V-Cache density scaled to N5 (2.0 MB/mm^2) that L2 is about 48 mm^2 of the 750 (6%), approximate. The +RX 9070 XT carries 64 MB of Infinity Cache plus 8 MB of L2 (vendor figures, approximate, bench-log "the 9070 XT on the +eGPU"). Reticle: the EUV field is 26 x 33 mm = 858 mm^2, about 830 mm^2 usable after scribe lanes (SemiAnalysis, "Die Size And Reticle Conundrum", https://newsletter.semianalysis.com/p/die-size-and-reticle-conundrum-cost ; WikiChip "Mask", @@ -66,45 +92,51 @@ Wafer prices (approximate; TSMC publishes none, every figure is supply-chain rep about $20,000 (Silicon Analysts, "Wafer Pricing by Node", September 2026, https://siliconanalysts.com/data/wafer-pricing); N2 about $30,000 (Tom's Hardware, https://www.tomshardware.com/tech-industry/semiconductors/tsmc-could-charge-up-to-usd45-000-for-1-6nm-wafers-rumors-allege-a-50-percent-increase-in-pricing-over-prior-gen-wafers). -## 4. Die area to mirror the cache, per node +## 4. Die area to mirror the cache, per node, two columns -Area = bits / (raw density x 0.70). The columns are the 256 MiB cache alone, the cache plus each hot-table size of -layer 5 (32, 64, 96 MB taken as MiB), and the larger caches of the options in section 6. +Headline = V-Cache density (41 mm^2 per 64 MiB at N7) scaled by the bit-cell ratio. Lower bound = bits / (raw +density x 0.70). Columns: the 256 MiB cache alone, the cache plus the 96 MB hot table of layer 5 (as MiB), and the +larger caches of the options in section 7. Area in mm^2; a figure over 830 is split into the dies shown. -| Node | Macro Mbit/mm^2 at 0.70 | 256 MiB | 256 + 32 | 256 + 64 | 256 + 96 | 512 MiB | 1 GiB | 2 GiB | 4 GiB | -|---|---|---|---|---|---|---|---|---|---| -| N7 | 25.9 | 83 mm^2 | 93 | 104 | 114 | 166 | 331 | 663 | 1,325 (2 dies) | -| N5 | 33.3 | 64 | 72 | 81 | 89 | 129 | 258 | 515 | 1,031 (2 dies) | -| N3B | 35.2 | 61 | 69 | 76 | 84 | 122 | 244 | 488 | 977 (2 dies) | -| N3E, Intel 18A | 33.3 | 64 | 72 | 81 | 89 | 129 | 258 | 515 | 1,031 (2 dies) | -| N2 | 40.0 | 54 | 60 | 67 | 74 | 107 | 215 | 429 | 859 (2 dies) | +| Node | 256 MiB, headline | 256 MiB, lower bound | 256 + 96, headline | 256 + 96, lower bound | 512 MiB, headline / lower | 1 GiB, headline / lower | 4 GiB, headline / lower | +|---|---|---|---|---|---|---|---| +| N7 | 164 | 83 | 226 | 114 | 328 / 166 | 656 / 331 | 2,624 (4 dies) / 1,325 (2 dies) | +| N5 | 128 | 64 | 175 | 89 | 255 / 129 | 510 / 258 | 2,041 (3 dies) / 1,031 (2 dies) | +| N3B | 121 | 61 | 166 | 84 | 242 / 122 | 483 / 244 | 1,934 (3 dies) / 977 (2 dies) | +| N3E, Intel 18A | 128 | 64 | 175 | 89 | 255 / 129 | 510 / 258 | 2,041 (3 dies) / 1,031 (2 dies) | +| N2 | 106 | 54 | 146 | 74 | 213 / 107 | 425 / 215 | 1,701 (3 dies) / 859 (2 dies) | -One reticle (830 mm^2) holds 2.5 GiB of SRAM at N7, 3.2 GiB at N5, N3E and 18A, 3.9 GiB at N2 (same arithmetic). +One reticle (830 mm^2) holds, at the headline density, 1.3 GiB of SRAM at N7, 1.6 GiB at N5, N3E and 18A, 1.9 GiB at +N2 (lower-bound column: 2.5, 3.2, 3.9 GiB). Against the figures the ledger carries: M16's "100 to 300 mm^2" (low end from a 0.02 um^2 cell with overhead, high -end from wafer-scale parts at about 1 MB/mm^2) and the plan's "about 45 mm^2 at a leading node" both bracket the -cited 54 to 64 mm^2; the wafer-scale high end is a different efficiency (Cerebras-class arrays sit beside logic) and -is not the right number for a pure SRAM die. The right figure for the ledger is 54 to 83 mm^2 depending on node, -cited above. +end from wafer-scale parts at about 1 MB per mm^2) brackets the headline 106 to 164 mm^2 well; the plan's "about +45 mm^2 at a leading node" is below even the lower bound and should be read as the bit-cell area with no overhead. +The right figures for the ledger are 106 to 164 mm^2 (shipped density) with 54 to 83 mm^2 as the floor. -## 5. Cost per good die +## 5. Cost per good die, two columns Dies per 300 mm wafer by the usual approximation pi x 150^2 / A minus the edge term pi x 300 / sqrt(2A); yield by Poisson exp(-A x D0) with D0 = 0.1 defects per cm^2 (an assumption, approximate; SRAM arrays carry redundancy so real yield is higher, which lowers these costs). Cost per good die = wafer price / (dies x yield). Packaging, test, the logic beside the SRAM and the design (masks at N5 and below run into the tens of millions of dollars, -approximate) are not in these numbers; they are per-die silicon only. +approximate) are not in these numbers; they are per-die silicon only. Headline / lower bound in each cell. -| Node, wafer price | 256 MiB | 256 + 96 MiB | 1 GiB | 4 GiB (2 dies) | +| Node, wafer price | 256 MiB | 256 + 96 MiB | 1 GiB | 4 GiB | |---|---|---|---|---| -| N7, $9,500 | 83 mm^2, 780 dies, yield 0.92, $13 | $19 | 331 mm^2, 177 dies, 0.72, $75 | $456 | -| N5, $20,000 | 64 mm^2, 1,014 dies, 0.94, $21 | $30 | 258 mm^2, 233 dies, 0.77, $111 | $621 | -| N3E, $20,000 | $21 | $30 | $111 | $621 | -| N2, $30,000 | 54 mm^2, 1,226 dies, 0.95, $26 | $37 | 215 mm^2, 284 dies, 0.81, $131 | $696 | +| N7, $9,500 | 164 mm^2, 379 dies, yield 0.85: $30 / $13 | $44 / $19 | $224 / $75 | $896 (4 dies) / $456 (2 dies) | +| N5, $20,000 | 128 mm^2, 495 dies, 0.88: $46 / $21 | $68 / $30 | $306 / $111 | $1,512 (3 dies) / $621 (2 dies) | +| N3B, $20,000 | 121 mm^2, 524 dies, 0.89: $43 / $20 | $63 / $28 | $280 / $103 | $1,371 (3 dies) / $569 (2 dies) | +| N3E, 18A, $20,000 | $46 / $21 | $68 / $30 | $306 / $111 | $1,512 / $621 | +| N2, $30,000 | 106 mm^2, 600 dies, 0.90: $56 / $26 | $81 / $37 | $343 / $131 | $1,641 (3 dies) / $696 (2 dies) | -Reading. The silicon for a 256 MiB mirror is tens of dollars per die on any node from N7 up. With the hot table it -is still under $40. It was never the SRAM that priced the recompute attacker out; the plan's premise for layer 6 -("the SRAM mirror stays unaffordable") does not hold for the cache as a mirror and did not hold at genesis either. +Reading. The silicon for a 256 MiB mirror is $30 to $56 per die at shipped density (2 to 3x the first draft's +figure), under $90 with the hot table. A funded chip programme pays that without noticing: it was never the SRAM +that priced the recompute attacker out, and the plan's premise for layer 6 ("the SRAM mirror stays unaffordable") +does not hold for the cache as a mirror and did not hold at genesis either. A 1 GiB cache is a 425 to 656 mm^2 die +($224 to $343), affordable too; 4 GiB is a 3 to 4 die part at about $900 to $1,600 of silicon, which is a different +product but not an impossible one (the attacker's problem at that size is the 1,024 dependent cross-die reads per +hash, section 6). ## 6. What the mirror buys the attacker, year by year @@ -112,37 +144,38 @@ From M16 (`docs/analysis/m16-recompute-attacker-2026-10-05.md`): with the cache items per hash at about 1,170 integer operations and 8 dependent 64-byte cache reads each, about 150,000 operations and 1,024 dependent SRAM reads per hash. At a 5090-class integer budget (about 50 T op/s, approximate) that is 0.33 Ghash/s against the honest 141 Mhash/s projected for version 2 programs: 2.4x at equal silicon before any -fixed-function factor, 3x to 6x with one (approximate). The SRAM is 54 to 83 mm^2 of that chip (7 to 11% of a -750 mm^2 die), so the mirror is cheap and the recompute route is bound by integer throughput, not by SRAM. +fixed-function factor, 3x to 6x with one (approximate). The SRAM is 106 to 164 mm^2 of that chip at the headline +density (14 to 22% of a 750 mm^2 die; the m16 model's 13 to 40% band holds), so the mirror is cheap and the recompute +route is bound by integer throughput, not by SRAM. The layer 5 hot table changes nothing in that arithmetic: the hot table is read-only and derived from the day key -like the cache, so a chip mirrors it in the same SRAM (another 32 to 96 MB, 7 to 24 mm^2 at N5) and reads it at -SRAM latency, which is exactly what a GPU's L2 does with it. Layer 5 taxes the DRAM-only chip (the one without SRAM); -it does not tax the SRAM chip. +like the cache, so a chip mirrors it in the same SRAM (another 32 to 96 MB, 24 to 48 mm^2 at N5 headline) and reads +it at SRAM latency, which is exactly what a GPU's L2 does with it. Layer 5 taxes the DRAM-only chip (the one without +SRAM); it does not tax the SRAM chip. Dataset growth does not touch the recompute attacker: the attacker never holds the dataset. It taxes the partial-store attacker (O-1.6, the time-memory curve, not drawn) and the honest card. Year by year under the schedule as it stands (flat 256 MiB), the mirror's area at the best node available that -year. Node years are approximate; the density trend from 2018 to 2025 is 37.0 to 57.1 Mbit/mm^2 raw, 1.54x in -7 years, about 6% per year, and it came in one step (N2); the extrapolation past 2026 assumes that average holds -(approximate, and optimistic for the attacker: A16 and A14 have no disclosed SRAM cell yet). +year, headline density. Node years are approximate; the density trend from 2018 to 2025 is 37.0 to 57.1 Mbit/mm^2 +raw, 1.54x in 7 years, about 6% per year, and it came in one step (N2); the extrapolation past 2026 assumes that +average holds (approximate, and optimistic for the attacker: A16 and A14 have no disclosed SRAM cell yet). -| Year | Calendar (approximate) | Dataset, GiB | Cache (spec) | Best node, raw Mbit/mm^2 | Mirror of the cache, mm^2 | With a 96 MiB hot table, mm^2 | Mirror as a share of a 750 mm^2 die | +| Year | Calendar (approximate) | Dataset, GiB | Cache (spec) | Best node | Mirror of the cache, headline (lower bound), mm^2 | With a 96 MiB hot table, headline, mm^2 | Mirror as a share of a 750 mm^2 die | |---|---|---|---|---|---|---|---| -| 0 | 2027 | 2.0 | 256 MiB | N2, 57.1 (cited) | 54 | 74 | 7% | -| 1 | 2028 | 2.5 | 256 MiB | N2 or A16, 57 to 61 | 50 to 54 | 69 to 74 | 7% | -| 2 | 2029 | 3.0 | 256 MiB | about 64 (trend) | 48 | 66 | 6% | -| 3 | 2030 | 3.5 | 256 MiB | about 68 | 45 | 62 | 6% | -| 4 | 2031 | 4.0 | 256 MiB | about 72 | 43 | 59 | 6% | -| 5 | 2032 | 4.5 | 256 MiB | about 76 | 40 | 55 | 5% | -| 6 | 2033 | 5.0 | 256 MiB | about 81 | 38 | 52 | 5% | -| 7 | 2034 | 5.5 | 256 MiB | about 86 | 36 | 49 | 5% | -| 8 | 2035 | 6.0 | 256 MiB | about 91 | 34 | 46 | 5% | -| 9 | 2036 | 6.5 | 256 MiB | about 97 | 32 | 44 | 4% | -| 10 | 2037 | 7.0 | 256 MiB | about 102 | 30 | 41 | 4% | +| 0 | 2027 | 2.0 | 256 MiB | N2 (cited) | 106 (54) | 146 | 14% | +| 1 | 2028 | 2.5 | 256 MiB | N2 or A16 | 103 (52) | 142 | 14% | +| 2 | 2029 | 3.0 | 256 MiB | trend | 95 (48) | 130 | 13% | +| 3 | 2030 | 3.5 | 256 MiB | trend | 89 (45) | 123 | 12% | +| 4 | 2031 | 4.0 | 256 MiB | trend | 84 (43) | 116 | 11% | +| 5 | 2032 | 4.5 | 256 MiB | trend | 79 (40) | 109 | 11% | +| 6 | 2033 | 5.0 | 256 MiB | trend | 75 (38) | 103 | 10% | +| 7 | 2034 | 5.5 | 256 MiB | trend | 71 (36) | 97 | 9% | +| 8 | 2035 | 6.0 | 256 MiB | trend | 67 (34) | 92 | 9% | +| 9 | 2036 | 6.5 | 256 MiB | trend | 63 (32) | 87 | 8% | +| 10 | 2037 | 7.0 | 256 MiB | trend | 59 (30) | 82 | 8% | -Reading. A flat cache's mirror shrinks from 7% to 4% of a large die over the decade, and a 5090-class consumer GPU +Reading. A flat cache's mirror shrinks from 14% to 8% of a large die over the decade, and a 5090-class consumer GPU already carries 96 MB of L2 on one die with the full GB202 at 128 MB; at the 2020 to 2025 pace of GPU L2 growth (6 MB, 72 MB, 96 MB on the three NVIDIA flagships in the whitepaper table) a consumer GPU could hold 256 MiB on die within the decade. The spec's own rule for the cache ("must exceed the largest on-chip cache of any card that @@ -152,8 +185,8 @@ DRAM-latency-bound and the recompute route stays a route only a custom chip can ## 7. Answer to the layer 6 question, and the options -Does the flat 256 MiB cache keep the SRAM mirror unaffordable through year 10? No. It is affordable at year 0 -(tens of dollars of silicon per die, section 5) and gets cheaper. What keeps the recompute attacker near 1x is +Does the flat 256 MiB cache keep the SRAM mirror unaffordable through year 10? No. It is affordable at year 0 ($30 to +$56 of silicon per die at shipped density, section 5) and gets cheaper. What keeps the recompute attacker near 1x is M16's integer arithmetic and the mixer-cost lever (4x the mixer cost puts the equal-silicon gain at 0.36x, bounded by the CPU verify gate), not the cache size. The cache size does one other job, keeping the cache larger than any GPU's L2, and that job needs growth. @@ -165,22 +198,23 @@ GPU fill: 0.67 ms per 256 MiB on the 5090 (section 1.8.3), linear. The GPU datas 5090, section 1.8.3) depends on the dataset size, not the cache size; a larger cache spreads the build's 8 dependent reads per item over more memory, which on a GPU means more of them miss L2 and the build slows by some factor between 1x and the L2-to-DRAM latency ratio, which is a measurement to take (approximate; owed). Mirror area is at N2 -(cited density), the node of the first years; at the trend's year-10 density divide by about 1.8. +headline density (lower bound in brackets), the node of the first years; at the trend's year-10 density divide by +about 1.8. -| Option | Rule | Cache at year 0 / 4 / 10 | Mirror at N2, year 0 / 4 / 10 (mm^2) | Dies at year 10 (830 mm^2 reticle) | Verifier fill, one core, year 0 / 10 | Verifier memory, year 10 | GPU cache fill (5090), year 10 | Keeps the cache above a 96 MB L2 at year 10 | Keeps it above a 256 MB L2 | +| Option | Rule | Cache at year 0 / 4 / 10 | Mirror at N2, headline (lower bound), year 0 / 4 / 10, mm^2 | Dies at year 10 (830 mm^2 reticle), headline | Verifier fill, one core, year 0 / 10 | Verifier memory, year 10 | GPU cache fill (5090), year 10 | Keeps the cache above a 96 MB L2 at year 10 | Keeps it above a 256 MB L2 | |---|---|---|---|---|---|---|---|---|---| -| A, as specified | flat 256 MiB | 256 / 256 / 256 MiB | 54 / 54 / 54 | 1 | 0.2 / 0.2 s | 256 MiB | 0.7 ms | yes, 2.7x | no | -| B | cache = dataset / 8 (today's ratio) | 256 / 512 / 896 MiB | 54 / 107 / 188 | 1 | 0.2 / 0.7 s | 896 MiB | 2.3 ms | yes, 9.3x | yes, 3.5x | -| C | cache doubles when the dataset doubles (the dataset's own clock: year 4, then year 12) | 256 / 512 / 512 MiB | 54 / 107 / 107 | 1 | 0.2 / 0.4 s | 512 MiB | 1.3 ms | yes, 5.3x | yes, 2x | -| D | cache = dataset / 4 | 512 / 1,024 / 1,792 MiB | 107 / 215 / 376 | 1 | 0.4 / 1.4 s | 1.75 GiB | 4.7 ms | yes | yes, 7x | -| E, one reticle | cache sized so the mirror exceeds one reticle at the node of the day: 4 GiB at N2 (section 4), growing with density | 4 GiB / about 4.5 / about 7 GiB | 859 / 860 / 860 (by construction) | 2 | 3.2 / 5.6 s | 7 GiB | 11 / 19 ms | yes | yes | +| A, as specified | flat 256 MiB | 256 / 256 / 256 MiB | 106 (54) / 106 / 106 | 1 | 0.2 / 0.2 s | 256 MiB | 0.7 ms | yes, 2.7x | no | +| B | cache = dataset / 8 (today's ratio) | 256 / 512 / 896 MiB | 106 (54) / 213 (107) / 372 (188) | 1 | 0.2 / 0.7 s | 896 MiB | 2.3 ms | yes, 9.3x | yes, 3.5x | +| C | cache doubles when the dataset doubles (the dataset's own clock: year 4, then year 12) | 256 / 512 / 512 MiB | 106 (54) / 213 (107) / 213 (107) | 1 | 0.2 / 0.4 s | 512 MiB | 1.3 ms | yes, 5.3x | yes, 2x | +| D | cache = dataset / 4 | 512 / 1,024 / 1,792 MiB | 213 (107) / 425 (215) / 744 (376) | 1 | 0.4 / 1.4 s | 1.75 GiB | 4.7 ms | yes | yes, 7x | +| E, one reticle | cache sized so the mirror exceeds one reticle at the node of the day: 2 GiB at N2 headline density (section 4; 4 GiB on the lower bound), growing with density | 2 GiB / about 2.3 / about 3.5 GiB | 850 / 850 / 850 (by construction) | 2 | 1.6 / 2.8 s | 3.5 GiB | 5.4 / 9.4 ms | yes | yes | Where the working set enters (coordinator's budget: 1 GiB table + hot table + scratch for every resident warp + buffers under 6 GB on an 8 GB card): the cache is not in the miner's working set at hash time (the dataset is built from it once a day and the cache can be dropped or kept), so options A to D do not move that budget; the dataset's own growth does (2 GiB at genesis, 4 GiB at year 4, 7 GiB at year 10, which is past an 8 GB card at about year 8 on its -own). Option E's 4 GiB cache would have to be built on the card and dropped, which is fine for a 16 GB card and tight -on an 8 GB one at build time (4 GiB cache + 2 GiB dataset + hot table). The per-warp scratch at 170 SMs x 64 warps +own). Option E's 2 GiB cache would have to be built on the card and dropped, which is fine for a 16 GB card and tight +on an 8 GB one at build time (2 GiB cache + 2 GiB dataset + hot table). The per-warp scratch at 170 SMs x 64 warps (approximate, readwidth) is 340 MB at 32 KB and 1.36 GB at 128 KB per warp; with the 1 GiB table, a 96 MB hot table and buffers that is 1.5 to 2.5 GB at the prototype dataset size, 2.5 to 3.5 GB at the 2 GiB genesis size, inside 6 GB either way. @@ -191,37 +225,59 @@ every step, which is the 1.13.3 option (b) argument again), costs the verifier 0 more until year 12, and keeps the cache 2x above a 256 MB GPU L2 if one appears. It does not price a chip out; nothing about cache size does (section 5). The lever that does is the mixer cost multiplier of M16, which is the gate 1 decision to take beside this one. Option B is the same idea in a smooth form and costs the verifier 0.7 s at year 10. -Option E is the only one that makes the mirror a multi-die part and it costs every verifier 3.2 s and 4 GiB at -genesis, which fails the spirit of the 10 ms verify gate (the fill is once a day, but a light node joining pays it on -every day it syncs across). +Option E is the only one that makes the mirror a multi-die part and it costs every verifier 1.6 s and 2 GiB at +genesis (at the headline density; the lower-bound density would ask for 4 GiB and 3.2 s), which fails the spirit of +the 10 ms verify gate (the fill is once a day, but a light node joining pays it on every day it syncs across). Decision for Josh, at gate 1: A, B, C, D or E above, together with M16's mixer multiplier. Nothing here changes a vector today: the cache size is a prototype value of spec 1.16 and the growth rule would be a new sentence in 1.13.3. -## 8. What is cited, what is approximate, what is owed +## 8. Why the latency bound is the property to lean on (citations behind the plan's rule) + +The plan's "what stays true" paragraph says DRAM latency is the same physics for everyone and bandwidth per watt is +what a custom memory chip buys. The sources behind that: + +| Claim | Figure | Source | +|---|---|---| +| Random-access DRAM latency is the same across memory types | Row cycle time 40 to 48 ns across DDR4, GDDR5 and HBM2 | Li, Reddy and Jacob, "A Performance and Power Comparison of Contemporary DRAM Architectures", MEMSYS 2018 (coordinator's chip-economics research, 5 October 2026) | +| Latency does not scale, bandwidth does | DRAM latency improved about 1.3x in two decades while bandwidth improved about 20x | K. Chang, "Understanding and Improving the Latency of DRAM-Based Memory Systems", PhD thesis, CMU, 2017 (same research) | +| No mining chip has bought latency with exotic memory | No shipped mining chip has used HBM or stacked memory; the Ethash chips used DDR3, GDDR6 and undisclosed types | same research; the Ethash chip gain of about 3x in the plan came from bandwidth per watt, not latency | +| The honest hash is latency-bound on every card measured | The hash runs within a few percent of 1/128 of each card's dependent random-read ceiling (5090, 9070 XT, M5 Max) | `docs/bench-log.md`, "the 9070 XT on the eGPU", 5 October 2026 (measured) | + +Reading for layer 6: an SRAM mirror beats DRAM latency by about 10x per read (a 64 MiB buffer inside the 9070 XT's +Infinity Cache chased at 9.2 G loads/s against 2.5 in GDDR6, the same bench-log entry; the 5090's L2 at 5.8x the +hash rate of its 1 GiB dataset, M16), which is why the recompute attacker is bound by the 1,024 dependent SRAM reads +and the 150,000 integer operations per hash and not by the SRAM's size or price. The cache size decides whether the +mirror is one die or several (section 4); it does not decide whether the mirror exists. + +## 9. What is cited, what is approximate, what is owed | Item | Status | |---|---| -| Bit cells for N7, N5, N3B, N3E, N2, Intel 18A | cited (section 3) | +| Bit cells for N7, N5, N3B, N3E, N2, Intel 18A | cited (section 3.1) | +| Shipped cache-die density (AMD V-Cache 64 MB on 41 mm^2 at 7 nm; Graphcore GC200; Groq TSP) | cited (section 3.2; V-Cache checked against Tom's Hardware's Hot Chips 33 report, 5 October 2026; the Graphcore and Groq rows are from the coordinator's research and were not re-checked tonight) | +| Scaling the V-Cache density to other nodes by the bit-cell ratio | approximate, stated | | Samsung SF2 or SF3 bit cell | not found; left out | -| Array efficiency 0.70 | WikiChip's convention, bracketed by two ISSCC 2025 macros (67 to 80%) | +| Array efficiency 0.70 | WikiChip's and SemiAnalysis's convention, bracketed by two ISSCC 2025 macros (67 to 80%); a macro figure, used only as the lower bound | | Wafer prices | approximate, supply-chain reporting, cited | | D0 = 0.1 per cm^2, Poisson yield | assumption, stated | | Node years and the 6% per year density trend past 2026 | approximate, extrapolated from cited 2018 to 2025 points | | GPU L2 sizes | cited (NVIDIA whitepaper); AMD Infinity Cache approximate | +| Latency citations (MEMSYS 2018, Chang 2017, mining-chip memory types) | from the coordinator's research, not re-read tonight | | Recompute attacker arithmetic | M16, which is itself arithmetic on measured rates, not a chip measurement | | Dataset-build slowdown at a larger cache on a GPU | owed, a measurement (5090 at a 512 MiB and 1 GiB cache) | | The on-die emulation of M16 (inline kernel with a 64 MiB cache inside the 5090's L2) | still a PC job (M16) | -## 9. The arithmetic +## 10. The arithmetic ``` MiB = 2^20; bits = cache_MiB * MiB * 8 -raw_Mbit_per_mm2 = 1 / cell_um2 (1e6 cells per mm2 per um2 of cell) -area_mm2 = bits / (raw * 0.70 * 1e6) +headline_mm2 = cache_MiB * (41 / 64) * (cell_um2 / 0.027) (V-Cache: 41 mm2 per 64 MiB at N7, scaled by cell) +raw_Mbit_per_mm2 = 1 / cell_um2 (1e6 cells per mm2 per um2 of cell) +lower_bound_mm2 = bits / (raw * 0.70 * 1e6) dies_per_wafer = pi * 150^2 / area - pi * 300 / sqrt(2 * area) -yield = exp(-area_mm2 * 0.001) (D0 = 0.1 per cm2) -cost_per_good_die = wafer_price / (dies * yield) -reticle_GiB = 830 * raw * 0.70 * 1e6 / 8 / 2^30 +yield = exp(-area_mm2 * 0.001) (D0 = 0.1 per cm2) +cost_per_good_die = wafer_price / (dies * yield); over 830 mm2: k = ceil(area / 830) dies of area / k, cost x k +reticle_GiB = 830 / (mm2 per MiB) / 1024 ``` Run on 5 October 2026 with Python 3 on the M5 Max; the printed tables are the ones above, rounded. From 65a6cfd83a71087447eb3af6cb24c53014fadba6 Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:08:28 +0100 Subject: [PATCH 010/311] Proving v1: prover on by default, the aggregated segment record, host modes chain, aggregate and verify-segment, the app's aggregator step, spec 7.8, the fast-time harness and the coverage tool Step 1: app/igneum-app/src/provedefault.rs decides once per install (NVIDIA 12 GB or more, WSL2 answering on Windows, Linux native, Apple silicon off until measured), never switching an explicit on back off; the Settings switch line and the tile line say why (5 unit tests). tools/proving-v1/pc2-prover-cost.ps1 is the PC 2 job (5 min mining alone, 5 min with the prover, GPU memory and host RAM peaks, the sp1-gpu-server's SM targets). Step 2: the host gains --mode chain (consecutive fixtures, each block aggregated with the previous block's proof by recursion), --mode aggregate (the live aggregator over shard proof files, a run of blocks in one process) and --mode verify-segment (the node's verifier against the pinned aggregator key); the app's prover loop gains aggregate_once (spec 7.8). Eight consecutive live fixtures (blocks 81046 to 81053, node 1's export at tip 81076) under proving/fixtures/chain/. tools/proving-v1/pc2-chain.ps1 is the PC 2 job (held). Steps 3 and 4: tools/proving-v1/coverage.mjs (the proven-block share and the on-chain latency from one node's RPC), tools/proving-v1/net.mjs (the fast-time 3-node harness on 29950+ with the known-finished and known-failed cases of the chain rule and the unproven rule), the four proving_v1 fields in infra/fast-time/override-60x.json. Spec 7.8, the 7.4 rows, the 5.3 sentence, docs/plans/proving-v1.md. Co-Authored-By: Claude Fable 5.1 --- app/igneum-app/src/config.rs | 6 +- app/igneum-app/src/engine.rs | 36 + app/igneum-app/src/main.rs | 1 + app/igneum-app/src/provedefault.rs | 105 ++ app/igneum-app/src/prover.rs | 125 +++ app/igneum-app/src/state.rs | 5 + app/igneum-app/src/wslhost.rs | 37 + app/igneum-app/ui/app.js | 3 +- app/igneum-app/ui/index.html | 2 +- docs/plans/proving-v1.md | 83 ++ docs/spec/05-fees-and-economics.md | 2 + docs/spec/07-execution.md | 24 + infra/fast-time/override-60x.json | 4 + proving/fixtures/chain/block-81046.json | 1325 ++++++++++++++++++++++ proving/fixtures/chain/block-81047.json | 1325 ++++++++++++++++++++++ proving/fixtures/chain/block-81048.json | 1343 +++++++++++++++++++++++ proving/fixtures/chain/block-81049.json | 1316 ++++++++++++++++++++++ proving/fixtures/chain/block-81050.json | 1325 ++++++++++++++++++++++ proving/fixtures/chain/block-81051.json | 1316 ++++++++++++++++++++++ proving/fixtures/chain/block-81052.json | 1334 ++++++++++++++++++++++ proving/fixtures/chain/block-81053.json | 1316 ++++++++++++++++++++++ proving/igneum-prove/host/src/main.rs | 327 +++++- proving/igneum-prove/host/src/pinned.rs | 3 +- tools/prove-fixtures/node_modules | 1 + tools/proving-v1/coverage.mjs | 89 ++ tools/proving-v1/net.mjs | 246 +++++ tools/proving-v1/pc2-chain.ps1 | 88 ++ tools/proving-v1/pc2-memory-sweep.ps1 | 76 ++ tools/proving-v1/pc2-prover-cost.ps1 | 111 ++ 29 files changed, 11967 insertions(+), 7 deletions(-) create mode 100644 app/igneum-app/src/provedefault.rs create mode 100644 docs/plans/proving-v1.md create mode 100644 proving/fixtures/chain/block-81046.json create mode 100644 proving/fixtures/chain/block-81047.json create mode 100644 proving/fixtures/chain/block-81048.json create mode 100644 proving/fixtures/chain/block-81049.json create mode 100644 proving/fixtures/chain/block-81050.json create mode 100644 proving/fixtures/chain/block-81051.json create mode 100644 proving/fixtures/chain/block-81052.json create mode 100644 proving/fixtures/chain/block-81053.json create mode 120000 tools/prove-fixtures/node_modules create mode 100644 tools/proving-v1/coverage.mjs create mode 100644 tools/proving-v1/net.mjs create mode 100644 tools/proving-v1/pc2-chain.ps1 create mode 100644 tools/proving-v1/pc2-memory-sweep.ps1 create mode 100644 tools/proving-v1/pc2-prover-cost.ps1 diff --git a/app/igneum-app/src/config.rs b/app/igneum-app/src/config.rs index f59fee072..23d633ed7 100644 --- a/app/igneum-app/src/config.rs +++ b/app/igneum-app/src/config.rs @@ -87,6 +87,10 @@ pub struct Settings { /// so it includes proof records it never verified (src/verifier.rs). Default off; a found verifier always wins. #[serde(default)] pub proof_verify_trust: bool, + /// Proving v1 step 1 (5 October 2026): the install-time default for `prove` has been applied once (src/provedefault.rs: + /// on when the machine can prove, never switching an explicit on back off). Older installs apply it at their next start. + #[serde(default)] + pub prove_default_applied: bool, } fn one() -> u32 { @@ -98,7 +102,7 @@ fn yes() -> bool { impl Default for Settings { fn default() -> Settings { - Settings { setup_done: false, address: String::new(), address_source: String::new(), key_saved: false, identities: 1, cards: HashMap::new(), display_name: String::new(), vote: true, paused: false, accepted_total: 0, auto_update: true, remote_jobs: true, prove: false, sweep: true, installed_at: 0, dev_fee: true, fee_total: 0, proof_verify_trust: false } + Settings { setup_done: false, address: String::new(), address_source: String::new(), key_saved: false, identities: 1, cards: HashMap::new(), display_name: String::new(), vote: true, paused: false, accepted_total: 0, auto_update: true, remote_jobs: true, prove: false, sweep: true, installed_at: 0, dev_fee: true, fee_total: 0, proof_verify_trust: false, prove_default_applied: false } } } diff --git a/app/igneum-app/src/engine.rs b/app/igneum-app/src/engine.rs index 3ad329d00..599902c82 100644 --- a/app/igneum-app/src/engine.rs +++ b/app/igneum-app/src/engine.rs @@ -253,6 +253,40 @@ impl Shared { Ok(json!({ "ok": true, "address": w.address, "display": keys::checksum(&w.address), "private_key": w.private_key, "wallet_file": self.wallet_path.display().to_string() })) } + /// Proving v1 step 1 (5 October 2026): once per install, after the cards are known, the prover goes on by itself + /// when this machine can prove (src/provedefault.rs: an NVIDIA card with 12 GB or more, WSL2 on Windows, Linux + /// native, Apple silicon off until measured). An explicit on is never switched off; the line goes to the log and + /// to the Proving tile. Older installs apply it at their first start on this version. + pub fn apply_prove_default(&self) { + let (applied, already_on) = { + let s = self.settings.lock().unwrap(); + (s.prove_default_applied, s.prove) + }; + if applied { + return; + } + let cards = self.state.lock().unwrap().mining.cards.clone(); + let wsl = if cfg!(windows) { Some(crate::wslhost::distro_answers()) } else { None }; + let d = crate::provedefault::decide(&cards, std::env::consts::OS, wsl); + let on = d.on || already_on; + { + let mut s = self.settings.lock().unwrap(); + s.prove = on; + s.prove_default_applied = true; + s.save(&self.settings_path); + } + { + let mut st = self.state.lock().unwrap(); + st.settings.prove = on; + st.proving.enabled = on; + st.proving.default_note = d.line.clone(); + if !on { + st.proving.status = "off".into(); + } + } + self.log(&format!("prover default: {}{}", d.line, if already_on && !d.on { " (left on: it was switched on by hand)" } else { "" })); + } + /// The prover service switch (src/prover.rs); the thread picks it up within 10 s. pub fn set_prove(&self, on: bool) -> Result { { @@ -711,6 +745,8 @@ impl Engine { self.detect_again = false; self.shared.send(Cmd::Detect); } + // proving v1 step 1: the install-time prover default, once the cards are known + self.shared.apply_prove_default(); } Cmd::ApplyCards(choices) => self.apply_cards(choices), Cmd::Start => self.start(), diff --git a/app/igneum-app/src/main.rs b/app/igneum-app/src/main.rs index dc37f2bb3..2e7164393 100644 --- a/app/igneum-app/src/main.rs +++ b/app/igneum-app/src/main.rs @@ -29,6 +29,7 @@ mod jobs; mod jobrun; mod jobbuild; mod prover; +mod provedefault; mod verifier; mod wslhost; mod sweep; diff --git a/app/igneum-app/src/provedefault.rs b/app/igneum-app/src/provedefault.rs new file mode 100644 index 000000000..3a72daae1 --- /dev/null +++ b/app/igneum-app/src/provedefault.rs @@ -0,0 +1,105 @@ +//! Proving v1 step 1 (5 October 2026, Josh: "open the proving round asap"): the prover is on by default on every +//! mining machine that can prove, decided once per install after the cards are detected (src/engine.rs +//! `apply_prove_default`). The rule, one line each: +//! +//! | Machine | Default | Why | +//! |---|---|---| +//! | NVIDIA card with 12 GB or more, Windows, WSL2 (Ubuntu-24.04) answers | on | the SP1 CUDA prover runs inside WSL2 (src/prover.rs) | +//! | NVIDIA card with 12 GB or more, Windows, WSL2 silent | off, with the Set up hint | nothing can prove until the distribution exists | +//! | NVIDIA card with 12 GB or more, Linux | on | the host runs next to the engine | +//! | Apple silicon | off | the CPU prover is minutes per shard; on until it is measured on the GPU | +//! | no NVIDIA card with 12 GB | off | the 12 GB gate (spec 5.1, ledger P1); measured on a 32 GB card only so far | +//! +//! The default never switches an explicit on back off, and Settings always wins afterwards. + +use crate::state::CardState; + +/// The gate card: 12 GB (spec 5.1, "a shard on a 12 GB card"). `nvidia-smi` reports MiB; 12 GB cards report +/// 12,288 MiB or a little under (the RTX 3060 12 GB reports 12,288), so the test is at 11.5 GB. +pub const MIN_VRAM_MB: u64 = 11_776; + +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct Decision { + pub on: bool, + /// One plain sentence for the log and the Proving tile. + pub line: String, +} + +fn gb(mb: u64) -> u64 { + (mb + 512) / 1024 +} + +/// `os` is `std::env::consts::OS` ("windows", "linux", "macos"); `wsl_answers` is read on Windows only. +pub fn decide(cards: &[CardState], os: &str, wsl_answers: Option) -> Decision { + let nvidia: Vec<&CardState> = cards.iter().filter(|c| c.vendor == "nvidia").collect(); + let able: Vec<&CardState> = nvidia.iter().copied().filter(|c| c.vram_mb >= MIN_VRAM_MB).collect(); + let off = |line: String| Decision { on: false, line }; + if os == "macos" { + return off("proving stays off on Apple silicon until the GPU prover is measured there; Settings switches it on (CPU, slow)".into()); + } + let Some(best) = able.iter().max_by_key(|c| c.vram_mb) else { + let seen = if nvidia.is_empty() { + "no NVIDIA card".to_string() + } else { + nvidia.iter().map(|c| format!("{} {} GB", c.name, gb(c.vram_mb))).collect::>().join(", ") + }; + return off(format!("proving off by default: no NVIDIA card with 12 GB or more ({seen}); Settings switches it on")); + }; + let card = format!("{} ({} GB)", best.name, gb(best.vram_mb)); + match os { + "windows" => match wsl_answers { + Some(true) => Decision { on: true, line: format!("proving on by default: {card} with WSL2 (Ubuntu-24.04 answers); Settings switches it off") }, + _ => off(format!("proving off: {card} qualifies but WSL2 (Ubuntu-24.04) did not answer; Set up installs it, then Settings switches proving on")), + }, + "linux" => Decision { on: true, line: format!("proving on by default: {card} on Linux (the host runs next to the engine); Settings switches it off") }, + other => off(format!("proving off: {card} on {other}, no prover path there; Settings switches it on")), + } +} + +#[cfg(test)] +mod tests { + use super::*; + + fn card(vendor: &str, name: &str, vram_mb: u64) -> CardState { + CardState { vendor: vendor.into(), name: name.into(), vram_mb, ..Default::default() } + } + + #[test] + fn a_5090_with_wsl2_on_windows_is_on() { + let d = decide(&[card("nvidia", "NVIDIA GeForce RTX 5090", 32_607), card("amd", "AMD Radeon(TM) Graphics", 512)], "windows", Some(true)); + assert!(d.on); + assert!(d.line.starts_with("proving on by default: NVIDIA GeForce RTX 5090 (32 GB) with WSL2"), "{}", d.line); + } + + #[test] + fn windows_without_wsl2_is_off_with_the_setup_hint() { + let d = decide(&[card("nvidia", "NVIDIA GeForce RTX 4090", 24_564)], "windows", Some(false)); + assert!(!d.on); + assert!(d.line.contains("did not answer") && d.line.contains("Set up"), "{}", d.line); + assert!(!decide(&[card("nvidia", "RTX 4090", 24_564)], "windows", None).on, "an unread probe is not an answer"); + } + + #[test] + fn linux_needs_no_wsl2_and_the_12_gb_gate_holds() { + assert!(decide(&[card("nvidia", "NVIDIA GeForce RTX 3060", 12_288)], "linux", None).on); + let d = decide(&[card("nvidia", "NVIDIA GeForce RTX 3080", 10_240)], "linux", None); + assert!(!d.on); + assert!(d.line.contains("no NVIDIA card with 12 GB or more (NVIDIA GeForce RTX 3080 10 GB)"), "{}", d.line); + assert!(!decide(&[card("amd", "Radeon RX 9070 XT", 16_384)], "linux", None).on, "no CUDA prover for AMD yet"); + assert!(decide(&[], "linux", None).line.contains("no NVIDIA card")); + } + + #[test] + fn apple_silicon_stays_off() { + let d = decide(&[card("apple", "Apple M5 Max", 65_536)], "macos", None); + assert!(!d.on); + assert!(d.line.contains("Apple silicon")); + assert!(!decide(&[card("nvidia", "RTX 5090", 32_607)], "macos", Some(true)).on, "the OS rule comes first"); + } + + #[test] + fn the_biggest_qualifying_card_is_named() { + let d = decide(&[card("nvidia", "RTX 3060", 12_288), card("nvidia", "RTX 5090", 32_607)], "linux", None); + assert!(d.line.contains("RTX 5090 (32 GB)"), "{}", d.line); + } +} diff --git a/app/igneum-app/src/prover.rs b/app/igneum-app/src/prover.rs index c5cfb1ca0..7b34b9b52 100644 --- a/app/igneum-app/src/prover.rs +++ b/app/igneum-app/src/prover.rs @@ -339,6 +339,7 @@ pub fn start(shared: Arc, bin_dir: PathBuf) { fn loop_forever(shared: Arc, bin_dir: PathBuf) { let mut attempted: HashSet<(String, u32)> = HashSet::new(); + let mut attempted_segments: HashSet = HashSet::new(); let mut tools: Option = None; let mut last_probe = Instant::now() - Duration::from_secs(600); let mut submitted: Vec<(u64, String, u32, u128)> = Vec::new(); @@ -466,6 +467,20 @@ fn loop_forever(shared: Arc, bin_dir: PathBuf) { p.assigned = assigned; p.keys = keys.len() as u32; }); + // proving v1 (spec 7.8): the aggregator step, when the node says v1 is active; one attempt a pass + if let Some((label0, _)) = keys.first() { + match aggregate_once(&shared, t, label0, &payout_address(&shared), &mut attempted_segments) { + Ok(Some(msg)) => { + shared.log(&format!("aggregator: {msg}")); + set(&shared, |p| p.segment_note = msg); + } + Ok(None) => {} + Err(e) => { + shared.log(&format!("aggregator: {e}")); + set(&shared, |p| p.segment_note = e); + } + } + } let Some(w) = choose(&work, &attempted) else { set(&shared, |p| { p.status = if submitted.is_empty() { "idle".into() } else { "submitted".into() }; @@ -558,6 +573,116 @@ fn loop_forever(shared: Arc, bin_dir: PathBuf) { } } +/// Proving v1 (spec 7.8): one aggregation attempt. When the node reports v1 active, takes the newest executed +/// segment that is still pending and not yet attempted here, needs one shard proof per shard of every block in +/// this node's pool (`igneum_getProofBytes`, a verified one when there is one) and, when the previous segment is +/// proven, its aggregated proof (`igneum_getSegmentProofBytes`); runs `igneum-prove-host --mode aggregate` over the +/// run of blocks (one process, one key setup), checks the public values against the node's native statement +/// (every field but `provers`), signs the record with the first key's label and submits it. Returns a line for +/// the log and the tile, or None when there is nothing to do. +fn aggregate_once(shared: &Shared, t: &Tools, label: &str, payout: &str, attempted: &mut HashSet) -> Result, String> { + let hexu = |x: &Value| x.as_str().and_then(|s| u64::from_str_radix(s.trim_start_matches("0x"), 16).ok()).unwrap_or(0); + let st = evm_rpc(shared, "igneum_getProvingStatus", json!([]), Duration::from_secs(10))?; + let v1 = &st["v1"]; + if !v1["active"].as_bool().unwrap_or(false) || v1["start"].is_null() { + return Ok(None); + } + let tip = hexu(&evm_rpc(shared, "eth_blockNumber", json!([]), Duration::from_secs(10))?); + let mut seg = evm_rpc(shared, "igneum_getSegmentStatement", json!([format!("{tip:#x}")]), Duration::from_secs(10))?; + if !seg["executed"].as_bool().unwrap_or(false) { + let first = hexu(&seg["first"]); + if first == 0 || first - 1 < hexu(&v1["start"]) { + return Ok(None); + } + seg = evm_rpc(shared, "igneum_getSegmentStatement", json!([format!("{:#x}", first - 1)]), Duration::from_secs(10))?; + } + let (first, last) = (hexu(&seg["first"]), hexu(&seg["last"])); + if seg["status"]["status"].as_str() != Some("pending") || attempted.contains(&first) || payout.len() != 42 { + return Ok(None); + } + let dir = shared.runtime.app_dir.join("proving").join(format!("seg-{first}")); + let _ = std::fs::create_dir_all(&dir); + let as_host_path = |p: &Path| if t.wsl { wsl_path(p) } else { p.display().to_string() }; + // the shard proofs, one per shard of every block, from this node's pool + let mut groups: Vec = Vec::new(); + let mut missing: Vec = Vec::new(); + for b in seg["blocks"].as_array().cloned().unwrap_or_default() { + let n = hexu(&b["number"]); + let shards = b["shards"].as_u64().unwrap_or(0) as u32; + let have = b["shardProofs"].as_array().cloned().unwrap_or_default(); + let mut files = Vec::new(); + for i in 0..shards { + let pick = have.iter().find(|e| e["shard"].as_u64() == Some(i as u64) && e["verified"] == json!(true)).or_else(|| have.iter().find(|e| e["shard"].as_u64() == Some(i as u64))); + let Some(e) = pick else { + missing.push(format!("{n}/{i}")); + continue; + }; + let got = evm_rpc(shared, "igneum_getProofBytes", json!([format!("{n:#x}"), i, e["keyHash"]]), Duration::from_secs(60))?; + let hex = got["proof"].as_str().ok_or("no proof bytes")?.trim_start_matches("0x").to_string(); + let bytes: Vec = (0..hex.len() / 2).map(|k| u8::from_str_radix(&hex[2 * k..2 * k + 2], 16).unwrap_or(0)).collect(); + let f = dir.join(format!("b{n}-s{i}.bin")); + std::fs::write(&f, bytes).map_err(|e| e.to_string())?; + files.push(as_host_path(&f)); + } + groups.push(files.join(",")); + } + if !missing.is_empty() { + return Ok(Some(format!("segment {first}..{last}: waiting for shard proofs {} in this node's pool", missing.join(" ")))); + } + // the previous segment's aggregated proof, when the chain continues + let prev = &seg["previous"]; + let (prev_file, expected_pv) = if prev.is_null() { + (None, seg["publicValuesFresh"].as_str().unwrap_or("").to_string()) + } else if prev["proofInPool"] == json!(true) { + let got = evm_rpc(shared, "igneum_getSegmentProofBytes", json!([prev["first"], prev["keyHash"]]), Duration::from_secs(60))?; + let hex = got["proof"].as_str().ok_or("no segment proof bytes")?.trim_start_matches("0x").to_string(); + let bytes: Vec = (0..hex.len() / 2).map(|k| u8::from_str_radix(&hex[2 * k..2 * k + 2], 16).unwrap_or(0)).collect(); + let f = dir.join("prev-aggregated.bin"); + std::fs::write(&f, bytes).map_err(|e| e.to_string())?; + (Some(as_host_path(&f)), seg["publicValuesContinuing"].as_str().unwrap_or("").to_string()) + } else { + return Ok(Some(format!("segment {first}..{last}: the previous segment's proof is not in this node's pool; waiting"))); + }; + attempted.insert(first); + let parent = seg["blocks"][0]["parentHash"].as_str().ok_or("no parent hash")?.to_string(); + let last_hash = seg["blocks"].as_array().and_then(|a| a.last()).and_then(|b| b["hash"].as_str()).ok_or("no last hash")?.to_string(); + let results = dir.join("results.json"); + let started = Instant::now(); + set(shared, |p| p.message = format!("aggregating segment {first}..{last} ({})", if t.cuda { "GPU" } else { "CPU, slow" })); + let mut args: Vec = vec!["--mode".into(), "aggregate".into(), "--proofs".into(), groups.join(";"), "--parent".into(), parent, "--out".into(), as_host_path(&results)]; + if let Some(pf) = prev_file { + args.push("--prev".into()); + args.push(pf); + } + let (ok, out) = run_tool(shared, t, &t.host, &args, &[("SP1_PROVER", if t.cuda { "cuda" } else { "cpu" }), ("RUST_LOG", "off")], Duration::from_secs(2 * 3600), &dir.join("aggregate.log")); + if !ok || !results.exists() { + return Err(format!("segment {first}..{last}: aggregator: {}", out.lines().rev().find(|l| l.contains("RESULT") || l.contains("rror")).unwrap_or("failed"))); + } + let res: Value = serde_json::from_str(&std::fs::read_to_string(&results).map_err(|e| e.to_string())?).map_err(|e| e.to_string())?; + let pv = res["segment_public_values"].as_str().ok_or("no public values in the results")?.to_string(); + let proof_sha = res["segment_proof_sha256"].as_str().ok_or("no proof hash in the results")?.to_string(); + let proof_file = res["segment_proof_file"].as_str().ok_or("no proof file in the results")?.to_string(); + // the node's native statement, every field but provers (bytes 236..268 of the 340) + let strip = |h: &str| { let h = h.trim_start_matches("0x"); if h.len() == 680 { format!("{}{}", &h[..472], &h[536..]) } else { h.to_string() } }; + if strip(&pv) != strip(&expected_pv) { + return Err(format!("segment {first}..{last}: the aggregated statement differs from the node's native statement (it would be vetoed); ours {} node {}", &pv[..66.min(pv.len())], &expected_pv[..66.min(expected_pv.len())])); + } + let proof_path = if t.wsl { PathBuf::from(proof_file.replace("/mnt/c/", "C:/")) } else { PathBuf::from(proof_file) }; + let sg = crate::detect::run_timeout(crate::platform::quiet(&mut Command::new(&t.miner)).args(["sign-segment-record", label, &chain_name(shared), &first.to_string(), &last.to_string(), &last_hash, payout, &pv, &proof_sha]), None, Duration::from_secs(20)).ok_or("sign-segment-record did not run")?; + let signed: Value = serde_json::from_str(sg.lines().last().unwrap_or("")).map_err(|_| format!("sign-segment-record: {}", sg.trim()))?; + let record = signed["record"].as_str().ok_or("sign-segment-record gave no record")?.to_string(); + let proof = std::fs::read(&proof_path).map_err(|e| format!("proof file {}: {e}", proof_path.display()))?; + let proof_hex = format!("0x{}", proof.iter().map(|b| format!("{b:02x}")).collect::()); + let r = evm_rpc(shared, "igneum_submitSegmentRecord", json!([{ "record": record, "proof": proof_hex }]), Duration::from_secs(60))?; + if !r["accepted"].as_bool().unwrap_or(false) { + return Err(format!("segment {first}..{last}: record refused: {}", r["reason"].as_str().unwrap_or("?"))); + } + let secs = started.elapsed().as_secs_f64(); + shared.event("proving", &format!("segment {first}..{last} aggregated and submitted in {secs:.0} s (chain_len {})", res["segment_chain_len"])); + set(shared, |p| p.aggregated += 1); + Ok(Some(format!("segment {first}..{last} aggregated in {secs:.0} s and submitted; paid when a block carries it"))) +} + /// Windows: runs the WSL2 setup from the payload (`wsl2/setup-wsl.sh` next to the engine) in a window of its own; /// the user watches it and reboots when it asks. Elsewhere there is nothing to set up. pub fn setup(shared: &Shared) -> Result { diff --git a/app/igneum-app/src/state.rs b/app/igneum-app/src/state.rs index 356b0cf8d..5cf32e57a 100644 --- a/app/igneum-app/src/state.rs +++ b/app/igneum-app/src/state.rs @@ -196,6 +196,11 @@ pub struct ProvingState { /// the pinned guests' ids (`igneum-prove-host --mode id`): the shard program and the aggregator; empty until read pub program_id: String, pub aggregator_id: String, + /// proving v1 step 1: the install-time default's one plain line (why proving is on or off on this machine) + pub default_note: String, + /// proving v1: segment records this machine aggregated and submitted, and the aggregator's last line + pub aggregated: u32, + pub segment_note: String, } #[derive(Clone, Serialize, Default)] diff --git a/app/igneum-app/src/wslhost.rs b/app/igneum-app/src/wslhost.rs index b01b928f3..2c3685b76 100644 --- a/app/igneum-app/src/wslhost.rs +++ b/app/igneum-app/src/wslhost.rs @@ -41,6 +41,43 @@ pub fn candidates(bin_dir: &Path) -> Vec { } /// The candidates as one line for a message. +/// Whether the distribution answers at all (`wsl.exe -d Ubuntu-24.04 -- echo ` within 30 s): the install-time +/// prover default (src/provedefault.rs) needs WSL2 on Windows before it switches proving on. Elsewhere: false. +#[allow(dead_code)] // also compiled into src/bin/prove-verify.rs, which does not call it +pub fn distro_answers() -> bool { + if !cfg!(windows) { + return false; + } + // self-contained (this file is also compiled into src/bin/prove-verify.rs, which has no detect or platform module) + let mut cmd = std::process::Command::new("wsl"); + cmd.args(["-d", DISTRO, "--", "echo", "igneum-wsl-answers"]).stdin(std::process::Stdio::null()).stdout(std::process::Stdio::piped()).stderr(std::process::Stdio::null()); + #[cfg(windows)] + { + use std::os::windows::process::CommandExt; + cmd.creation_flags(0x0800_0000); // CREATE_NO_WINDOW + } + let Ok(mut child) = cmd.spawn() else { return false }; + let Some(out) = child.stdout.take() else { return false }; + let reader = std::thread::spawn(move || { + let mut s = String::new(); + let _ = std::io::Read::read_to_string(&mut std::io::BufReader::new(out), &mut s); + s + }); + let deadline = std::time::Instant::now() + std::time::Duration::from_secs(30); + loop { + match child.try_wait() { + Ok(Some(_)) => break, + Ok(None) if std::time::Instant::now() < deadline => std::thread::sleep(std::time::Duration::from_millis(100)), + _ => { + let _ = child.kill(); + let _ = child.wait(); + break; + } + } + } + reader.join().map(|o| o.replace('\0', "").contains("igneum-wsl-answers")).unwrap_or(false) +} + pub fn candidates_text(bin_dir: &Path) -> String { candidates(bin_dir).join(", ") } diff --git a/app/igneum-app/ui/app.js b/app/igneum-app/ui/app.js index 92b396ea2..af1bb188b 100644 --- a/app/igneum-app/ui/app.js +++ b/app/igneum-app/ui/app.js @@ -1005,7 +1005,8 @@ if (typeof document !== 'undefined') (function () { setText('pv-state', w.word + (pv.backend && enabled && pv.available ? ' (' + pv.backend.toUpperCase() + ')' : '')); setText('pv-state-sub', w.sub); var cell = $('pv-state').parentNode; cell.classList.toggle('ok', w.tone === 'on' || w.tone === 'ok'); cell.classList.toggle('bad', w.tone === 'bad'); - setText('pv-note', w.note); + // proving v1: when off by the install-time default, the default's own line says why (src/provedefault.rs) + setText('pv-note', (!enabled && pv.default_note) ? 'Off. ' + pv.default_note + '.' : w.note); setText('pv-assigned', String(pv.assigned || 0)); setText('pv-submitted', String(pv.submitted || 0)); setText('pv-paid', String(pv.paid || 0)); diff --git a/app/igneum-app/ui/index.html b/app/igneum-app/ui/index.html index ec32a9637..9c98989a2 100644 --- a/app/igneum-app/ui/index.html +++ b/app/igneum-app/ui/index.html @@ -206,7 +206,7 @@

Prove shards on this machine

-

Every block on Igneum is turned into a short mathematical proof, in pieces called shards. The chain assigns shards to your keys; this machine proves them and is paid for each one.

+

Every block on Igneum is turned into a short mathematical proof, in pieces called shards. The chain assigns shards to your keys; this machine proves them and is paid for each one. On by default on an NVIDIA card with enough memory (the line below says which); off on a Mac, whose CPU prover is slow.

diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md new file mode 100644 index 000000000..2f4ef493d --- /dev/null +++ b/docs/plans/proving-v1.md @@ -0,0 +1,83 @@ +# Proving v1: segment records, the chain rule, the unproven rule; the 0.3.11 rollout + +5 October 2026, from 18:55 UTC (Josh: "open the proving round asap"). Branches `proving-v1` in the main repository +(worktree `/Users/joshm/Projects/igneum-wt-proving-v1`, from master a93199a) and in the fork +(`vendor/igneum-node-pv1`, from release-0.3.6 a24ab01a; to be rebased onto the 0.3.10 tip when it lands on +release-0.3.6). Status words follow `docs/spec/00-overview.md` 0.2. Every number here is in `docs/bench-log.md` +with its command. Nothing ships from this plan: it delivers branches, numbers and the rollout for 0.3.11. + +## The gap this closes + +The litepaper says every block is proven within about a minute. Proving v0 (`proving-v0.md`, spec 7.7) proves +some shards: one prover (PC 2's RTX 5090) takes the newest shard assigned to it, about one shard every 30 s, so +under a tenth of blocks carry a proof; consensus does not require one; the aggregator guest (design 5.3) runs +on fixtures only. Proving v1 adds the aggregated segment record on chain (spec 7.8), the chain rule (segment N's +record verifies N-1, inside the proof by recursion), the unproven rule (a segment nobody proves in T seconds pays +nothing and may be skipped), the prover on by default on every machine that can prove, and the measurements +that say how many cards cover the chain. + +## The round, step by step + +| Step | What | State | +|---|---|---| +| 1 | Prover on by default (`app/igneum-app/src/provedefault.rs`, `engine.rs apply_prove_default`): on at install when the machine can prove (NVIDIA card with 12 GB or more; WSL2 answering on Windows; Linux native; Apple silicon off until measured), never switching an explicit on back off; the Settings switch line and the tile line say why. Unit tests (5). The cost of proving on a mining machine: PC 2 job `prover-cost-pc2-pv1` (5 min mining alone, 5 min with the prover, the GPU memory peak and the host RAM peak, the sp1-gpu-server's compiled SM targets) | Implemented; the measurement is HELD (coordinator, 19:00Z): PC 2's RTX 5090 worker has been exiting on a pack seed mismatch since 18:35Z, so the first run's "mining alone" is 0 MH/s and void; re-run after the go | +| 2 | Segment aggregation: `SegmentRecord` (586 bytes, the aggregator guest's 340-byte statement inline), section `IGNS` before the shard section, p2p message 72 at protocol 14, the native block statement and the veto, the credit split (`split_pool_credit`), the payout at the carrier, `igneum-miner sign-segment-record`, the RPCs; host modes `chain` (consecutive fixtures), `aggregate` (live shard proofs from the pool, a run of blocks in one process) and `verify-segment` (the node's verifier, pinned aggregator key); the app's aggregator step (`prover.rs aggregate_once`) | Implemented, unit-tested (consensus core 2 new tests, exec 2, params 1); the GPU measurement (N = 2, 4, 8 blocks on PC 2) is HELD with step 1; the Mac CPU run of `--mode chain` over 2 live blocks is the known-finished case | +| 3 | Coverage: `tools/proving-v1/coverage.mjs` (the proven-block share and the on-chain proof latency over a window from one node's RPC, the live page beside it) | Implemented and run for 3 min (below); the 30-min window waits for the fleet | +| 4 | The chain rule and the unproven rule in consensus behind `proving_v1_activation_daa` (spec 7.8 items 2, 6, 7); unit tests; the fast-time 3-node harness `tools/proving-v1/net.mjs` (ports 29950+, suffix 956, trust mode) with the known-finished and known-failed cases | Implemented; the harness run waits for the Mac build of the fork (`vendor/igneum-node/target-pv1`) | +| 5 | This plan: the rollout for 0.3.11 and Josh's decisions | Written below | + +## Numbers (every one from `docs/bench-log.md`, "proving v1") + +(filled as the measurements land; see the bench log entries of 5 October 2026 named "proving v1 ...") + +## The rule, in one paragraph (spec 7.8) + +From the first chain block `A` at or above `proving_v1_activation_daa`, chain blocks form fixed segments of `N` +(`proving_v1_segment_blocks`). A segment record carries the aggregated proof of the segment's last block, whose +`chain_len` says how many consecutive blocks the recursion attests. Every node checks the record's statement +against its own native block statement (every field but the provers commitment and `chain_len`), the chain rule +(a proof that does not chain to the previous segment, `chain_len = N`, is valid only for the first segment or +after an unproven one) and the deadline (`T = proving_v1_unproven_daa` DAA seconds after the segment's last block; +a record carried later pays nothing). The pool credit of every attested block splits: `proving_v1_aggregator_share_bps` +to the aggregator, the rest to the shards as v0. A block is never invalid for lack of a proof; the mandatory rule +(spec 7.8 item 10) is Designed and off, with no switch yet. + +## Rollout for 0.3.11 (the digest handshake pattern of 0.3.9 and tonight's switches) + +The switch moves the consensus digest only once it is set (`consensus_digest`: the four v1 fields enter the hash +when `proving_v1_activation_daa != never`), so a 0.3.11 node on the unswitched devnet keeps the 0.3.10 digest and +the rolling upgrade does not partition the network. The order, each step with its check: + +1. **Rebase and build.** Fork `proving-v1` rebased onto the 0.3.10 tip on `release-0.3.6`; the six node suites and + the app tests as PC 2 build jobs; the Mac node and the Windows exes by the Mac cross-build; the Linux node by + PC 1; the HiveOS package republished from the same fork commit (rule: the HiveOS package carries the node of + the release commit and the same override object, `infra/hive`, as 0.3.9's `hive-sync-039o` checked it). +2. **Pinned guests.** The guest ids do not change in this round (shard `0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a`, + aggregator `0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896`, pinned 2026-10-05T16:20:38Z): + the aggregator guest already carried the chain rule and only the HOST gained modes. So no provers-off drain is + needed for the guests; `--mode id` on every machine after the update must print the same two ids, and the node + now reads them at start (`program_ids`: `IGNEUM_PROOF_PROGRAM_IDS` or the verifier's `--mode id`) and names them + in the native statement. +3. **Hand nodes and the seed first**, with the UNCHANGED override object (the digest stays): observer, node 1, the + seed on the 0.3.11 node; peers back within 20 s; the `proving v1: segment records from DAA score never` line in + each log. +4. **Manifest and apps.** `publish-manifest.sh --version 0.3.11` with the unchanged object; `update-now` to every + app; every machine on 0.3.11 with a DAA score and a hash rate (the watcher takes the commit as an argument). + The app's prover default applies at the first start on 0.3.11: every NVIDIA machine with WSL2 goes on; the + log line `prover default: ...` on each. +5. **The switch.** When every node runs 0.3.11: publish the object with `proving_v1_activation_daa` = H (24 h + ahead, the rule of `fee-switch-devnet.md`) and the three parameters; read the expected digest on a scratch node + first; `update-now`; the hand nodes and the seed with the same object; the digest sweep; the first paid segment + record (`igneum_getSegmentRecords`) and `igneum_getProvingStatus.v1.segmentsInWindow` after H. +6. **The mandatory rule** stays off: no switch exists for it yet; it gets one when the measured share is one. + +## What Josh must decide + +| Decision | Proposed | Why | +|---|---|---| +| `proving_v1_segment_blocks` (N) | 4 | the N = 2, 4, 8 measurement on the 5090 decides; 4 keeps a record every 4 s at 1 block/s and the recursion cost per block constant | +| `proving_v1_unproven_daa` (T) | 600 | equals the 600-block record window of v0: nothing is payable for a segment after it either way; the fleet's measured latency (p99) must sit well inside it | +| `proving_v1_aggregator_share_bps` | 1,000 (a tenth) | the aggregation is one recursion per block, far cheaper than the shards; a tenth pays a second role without starving the shard provers; it is a consensus parameter in the digest | +| `proving_v1_activation_daa` (H) | 24 h after the 0.3.11 publish | the fee-switch rule | +| Aggregator sortition | none in v1 (first valid record wins) | design 5.3's VRF draw is O-7.3; with one or two aggregators on the devnet a draw changes nothing yet | +| Apple silicon default | off | until the GPU prover is measured on a Mac | diff --git a/docs/spec/05-fees-and-economics.md b/docs/spec/05-fees-and-economics.md index 18fdb58e1..5e2692913 100644 --- a/docs/spec/05-fees-and-economics.md +++ b/docs/spec/05-fees-and-economics.md @@ -36,6 +36,8 @@ Self-dealing: a developer who also mines the including block collects 80% plus 2 Designed. The 20% emission share (section 2.5) and the provers' part of the 80% tip share are paid per block as a fixed amount for that block, divided among the block's shards by consensus proving cost, so a stuffed block earns no more than an honest one. Shards are not claimed first-come and carry no bond: each shard is assigned by sortition to 8 eligible provers for a 10-s exclusive window, then open to anyone, and the first valid proof included in a block is paid (section 7.2, decided 3 October 2026, ledger P8, C9). The parameters 8 and 10 s are set on the phase 4 devnet (O-5.1). A withheld shard costs nothing to bond against because nothing waits on an assigned prover: an unproven block delays only its proof; execution and the 30-s lock do not wait for it (ledger P9). The bond, slashed on a bad or late proof, remains in the external job market (5.4), where a customer does wait; its size and timeout are Open (O-5.6). +Proving v1 (section 7.8, 5 October 2026, Implemented behind `proving_v1_activation_daa`): from the switch, `proving_v1_aggregator_share_bps` of a block's fixed amount (a tenth, Josh's decision at 0.3.11) goes to the aggregator whose segment record attests the block, the rest to the shards as before; a block in a segment that stays unproven past `proving_v1_unproven_daa` pays no aggregator share. + ## 5.4 External job market Designed. At launch external proving jobs are paid on the customer's chain, in the customer's currency, to a payout contract keyed by miner address, because Igneum cannot yet see Ethereum; the customer chain's own bond and slashing apply (design document, "The first six months"). When the job market settles on Igneum, which needs the proof bridge (phase 2 consensus proof, ledger P4, E7), every job fee paid in IGN splits: diff --git a/docs/spec/07-execution.md b/docs/spec/07-execution.md index cfcafa954..ea023e77d 100644 --- a/docs/spec/07-execution.md +++ b/docs/spec/07-execution.md @@ -78,6 +78,13 @@ Nothing in consensus changes for any of this: the segment claim already commits | Records per block | 8 | Implemented, 7.7 (Designed value) | | Payout per shard | the segment's pool credit in equal parts, remainder to shard 0, paid by the carrying segment | Implemented, 7.7 | | Proving v0 activation | `proving_v0_activation_daa`, default never | Implemented, 7.7 | +| Segment record (v1) | per segment of `proving_v1_segment_blocks` chain blocks, 586 bytes (the aggregator guest's 340-byte statement inline), BLS-signed by the aggregator's vote key, in the coinbase extra data before the shard record section (`IGNS`); proof bytes on p2p message 72 (protocol 14) | Implemented, 7.8 (branch `proving-v1`, 5 October 2026) | +| Segment records per block | 2 | Implemented, 7.8 (Designed value) | +| Segment length `N` | `proving_v1_segment_blocks`, 4 | Implemented, value Designed (Josh decides at 0.3.11) | +| Unproven deadline `T` | `proving_v1_unproven_daa`, 600 DAA s after the segment's last chain block | Implemented, value Designed (Josh decides at 0.3.11) | +| Aggregator share | `proving_v1_aggregator_share_bps`, 1,000 (a tenth of every attested block's pool credit; the shards share the rest) | Implemented, value Designed (Josh decides at 0.3.11) | +| Proving v1 activation | `proving_v1_activation_daa`, default never; the segment grid starts at the first chain block at or above it | Implemented, 7.8 | +| Mandatory proofs | the rule of 7.8 item 9, no switch yet, off | Designed | ## 7.5 Proving gas per transaction: the cap and the abort @@ -112,3 +119,20 @@ Added 4 October 2026 because the implementation (`vendor/igneum-node-proving`, b 8. **The plan.** Every segment has at least one shard; an empty segment is one shard whose statement applies the rewards and payouts only. The node cuts from its own per-transaction boundaries (`TxBoundary`: the carry of 7.6, the state root after every transaction) with the cut of `igneum_prove_core::plan`; `igneum-prove-export` must reproduce the node's plan on the same export, and the test network checks that it does. RPCs (the execution layer's JSON-RPC): `igneum_getShardPlan(block)`, `igneum_getProofRecords(block)`, `igneum_submitProofRecord({record, proof})`, `igneum_getAssignedShards([keyHash...], lookback)` (the prover's work list), `igneum_getProvingStatus()`. Signing without the BLS key material in the prover process: `igneum-miner sign-record` and `key-hash`. + +## 7.8 Segment records, the chain rule and the unproven rule, as implemented (proving v1) + +Added 5 October 2026 (Josh: "open the proving round asap"; branch `proving-v1` of the fork and of the main repository, `docs/plans/proving-v1.md`). Every item is Implemented on the branch and behind `proving_v1_activation_daa` (default never); nothing here changes the devnet until the 0.3.11 rollout sets the switch. Item 9 is Designed and off. Proving v0 (7.7) keeps running underneath: per-shard records stay valid and paid. + +1. **The aggregated proof.** The aggregator guest of 7.6 (pinned, `elf/igneum-prove-aggregator`) verifies every shard proof of one chain block and, by recursion, the previous chain block's aggregated proof (`AggInput.prev`): its public values (`BlockOutput`, 340 bytes, mirrored in consensus as `BlockStatement`) carry `chain_len`, the number of consecutive chain blocks the proof attests, and `agg_vk`, the aggregator's own id whenever `chain_len > 1`. One compressed proof of constant size therefore attests any run of consecutive chain blocks (measured: `docs/bench-log.md`, "proving v1, aggregated chains on the RTX 5090"). This is design 5.3's "segment N verifies N-1" realised inside the proof rather than beside it. +2. **Segments.** From the first chain block `A` whose DAA score reaches `proving_v1_activation_daa`, chain blocks are grouped in fixed segments of `N = proving_v1_segment_blocks`: segment `k` is `A + kN ..= A + kN + N - 1`. The grid is a pure function of the chain, so every node names the same segments. +3. **The record.** One `SegmentRecord` (`kaspa_consensus_core::proving`): version 2, `first`, `last`, the hash of chain block `last`, the aggregator's BLS vote key, the payout address, the aggregated proof's 340 public values inline, the SHA-256 of the proof bytes, and a BLS signature over all of it under `IGNEUM_SEGMENT_RECORD_V1`, domain-separated with the network name. 586 bytes. The statement the verifier checks is keccak256 of the public values. +4. **Carriage.** Records ride in the coinbase extra data as a section `records || len_le32 || "IGNS"` placed before the shard record section (`IGNP`), which sits before the finality section; at most 2 per block. A node that does not read the section sees miner bytes. The proof bytes travel beside the record on p2p message `IgneumSegmentRecordMessage` (type 72, protocol version 14; peers at 13 never receive it) and through `igneum_submitSegmentRecord`. +5. **What every node checks on a carried record (consensus).** The segment is aligned on the grid and its last block is on the executor's own chain, at most 600 chain blocks behind the carrier; the signature verifies; the public values equal the node's native block statement for chain block `last` in every field but `provers` and `chain_len` (`chain_id`, `number`, `block_hash`, `parent_hash`, `shard_count`, `tx_commitment`, `pre_root`, `post_root`, the keccak of the shard receipts roots, gas, pgas, executed, skipped, `shard_vk` = the pinned shard program id, `agg_vk` = the pinned aggregator id when `chain_len > 1` and zero otherwise); `chain_len` is at least `N` and at most the chain height; the chain rule of item 6; and the deadline of item 7. A record that fails is ignored, not a fault of the block: it pays nothing. The SP1 proof is not verified on this path (item 8). +6. **The chain rule.** A record whose `chain_len > N` chains to the previous segment's proof by recursion (the proof itself verified it), and is valid whatever the chain says about that segment. A record whose `chain_len = N` starts a fresh chain and is valid only for the first segment of the grid, or when the previous segment is unproven (item 7). So a proven segment is followed only by records that verify it; an unproven one may be skipped. +7. **The unproven rule.** A segment whose last chain block is more than `T = proving_v1_unproven_daa` DAA seconds old at the carrier, with no paid record, is unproven: the pool pays nothing for it (its aggregator share stays in the escrow), a record for it carried after the deadline is invalid, and the next segment may start a fresh chain. A segment with a paid record is proven; before its deadline it is pending. `igneum_getProvingStatus` reports the three counts over the record window. +8. **Verification is off the consensus path**, as 7.7 item 4: the proof pool verifies segment proofs through `igneum-prove-host --mode verify-segment` (SP1's light verifier against the pinned aggregator key; the shard id and the aggregator id inside the statement checked against the pinned ids; keccak of the public values against the statement) and a producer offers only verified records to its templates. The native statement bounds what an unverified record can do to the aggregator's payout, never to state. +9. **Payout.** From the activation, a chain block's pool credit (5.3) splits: `proving_v1_aggregator_share_bps` of it to the aggregator of the segment record that attests the block, the rest to its shards in equal parts as 7.7 item 6 (`split_pool_credit`). At the carrying chain block, for every segment record its blocks carry in sequence order, the first valid record per segment pays the sum of the aggregator shares of the segment's blocks to the record's payout address, from the escrow, after the shard payouts and before the transactions; a reorg unwinds it with the carrier. There is no aggregator sortition yet: the first valid record carried wins (Open, O-7.3: the VRF draw of design 5.3). +10. **Mandatory proofs (Designed, off).** The rule that makes a proof a condition of validity: a chain block is invalid if the segment ending `T` DAA seconds before it is unproven. It needs an activation height of its own (not yet a parameter) and a consensus-level check in the block validator; it is written here so the devnet measures the coverage the fleet can meet first (`docs/plans/proving-v1.md`, step 3) and Josh sets the height when the share is one. + +RPCs: `igneum_submitSegmentRecord({record, proof})`, `igneum_getSegmentStatement(block)` (the segment, its status, the native public values for a fresh and a continuing chain, the shard proofs the pool holds per block, the previous paid record), `igneum_getSegmentRecords(block)`, `igneum_getProofBytes(block, shard, keyHash)` and `igneum_getSegmentProofBytes(first, keyHash)` (the aggregator's inputs), the `v1` object of `igneum_getProvingStatus`. Signing: `igneum-miner sign-segment-record`. Host modes: `chain`, `aggregate`, `verify-segment`. diff --git a/infra/fast-time/override-60x.json b/infra/fast-time/override-60x.json index d5db07547..1b5a2521a 100644 --- a/infra/fast-time/override-60x.json +++ b/infra/fast-time/override-60x.json @@ -55,5 +55,9 @@ "proving_v0_activation_daa": 18446744073709551615, "finality_v3_activation_daa": 18446744073709551615, "fees_v1_activation_daa": 0, + "proving_v1_activation_daa": 18446744073709551615, + "proving_v1_segment_blocks": 4, + "proving_v1_unproven_daa": 60, + "proving_v1_aggregator_share_bps": 1000, "fees": {"pgas": {"version": 1, "cycles_per_pgas": 1000, "intrinsic_pgas_per_tx": 300, "modexp_base": 10, "modexp_per_byte_numer": 1, "modexp_per_byte_denom": 10}, "block_proving_gas_limit": 120000, "shard_proving_gas_budget": 30000, "min_execution_base_fee_wei": 100000000000, "min_proving_base_fee_wei": 10000000000000, "initial_execution_base_fee_wei": 100000000000, "initial_proving_base_fee_wei": 10000000000000, "base_fee_change_denominator": 8} } diff --git a/proving/fixtures/chain/block-81046.json b/proving/fixtures/chain/block-81046.json new file mode 100644 index 000000000..219769281 --- /dev/null +++ b/proving/fixtures/chain/block-81046.json @@ -0,0 +1,1325 @@ +{ + "format": "igneum-prove-fixture-v1", + "source": "live devnet export from node 1 at tip 81076, 5 October 2026 20:06 BST, proving v1 chain fixtures", + "block": { + "chain_id": 4463, + "env": { + "number": 81046, + "hash": "0xae10dd9883739a2cd9e95cafa4f9cc7b2e38e477207ec49fb3fe64e3fcf21613", + "parent_hash": "0x4e420f14ba78477c090eef99a098b62dfdadc3a61f7c0a5499797380763693db", + "timestamp": 1791227122, + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "prevrandao": "0x63d5cfecd24e23244cb2c82def5d7a2102444d1a0f06a2d2239045d72498a2eb", + "base_fee_exec": 1000000000, + "base_fee_proving": 1000000000, + "daa_score": 0 + }, + "block_hashes": [ + [ + 80790, + "0xc1dd108438deeb4865751fa050cf2c5bce2160c2e8d306179297d2273ee7a2ac" + ], + [ + 80791, + "0xd89b9e7a2c04f5067579d93e7b2174011b2efb213dc666c0894686bcd7abd7ed" + ], + [ + 80792, + "0xb15e9df888384729a98f0dc1bc33017e71bfa7d40677326f2461a950095804e7" + ], + [ + 80793, + "0x1a75bb3615a32dad3eb935c2a1a3f4d2791cb266be47ef17bd4285962108ff15" + ], + [ + 80794, + "0xf657d95990ce07d72922d34f07dbc66bb681797076a3e9819b2b3349218efaec" + ], + [ + 80795, + "0x4bf9d010c0798a5340983b27eda6a9876e6fefb925104fdc46291a5789f1f70a" + ], + [ + 80796, + "0xb09af3831aa1d8f4945d50fa16a602848a2da22b14dbc1f7af2844cf19e90040" + ], + [ + 80797, + "0xe099536cfd64654690afc152b97866b32da6bf81f959a6c61f9e43bfd2ca2d9a" + ], + [ + 80798, + "0xd6af2cf6184f15ecca77ee9e7e8d2ca718d88f9cea39b475713ceb0e151ddc36" + ], + [ + 80799, + "0x9c76f9f6036c7035fd52a7a4e3fca99451715e729fcc9fc424487faee9a6134d" + ], + [ + 80800, + "0x9a046cd99eb89c551dd0c698f1969069c02ae36cde4140ff4b864d624fd9dff7" + ], + [ + 80801, + "0xc9bffa329d649c6540518cf249e25e845743dc01fcbcfdfb8c798f1455efeae6" + ], + [ + 80802, + "0xb124f2de3c17a985f10c0b58d2283631bcc12524de8d79b0955aecfd2d7f78aa" + ], + [ + 80803, + "0x97b21d6cd3f1b321d2c9729062c305a7955b97646d76910714f899c75d275568" + ], + [ + 80804, + "0xaa3f0e114e156f706e11a701fb940bf54e89b09d83953be7f6733478c7a33093" + ], + [ + 80805, + "0xe5a5868b02ec30304079652a8d5b8af4a9da0a5a676615c1a10d833feb4eb0f7" + ], + [ + 80806, + "0xd97c5c4288f56955e1a3e134eb63250e1becb907b0fdde9ce1383d8d2c8a3939" + ], + [ + 80807, + "0x8f7bf1128e5d03607cca246b86e63cfb4bfe73e74395c8660d027172376b5c36" + ], + [ + 80808, + "0x5d5fe4c3bee7d758986da5a433cfe5221be392419ac22e5acddd701cd12c8142" + ], + [ + 80809, + "0x9b0d0a871ce8eaab11a185a4344f9f48a58b7111f94a8f5221d557722cc5ca5a" + ], + [ + 80810, + "0x58cad4ab776ba7eca94e813323e28a6467a7fcf39d5aa2069e9dfef6f8058e5c" + ], + [ + 80811, + "0x706788247d6c05475eb3bdbc6f181e76429b57a97f168004832fe5ba4b20c7b7" + ], + [ + 80812, + "0xdf26fb97f999b1016865b3fee539e72e3f66d6d228b2aec7e2f59ab93e5a9ee6" + ], + [ + 80813, + "0x8e612c07fdb693a44353a08899e8beae8c93fb1b85845ad965f96e7d2f77fed4" + ], + [ + 80814, + "0x806f4c29ea28af1a3ce0c3ba775c776006ba8cf4b308ee617dc207dc3f740b29" + ], + [ + 80815, + "0xebd8e602049ddc2f8d11d5edc8ee837eefac77418e73e9bad97f6f1c7a85e8aa" + ], + [ + 80816, + "0xe2b069b77ac05cb0a95ae419966feeb20660e4ff3b9a68d044a5721e1525dd86" + ], + [ + 80817, + "0x1ca5b79e26dc425a94c3a6bfd339516caa0a7f8e5e2a62389bfa466186bca495" + ], + [ + 80818, + "0x7cf382f956597cbd36917036dd15b774500f97714e50587a35ec9617662052d6" + ], + [ + 80819, + "0xe70494836c5b5c68b6a167c868db166f9630140499f96e6660d3e3928cad6a14" + ], + [ + 80820, + "0x15f73eefc39e00f9d09390764f9ea0069908062c563686e73b49ec3919c24be3" + ], + [ + 80821, + "0x7f2bd2cf68fc14e6a5debaac39f368534657c41cc587402fac1156d24998e55b" + ], + [ + 80822, + "0x2641ec1c3397b06ab5855dcf6ec2e23de301f5831015a32ffa9f29e18cc2f084" + ], + [ + 80823, + "0xf643e963f5a62a49e935b4ef51aab9bf2689b0b9ec483924d7f42934c8ce5e07" + ], + [ + 80824, + "0x867ee0219bd30fc3117c79295cb1075ca21281fdd285467ae128aece1c1dd9a7" + ], + [ + 80825, + "0xab6bb19259a039d30c3022c9da3494af8c9a541dce30895df083252894c433df" + ], + [ + 80826, + "0x155511de8d34bef351ba0ae7af5bccf515e6e490135130aad5ce2eb02e103d75" + ], + [ + 80827, + "0x4e297a53901da9b364f4dce686f016f4039658ea8832d591dcb21aa44acd197b" + ], + [ + 80828, + "0x7d07b8bc9b06b20e7b56f96fff9051590caac01159d39b3f95bd592a4627daf4" + ], + [ + 80829, + "0x8392a1875388baec163e64854d6ee7664a31017402ee8e269ddc8aa42ca06647" + ], + [ + 80830, + "0x5ab4e5c586d2fb4272941bf26c15288cc129ca1092d3aaab82c478edd1942e29" + ], + [ + 80831, + "0xb0100f0a559daa435128cae74f47f9353ca2a93c2bdac8dd9d12a3130ce6a818" + ], + [ + 80832, + "0x326d1388d0911871c5087925500c4ee5091624331d07659b9d2fa09e83f5884a" + ], + [ + 80833, + "0xc01c186cfcebb8d53953589ce595afef213d46e481b8b5dee7e0a71abd773f6a" + ], + [ + 80834, + "0xe845934e79eaeb8a8ee0b9396b3a13f13c8ac45a0d4e81017c73042789013c38" + ], + [ + 80835, + "0xa45f8f03249624a2d35b54eb83f4edf54e8d1acd061fee3965d9164a132c72aa" + ], + [ + 80836, + "0xe72a59a64af9aa223b5341819a5872f8626908b5c1857d0617706f6999632e01" + ], + [ + 80837, + "0x94063f426663347bb4a3e79f4dedeb76fbce6fd69117a2c74ad1fb4714b856dd" + ], + [ + 80838, + "0x157ab5f9e7fbf98b48af049adcfb4a8e071253ec95b991d031a91668ff3b8d57" + ], + [ + 80839, + "0x5bff632aa54b1a71a7d9a2a57d9c42ef223eb6efe86c72f9aa8dfd67ab7b218a" + ], + [ + 80840, + "0x7f5d1952ad8ad27dd17bed3354395b646468985bae0341b4048376a425e43076" + ], + [ + 80841, + "0x8e69df81fd1b0f186813304260dda1d0c9dad5792126e69d021e679b837b01a5" + ], + [ + 80842, + "0x6733c5472e38830dfb273a0e0802924c1b60a2cc916d6efaf82785e1bf960724" + ], + [ + 80843, + "0x0838513dc8264efffb2c11f8c00ebbc9c57aeb67dd7be93f18e1c613872bc929" + ], + [ + 80844, + "0x8c2f26b17db95bb20983ac41df7a5fca3e3826d2e57831eba13bc3175daeee34" + ], + [ + 80845, + "0xe3edcab823e03598c66e16c5743f7a0917735c975ac946db4cb3ad862db40a21" + ], + [ + 80846, + "0xd075510e75159d941946e666e2c75604a3f3c959523cbe2c0f025427ae85d4a2" + ], + [ + 80847, + "0x57b4b6c0a6ca2532515653445bc872dd44e621f9774f3dd4e2000bf82a0f510a" + ], + [ + 80848, + "0x38a85c99bc3e489a4431ba9a125e74f4d807c630a239419287e039f05cc91ead" + ], + [ + 80849, + "0x3c95d08c99482d57bbcb9fd333cae4010628fc8b1c96b3ad4a68d8f3dd21695f" + ], + [ + 80850, + "0xe767527430f19da2c2b21bc648e33ed2b241f5b7c9413649d9c5a555491a9253" + ], + [ + 80851, + "0xe9941baa9f3378854aaff5563340d130700f0753197c7373229542c56ac91a8f" + ], + [ + 80852, + "0xb92e88ec7cc8c340959d8d0cd5434eec5c1033b389b953b02cda103d53292df0" + ], + [ + 80853, + "0x39f01f35ba2ea7b45e3eec1681c44bd85b066a6a8521a10fe6cec5d4d98aace8" + ], + [ + 80854, + "0x36456a5e36a9a66488656bc0a92705632326e59ca40bd09f6d5f99ccc976972a" + ], + [ + 80855, + "0x12e52067750279b8bc34b03cd21bdb2809ea18a3aa68739623780aaf1150129b" + ], + [ + 80856, + "0x64951c693d7f246647fac0df504f35ff9812416f966afaf4fecf7ca92c0e0b24" + ], + [ + 80857, + "0x3918b797fe9d8cd6c34a0b315aaef550de15f4d30dce932f3be8cdfdc27343a2" + ], + [ + 80858, + "0x86ec8713a68034acb9f34b015a48644c7562766c1760927a299e517ae8ef6f8a" + ], + [ + 80859, + "0xfe75a4c99fc17131b7aae61dd0a588b7d7019e6c662aab8c7d6a8c56370fe07f" + ], + [ + 80860, + "0xf4afa8f14a0f814639c88dea304a8ccf76450ee56f4fbb89caf29dbeef80399b" + ], + [ + 80861, + "0xd4c72bc192643e20658ca54b730ef865d1c29ebfb6d1f78263c796a23b6965b4" + ], + [ + 80862, + "0x7a73b7302ab47b7acc1c5359620df9026fdc08d78c4173bb4e45e73a66651767" + ], + [ + 80863, + "0x9033ce568146d2c69139d3674fa7e18cb8319abae5e6cd24f924f750d2ca3446" + ], + [ + 80864, + "0xe4d8da84ba631f233ea7dd1336a264f499083b08290c82b717f79bfb7ffa09a6" + ], + [ + 80865, + "0xc53ec9d6b11d592921f5cf850584ebb562119ea968a5fa7bb00af381d6eaef46" + ], + [ + 80866, + "0xedcbb7d2f9626af53341464fbfc54e838e6417b7ffcfb0c2f9758641c98d5f9e" + ], + [ + 80867, + "0xbe5f726816d43ca5bbcb70f13901e60ef9040debdfaf2a73a0692c3cbef578ec" + ], + [ + 80868, + "0xb7cf72378b14c086afb4094b9b0121fcf7c6929bf2c2134b5189862644e2e506" + ], + [ + 80869, + "0xa29e21de7e54defda356f088dea38a06cf1a1a75bf93c14e30afdba213df7ef3" + ], + [ + 80870, + "0x504d3f83a27c37c5a15cdbbdfc03285f53071ad01d82bf9be2d6f862548a9a8e" + ], + [ + 80871, + "0xe02107fefe75bc8d1abb0443bff2e7a658e3a9b2fccd5157c1a7c7ccc995fe17" + ], + [ + 80872, + "0xdb92e1b449e7670ea719a6d68f496965620656ff5d25a0521f1873eb8d7ae1ab" + ], + [ + 80873, + "0xf12689607f86ce8f762df5d67054a00b8b580300511a221ff5addcf732393988" + ], + [ + 80874, + "0x0e0fe961c78fe52f44f6742795bae3e96713f4618e4a95623045056eb9e046f6" + ], + [ + 80875, + "0x7fd5bb64e800686ae28c3b55613d27f181dbd91e35d342d33d7b3e5d4df17d00" + ], + [ + 80876, + "0x83c2feae4eacf2ada1c9a630a97614a6296df6b89bfb5f954c3e306b16f0ab2e" + ], + [ + 80877, + "0xd1f3af4d6464c5bc9416d73a291bcf3e1df4fa1a396bb3e9dff92dcd451f75d9" + ], + [ + 80878, + "0xa4cd4b49e1d475a5561a91a0de592ba0c00f63b39a61eab150c6486d1bf2b916" + ], + [ + 80879, + "0x714904f162c3b9013a9defd79d4552b58c5ca8747cc94187079f2ebefee2f0d9" + ], + [ + 80880, + "0xb2263d12431e74591e2843047d79933820d2dd352d5b5c977698b328cd59f46d" + ], + [ + 80881, + "0x67977d72a573f121d9a818dd035959cd502e5bec07a4a5c30a28fc66d4fdf964" + ], + [ + 80882, + "0xf1d5f1b271cd23a79f1ab32655418e4b664663c4a77ebaf0a213fc2456e59c5a" + ], + [ + 80883, + "0x91a9f94d339407747a77d79a9c2be608744d4ecc8b84699d0120814f55a5060b" + ], + [ + 80884, + "0x8f709bf4582d3a912e6006f6c417459bbbf182af2b500b4376cd8dff78201c25" + ], + [ + 80885, + "0xbc66848a66fc038200186f09665540b5905a350a3149bfdddf9575eabb2f0a98" + ], + [ + 80886, + "0x8827e43b9a02e7424513e6d0fb6db7e93cc5555fac9b832584570784708a8f90" + ], + [ + 80887, + "0x7c76b6619ea66b8e7fa2791ad1a3dcafe863388d14db50dd022f63748e2b6ccb" + ], + [ + 80888, + "0xace6e76c69db6530277c9749ff3e77a64c6e7123790fdf5f248cc67b761e10bd" + ], + [ + 80889, + "0x8f96a0d2f9524f0b2bf9f520f35516969fba795e005f003dac5721cc37c6b5ce" + ], + [ + 80890, + "0x68a635c0d7ba10b3fff9108deecc4e9af70d760d688c5614a000503b1f2f006c" + ], + [ + 80891, + "0x47dc91daaf1758699e90f8906c61bcb6dd020b3726a1e0d15a7fe5177bfab696" + ], + [ + 80892, + "0x79f1ff991dd34c0384b9ddb64a7f3791116891447f2c960de5c6d80f9a705b5f" + ], + [ + 80893, + "0xf41832c4af88cdd31eec42b233ccfae51a150cd41baf59e9dc6b64dba2b63c7a" + ], + [ + 80894, + "0xfaf2c58e9ba01047e4cbe2c147a4614d46ad39c1d64cbb43d1de60f9c75286b7" + ], + [ + 80895, + "0x65a671bcbc9b6f905811ce62594feda18a12b8c6ca1c72081c1dde3f6227d1a7" + ], + [ + 80896, + "0x815d16976077a9b599fd5d6a82eda30a0bd1d91c66718528aaea9d4dcd12274b" + ], + [ + 80897, + "0x710c0c1d4d44ed7925343b10c7ace3216741b31a91506fa0f1ad80c96c049b14" + ], + [ + 80898, + "0x5fff52604275c206678c800d95cd0569325ad67413b50ac15252c3440856cc78" + ], + [ + 80899, + "0xe67fe87df50a63d348d300de16353e560c1b86f1a79a0bd09e2a7771ae7e92ba" + ], + [ + 80900, + "0x5d854ceaef9de086a2361cfa1f829f843aac00181fa5ae27e69171c198853895" + ], + [ + 80901, + "0xca45572516a2d4a948cfe518d0bb5378a4d23b800522f515e59176c31c069160" + ], + [ + 80902, + "0xa724d3e05f16d3d8c97296f59802c7f9ebba4284808126d955d6fc96dc1c4729" + ], + [ + 80903, + "0x552061794bddc5185033e936ac8ad07d3a315c7a3d876989b2d61328ad0a3128" + ], + [ + 80904, + "0x23c1ee3ff43ae0d42b9b3f0d2c6c2fef6c8ab8fdec9772540b563cfebbb77a2c" + ], + [ + 80905, + "0xc7e33046cb4819a17a21a00e9c0efa55765a974d8a2364f4bed9a16034ec2828" + ], + [ + 80906, + "0xa7b5712f7d22449a6f6a8a55a33c2f76cb54be86faa2f0d892c8bf6aa9fe59d9" + ], + [ + 80907, + "0xbd8696f54d266714d694b41d3b6466ef242e999938b9effd2de035c3f2c5677b" + ], + [ + 80908, + "0x7ca324c51788d4cc173bd3a013702f80e68524b3b005299be9012621dc8e82c4" + ], + [ + 80909, + "0xaea0e3a560abbe0de857592d37bfd43bd986d1cd461f43139d951d123140d495" + ], + [ + 80910, + "0x640f0ed5ec76969f53d90d5e502f6fce0ec65ab28a728e92a81dd27c466dad09" + ], + [ + 80911, + "0xf4cb4beab37872930f84e37b38c15a0a5e5f5d95927c91284ae880337abcfd9a" + ], + [ + 80912, + "0x65121e870e1668a0a55c9a50d31166294cb8e57bc007897640a7414f0866ec3b" + ], + [ + 80913, + "0xb3ea0a195e00db811104db7362b555fd1ab8401814ee841ca7f4d3344843708c" + ], + [ + 80914, + "0xf92735fd64122d7f6f6df6f20f83fc0793675c81e25e54dc64b3ca027d8ef9a4" + ], + [ + 80915, + "0x78e07fd2fcedb51175785be4657b8d2a554c16c1d2cc396ca3738dbf6f9373f7" + ], + [ + 80916, + "0x0ce35f9c3a736ff2937e6b7addcb632855e347ab3c5c08b74d2dca5f82265426" + ], + [ + 80917, + "0x007b1e317c0d9e7565260e1fe7ebfe031de4291f6ddad75210f8563479e84178" + ], + [ + 80918, + "0xbdb4b040d7942b8bcd206e4e12632892d9cfb87c905f760f0c918dcb682d4eca" + ], + [ + 80919, + "0x1eb8ee38b4d246a0502cc56c8948ec6e17fa55b02002c9916e1a3d9867e3c852" + ], + [ + 80920, + "0x2bd9fea5702cc9c0b96fa3b34559c4a72c5ff4f1d572d9e702390a13b0b1c145" + ], + [ + 80921, + "0x12ae91b4df38f549004ab8c4482d9ed1a2841e43f649ee1a89156f9accea51ea" + ], + [ + 80922, + "0xb960711112f51fc11689061a66a9b4ed3c0bccb3a6f027bd2cceeb2d3786e780" + ], + [ + 80923, + "0x4d474df4920e9136e7564cf0e9ebd03bbcd8cb8bdb38c68b8a9c35120888bfea" + ], + [ + 80924, + "0x73d84f1013038f10bebf6606486c1bb6c6cf530a4d35f482be91df74fe15c93e" + ], + [ + 80925, + "0xf08dc8a601c46c43bb6f94edc1b0734a61e7f8efc82d915b33136fac63d5a407" + ], + [ + 80926, + "0x2b695a3c4d05e233cbe18213d50b10fb9d83914aca9a935aa6f6639b1f54c32e" + ], + [ + 80927, + "0x893a6295ba2f66f083feaa39e0570ccbd446166c2e15238d59a911f2826a47d8" + ], + [ + 80928, + "0xc2b2870e39e8344e18296012c4d846b78cd3bdc80145de6b595527c8845ac57b" + ], + [ + 80929, + "0xc32bf651e3ea72bd6808db5b4dbe4577078d418404221809ac12897082ab8bfa" + ], + [ + 80930, + "0xf1c2f5ed0a7d6b7f083c1a68f75004fbfc929ef3f6ab3b46acbba373feecfdde" + ], + [ + 80931, + "0x51638cd9a09c8ee630706d736c81ea0dde7d97234a515c6064be20d0b92b8418" + ], + [ + 80932, + "0xf3cb85e3f551c4e21d0c6c28e940c6bd965c2a20b5eb3375a9d8d32a65c58c4c" + ], + [ + 80933, + "0x65f05787c7f239dcd4dd9cac2e5c516e25f8a8b898eec284531c16a8da83ed2d" + ], + [ + 80934, + "0x58567dd84c5ce96fdfb6ea5794201199d259374e4879e76df9a6622694de95f3" + ], + [ + 80935, + "0x2031a54f25cefcf00c464f36897691a73616f777582bfcd9c03e5649d3339821" + ], + [ + 80936, + "0x0fc1b95d2579a6cb03f21083da53a989db2491f37bcc339195864baf21932823" + ], + [ + 80937, + "0xfc501a03f2666c7c7db743b6a7296692789824828c033f2254e3565de63bccec" + ], + [ + 80938, + "0xe752390723b297a6d326cc9d10f5a3219b5eede3159f9ea1c7fc7f26401b02f8" + ], + [ + 80939, + "0xfdba580d13c6964b46e7d2a66f2aeee150174facaf1cbff66016832b02d9f796" + ], + [ + 80940, + "0x2221d3b867caf0c1d4993845a60ba8389558ccd9d48bf97f51331e87ff51ce84" + ], + [ + 80941, + "0xb4ba8a208861ab1db5302570f7d589adc5f4826812b4759d625cabe2b68231a0" + ], + [ + 80942, + "0x0bc24048f15b64e0c836c387539686906198ed0b3a7adbe2aed37f41eaee5577" + ], + [ + 80943, + "0x5c3230d24716dfeb57c624bc44f7d80b65c95f4b810aa945f8e5309e45327104" + ], + [ + 80944, + "0x798ef63bf3ed33e8608b899ee956fc3a0a0b359a4466141728589ea914d42453" + ], + [ + 80945, + "0xb158a6b683a60e8d3debc803a89bd18af0f1307941039304a2e33a02f237325a" + ], + [ + 80946, + "0xd420374f3f533e0136b3fb1211357adbdde913776fbb851492ca80980e77ce2e" + ], + [ + 80947, + "0xbf7fea59019b30fec39a94c29b026285e0c2d6fe034da2555095999ca694914b" + ], + [ + 80948, + "0x3d232c12ca3a428474f8c8e21992efc614775842c365a29ef44d9b7f24caafc8" + ], + [ + 80949, + "0x87d17a79601cf126d2b5c95f1ac989f8a90f2b4e9f3b13d1d721776ee2a53c24" + ], + [ + 80950, + "0xd23290b4b63c19aafd1edefd1cfcf34829bdc6efa723df17cce5eb7612b17a0d" + ], + [ + 80951, + "0xa8d8db2850de878732f5a7b95bc08ddbb944e1e8e8884a2a2a6a98275bb7663d" + ], + [ + 80952, + "0xb2c005429b5d1f35d006693509651726e67fd19d34ac04a3473df90d8bed52ab" + ], + [ + 80953, + "0xfcdb4ee3a4fa1273ec429c49704ac1afe6c8c6f858640155276661b19a5c4ca0" + ], + [ + 80954, + "0x9b988ccb0403ebfb0c15486ed6b63ea070121d7754f88d1f14f85651c083f240" + ], + [ + 80955, + "0x39c3d647ae0dd09b1bbf1699803e3c1e76448787baf12cfced826b8d75dd4b7c" + ], + [ + 80956, + "0xb24d2ff081df7d34392e460ddf8fdb8c7cc5ee502e8cae2a279474f329fe7d05" + ], + [ + 80957, + "0xf1f1daa9a70d68fea94a204260b298285ca15469cdc1855be3ef731ae44707bb" + ], + [ + 80958, + "0xeadcfd836598031d7d1a6f5c14028edeffc5baed420b2fbd343a5e21cfdd1fd9" + ], + [ + 80959, + "0xe1391944a4bd9d2fe474c8b12de8b3d151bf3afc97addb223ebaa59592ddcd94" + ], + [ + 80960, + "0xdef9733fd939d6b260fe671867ee7cdcd802332b9afdadbba0ac24b964126d29" + ], + [ + 80961, + "0x641d150032ebcce0d0d7616b010eb29df976191edfad5c8751da8533f78f7946" + ], + [ + 80962, + "0x7480f68a44e03c35a1ece1cec7402fe9fb5b4d34d4f246c967c5178e1a7fa330" + ], + [ + 80963, + "0xefae600a6e868a4213011b106f295f79e5c86186b8382143c0a402761226a9ec" + ], + [ + 80964, + "0xc6ce05cbc32d8a1d786baa6473cc65c29892a5a47e7474f384b65c567d425748" + ], + [ + 80965, + "0xd3a3e3c3c964a77a0821984903c958079b294b195b7b13e109f26f3bafcfa092" + ], + [ + 80966, + "0xd42712815a77eaca169de5c74653b4c007159f47fb5d50d63f57cfd3f2c5868e" + ], + [ + 80967, + "0x9b309f8c13534ed43e399d45a5bcfdb38f49c5633b458f6db71d39571e9d8d41" + ], + [ + 80968, + "0x74729f1a9de838f15def5f5c183b7d3596b03afde1ed6302cabacf8a677f82a2" + ], + [ + 80969, + "0x4c07a8677c028eda23e1665416eb54d364959f3210e682978082d560ebf2bea4" + ], + [ + 80970, + "0xe3f992775c7dfbd2485ccecf516ca7ee9961c897296190c45fd7abb1431ee04b" + ], + [ + 80971, + "0xa8492953412a007f2f7696922b661a27fe2ec42e566504b3aacae2f97a24fab0" + ], + [ + 80972, + "0xe14d6f60ece95189dc8eaed8a2ea5c2da62cbdf988588222a4e99d242ea3beb8" + ], + [ + 80973, + "0x5c7ca844c8eec0cce4e4640ba6fc486c1739eb44aa50de86f55008a4d597357c" + ], + [ + 80974, + "0x2b33e852101ee35f2c61ee7b35152521eefb11364505bf6903a964c93d4001b6" + ], + [ + 80975, + "0x5516b628a8ec4328bcfdf90785fb5d96d22d60b4db638ea1b93bd837d7c739a3" + ], + [ + 80976, + "0x9959b993b66dc1285d0ecd8c95b74fe70949c50d5a97d822ce866291f94f2926" + ], + [ + 80977, + "0x04eb959364eae151a41d4a8d970795da0bf8a0ef2f64bde753cbca4ea3d745c4" + ], + [ + 80978, + "0x22907520934a83b056017bfaa72a0ce92fd8658c81dc8236f3b78f1f7e055aa6" + ], + [ + 80979, + "0x59fae978e8063799b4ab6335ca7ac097a25cbb4d60f9af04eb2f647a92ead4bc" + ], + [ + 80980, + "0x9ef7533f2e9c60f4acf9618efd1c47b147fa95bbb27caac54c1bb19540befeab" + ], + [ + 80981, + "0x9986f742ffa9d17f641293e788e53cd904b6edba3e971603470186f61f934e77" + ], + [ + 80982, + "0xbfb394920e4bd707342ddd41644a5733b7041e517d4e88863924046cd512ec19" + ], + [ + 80983, + "0xa3fef379a91df7a254bc4897cf2ef61ff25b380b2142374ff40d91d5140ed21c" + ], + [ + 80984, + "0x71626767d14a0735e3716b45b847611dc7fe04c6d7a4b22e3a6031d621b6d560" + ], + [ + 80985, + "0xb2c03912458086a6b421a8c83cee657c9276a1d38db128799a3170cba30e3266" + ], + [ + 80986, + "0x8f8000a6e85c09716e8766c9bbb9df1287d46f54625ff139d5e98142a58c9253" + ], + [ + 80987, + "0xfc88dba1647b7d43ef19a1a38780dd238228dd46beeb50ea0e00853f06c78471" + ], + [ + 80988, + "0x7e383c88765b3c499b05b60cdd4470cf9b23add0f6bd33990f207a43a51e431f" + ], + [ + 80989, + "0xa133fe8b34c9e839bcd4a75c2a5858874fa19d8bda43cd1c1af781cd2689753d" + ], + [ + 80990, + "0x466a4d175cb847a9899d98a6863900be93ecfec2f1681c8595db5a62b843478b" + ], + [ + 80991, + "0xb9b1bf393868e53ec5da954cb87fb65eef79ab9caf59cfb636f6cfd9c5c92f8e" + ], + [ + 80992, + "0x2f9eb9f04e16f420d5b98d6a28e25494207bc05affb6b914a1b98fb2244bba75" + ], + [ + 80993, + "0xc73829f98810dd017f45a0bd71ad7db3d07a1f5ff0a57c07994fa095989e6ef2" + ], + [ + 80994, + "0x9cdecf07b8f03979fdcb8f7d4d2c5ef4a13adc19edc0bb0c87a8e0beac6ff3be" + ], + [ + 80995, + "0x978c9e3471a4c37fb6aebab56d41d660720681073dfc68e5bef41d594d67ac8b" + ], + [ + 80996, + "0x739dad0614f0e852037aafbab92b4072f0f9dc9fc30c489465d67a18023f80a4" + ], + [ + 80997, + "0x1fbba34f589373e32f170a1cc3bf5c69083e8e2af5167e6ff38eab89ca2c0a72" + ], + [ + 80998, + "0xb58b192d29e3605fa4006c019698a7fbdc14639784ceeac087ae09148514e68e" + ], + [ + 80999, + "0x77f67acdb2afb4754c6e31005b27aea691cd3edce3094f87238df3f887e1bebe" + ], + [ + 81000, + "0x50e198025598c36879156ee586cdaa18bfa620f117c49baba1b2093d77b83831" + ], + [ + 81001, + "0x416e7d689843b2ddabbc9efd9b6f183b883dcb28e857977cd130f989ff7c72d6" + ], + [ + 81002, + "0xd639747fba2f48c62f7aeb91ead4b7ffa1292d55bd15ebec58879bc51b05863c" + ], + [ + 81003, + "0xc43806d1ab9e6d84be147ac3b74130fc2dca01254b4a847eb2713c3e7c1cd11c" + ], + [ + 81004, + "0xd854d522a33e98313a872fb4c47675f54eeed084d99794c5c72aab4464ee70b4" + ], + [ + 81005, + "0x752458c4e11fe4d564250bd4bb0122c1a1a4fabb124942ed9c3981887d612854" + ], + [ + 81006, + "0xe3eef7ded5487286de24e9802f5dfa1afe7c477666f2a26fca7606d3b9b8694c" + ], + [ + 81007, + "0xc525b44a5e888c3bdb0f66edc5bb4dbe9ce9116279b01fa510037760319ca15f" + ], + [ + 81008, + "0x9552445c57b6e2cd01d95bf4013bb913ce1611e5d707f71e528ad7522e830806" + ], + [ + 81009, + "0xd0eba42e3dd5cb37d42d6891eb47e5e2bdaeccbf9ddabeba06a09cd2b4b15f55" + ], + [ + 81010, + "0xcc6fd1b1758929e8a94d3bfc1b4c5444c55023f8939d4b4fb539b523219b00fd" + ], + [ + 81011, + "0x826ee09096704d3798aa4fccfc02cdedc59bb7472332b624d3716fa3aa14e612" + ], + [ + 81012, + "0x3f82182c620ce4d5ac0a776b0ec52b590f7d30c3bbea039d9a6ec3c6ec8a4aeb" + ], + [ + 81013, + "0xbdc7e4e07b4fc9245d8f5e7d7b207af203ce271b96912963aae2012de4d0dfd4" + ], + [ + 81014, + "0x797fd61f65fdc95f190d0dec1ad3eee5c98211c4529c83d013236838eda13d67" + ], + [ + 81015, + "0x7928f9107a62bd4d597894701cf4bca71365052e3560c23af7fdfa3374105b45" + ], + [ + 81016, + "0x78dc4dfd56d144d629278047ce28d8649e02b55768a137cf98260e40e6c35989" + ], + [ + 81017, + "0x3cd444ed1bd6f3012197839db12d566729ca5fb8c133af4353afbd756104c03c" + ], + [ + 81018, + "0x26350e310daef399afac9ac19b106d5f160d8dc6921fd289b196b4e979742bef" + ], + [ + 81019, + "0x99e1715d9ba9a3d125e77357595bd918766a4d46c20d9db1dfa9ebbd002a7bfa" + ], + [ + 81020, + "0xdf16ba27999f8dacfde10d403cf9982fd3b7545793e678219f09818e04ee64d6" + ], + [ + 81021, + "0xc4690b3c4b89b7282988cd2f71c9e43d3c17d081f0c735d04fbd01247a74bd6f" + ], + [ + 81022, + "0x94a1951db605f9a7dff70e2fd919e154f1291d31b4baa6f9e88d8f904c74fd54" + ], + [ + 81023, + "0x9846feb46081eafd35bd5f134cdf98f7a2a270b52a36779de7fb132250823a6b" + ], + [ + 81024, + "0x3716fe1bc70f07f53358e7936ff4b652a035eabe199265f99809dc4aac49d002" + ], + [ + 81025, + "0x841c89865fcece06f48c90ca283a1dbafb49e3fc422ff4cc5d3ec7fa68ac0f14" + ], + [ + 81026, + "0xcf2be45d7a88f2ffa7fa45366404b8b0af131758b7706f099ee8daa6b5d17b94" + ], + [ + 81027, + "0xf3c7dd48103bf10e13bea5800814475a543f4bf53958e1ad464b112d0c763d00" + ], + [ + 81028, + "0x6ebafde3e8fe7067903074cbeb854b9957bbc45859143f618ada972fccdc2c0d" + ], + [ + 81029, + "0x05146b0d5efdf34a927593169ecce2d2d112a3faa3b79cb2df1a0a2e6a4bd7c0" + ], + [ + 81030, + "0x33def97a4f3ca92c7af02bd19bb9703c516f6df8917928734953ec9fddae68ce" + ], + [ + 81031, + "0x18bbaeb286265fc8edb34abbfd3fade65196f09125dae4d31347401eb635b63f" + ], + [ + 81032, + "0x56aacb0d332a22fabb3bf8ac90b030fb82337b3b22b5be334d9a4f633847480c" + ], + [ + 81033, + "0x8d5d77464c966282353e2019598e0f1d77b5afb938475275f4a3b0d0c29d9c5e" + ], + [ + 81034, + "0xc8ccc7baba52e4fa56d488c78abf451919aeba4a029068fd364059633446c80d" + ], + [ + 81035, + "0xc295e160d3378e12e6b96de0eac43d903aa656f9255e58a3b0222bc229b18bbb" + ], + [ + 81036, + "0xd4cacd3eb6e8eebf312225c23284c15e347e73539816fe91954b482a9acb1753" + ], + [ + 81037, + "0x096f15eb2cfa8a6811e08ee533e3a9bfaf3447225c214157d33f27d532423598" + ], + [ + 81038, + "0xd2464a278a262d447f8d67feb035e8b9eaaaa849eb87fb330e1ac868e6b1d063" + ], + [ + 81039, + "0xfea1ffb0177d362f9e041268c395858091da08fb55de5c922493bf0f3dd19476" + ], + [ + 81040, + "0x5ab2d5c6f5789372abc97b4792b2c4e3476c70bd538764ba74d684ed102fdd0e" + ], + [ + 81041, + "0x2f5be38f279790c0e69e554fcaa28f82bd17ef645338539416db0c9252040f91" + ], + [ + 81042, + "0x844ebc844ba1c0c301a6a9189779e17e557f824d85d01247b4a0b1af49f50d56" + ], + [ + 81043, + "0x944855a63daa3b755a70f2ba5e6a956b037eeb210dd9e80cef73b7c814645adc" + ], + [ + 81044, + "0xf046f77959a6aa199f3d999b2c16dbeb1e9038c308cb1dcc9bba26141e0e5f35" + ], + [ + 81045, + "0x4e420f14ba78477c090eef99a098b62dfdadc3a61f7c0a5499797380763693db" + ] + ], + "rewards": [ + [ + "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "0x3268dd8daacd1800" + ], + [ + "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "0x3268dd8daacd1800" + ] + ], + "proving_pool_credit": "0x19346ec4815aa800", + "payouts": [], + "blocks": [ + { + "miner": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "blue": true, + "txs": [] + }, + { + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "blue": true, + "txs": [] + } + ], + "pre_state": [ + { + "address": "0x0000000000000000000000000000000000000210", + "nonce": 1, + "balance": "0x0", + "code": "0x608060405234801561000f575f5ffd5b506004361061003f575f3560e01c8063aa67735414610043578063dea5c2e014610058578063fe7e05d51461009f575b5f5ffd5b6100566100513660046101e0565b6100ca565b005b610083610066366004610211565b6001600160a01b039081165f908152600160205260409020541690565b6040516001600160a01b03909116815260200160405180910390f35b6100836100ad366004610211565b6001600160a01b039081165f908152602081905260409020541690565b336001600160a01b03831614806100f957506001600160a01b038281165f908152600160205260409020541633145b6101635760405162461bcd60e51b815260206004820152603160248201527f446576656c6f70657252656769737472793a206e6f7420746865206163636f75604482015270373a1037b91034ba399031b932b0ba37b960791b606482015260840160405180910390fd5b6001600160a01b038281165f818152602081815260409182902080546001600160a01b031916948616948517905590513381527fa47563c41dab010f91a8ef9dc7ac2bcdfa0ef2af697e575048e71e6eec60dda3910160405180910390a35050565b80356001600160a01b03811681146101db575f5ffd5b919050565b5f5f604083850312156101f1575f5ffd5b6101fa836101c5565b9150610208602084016101c5565b90509250929050565b5f60208284031215610221575f5ffd5b61022a826101c5565b939250505056fea2646970667358221220cbf48f5aa6f911f83b3c2b09adf8c418f5da6fc530c3d5db9d4ba0be24c4bf5864736f6c63430008250033", + "storage": [] + }, + { + "address": "0x0000000000000000000000000000000000000220", + "nonce": 0, + "balance": "0x13cb303d3f5c73301000", + "code": "0x", + "storage": [] + }, + { + "address": "0x0f002c928c363c7041cafe788f099cf7d35452a1", + "nonce": 243, + "balance": "0x1bc36179476e9a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x18524811fa2e76dd0770d29308b51bcbdc0499f9", + "nonce": 242, + "balance": "0x1bb8aa483e652c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x1aca71f7872aebc85b4d4a9ad47a09abcc624d65", + "nonce": 0, + "balance": "0x2860cb14365c048800", + "code": "0x", + "storage": [] + }, + { + "address": "0x233d639f53225ea5012dc01ceb0a5a30891021cd", + "nonce": 264, + "balance": "0x1b8eea59c54b4200", + "code": "0x", + "storage": [] + }, + { + "address": "0x27fd8475c3db352fdddeaf29823e00c7d33ca39a", + "nonce": 248, + "balance": "0x1bccb45d83401000", + "code": "0x", + "storage": [] + }, + { + "address": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "nonce": 0, + "balance": "0x1dfcc3696501217d400", + "code": "0x", + "storage": [] + }, + { + "address": "0x46681948060140945958068b373f94581e049d4f", + "nonce": 240, + "balance": "0x1baadca68fcb1a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4cb00bc539538d2ee7172474f2b06f021944ece4", + "nonce": 258, + "balance": "0x1b3e54c72c629c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4d7c0d6f3ad5466691649e98ebb59cd5b5194687", + "nonce": 0, + "balance": "0x1b598fbe8e08549c6000", + "code": "0x", + "storage": [] + }, + { + "address": "0x5a5e606bda0fe1b2b4298aadd55cc6ec57ff698a", + "nonce": 0, + "balance": "0x648cba8883ddb562c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x6613ca66c14338db771d01c238ae140c71310535", + "nonce": 243, + "balance": "0x1b98d3921da34800", + "code": "0x", + "storage": [] + }, + { + "address": "0x66577de387feeec96c10f7f8748f13912589c78c", + "nonce": 240, + "balance": "0x1bc37e3123b20a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x7e7b1db26094aa913933f65127b328c1861ff5b0", + "nonce": 256, + "balance": "0x1b89d2467bcd2600", + "code": "0x", + "storage": [] + }, + { + "address": "0x90acb15171deb958d3191d5c02d807afc367f0e8", + "nonce": 240, + "balance": "0x1b9e3c5e84fb6c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x9f296bf64eb7051912c9c8112da3e3ea6cf546e1", + "nonce": 266, + "balance": "0x1b69b11e00d88c00", + "code": "0x", + "storage": [] + }, + { + "address": "0xaa19223d63c82bf9be5543ec96e5504834d17bc8", + "nonce": 246, + "balance": "0x1b9abb1eb524dc00", + "code": "0x", + "storage": [] + }, + { + "address": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "nonce": 0, + "balance": "0x5490f60ff224484c00", + "code": "0x", + "storage": [] + }, + { + "address": "0xc2faa4a2866422a865c9b3b3cb5988bf4d07f0b4", + "nonce": 0, + "balance": "0xc05b9fc73002b32800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc45d23f49451faad9afadf877feead626ea8d083", + "nonce": 0, + "balance": "0x9ee8f806428adc9800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc8621de5921418701619ba715044575073d2ca62", + "nonce": 243, + "balance": "0x1bbb8989a888a800", + "code": "0x", + "storage": [] + }, + { + "address": "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "nonce": 0, + "balance": "0x1f299e7e965325154400", + "code": "0x", + "storage": [] + }, + { + "address": "0xd958300657e7931b06fc6f60dc4bfe7d8e8ea7b8", + "nonce": 252, + "balance": "0x1b79a326be219400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdadb11b01f6004eba6da3846a0e1bf6f4f71fb82", + "nonce": 244, + "balance": "0x1b900b01183e5600", + "code": "0x", + "storage": [] + }, + { + "address": "0xdd442fcbb964a3afdc90d49b408e8dd296fa86e8", + "nonce": 0, + "balance": "0xafd0c09109be7410400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdfaea67368f3e3753397d878f97efe6aa8020c2e", + "nonce": 16, + "balance": "0x4f20fde2ac0f98ec00", + "code": "0x", + "storage": [] + }, + { + "address": "0xfd4fca79b266a73a7defa776b04c7e15bdc70e86", + "nonce": 238, + "balance": "0x1bb5cf9909338400", + "code": "0x", + "storage": [] + } + ], + "fees": { + "base": { + "pgas": { + "version": 0, + "cycles_per_pgas": 1000, + "intrinsic_pgas_per_tx": 200, + "modexp_base": 1000, + "modexp_per_byte_numer": 10, + "modexp_per_byte_denom": 1 + }, + "block_proving_gas_limit": 30000000, + "shard_proving_gas_budget": 7500000, + "min_execution_base_fee_wei": 1000000000, + "min_proving_base_fee_wei": 1000000000, + "initial_execution_base_fee_wei": 1000000000, + "initial_proving_base_fee_wei": 1000000000, + "base_fee_change_denominator": 8 + }, + "v1_activation_daa": 18446744073709551615 + } + }, + "plan": { + "shard_budget": 7500000, + "consensus": true, + "shards": [ + { + "index": 0, + "tx_start": 0, + "tx_end": 0, + "over_budget": false, + "pre_root": "0xee8ee30985ffd378b51927f2c99e4c66f0bfd7c066a994b95dd0fb89eed98586", + "post_root": "0x5ef730d673038ebbd350644d6562558962c25155ae4e37024c6a32c8dd2ecdbc", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "link_in": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "link_out": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "witness": [ + 3, + 0, + 4, + 15, + 14398 + ] + } + ] + }, + "expected": { + "pre_state_root": "0xee8ee30985ffd378b51927f2c99e4c66f0bfd7c066a994b95dd0fb89eed98586", + "post_state_root": "0x5ef730d673038ebbd350644d6562558962c25155ae4e37024c6a32c8dd2ecdbc", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "tx_commitment": "0x0000000000000000000000000000000000000000000000000000000000000000", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "node_state_root": "0x5ef730d673038ebbd350644d6562558962c25155ae4e37024c6a32c8dd2ecdbc" + } +} \ No newline at end of file diff --git a/proving/fixtures/chain/block-81047.json b/proving/fixtures/chain/block-81047.json new file mode 100644 index 000000000..0932ca816 --- /dev/null +++ b/proving/fixtures/chain/block-81047.json @@ -0,0 +1,1325 @@ +{ + "format": "igneum-prove-fixture-v1", + "source": "live devnet export from node 1 at tip 81076, 5 October 2026 20:06 BST, proving v1 chain fixtures", + "block": { + "chain_id": 4463, + "env": { + "number": 81047, + "hash": "0xa9e686e36215cf1053ac16cb445e537ef8a633aac308ad755716f6374c3c1238", + "parent_hash": "0xae10dd9883739a2cd9e95cafa4f9cc7b2e38e477207ec49fb3fe64e3fcf21613", + "timestamp": 1791227124, + "miner": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "prevrandao": "0xd4e117a69f0cfd628e6083fbb780198705f751d7c0ac069ee96ea4d9b2de8bfa", + "base_fee_exec": 1000000000, + "base_fee_proving": 1000000000, + "daa_score": 0 + }, + "block_hashes": [ + [ + 80791, + "0xd89b9e7a2c04f5067579d93e7b2174011b2efb213dc666c0894686bcd7abd7ed" + ], + [ + 80792, + "0xb15e9df888384729a98f0dc1bc33017e71bfa7d40677326f2461a950095804e7" + ], + [ + 80793, + "0x1a75bb3615a32dad3eb935c2a1a3f4d2791cb266be47ef17bd4285962108ff15" + ], + [ + 80794, + "0xf657d95990ce07d72922d34f07dbc66bb681797076a3e9819b2b3349218efaec" + ], + [ + 80795, + "0x4bf9d010c0798a5340983b27eda6a9876e6fefb925104fdc46291a5789f1f70a" + ], + [ + 80796, + "0xb09af3831aa1d8f4945d50fa16a602848a2da22b14dbc1f7af2844cf19e90040" + ], + [ + 80797, + "0xe099536cfd64654690afc152b97866b32da6bf81f959a6c61f9e43bfd2ca2d9a" + ], + [ + 80798, + "0xd6af2cf6184f15ecca77ee9e7e8d2ca718d88f9cea39b475713ceb0e151ddc36" + ], + [ + 80799, + "0x9c76f9f6036c7035fd52a7a4e3fca99451715e729fcc9fc424487faee9a6134d" + ], + [ + 80800, + "0x9a046cd99eb89c551dd0c698f1969069c02ae36cde4140ff4b864d624fd9dff7" + ], + [ + 80801, + "0xc9bffa329d649c6540518cf249e25e845743dc01fcbcfdfb8c798f1455efeae6" + ], + [ + 80802, + "0xb124f2de3c17a985f10c0b58d2283631bcc12524de8d79b0955aecfd2d7f78aa" + ], + [ + 80803, + "0x97b21d6cd3f1b321d2c9729062c305a7955b97646d76910714f899c75d275568" + ], + [ + 80804, + "0xaa3f0e114e156f706e11a701fb940bf54e89b09d83953be7f6733478c7a33093" + ], + [ + 80805, + "0xe5a5868b02ec30304079652a8d5b8af4a9da0a5a676615c1a10d833feb4eb0f7" + ], + [ + 80806, + "0xd97c5c4288f56955e1a3e134eb63250e1becb907b0fdde9ce1383d8d2c8a3939" + ], + [ + 80807, + "0x8f7bf1128e5d03607cca246b86e63cfb4bfe73e74395c8660d027172376b5c36" + ], + [ + 80808, + "0x5d5fe4c3bee7d758986da5a433cfe5221be392419ac22e5acddd701cd12c8142" + ], + [ + 80809, + "0x9b0d0a871ce8eaab11a185a4344f9f48a58b7111f94a8f5221d557722cc5ca5a" + ], + [ + 80810, + "0x58cad4ab776ba7eca94e813323e28a6467a7fcf39d5aa2069e9dfef6f8058e5c" + ], + [ + 80811, + "0x706788247d6c05475eb3bdbc6f181e76429b57a97f168004832fe5ba4b20c7b7" + ], + [ + 80812, + "0xdf26fb97f999b1016865b3fee539e72e3f66d6d228b2aec7e2f59ab93e5a9ee6" + ], + [ + 80813, + "0x8e612c07fdb693a44353a08899e8beae8c93fb1b85845ad965f96e7d2f77fed4" + ], + [ + 80814, + "0x806f4c29ea28af1a3ce0c3ba775c776006ba8cf4b308ee617dc207dc3f740b29" + ], + [ + 80815, + "0xebd8e602049ddc2f8d11d5edc8ee837eefac77418e73e9bad97f6f1c7a85e8aa" + ], + [ + 80816, + "0xe2b069b77ac05cb0a95ae419966feeb20660e4ff3b9a68d044a5721e1525dd86" + ], + [ + 80817, + "0x1ca5b79e26dc425a94c3a6bfd339516caa0a7f8e5e2a62389bfa466186bca495" + ], + [ + 80818, + "0x7cf382f956597cbd36917036dd15b774500f97714e50587a35ec9617662052d6" + ], + [ + 80819, + "0xe70494836c5b5c68b6a167c868db166f9630140499f96e6660d3e3928cad6a14" + ], + [ + 80820, + "0x15f73eefc39e00f9d09390764f9ea0069908062c563686e73b49ec3919c24be3" + ], + [ + 80821, + "0x7f2bd2cf68fc14e6a5debaac39f368534657c41cc587402fac1156d24998e55b" + ], + [ + 80822, + "0x2641ec1c3397b06ab5855dcf6ec2e23de301f5831015a32ffa9f29e18cc2f084" + ], + [ + 80823, + "0xf643e963f5a62a49e935b4ef51aab9bf2689b0b9ec483924d7f42934c8ce5e07" + ], + [ + 80824, + "0x867ee0219bd30fc3117c79295cb1075ca21281fdd285467ae128aece1c1dd9a7" + ], + [ + 80825, + "0xab6bb19259a039d30c3022c9da3494af8c9a541dce30895df083252894c433df" + ], + [ + 80826, + "0x155511de8d34bef351ba0ae7af5bccf515e6e490135130aad5ce2eb02e103d75" + ], + [ + 80827, + "0x4e297a53901da9b364f4dce686f016f4039658ea8832d591dcb21aa44acd197b" + ], + [ + 80828, + "0x7d07b8bc9b06b20e7b56f96fff9051590caac01159d39b3f95bd592a4627daf4" + ], + [ + 80829, + "0x8392a1875388baec163e64854d6ee7664a31017402ee8e269ddc8aa42ca06647" + ], + [ + 80830, + "0x5ab4e5c586d2fb4272941bf26c15288cc129ca1092d3aaab82c478edd1942e29" + ], + [ + 80831, + "0xb0100f0a559daa435128cae74f47f9353ca2a93c2bdac8dd9d12a3130ce6a818" + ], + [ + 80832, + "0x326d1388d0911871c5087925500c4ee5091624331d07659b9d2fa09e83f5884a" + ], + [ + 80833, + "0xc01c186cfcebb8d53953589ce595afef213d46e481b8b5dee7e0a71abd773f6a" + ], + [ + 80834, + "0xe845934e79eaeb8a8ee0b9396b3a13f13c8ac45a0d4e81017c73042789013c38" + ], + [ + 80835, + "0xa45f8f03249624a2d35b54eb83f4edf54e8d1acd061fee3965d9164a132c72aa" + ], + [ + 80836, + "0xe72a59a64af9aa223b5341819a5872f8626908b5c1857d0617706f6999632e01" + ], + [ + 80837, + "0x94063f426663347bb4a3e79f4dedeb76fbce6fd69117a2c74ad1fb4714b856dd" + ], + [ + 80838, + "0x157ab5f9e7fbf98b48af049adcfb4a8e071253ec95b991d031a91668ff3b8d57" + ], + [ + 80839, + "0x5bff632aa54b1a71a7d9a2a57d9c42ef223eb6efe86c72f9aa8dfd67ab7b218a" + ], + [ + 80840, + "0x7f5d1952ad8ad27dd17bed3354395b646468985bae0341b4048376a425e43076" + ], + [ + 80841, + "0x8e69df81fd1b0f186813304260dda1d0c9dad5792126e69d021e679b837b01a5" + ], + [ + 80842, + "0x6733c5472e38830dfb273a0e0802924c1b60a2cc916d6efaf82785e1bf960724" + ], + [ + 80843, + "0x0838513dc8264efffb2c11f8c00ebbc9c57aeb67dd7be93f18e1c613872bc929" + ], + [ + 80844, + "0x8c2f26b17db95bb20983ac41df7a5fca3e3826d2e57831eba13bc3175daeee34" + ], + [ + 80845, + "0xe3edcab823e03598c66e16c5743f7a0917735c975ac946db4cb3ad862db40a21" + ], + [ + 80846, + "0xd075510e75159d941946e666e2c75604a3f3c959523cbe2c0f025427ae85d4a2" + ], + [ + 80847, + "0x57b4b6c0a6ca2532515653445bc872dd44e621f9774f3dd4e2000bf82a0f510a" + ], + [ + 80848, + "0x38a85c99bc3e489a4431ba9a125e74f4d807c630a239419287e039f05cc91ead" + ], + [ + 80849, + "0x3c95d08c99482d57bbcb9fd333cae4010628fc8b1c96b3ad4a68d8f3dd21695f" + ], + [ + 80850, + "0xe767527430f19da2c2b21bc648e33ed2b241f5b7c9413649d9c5a555491a9253" + ], + [ + 80851, + "0xe9941baa9f3378854aaff5563340d130700f0753197c7373229542c56ac91a8f" + ], + [ + 80852, + "0xb92e88ec7cc8c340959d8d0cd5434eec5c1033b389b953b02cda103d53292df0" + ], + [ + 80853, + "0x39f01f35ba2ea7b45e3eec1681c44bd85b066a6a8521a10fe6cec5d4d98aace8" + ], + [ + 80854, + "0x36456a5e36a9a66488656bc0a92705632326e59ca40bd09f6d5f99ccc976972a" + ], + [ + 80855, + "0x12e52067750279b8bc34b03cd21bdb2809ea18a3aa68739623780aaf1150129b" + ], + [ + 80856, + "0x64951c693d7f246647fac0df504f35ff9812416f966afaf4fecf7ca92c0e0b24" + ], + [ + 80857, + "0x3918b797fe9d8cd6c34a0b315aaef550de15f4d30dce932f3be8cdfdc27343a2" + ], + [ + 80858, + "0x86ec8713a68034acb9f34b015a48644c7562766c1760927a299e517ae8ef6f8a" + ], + [ + 80859, + "0xfe75a4c99fc17131b7aae61dd0a588b7d7019e6c662aab8c7d6a8c56370fe07f" + ], + [ + 80860, + "0xf4afa8f14a0f814639c88dea304a8ccf76450ee56f4fbb89caf29dbeef80399b" + ], + [ + 80861, + "0xd4c72bc192643e20658ca54b730ef865d1c29ebfb6d1f78263c796a23b6965b4" + ], + [ + 80862, + "0x7a73b7302ab47b7acc1c5359620df9026fdc08d78c4173bb4e45e73a66651767" + ], + [ + 80863, + "0x9033ce568146d2c69139d3674fa7e18cb8319abae5e6cd24f924f750d2ca3446" + ], + [ + 80864, + "0xe4d8da84ba631f233ea7dd1336a264f499083b08290c82b717f79bfb7ffa09a6" + ], + [ + 80865, + "0xc53ec9d6b11d592921f5cf850584ebb562119ea968a5fa7bb00af381d6eaef46" + ], + [ + 80866, + "0xedcbb7d2f9626af53341464fbfc54e838e6417b7ffcfb0c2f9758641c98d5f9e" + ], + [ + 80867, + "0xbe5f726816d43ca5bbcb70f13901e60ef9040debdfaf2a73a0692c3cbef578ec" + ], + [ + 80868, + "0xb7cf72378b14c086afb4094b9b0121fcf7c6929bf2c2134b5189862644e2e506" + ], + [ + 80869, + "0xa29e21de7e54defda356f088dea38a06cf1a1a75bf93c14e30afdba213df7ef3" + ], + [ + 80870, + "0x504d3f83a27c37c5a15cdbbdfc03285f53071ad01d82bf9be2d6f862548a9a8e" + ], + [ + 80871, + "0xe02107fefe75bc8d1abb0443bff2e7a658e3a9b2fccd5157c1a7c7ccc995fe17" + ], + [ + 80872, + "0xdb92e1b449e7670ea719a6d68f496965620656ff5d25a0521f1873eb8d7ae1ab" + ], + [ + 80873, + "0xf12689607f86ce8f762df5d67054a00b8b580300511a221ff5addcf732393988" + ], + [ + 80874, + "0x0e0fe961c78fe52f44f6742795bae3e96713f4618e4a95623045056eb9e046f6" + ], + [ + 80875, + "0x7fd5bb64e800686ae28c3b55613d27f181dbd91e35d342d33d7b3e5d4df17d00" + ], + [ + 80876, + "0x83c2feae4eacf2ada1c9a630a97614a6296df6b89bfb5f954c3e306b16f0ab2e" + ], + [ + 80877, + "0xd1f3af4d6464c5bc9416d73a291bcf3e1df4fa1a396bb3e9dff92dcd451f75d9" + ], + [ + 80878, + "0xa4cd4b49e1d475a5561a91a0de592ba0c00f63b39a61eab150c6486d1bf2b916" + ], + [ + 80879, + "0x714904f162c3b9013a9defd79d4552b58c5ca8747cc94187079f2ebefee2f0d9" + ], + [ + 80880, + "0xb2263d12431e74591e2843047d79933820d2dd352d5b5c977698b328cd59f46d" + ], + [ + 80881, + "0x67977d72a573f121d9a818dd035959cd502e5bec07a4a5c30a28fc66d4fdf964" + ], + [ + 80882, + "0xf1d5f1b271cd23a79f1ab32655418e4b664663c4a77ebaf0a213fc2456e59c5a" + ], + [ + 80883, + "0x91a9f94d339407747a77d79a9c2be608744d4ecc8b84699d0120814f55a5060b" + ], + [ + 80884, + "0x8f709bf4582d3a912e6006f6c417459bbbf182af2b500b4376cd8dff78201c25" + ], + [ + 80885, + "0xbc66848a66fc038200186f09665540b5905a350a3149bfdddf9575eabb2f0a98" + ], + [ + 80886, + "0x8827e43b9a02e7424513e6d0fb6db7e93cc5555fac9b832584570784708a8f90" + ], + [ + 80887, + "0x7c76b6619ea66b8e7fa2791ad1a3dcafe863388d14db50dd022f63748e2b6ccb" + ], + [ + 80888, + "0xace6e76c69db6530277c9749ff3e77a64c6e7123790fdf5f248cc67b761e10bd" + ], + [ + 80889, + "0x8f96a0d2f9524f0b2bf9f520f35516969fba795e005f003dac5721cc37c6b5ce" + ], + [ + 80890, + "0x68a635c0d7ba10b3fff9108deecc4e9af70d760d688c5614a000503b1f2f006c" + ], + [ + 80891, + "0x47dc91daaf1758699e90f8906c61bcb6dd020b3726a1e0d15a7fe5177bfab696" + ], + [ + 80892, + "0x79f1ff991dd34c0384b9ddb64a7f3791116891447f2c960de5c6d80f9a705b5f" + ], + [ + 80893, + "0xf41832c4af88cdd31eec42b233ccfae51a150cd41baf59e9dc6b64dba2b63c7a" + ], + [ + 80894, + "0xfaf2c58e9ba01047e4cbe2c147a4614d46ad39c1d64cbb43d1de60f9c75286b7" + ], + [ + 80895, + "0x65a671bcbc9b6f905811ce62594feda18a12b8c6ca1c72081c1dde3f6227d1a7" + ], + [ + 80896, + "0x815d16976077a9b599fd5d6a82eda30a0bd1d91c66718528aaea9d4dcd12274b" + ], + [ + 80897, + "0x710c0c1d4d44ed7925343b10c7ace3216741b31a91506fa0f1ad80c96c049b14" + ], + [ + 80898, + "0x5fff52604275c206678c800d95cd0569325ad67413b50ac15252c3440856cc78" + ], + [ + 80899, + "0xe67fe87df50a63d348d300de16353e560c1b86f1a79a0bd09e2a7771ae7e92ba" + ], + [ + 80900, + "0x5d854ceaef9de086a2361cfa1f829f843aac00181fa5ae27e69171c198853895" + ], + [ + 80901, + "0xca45572516a2d4a948cfe518d0bb5378a4d23b800522f515e59176c31c069160" + ], + [ + 80902, + "0xa724d3e05f16d3d8c97296f59802c7f9ebba4284808126d955d6fc96dc1c4729" + ], + [ + 80903, + "0x552061794bddc5185033e936ac8ad07d3a315c7a3d876989b2d61328ad0a3128" + ], + [ + 80904, + "0x23c1ee3ff43ae0d42b9b3f0d2c6c2fef6c8ab8fdec9772540b563cfebbb77a2c" + ], + [ + 80905, + "0xc7e33046cb4819a17a21a00e9c0efa55765a974d8a2364f4bed9a16034ec2828" + ], + [ + 80906, + "0xa7b5712f7d22449a6f6a8a55a33c2f76cb54be86faa2f0d892c8bf6aa9fe59d9" + ], + [ + 80907, + "0xbd8696f54d266714d694b41d3b6466ef242e999938b9effd2de035c3f2c5677b" + ], + [ + 80908, + "0x7ca324c51788d4cc173bd3a013702f80e68524b3b005299be9012621dc8e82c4" + ], + [ + 80909, + "0xaea0e3a560abbe0de857592d37bfd43bd986d1cd461f43139d951d123140d495" + ], + [ + 80910, + "0x640f0ed5ec76969f53d90d5e502f6fce0ec65ab28a728e92a81dd27c466dad09" + ], + [ + 80911, + "0xf4cb4beab37872930f84e37b38c15a0a5e5f5d95927c91284ae880337abcfd9a" + ], + [ + 80912, + "0x65121e870e1668a0a55c9a50d31166294cb8e57bc007897640a7414f0866ec3b" + ], + [ + 80913, + "0xb3ea0a195e00db811104db7362b555fd1ab8401814ee841ca7f4d3344843708c" + ], + [ + 80914, + "0xf92735fd64122d7f6f6df6f20f83fc0793675c81e25e54dc64b3ca027d8ef9a4" + ], + [ + 80915, + "0x78e07fd2fcedb51175785be4657b8d2a554c16c1d2cc396ca3738dbf6f9373f7" + ], + [ + 80916, + "0x0ce35f9c3a736ff2937e6b7addcb632855e347ab3c5c08b74d2dca5f82265426" + ], + [ + 80917, + "0x007b1e317c0d9e7565260e1fe7ebfe031de4291f6ddad75210f8563479e84178" + ], + [ + 80918, + "0xbdb4b040d7942b8bcd206e4e12632892d9cfb87c905f760f0c918dcb682d4eca" + ], + [ + 80919, + "0x1eb8ee38b4d246a0502cc56c8948ec6e17fa55b02002c9916e1a3d9867e3c852" + ], + [ + 80920, + "0x2bd9fea5702cc9c0b96fa3b34559c4a72c5ff4f1d572d9e702390a13b0b1c145" + ], + [ + 80921, + "0x12ae91b4df38f549004ab8c4482d9ed1a2841e43f649ee1a89156f9accea51ea" + ], + [ + 80922, + "0xb960711112f51fc11689061a66a9b4ed3c0bccb3a6f027bd2cceeb2d3786e780" + ], + [ + 80923, + "0x4d474df4920e9136e7564cf0e9ebd03bbcd8cb8bdb38c68b8a9c35120888bfea" + ], + [ + 80924, + "0x73d84f1013038f10bebf6606486c1bb6c6cf530a4d35f482be91df74fe15c93e" + ], + [ + 80925, + "0xf08dc8a601c46c43bb6f94edc1b0734a61e7f8efc82d915b33136fac63d5a407" + ], + [ + 80926, + "0x2b695a3c4d05e233cbe18213d50b10fb9d83914aca9a935aa6f6639b1f54c32e" + ], + [ + 80927, + "0x893a6295ba2f66f083feaa39e0570ccbd446166c2e15238d59a911f2826a47d8" + ], + [ + 80928, + "0xc2b2870e39e8344e18296012c4d846b78cd3bdc80145de6b595527c8845ac57b" + ], + [ + 80929, + "0xc32bf651e3ea72bd6808db5b4dbe4577078d418404221809ac12897082ab8bfa" + ], + [ + 80930, + "0xf1c2f5ed0a7d6b7f083c1a68f75004fbfc929ef3f6ab3b46acbba373feecfdde" + ], + [ + 80931, + "0x51638cd9a09c8ee630706d736c81ea0dde7d97234a515c6064be20d0b92b8418" + ], + [ + 80932, + "0xf3cb85e3f551c4e21d0c6c28e940c6bd965c2a20b5eb3375a9d8d32a65c58c4c" + ], + [ + 80933, + "0x65f05787c7f239dcd4dd9cac2e5c516e25f8a8b898eec284531c16a8da83ed2d" + ], + [ + 80934, + "0x58567dd84c5ce96fdfb6ea5794201199d259374e4879e76df9a6622694de95f3" + ], + [ + 80935, + "0x2031a54f25cefcf00c464f36897691a73616f777582bfcd9c03e5649d3339821" + ], + [ + 80936, + "0x0fc1b95d2579a6cb03f21083da53a989db2491f37bcc339195864baf21932823" + ], + [ + 80937, + "0xfc501a03f2666c7c7db743b6a7296692789824828c033f2254e3565de63bccec" + ], + [ + 80938, + "0xe752390723b297a6d326cc9d10f5a3219b5eede3159f9ea1c7fc7f26401b02f8" + ], + [ + 80939, + "0xfdba580d13c6964b46e7d2a66f2aeee150174facaf1cbff66016832b02d9f796" + ], + [ + 80940, + "0x2221d3b867caf0c1d4993845a60ba8389558ccd9d48bf97f51331e87ff51ce84" + ], + [ + 80941, + "0xb4ba8a208861ab1db5302570f7d589adc5f4826812b4759d625cabe2b68231a0" + ], + [ + 80942, + "0x0bc24048f15b64e0c836c387539686906198ed0b3a7adbe2aed37f41eaee5577" + ], + [ + 80943, + "0x5c3230d24716dfeb57c624bc44f7d80b65c95f4b810aa945f8e5309e45327104" + ], + [ + 80944, + "0x798ef63bf3ed33e8608b899ee956fc3a0a0b359a4466141728589ea914d42453" + ], + [ + 80945, + "0xb158a6b683a60e8d3debc803a89bd18af0f1307941039304a2e33a02f237325a" + ], + [ + 80946, + "0xd420374f3f533e0136b3fb1211357adbdde913776fbb851492ca80980e77ce2e" + ], + [ + 80947, + "0xbf7fea59019b30fec39a94c29b026285e0c2d6fe034da2555095999ca694914b" + ], + [ + 80948, + "0x3d232c12ca3a428474f8c8e21992efc614775842c365a29ef44d9b7f24caafc8" + ], + [ + 80949, + "0x87d17a79601cf126d2b5c95f1ac989f8a90f2b4e9f3b13d1d721776ee2a53c24" + ], + [ + 80950, + "0xd23290b4b63c19aafd1edefd1cfcf34829bdc6efa723df17cce5eb7612b17a0d" + ], + [ + 80951, + "0xa8d8db2850de878732f5a7b95bc08ddbb944e1e8e8884a2a2a6a98275bb7663d" + ], + [ + 80952, + "0xb2c005429b5d1f35d006693509651726e67fd19d34ac04a3473df90d8bed52ab" + ], + [ + 80953, + "0xfcdb4ee3a4fa1273ec429c49704ac1afe6c8c6f858640155276661b19a5c4ca0" + ], + [ + 80954, + "0x9b988ccb0403ebfb0c15486ed6b63ea070121d7754f88d1f14f85651c083f240" + ], + [ + 80955, + "0x39c3d647ae0dd09b1bbf1699803e3c1e76448787baf12cfced826b8d75dd4b7c" + ], + [ + 80956, + "0xb24d2ff081df7d34392e460ddf8fdb8c7cc5ee502e8cae2a279474f329fe7d05" + ], + [ + 80957, + "0xf1f1daa9a70d68fea94a204260b298285ca15469cdc1855be3ef731ae44707bb" + ], + [ + 80958, + "0xeadcfd836598031d7d1a6f5c14028edeffc5baed420b2fbd343a5e21cfdd1fd9" + ], + [ + 80959, + "0xe1391944a4bd9d2fe474c8b12de8b3d151bf3afc97addb223ebaa59592ddcd94" + ], + [ + 80960, + "0xdef9733fd939d6b260fe671867ee7cdcd802332b9afdadbba0ac24b964126d29" + ], + [ + 80961, + "0x641d150032ebcce0d0d7616b010eb29df976191edfad5c8751da8533f78f7946" + ], + [ + 80962, + "0x7480f68a44e03c35a1ece1cec7402fe9fb5b4d34d4f246c967c5178e1a7fa330" + ], + [ + 80963, + "0xefae600a6e868a4213011b106f295f79e5c86186b8382143c0a402761226a9ec" + ], + [ + 80964, + "0xc6ce05cbc32d8a1d786baa6473cc65c29892a5a47e7474f384b65c567d425748" + ], + [ + 80965, + "0xd3a3e3c3c964a77a0821984903c958079b294b195b7b13e109f26f3bafcfa092" + ], + [ + 80966, + "0xd42712815a77eaca169de5c74653b4c007159f47fb5d50d63f57cfd3f2c5868e" + ], + [ + 80967, + "0x9b309f8c13534ed43e399d45a5bcfdb38f49c5633b458f6db71d39571e9d8d41" + ], + [ + 80968, + "0x74729f1a9de838f15def5f5c183b7d3596b03afde1ed6302cabacf8a677f82a2" + ], + [ + 80969, + "0x4c07a8677c028eda23e1665416eb54d364959f3210e682978082d560ebf2bea4" + ], + [ + 80970, + "0xe3f992775c7dfbd2485ccecf516ca7ee9961c897296190c45fd7abb1431ee04b" + ], + [ + 80971, + "0xa8492953412a007f2f7696922b661a27fe2ec42e566504b3aacae2f97a24fab0" + ], + [ + 80972, + "0xe14d6f60ece95189dc8eaed8a2ea5c2da62cbdf988588222a4e99d242ea3beb8" + ], + [ + 80973, + "0x5c7ca844c8eec0cce4e4640ba6fc486c1739eb44aa50de86f55008a4d597357c" + ], + [ + 80974, + "0x2b33e852101ee35f2c61ee7b35152521eefb11364505bf6903a964c93d4001b6" + ], + [ + 80975, + "0x5516b628a8ec4328bcfdf90785fb5d96d22d60b4db638ea1b93bd837d7c739a3" + ], + [ + 80976, + "0x9959b993b66dc1285d0ecd8c95b74fe70949c50d5a97d822ce866291f94f2926" + ], + [ + 80977, + "0x04eb959364eae151a41d4a8d970795da0bf8a0ef2f64bde753cbca4ea3d745c4" + ], + [ + 80978, + "0x22907520934a83b056017bfaa72a0ce92fd8658c81dc8236f3b78f1f7e055aa6" + ], + [ + 80979, + "0x59fae978e8063799b4ab6335ca7ac097a25cbb4d60f9af04eb2f647a92ead4bc" + ], + [ + 80980, + "0x9ef7533f2e9c60f4acf9618efd1c47b147fa95bbb27caac54c1bb19540befeab" + ], + [ + 80981, + "0x9986f742ffa9d17f641293e788e53cd904b6edba3e971603470186f61f934e77" + ], + [ + 80982, + "0xbfb394920e4bd707342ddd41644a5733b7041e517d4e88863924046cd512ec19" + ], + [ + 80983, + "0xa3fef379a91df7a254bc4897cf2ef61ff25b380b2142374ff40d91d5140ed21c" + ], + [ + 80984, + "0x71626767d14a0735e3716b45b847611dc7fe04c6d7a4b22e3a6031d621b6d560" + ], + [ + 80985, + "0xb2c03912458086a6b421a8c83cee657c9276a1d38db128799a3170cba30e3266" + ], + [ + 80986, + "0x8f8000a6e85c09716e8766c9bbb9df1287d46f54625ff139d5e98142a58c9253" + ], + [ + 80987, + "0xfc88dba1647b7d43ef19a1a38780dd238228dd46beeb50ea0e00853f06c78471" + ], + [ + 80988, + "0x7e383c88765b3c499b05b60cdd4470cf9b23add0f6bd33990f207a43a51e431f" + ], + [ + 80989, + "0xa133fe8b34c9e839bcd4a75c2a5858874fa19d8bda43cd1c1af781cd2689753d" + ], + [ + 80990, + "0x466a4d175cb847a9899d98a6863900be93ecfec2f1681c8595db5a62b843478b" + ], + [ + 80991, + "0xb9b1bf393868e53ec5da954cb87fb65eef79ab9caf59cfb636f6cfd9c5c92f8e" + ], + [ + 80992, + "0x2f9eb9f04e16f420d5b98d6a28e25494207bc05affb6b914a1b98fb2244bba75" + ], + [ + 80993, + "0xc73829f98810dd017f45a0bd71ad7db3d07a1f5ff0a57c07994fa095989e6ef2" + ], + [ + 80994, + "0x9cdecf07b8f03979fdcb8f7d4d2c5ef4a13adc19edc0bb0c87a8e0beac6ff3be" + ], + [ + 80995, + "0x978c9e3471a4c37fb6aebab56d41d660720681073dfc68e5bef41d594d67ac8b" + ], + [ + 80996, + "0x739dad0614f0e852037aafbab92b4072f0f9dc9fc30c489465d67a18023f80a4" + ], + [ + 80997, + "0x1fbba34f589373e32f170a1cc3bf5c69083e8e2af5167e6ff38eab89ca2c0a72" + ], + [ + 80998, + "0xb58b192d29e3605fa4006c019698a7fbdc14639784ceeac087ae09148514e68e" + ], + [ + 80999, + "0x77f67acdb2afb4754c6e31005b27aea691cd3edce3094f87238df3f887e1bebe" + ], + [ + 81000, + "0x50e198025598c36879156ee586cdaa18bfa620f117c49baba1b2093d77b83831" + ], + [ + 81001, + "0x416e7d689843b2ddabbc9efd9b6f183b883dcb28e857977cd130f989ff7c72d6" + ], + [ + 81002, + "0xd639747fba2f48c62f7aeb91ead4b7ffa1292d55bd15ebec58879bc51b05863c" + ], + [ + 81003, + "0xc43806d1ab9e6d84be147ac3b74130fc2dca01254b4a847eb2713c3e7c1cd11c" + ], + [ + 81004, + "0xd854d522a33e98313a872fb4c47675f54eeed084d99794c5c72aab4464ee70b4" + ], + [ + 81005, + "0x752458c4e11fe4d564250bd4bb0122c1a1a4fabb124942ed9c3981887d612854" + ], + [ + 81006, + "0xe3eef7ded5487286de24e9802f5dfa1afe7c477666f2a26fca7606d3b9b8694c" + ], + [ + 81007, + "0xc525b44a5e888c3bdb0f66edc5bb4dbe9ce9116279b01fa510037760319ca15f" + ], + [ + 81008, + "0x9552445c57b6e2cd01d95bf4013bb913ce1611e5d707f71e528ad7522e830806" + ], + [ + 81009, + "0xd0eba42e3dd5cb37d42d6891eb47e5e2bdaeccbf9ddabeba06a09cd2b4b15f55" + ], + [ + 81010, + "0xcc6fd1b1758929e8a94d3bfc1b4c5444c55023f8939d4b4fb539b523219b00fd" + ], + [ + 81011, + "0x826ee09096704d3798aa4fccfc02cdedc59bb7472332b624d3716fa3aa14e612" + ], + [ + 81012, + "0x3f82182c620ce4d5ac0a776b0ec52b590f7d30c3bbea039d9a6ec3c6ec8a4aeb" + ], + [ + 81013, + "0xbdc7e4e07b4fc9245d8f5e7d7b207af203ce271b96912963aae2012de4d0dfd4" + ], + [ + 81014, + "0x797fd61f65fdc95f190d0dec1ad3eee5c98211c4529c83d013236838eda13d67" + ], + [ + 81015, + "0x7928f9107a62bd4d597894701cf4bca71365052e3560c23af7fdfa3374105b45" + ], + [ + 81016, + "0x78dc4dfd56d144d629278047ce28d8649e02b55768a137cf98260e40e6c35989" + ], + [ + 81017, + "0x3cd444ed1bd6f3012197839db12d566729ca5fb8c133af4353afbd756104c03c" + ], + [ + 81018, + "0x26350e310daef399afac9ac19b106d5f160d8dc6921fd289b196b4e979742bef" + ], + [ + 81019, + "0x99e1715d9ba9a3d125e77357595bd918766a4d46c20d9db1dfa9ebbd002a7bfa" + ], + [ + 81020, + "0xdf16ba27999f8dacfde10d403cf9982fd3b7545793e678219f09818e04ee64d6" + ], + [ + 81021, + "0xc4690b3c4b89b7282988cd2f71c9e43d3c17d081f0c735d04fbd01247a74bd6f" + ], + [ + 81022, + "0x94a1951db605f9a7dff70e2fd919e154f1291d31b4baa6f9e88d8f904c74fd54" + ], + [ + 81023, + "0x9846feb46081eafd35bd5f134cdf98f7a2a270b52a36779de7fb132250823a6b" + ], + [ + 81024, + "0x3716fe1bc70f07f53358e7936ff4b652a035eabe199265f99809dc4aac49d002" + ], + [ + 81025, + "0x841c89865fcece06f48c90ca283a1dbafb49e3fc422ff4cc5d3ec7fa68ac0f14" + ], + [ + 81026, + "0xcf2be45d7a88f2ffa7fa45366404b8b0af131758b7706f099ee8daa6b5d17b94" + ], + [ + 81027, + "0xf3c7dd48103bf10e13bea5800814475a543f4bf53958e1ad464b112d0c763d00" + ], + [ + 81028, + "0x6ebafde3e8fe7067903074cbeb854b9957bbc45859143f618ada972fccdc2c0d" + ], + [ + 81029, + "0x05146b0d5efdf34a927593169ecce2d2d112a3faa3b79cb2df1a0a2e6a4bd7c0" + ], + [ + 81030, + "0x33def97a4f3ca92c7af02bd19bb9703c516f6df8917928734953ec9fddae68ce" + ], + [ + 81031, + "0x18bbaeb286265fc8edb34abbfd3fade65196f09125dae4d31347401eb635b63f" + ], + [ + 81032, + "0x56aacb0d332a22fabb3bf8ac90b030fb82337b3b22b5be334d9a4f633847480c" + ], + [ + 81033, + "0x8d5d77464c966282353e2019598e0f1d77b5afb938475275f4a3b0d0c29d9c5e" + ], + [ + 81034, + "0xc8ccc7baba52e4fa56d488c78abf451919aeba4a029068fd364059633446c80d" + ], + [ + 81035, + "0xc295e160d3378e12e6b96de0eac43d903aa656f9255e58a3b0222bc229b18bbb" + ], + [ + 81036, + "0xd4cacd3eb6e8eebf312225c23284c15e347e73539816fe91954b482a9acb1753" + ], + [ + 81037, + "0x096f15eb2cfa8a6811e08ee533e3a9bfaf3447225c214157d33f27d532423598" + ], + [ + 81038, + "0xd2464a278a262d447f8d67feb035e8b9eaaaa849eb87fb330e1ac868e6b1d063" + ], + [ + 81039, + "0xfea1ffb0177d362f9e041268c395858091da08fb55de5c922493bf0f3dd19476" + ], + [ + 81040, + "0x5ab2d5c6f5789372abc97b4792b2c4e3476c70bd538764ba74d684ed102fdd0e" + ], + [ + 81041, + "0x2f5be38f279790c0e69e554fcaa28f82bd17ef645338539416db0c9252040f91" + ], + [ + 81042, + "0x844ebc844ba1c0c301a6a9189779e17e557f824d85d01247b4a0b1af49f50d56" + ], + [ + 81043, + "0x944855a63daa3b755a70f2ba5e6a956b037eeb210dd9e80cef73b7c814645adc" + ], + [ + 81044, + "0xf046f77959a6aa199f3d999b2c16dbeb1e9038c308cb1dcc9bba26141e0e5f35" + ], + [ + 81045, + "0x4e420f14ba78477c090eef99a098b62dfdadc3a61f7c0a5499797380763693db" + ], + [ + 81046, + "0xae10dd9883739a2cd9e95cafa4f9cc7b2e38e477207ec49fb3fe64e3fcf21613" + ] + ], + "rewards": [ + [ + "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "0x3268ed91d0987c00" + ], + [ + "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "0x3268ed91d0987c00" + ] + ], + "proving_pool_credit": "0x193476c56a3a6800", + "payouts": [], + "blocks": [ + { + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "blue": true, + "txs": [] + }, + { + "miner": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "blue": true, + "txs": [] + } + ], + "pre_state": [ + { + "address": "0x0000000000000000000000000000000000000210", + "nonce": 1, + "balance": "0x0", + "code": "0x608060405234801561000f575f5ffd5b506004361061003f575f3560e01c8063aa67735414610043578063dea5c2e014610058578063fe7e05d51461009f575b5f5ffd5b6100566100513660046101e0565b6100ca565b005b610083610066366004610211565b6001600160a01b039081165f908152600160205260409020541690565b6040516001600160a01b03909116815260200160405180910390f35b6100836100ad366004610211565b6001600160a01b039081165f908152602081905260409020541690565b336001600160a01b03831614806100f957506001600160a01b038281165f908152600160205260409020541633145b6101635760405162461bcd60e51b815260206004820152603160248201527f446576656c6f70657252656769737472793a206e6f7420746865206163636f75604482015270373a1037b91034ba399031b932b0ba37b960791b606482015260840160405180910390fd5b6001600160a01b038281165f818152602081815260409182902080546001600160a01b031916948616948517905590513381527fa47563c41dab010f91a8ef9dc7ac2bcdfa0ef2af697e575048e71e6eec60dda3910160405180910390a35050565b80356001600160a01b03811681146101db575f5ffd5b919050565b5f5f604083850312156101f1575f5ffd5b6101fa836101c5565b9150610208602084016101c5565b90509250929050565b5f60208284031215610221575f5ffd5b61022a826101c5565b939250505056fea2646970667358221220cbf48f5aa6f911f83b3c2b09adf8c418f5da6fc530c3d5db9d4ba0be24c4bf5864736f6c63430008250033", + "storage": [] + }, + { + "address": "0x0000000000000000000000000000000000000220", + "nonce": 0, + "balance": "0x13cb4971ae20f48ab800", + "code": "0x", + "storage": [] + }, + { + "address": "0x0f002c928c363c7041cafe788f099cf7d35452a1", + "nonce": 243, + "balance": "0x1bc36179476e9a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x18524811fa2e76dd0770d29308b51bcbdc0499f9", + "nonce": 242, + "balance": "0x1bb8aa483e652c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x1aca71f7872aebc85b4d4a9ad47a09abcc624d65", + "nonce": 0, + "balance": "0x2860cb14365c048800", + "code": "0x", + "storage": [] + }, + { + "address": "0x233d639f53225ea5012dc01ceb0a5a30891021cd", + "nonce": 264, + "balance": "0x1b8eea59c54b4200", + "code": "0x", + "storage": [] + }, + { + "address": "0x27fd8475c3db352fdddeaf29823e00c7d33ca39a", + "nonce": 248, + "balance": "0x1bccb45d83401000", + "code": "0x", + "storage": [] + }, + { + "address": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "nonce": 0, + "balance": "0x1dffe9f73ddbce4ec00", + "code": "0x", + "storage": [] + }, + { + "address": "0x46681948060140945958068b373f94581e049d4f", + "nonce": 240, + "balance": "0x1baadca68fcb1a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4cb00bc539538d2ee7172474f2b06f021944ece4", + "nonce": 258, + "balance": "0x1b3e54c72c629c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4d7c0d6f3ad5466691649e98ebb59cd5b5194687", + "nonce": 0, + "balance": "0x1b598fbe8e08549c6000", + "code": "0x", + "storage": [] + }, + { + "address": "0x5a5e606bda0fe1b2b4298aadd55cc6ec57ff698a", + "nonce": 0, + "balance": "0x648cba8883ddb562c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x6613ca66c14338db771d01c238ae140c71310535", + "nonce": 243, + "balance": "0x1b98d3921da34800", + "code": "0x", + "storage": [] + }, + { + "address": "0x66577de387feeec96c10f7f8748f13912589c78c", + "nonce": 240, + "balance": "0x1bc37e3123b20a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x7e7b1db26094aa913933f65127b328c1861ff5b0", + "nonce": 256, + "balance": "0x1b89d2467bcd2600", + "code": "0x", + "storage": [] + }, + { + "address": "0x90acb15171deb958d3191d5c02d807afc367f0e8", + "nonce": 240, + "balance": "0x1b9e3c5e84fb6c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x9f296bf64eb7051912c9c8112da3e3ea6cf546e1", + "nonce": 266, + "balance": "0x1b69b11e00d88c00", + "code": "0x", + "storage": [] + }, + { + "address": "0xaa19223d63c82bf9be5543ec96e5504834d17bc8", + "nonce": 246, + "balance": "0x1b9abb1eb524dc00", + "code": "0x", + "storage": [] + }, + { + "address": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "nonce": 0, + "balance": "0x54c35eed7fcf156400", + "code": "0x", + "storage": [] + }, + { + "address": "0xc2faa4a2866422a865c9b3b3cb5988bf4d07f0b4", + "nonce": 0, + "balance": "0xc05b9fc73002b32800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc45d23f49451faad9afadf877feead626ea8d083", + "nonce": 0, + "balance": "0x9ee8f806428adc9800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc8621de5921418701619ba715044575073d2ca62", + "nonce": 243, + "balance": "0x1bbb8989a888a800", + "code": "0x", + "storage": [] + }, + { + "address": "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "nonce": 0, + "balance": "0x1f299e7e965325154400", + "code": "0x", + "storage": [] + }, + { + "address": "0xd958300657e7931b06fc6f60dc4bfe7d8e8ea7b8", + "nonce": 252, + "balance": "0x1b79a326be219400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdadb11b01f6004eba6da3846a0e1bf6f4f71fb82", + "nonce": 244, + "balance": "0x1b900b01183e5600", + "code": "0x", + "storage": [] + }, + { + "address": "0xdd442fcbb964a3afdc90d49b408e8dd296fa86e8", + "nonce": 0, + "balance": "0xafd0c09109be7410400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdfaea67368f3e3753397d878f97efe6aa8020c2e", + "nonce": 16, + "balance": "0x4f20fde2ac0f98ec00", + "code": "0x", + "storage": [] + }, + { + "address": "0xfd4fca79b266a73a7defa776b04c7e15bdc70e86", + "nonce": 238, + "balance": "0x1bb5cf9909338400", + "code": "0x", + "storage": [] + } + ], + "fees": { + "base": { + "pgas": { + "version": 0, + "cycles_per_pgas": 1000, + "intrinsic_pgas_per_tx": 200, + "modexp_base": 1000, + "modexp_per_byte_numer": 10, + "modexp_per_byte_denom": 1 + }, + "block_proving_gas_limit": 30000000, + "shard_proving_gas_budget": 7500000, + "min_execution_base_fee_wei": 1000000000, + "min_proving_base_fee_wei": 1000000000, + "initial_execution_base_fee_wei": 1000000000, + "initial_proving_base_fee_wei": 1000000000, + "base_fee_change_denominator": 8 + }, + "v1_activation_daa": 18446744073709551615 + } + }, + "plan": { + "shard_budget": 7500000, + "consensus": true, + "shards": [ + { + "index": 0, + "tx_start": 0, + "tx_end": 0, + "over_budget": false, + "pre_root": "0x5ef730d673038ebbd350644d6562558962c25155ae4e37024c6a32c8dd2ecdbc", + "post_root": "0xf607b0098f019577de100167e86771c9eec5cbc0193508abcf460385532aae09", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "link_in": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "link_out": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "witness": [ + 3, + 0, + 4, + 15, + 14398 + ] + } + ] + }, + "expected": { + "pre_state_root": "0x5ef730d673038ebbd350644d6562558962c25155ae4e37024c6a32c8dd2ecdbc", + "post_state_root": "0xf607b0098f019577de100167e86771c9eec5cbc0193508abcf460385532aae09", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "tx_commitment": "0x0000000000000000000000000000000000000000000000000000000000000000", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "node_state_root": "0xf607b0098f019577de100167e86771c9eec5cbc0193508abcf460385532aae09" + } +} \ No newline at end of file diff --git a/proving/fixtures/chain/block-81048.json b/proving/fixtures/chain/block-81048.json new file mode 100644 index 000000000..dcf11d228 --- /dev/null +++ b/proving/fixtures/chain/block-81048.json @@ -0,0 +1,1343 @@ +{ + "format": "igneum-prove-fixture-v1", + "source": "live devnet export from node 1 at tip 81076, 5 October 2026 20:06 BST, proving v1 chain fixtures", + "block": { + "chain_id": 4463, + "env": { + "number": 81048, + "hash": "0xa7ab60b451ae4a7fa639faa0c8b1f2ad7bb6bcb853242797f4aca608b2e69193", + "parent_hash": "0xa9e686e36215cf1053ac16cb445e537ef8a633aac308ad755716f6374c3c1238", + "timestamp": 1791227125, + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "prevrandao": "0xf59bfba5cc2092016d4c609cf22c3de671093cd598a52e57df98f90122e97d59", + "base_fee_exec": 1000000000, + "base_fee_proving": 1000000000, + "daa_score": 0 + }, + "block_hashes": [ + [ + 80792, + "0xb15e9df888384729a98f0dc1bc33017e71bfa7d40677326f2461a950095804e7" + ], + [ + 80793, + "0x1a75bb3615a32dad3eb935c2a1a3f4d2791cb266be47ef17bd4285962108ff15" + ], + [ + 80794, + "0xf657d95990ce07d72922d34f07dbc66bb681797076a3e9819b2b3349218efaec" + ], + [ + 80795, + "0x4bf9d010c0798a5340983b27eda6a9876e6fefb925104fdc46291a5789f1f70a" + ], + [ + 80796, + "0xb09af3831aa1d8f4945d50fa16a602848a2da22b14dbc1f7af2844cf19e90040" + ], + [ + 80797, + "0xe099536cfd64654690afc152b97866b32da6bf81f959a6c61f9e43bfd2ca2d9a" + ], + [ + 80798, + "0xd6af2cf6184f15ecca77ee9e7e8d2ca718d88f9cea39b475713ceb0e151ddc36" + ], + [ + 80799, + "0x9c76f9f6036c7035fd52a7a4e3fca99451715e729fcc9fc424487faee9a6134d" + ], + [ + 80800, + "0x9a046cd99eb89c551dd0c698f1969069c02ae36cde4140ff4b864d624fd9dff7" + ], + [ + 80801, + "0xc9bffa329d649c6540518cf249e25e845743dc01fcbcfdfb8c798f1455efeae6" + ], + [ + 80802, + "0xb124f2de3c17a985f10c0b58d2283631bcc12524de8d79b0955aecfd2d7f78aa" + ], + [ + 80803, + "0x97b21d6cd3f1b321d2c9729062c305a7955b97646d76910714f899c75d275568" + ], + [ + 80804, + "0xaa3f0e114e156f706e11a701fb940bf54e89b09d83953be7f6733478c7a33093" + ], + [ + 80805, + "0xe5a5868b02ec30304079652a8d5b8af4a9da0a5a676615c1a10d833feb4eb0f7" + ], + [ + 80806, + "0xd97c5c4288f56955e1a3e134eb63250e1becb907b0fdde9ce1383d8d2c8a3939" + ], + [ + 80807, + "0x8f7bf1128e5d03607cca246b86e63cfb4bfe73e74395c8660d027172376b5c36" + ], + [ + 80808, + "0x5d5fe4c3bee7d758986da5a433cfe5221be392419ac22e5acddd701cd12c8142" + ], + [ + 80809, + "0x9b0d0a871ce8eaab11a185a4344f9f48a58b7111f94a8f5221d557722cc5ca5a" + ], + [ + 80810, + "0x58cad4ab776ba7eca94e813323e28a6467a7fcf39d5aa2069e9dfef6f8058e5c" + ], + [ + 80811, + "0x706788247d6c05475eb3bdbc6f181e76429b57a97f168004832fe5ba4b20c7b7" + ], + [ + 80812, + "0xdf26fb97f999b1016865b3fee539e72e3f66d6d228b2aec7e2f59ab93e5a9ee6" + ], + [ + 80813, + "0x8e612c07fdb693a44353a08899e8beae8c93fb1b85845ad965f96e7d2f77fed4" + ], + [ + 80814, + "0x806f4c29ea28af1a3ce0c3ba775c776006ba8cf4b308ee617dc207dc3f740b29" + ], + [ + 80815, + "0xebd8e602049ddc2f8d11d5edc8ee837eefac77418e73e9bad97f6f1c7a85e8aa" + ], + [ + 80816, + "0xe2b069b77ac05cb0a95ae419966feeb20660e4ff3b9a68d044a5721e1525dd86" + ], + [ + 80817, + "0x1ca5b79e26dc425a94c3a6bfd339516caa0a7f8e5e2a62389bfa466186bca495" + ], + [ + 80818, + "0x7cf382f956597cbd36917036dd15b774500f97714e50587a35ec9617662052d6" + ], + [ + 80819, + "0xe70494836c5b5c68b6a167c868db166f9630140499f96e6660d3e3928cad6a14" + ], + [ + 80820, + "0x15f73eefc39e00f9d09390764f9ea0069908062c563686e73b49ec3919c24be3" + ], + [ + 80821, + "0x7f2bd2cf68fc14e6a5debaac39f368534657c41cc587402fac1156d24998e55b" + ], + [ + 80822, + "0x2641ec1c3397b06ab5855dcf6ec2e23de301f5831015a32ffa9f29e18cc2f084" + ], + [ + 80823, + "0xf643e963f5a62a49e935b4ef51aab9bf2689b0b9ec483924d7f42934c8ce5e07" + ], + [ + 80824, + "0x867ee0219bd30fc3117c79295cb1075ca21281fdd285467ae128aece1c1dd9a7" + ], + [ + 80825, + "0xab6bb19259a039d30c3022c9da3494af8c9a541dce30895df083252894c433df" + ], + [ + 80826, + "0x155511de8d34bef351ba0ae7af5bccf515e6e490135130aad5ce2eb02e103d75" + ], + [ + 80827, + "0x4e297a53901da9b364f4dce686f016f4039658ea8832d591dcb21aa44acd197b" + ], + [ + 80828, + "0x7d07b8bc9b06b20e7b56f96fff9051590caac01159d39b3f95bd592a4627daf4" + ], + [ + 80829, + "0x8392a1875388baec163e64854d6ee7664a31017402ee8e269ddc8aa42ca06647" + ], + [ + 80830, + "0x5ab4e5c586d2fb4272941bf26c15288cc129ca1092d3aaab82c478edd1942e29" + ], + [ + 80831, + "0xb0100f0a559daa435128cae74f47f9353ca2a93c2bdac8dd9d12a3130ce6a818" + ], + [ + 80832, + "0x326d1388d0911871c5087925500c4ee5091624331d07659b9d2fa09e83f5884a" + ], + [ + 80833, + "0xc01c186cfcebb8d53953589ce595afef213d46e481b8b5dee7e0a71abd773f6a" + ], + [ + 80834, + "0xe845934e79eaeb8a8ee0b9396b3a13f13c8ac45a0d4e81017c73042789013c38" + ], + [ + 80835, + "0xa45f8f03249624a2d35b54eb83f4edf54e8d1acd061fee3965d9164a132c72aa" + ], + [ + 80836, + "0xe72a59a64af9aa223b5341819a5872f8626908b5c1857d0617706f6999632e01" + ], + [ + 80837, + "0x94063f426663347bb4a3e79f4dedeb76fbce6fd69117a2c74ad1fb4714b856dd" + ], + [ + 80838, + "0x157ab5f9e7fbf98b48af049adcfb4a8e071253ec95b991d031a91668ff3b8d57" + ], + [ + 80839, + "0x5bff632aa54b1a71a7d9a2a57d9c42ef223eb6efe86c72f9aa8dfd67ab7b218a" + ], + [ + 80840, + "0x7f5d1952ad8ad27dd17bed3354395b646468985bae0341b4048376a425e43076" + ], + [ + 80841, + "0x8e69df81fd1b0f186813304260dda1d0c9dad5792126e69d021e679b837b01a5" + ], + [ + 80842, + "0x6733c5472e38830dfb273a0e0802924c1b60a2cc916d6efaf82785e1bf960724" + ], + [ + 80843, + "0x0838513dc8264efffb2c11f8c00ebbc9c57aeb67dd7be93f18e1c613872bc929" + ], + [ + 80844, + "0x8c2f26b17db95bb20983ac41df7a5fca3e3826d2e57831eba13bc3175daeee34" + ], + [ + 80845, + "0xe3edcab823e03598c66e16c5743f7a0917735c975ac946db4cb3ad862db40a21" + ], + [ + 80846, + "0xd075510e75159d941946e666e2c75604a3f3c959523cbe2c0f025427ae85d4a2" + ], + [ + 80847, + "0x57b4b6c0a6ca2532515653445bc872dd44e621f9774f3dd4e2000bf82a0f510a" + ], + [ + 80848, + "0x38a85c99bc3e489a4431ba9a125e74f4d807c630a239419287e039f05cc91ead" + ], + [ + 80849, + "0x3c95d08c99482d57bbcb9fd333cae4010628fc8b1c96b3ad4a68d8f3dd21695f" + ], + [ + 80850, + "0xe767527430f19da2c2b21bc648e33ed2b241f5b7c9413649d9c5a555491a9253" + ], + [ + 80851, + "0xe9941baa9f3378854aaff5563340d130700f0753197c7373229542c56ac91a8f" + ], + [ + 80852, + "0xb92e88ec7cc8c340959d8d0cd5434eec5c1033b389b953b02cda103d53292df0" + ], + [ + 80853, + "0x39f01f35ba2ea7b45e3eec1681c44bd85b066a6a8521a10fe6cec5d4d98aace8" + ], + [ + 80854, + "0x36456a5e36a9a66488656bc0a92705632326e59ca40bd09f6d5f99ccc976972a" + ], + [ + 80855, + "0x12e52067750279b8bc34b03cd21bdb2809ea18a3aa68739623780aaf1150129b" + ], + [ + 80856, + "0x64951c693d7f246647fac0df504f35ff9812416f966afaf4fecf7ca92c0e0b24" + ], + [ + 80857, + "0x3918b797fe9d8cd6c34a0b315aaef550de15f4d30dce932f3be8cdfdc27343a2" + ], + [ + 80858, + "0x86ec8713a68034acb9f34b015a48644c7562766c1760927a299e517ae8ef6f8a" + ], + [ + 80859, + "0xfe75a4c99fc17131b7aae61dd0a588b7d7019e6c662aab8c7d6a8c56370fe07f" + ], + [ + 80860, + "0xf4afa8f14a0f814639c88dea304a8ccf76450ee56f4fbb89caf29dbeef80399b" + ], + [ + 80861, + "0xd4c72bc192643e20658ca54b730ef865d1c29ebfb6d1f78263c796a23b6965b4" + ], + [ + 80862, + "0x7a73b7302ab47b7acc1c5359620df9026fdc08d78c4173bb4e45e73a66651767" + ], + [ + 80863, + "0x9033ce568146d2c69139d3674fa7e18cb8319abae5e6cd24f924f750d2ca3446" + ], + [ + 80864, + "0xe4d8da84ba631f233ea7dd1336a264f499083b08290c82b717f79bfb7ffa09a6" + ], + [ + 80865, + "0xc53ec9d6b11d592921f5cf850584ebb562119ea968a5fa7bb00af381d6eaef46" + ], + [ + 80866, + "0xedcbb7d2f9626af53341464fbfc54e838e6417b7ffcfb0c2f9758641c98d5f9e" + ], + [ + 80867, + "0xbe5f726816d43ca5bbcb70f13901e60ef9040debdfaf2a73a0692c3cbef578ec" + ], + [ + 80868, + "0xb7cf72378b14c086afb4094b9b0121fcf7c6929bf2c2134b5189862644e2e506" + ], + [ + 80869, + "0xa29e21de7e54defda356f088dea38a06cf1a1a75bf93c14e30afdba213df7ef3" + ], + [ + 80870, + "0x504d3f83a27c37c5a15cdbbdfc03285f53071ad01d82bf9be2d6f862548a9a8e" + ], + [ + 80871, + "0xe02107fefe75bc8d1abb0443bff2e7a658e3a9b2fccd5157c1a7c7ccc995fe17" + ], + [ + 80872, + "0xdb92e1b449e7670ea719a6d68f496965620656ff5d25a0521f1873eb8d7ae1ab" + ], + [ + 80873, + "0xf12689607f86ce8f762df5d67054a00b8b580300511a221ff5addcf732393988" + ], + [ + 80874, + "0x0e0fe961c78fe52f44f6742795bae3e96713f4618e4a95623045056eb9e046f6" + ], + [ + 80875, + "0x7fd5bb64e800686ae28c3b55613d27f181dbd91e35d342d33d7b3e5d4df17d00" + ], + [ + 80876, + "0x83c2feae4eacf2ada1c9a630a97614a6296df6b89bfb5f954c3e306b16f0ab2e" + ], + [ + 80877, + "0xd1f3af4d6464c5bc9416d73a291bcf3e1df4fa1a396bb3e9dff92dcd451f75d9" + ], + [ + 80878, + "0xa4cd4b49e1d475a5561a91a0de592ba0c00f63b39a61eab150c6486d1bf2b916" + ], + [ + 80879, + "0x714904f162c3b9013a9defd79d4552b58c5ca8747cc94187079f2ebefee2f0d9" + ], + [ + 80880, + "0xb2263d12431e74591e2843047d79933820d2dd352d5b5c977698b328cd59f46d" + ], + [ + 80881, + "0x67977d72a573f121d9a818dd035959cd502e5bec07a4a5c30a28fc66d4fdf964" + ], + [ + 80882, + "0xf1d5f1b271cd23a79f1ab32655418e4b664663c4a77ebaf0a213fc2456e59c5a" + ], + [ + 80883, + "0x91a9f94d339407747a77d79a9c2be608744d4ecc8b84699d0120814f55a5060b" + ], + [ + 80884, + "0x8f709bf4582d3a912e6006f6c417459bbbf182af2b500b4376cd8dff78201c25" + ], + [ + 80885, + "0xbc66848a66fc038200186f09665540b5905a350a3149bfdddf9575eabb2f0a98" + ], + [ + 80886, + "0x8827e43b9a02e7424513e6d0fb6db7e93cc5555fac9b832584570784708a8f90" + ], + [ + 80887, + "0x7c76b6619ea66b8e7fa2791ad1a3dcafe863388d14db50dd022f63748e2b6ccb" + ], + [ + 80888, + "0xace6e76c69db6530277c9749ff3e77a64c6e7123790fdf5f248cc67b761e10bd" + ], + [ + 80889, + "0x8f96a0d2f9524f0b2bf9f520f35516969fba795e005f003dac5721cc37c6b5ce" + ], + [ + 80890, + "0x68a635c0d7ba10b3fff9108deecc4e9af70d760d688c5614a000503b1f2f006c" + ], + [ + 80891, + "0x47dc91daaf1758699e90f8906c61bcb6dd020b3726a1e0d15a7fe5177bfab696" + ], + [ + 80892, + "0x79f1ff991dd34c0384b9ddb64a7f3791116891447f2c960de5c6d80f9a705b5f" + ], + [ + 80893, + "0xf41832c4af88cdd31eec42b233ccfae51a150cd41baf59e9dc6b64dba2b63c7a" + ], + [ + 80894, + "0xfaf2c58e9ba01047e4cbe2c147a4614d46ad39c1d64cbb43d1de60f9c75286b7" + ], + [ + 80895, + "0x65a671bcbc9b6f905811ce62594feda18a12b8c6ca1c72081c1dde3f6227d1a7" + ], + [ + 80896, + "0x815d16976077a9b599fd5d6a82eda30a0bd1d91c66718528aaea9d4dcd12274b" + ], + [ + 80897, + "0x710c0c1d4d44ed7925343b10c7ace3216741b31a91506fa0f1ad80c96c049b14" + ], + [ + 80898, + "0x5fff52604275c206678c800d95cd0569325ad67413b50ac15252c3440856cc78" + ], + [ + 80899, + "0xe67fe87df50a63d348d300de16353e560c1b86f1a79a0bd09e2a7771ae7e92ba" + ], + [ + 80900, + "0x5d854ceaef9de086a2361cfa1f829f843aac00181fa5ae27e69171c198853895" + ], + [ + 80901, + "0xca45572516a2d4a948cfe518d0bb5378a4d23b800522f515e59176c31c069160" + ], + [ + 80902, + "0xa724d3e05f16d3d8c97296f59802c7f9ebba4284808126d955d6fc96dc1c4729" + ], + [ + 80903, + "0x552061794bddc5185033e936ac8ad07d3a315c7a3d876989b2d61328ad0a3128" + ], + [ + 80904, + "0x23c1ee3ff43ae0d42b9b3f0d2c6c2fef6c8ab8fdec9772540b563cfebbb77a2c" + ], + [ + 80905, + "0xc7e33046cb4819a17a21a00e9c0efa55765a974d8a2364f4bed9a16034ec2828" + ], + [ + 80906, + "0xa7b5712f7d22449a6f6a8a55a33c2f76cb54be86faa2f0d892c8bf6aa9fe59d9" + ], + [ + 80907, + "0xbd8696f54d266714d694b41d3b6466ef242e999938b9effd2de035c3f2c5677b" + ], + [ + 80908, + "0x7ca324c51788d4cc173bd3a013702f80e68524b3b005299be9012621dc8e82c4" + ], + [ + 80909, + "0xaea0e3a560abbe0de857592d37bfd43bd986d1cd461f43139d951d123140d495" + ], + [ + 80910, + "0x640f0ed5ec76969f53d90d5e502f6fce0ec65ab28a728e92a81dd27c466dad09" + ], + [ + 80911, + "0xf4cb4beab37872930f84e37b38c15a0a5e5f5d95927c91284ae880337abcfd9a" + ], + [ + 80912, + "0x65121e870e1668a0a55c9a50d31166294cb8e57bc007897640a7414f0866ec3b" + ], + [ + 80913, + "0xb3ea0a195e00db811104db7362b555fd1ab8401814ee841ca7f4d3344843708c" + ], + [ + 80914, + "0xf92735fd64122d7f6f6df6f20f83fc0793675c81e25e54dc64b3ca027d8ef9a4" + ], + [ + 80915, + "0x78e07fd2fcedb51175785be4657b8d2a554c16c1d2cc396ca3738dbf6f9373f7" + ], + [ + 80916, + "0x0ce35f9c3a736ff2937e6b7addcb632855e347ab3c5c08b74d2dca5f82265426" + ], + [ + 80917, + "0x007b1e317c0d9e7565260e1fe7ebfe031de4291f6ddad75210f8563479e84178" + ], + [ + 80918, + "0xbdb4b040d7942b8bcd206e4e12632892d9cfb87c905f760f0c918dcb682d4eca" + ], + [ + 80919, + "0x1eb8ee38b4d246a0502cc56c8948ec6e17fa55b02002c9916e1a3d9867e3c852" + ], + [ + 80920, + "0x2bd9fea5702cc9c0b96fa3b34559c4a72c5ff4f1d572d9e702390a13b0b1c145" + ], + [ + 80921, + "0x12ae91b4df38f549004ab8c4482d9ed1a2841e43f649ee1a89156f9accea51ea" + ], + [ + 80922, + "0xb960711112f51fc11689061a66a9b4ed3c0bccb3a6f027bd2cceeb2d3786e780" + ], + [ + 80923, + "0x4d474df4920e9136e7564cf0e9ebd03bbcd8cb8bdb38c68b8a9c35120888bfea" + ], + [ + 80924, + "0x73d84f1013038f10bebf6606486c1bb6c6cf530a4d35f482be91df74fe15c93e" + ], + [ + 80925, + "0xf08dc8a601c46c43bb6f94edc1b0734a61e7f8efc82d915b33136fac63d5a407" + ], + [ + 80926, + "0x2b695a3c4d05e233cbe18213d50b10fb9d83914aca9a935aa6f6639b1f54c32e" + ], + [ + 80927, + "0x893a6295ba2f66f083feaa39e0570ccbd446166c2e15238d59a911f2826a47d8" + ], + [ + 80928, + "0xc2b2870e39e8344e18296012c4d846b78cd3bdc80145de6b595527c8845ac57b" + ], + [ + 80929, + "0xc32bf651e3ea72bd6808db5b4dbe4577078d418404221809ac12897082ab8bfa" + ], + [ + 80930, + "0xf1c2f5ed0a7d6b7f083c1a68f75004fbfc929ef3f6ab3b46acbba373feecfdde" + ], + [ + 80931, + "0x51638cd9a09c8ee630706d736c81ea0dde7d97234a515c6064be20d0b92b8418" + ], + [ + 80932, + "0xf3cb85e3f551c4e21d0c6c28e940c6bd965c2a20b5eb3375a9d8d32a65c58c4c" + ], + [ + 80933, + "0x65f05787c7f239dcd4dd9cac2e5c516e25f8a8b898eec284531c16a8da83ed2d" + ], + [ + 80934, + "0x58567dd84c5ce96fdfb6ea5794201199d259374e4879e76df9a6622694de95f3" + ], + [ + 80935, + "0x2031a54f25cefcf00c464f36897691a73616f777582bfcd9c03e5649d3339821" + ], + [ + 80936, + "0x0fc1b95d2579a6cb03f21083da53a989db2491f37bcc339195864baf21932823" + ], + [ + 80937, + "0xfc501a03f2666c7c7db743b6a7296692789824828c033f2254e3565de63bccec" + ], + [ + 80938, + "0xe752390723b297a6d326cc9d10f5a3219b5eede3159f9ea1c7fc7f26401b02f8" + ], + [ + 80939, + "0xfdba580d13c6964b46e7d2a66f2aeee150174facaf1cbff66016832b02d9f796" + ], + [ + 80940, + "0x2221d3b867caf0c1d4993845a60ba8389558ccd9d48bf97f51331e87ff51ce84" + ], + [ + 80941, + "0xb4ba8a208861ab1db5302570f7d589adc5f4826812b4759d625cabe2b68231a0" + ], + [ + 80942, + "0x0bc24048f15b64e0c836c387539686906198ed0b3a7adbe2aed37f41eaee5577" + ], + [ + 80943, + "0x5c3230d24716dfeb57c624bc44f7d80b65c95f4b810aa945f8e5309e45327104" + ], + [ + 80944, + "0x798ef63bf3ed33e8608b899ee956fc3a0a0b359a4466141728589ea914d42453" + ], + [ + 80945, + "0xb158a6b683a60e8d3debc803a89bd18af0f1307941039304a2e33a02f237325a" + ], + [ + 80946, + "0xd420374f3f533e0136b3fb1211357adbdde913776fbb851492ca80980e77ce2e" + ], + [ + 80947, + "0xbf7fea59019b30fec39a94c29b026285e0c2d6fe034da2555095999ca694914b" + ], + [ + 80948, + "0x3d232c12ca3a428474f8c8e21992efc614775842c365a29ef44d9b7f24caafc8" + ], + [ + 80949, + "0x87d17a79601cf126d2b5c95f1ac989f8a90f2b4e9f3b13d1d721776ee2a53c24" + ], + [ + 80950, + "0xd23290b4b63c19aafd1edefd1cfcf34829bdc6efa723df17cce5eb7612b17a0d" + ], + [ + 80951, + "0xa8d8db2850de878732f5a7b95bc08ddbb944e1e8e8884a2a2a6a98275bb7663d" + ], + [ + 80952, + "0xb2c005429b5d1f35d006693509651726e67fd19d34ac04a3473df90d8bed52ab" + ], + [ + 80953, + "0xfcdb4ee3a4fa1273ec429c49704ac1afe6c8c6f858640155276661b19a5c4ca0" + ], + [ + 80954, + "0x9b988ccb0403ebfb0c15486ed6b63ea070121d7754f88d1f14f85651c083f240" + ], + [ + 80955, + "0x39c3d647ae0dd09b1bbf1699803e3c1e76448787baf12cfced826b8d75dd4b7c" + ], + [ + 80956, + "0xb24d2ff081df7d34392e460ddf8fdb8c7cc5ee502e8cae2a279474f329fe7d05" + ], + [ + 80957, + "0xf1f1daa9a70d68fea94a204260b298285ca15469cdc1855be3ef731ae44707bb" + ], + [ + 80958, + "0xeadcfd836598031d7d1a6f5c14028edeffc5baed420b2fbd343a5e21cfdd1fd9" + ], + [ + 80959, + "0xe1391944a4bd9d2fe474c8b12de8b3d151bf3afc97addb223ebaa59592ddcd94" + ], + [ + 80960, + "0xdef9733fd939d6b260fe671867ee7cdcd802332b9afdadbba0ac24b964126d29" + ], + [ + 80961, + "0x641d150032ebcce0d0d7616b010eb29df976191edfad5c8751da8533f78f7946" + ], + [ + 80962, + "0x7480f68a44e03c35a1ece1cec7402fe9fb5b4d34d4f246c967c5178e1a7fa330" + ], + [ + 80963, + "0xefae600a6e868a4213011b106f295f79e5c86186b8382143c0a402761226a9ec" + ], + [ + 80964, + "0xc6ce05cbc32d8a1d786baa6473cc65c29892a5a47e7474f384b65c567d425748" + ], + [ + 80965, + "0xd3a3e3c3c964a77a0821984903c958079b294b195b7b13e109f26f3bafcfa092" + ], + [ + 80966, + "0xd42712815a77eaca169de5c74653b4c007159f47fb5d50d63f57cfd3f2c5868e" + ], + [ + 80967, + "0x9b309f8c13534ed43e399d45a5bcfdb38f49c5633b458f6db71d39571e9d8d41" + ], + [ + 80968, + "0x74729f1a9de838f15def5f5c183b7d3596b03afde1ed6302cabacf8a677f82a2" + ], + [ + 80969, + "0x4c07a8677c028eda23e1665416eb54d364959f3210e682978082d560ebf2bea4" + ], + [ + 80970, + "0xe3f992775c7dfbd2485ccecf516ca7ee9961c897296190c45fd7abb1431ee04b" + ], + [ + 80971, + "0xa8492953412a007f2f7696922b661a27fe2ec42e566504b3aacae2f97a24fab0" + ], + [ + 80972, + "0xe14d6f60ece95189dc8eaed8a2ea5c2da62cbdf988588222a4e99d242ea3beb8" + ], + [ + 80973, + "0x5c7ca844c8eec0cce4e4640ba6fc486c1739eb44aa50de86f55008a4d597357c" + ], + [ + 80974, + "0x2b33e852101ee35f2c61ee7b35152521eefb11364505bf6903a964c93d4001b6" + ], + [ + 80975, + "0x5516b628a8ec4328bcfdf90785fb5d96d22d60b4db638ea1b93bd837d7c739a3" + ], + [ + 80976, + "0x9959b993b66dc1285d0ecd8c95b74fe70949c50d5a97d822ce866291f94f2926" + ], + [ + 80977, + "0x04eb959364eae151a41d4a8d970795da0bf8a0ef2f64bde753cbca4ea3d745c4" + ], + [ + 80978, + "0x22907520934a83b056017bfaa72a0ce92fd8658c81dc8236f3b78f1f7e055aa6" + ], + [ + 80979, + "0x59fae978e8063799b4ab6335ca7ac097a25cbb4d60f9af04eb2f647a92ead4bc" + ], + [ + 80980, + "0x9ef7533f2e9c60f4acf9618efd1c47b147fa95bbb27caac54c1bb19540befeab" + ], + [ + 80981, + "0x9986f742ffa9d17f641293e788e53cd904b6edba3e971603470186f61f934e77" + ], + [ + 80982, + "0xbfb394920e4bd707342ddd41644a5733b7041e517d4e88863924046cd512ec19" + ], + [ + 80983, + "0xa3fef379a91df7a254bc4897cf2ef61ff25b380b2142374ff40d91d5140ed21c" + ], + [ + 80984, + "0x71626767d14a0735e3716b45b847611dc7fe04c6d7a4b22e3a6031d621b6d560" + ], + [ + 80985, + "0xb2c03912458086a6b421a8c83cee657c9276a1d38db128799a3170cba30e3266" + ], + [ + 80986, + "0x8f8000a6e85c09716e8766c9bbb9df1287d46f54625ff139d5e98142a58c9253" + ], + [ + 80987, + "0xfc88dba1647b7d43ef19a1a38780dd238228dd46beeb50ea0e00853f06c78471" + ], + [ + 80988, + "0x7e383c88765b3c499b05b60cdd4470cf9b23add0f6bd33990f207a43a51e431f" + ], + [ + 80989, + "0xa133fe8b34c9e839bcd4a75c2a5858874fa19d8bda43cd1c1af781cd2689753d" + ], + [ + 80990, + "0x466a4d175cb847a9899d98a6863900be93ecfec2f1681c8595db5a62b843478b" + ], + [ + 80991, + "0xb9b1bf393868e53ec5da954cb87fb65eef79ab9caf59cfb636f6cfd9c5c92f8e" + ], + [ + 80992, + "0x2f9eb9f04e16f420d5b98d6a28e25494207bc05affb6b914a1b98fb2244bba75" + ], + [ + 80993, + "0xc73829f98810dd017f45a0bd71ad7db3d07a1f5ff0a57c07994fa095989e6ef2" + ], + [ + 80994, + "0x9cdecf07b8f03979fdcb8f7d4d2c5ef4a13adc19edc0bb0c87a8e0beac6ff3be" + ], + [ + 80995, + "0x978c9e3471a4c37fb6aebab56d41d660720681073dfc68e5bef41d594d67ac8b" + ], + [ + 80996, + "0x739dad0614f0e852037aafbab92b4072f0f9dc9fc30c489465d67a18023f80a4" + ], + [ + 80997, + "0x1fbba34f589373e32f170a1cc3bf5c69083e8e2af5167e6ff38eab89ca2c0a72" + ], + [ + 80998, + "0xb58b192d29e3605fa4006c019698a7fbdc14639784ceeac087ae09148514e68e" + ], + [ + 80999, + "0x77f67acdb2afb4754c6e31005b27aea691cd3edce3094f87238df3f887e1bebe" + ], + [ + 81000, + "0x50e198025598c36879156ee586cdaa18bfa620f117c49baba1b2093d77b83831" + ], + [ + 81001, + "0x416e7d689843b2ddabbc9efd9b6f183b883dcb28e857977cd130f989ff7c72d6" + ], + [ + 81002, + "0xd639747fba2f48c62f7aeb91ead4b7ffa1292d55bd15ebec58879bc51b05863c" + ], + [ + 81003, + "0xc43806d1ab9e6d84be147ac3b74130fc2dca01254b4a847eb2713c3e7c1cd11c" + ], + [ + 81004, + "0xd854d522a33e98313a872fb4c47675f54eeed084d99794c5c72aab4464ee70b4" + ], + [ + 81005, + "0x752458c4e11fe4d564250bd4bb0122c1a1a4fabb124942ed9c3981887d612854" + ], + [ + 81006, + "0xe3eef7ded5487286de24e9802f5dfa1afe7c477666f2a26fca7606d3b9b8694c" + ], + [ + 81007, + "0xc525b44a5e888c3bdb0f66edc5bb4dbe9ce9116279b01fa510037760319ca15f" + ], + [ + 81008, + "0x9552445c57b6e2cd01d95bf4013bb913ce1611e5d707f71e528ad7522e830806" + ], + [ + 81009, + "0xd0eba42e3dd5cb37d42d6891eb47e5e2bdaeccbf9ddabeba06a09cd2b4b15f55" + ], + [ + 81010, + "0xcc6fd1b1758929e8a94d3bfc1b4c5444c55023f8939d4b4fb539b523219b00fd" + ], + [ + 81011, + "0x826ee09096704d3798aa4fccfc02cdedc59bb7472332b624d3716fa3aa14e612" + ], + [ + 81012, + "0x3f82182c620ce4d5ac0a776b0ec52b590f7d30c3bbea039d9a6ec3c6ec8a4aeb" + ], + [ + 81013, + "0xbdc7e4e07b4fc9245d8f5e7d7b207af203ce271b96912963aae2012de4d0dfd4" + ], + [ + 81014, + "0x797fd61f65fdc95f190d0dec1ad3eee5c98211c4529c83d013236838eda13d67" + ], + [ + 81015, + "0x7928f9107a62bd4d597894701cf4bca71365052e3560c23af7fdfa3374105b45" + ], + [ + 81016, + "0x78dc4dfd56d144d629278047ce28d8649e02b55768a137cf98260e40e6c35989" + ], + [ + 81017, + "0x3cd444ed1bd6f3012197839db12d566729ca5fb8c133af4353afbd756104c03c" + ], + [ + 81018, + "0x26350e310daef399afac9ac19b106d5f160d8dc6921fd289b196b4e979742bef" + ], + [ + 81019, + "0x99e1715d9ba9a3d125e77357595bd918766a4d46c20d9db1dfa9ebbd002a7bfa" + ], + [ + 81020, + "0xdf16ba27999f8dacfde10d403cf9982fd3b7545793e678219f09818e04ee64d6" + ], + [ + 81021, + "0xc4690b3c4b89b7282988cd2f71c9e43d3c17d081f0c735d04fbd01247a74bd6f" + ], + [ + 81022, + "0x94a1951db605f9a7dff70e2fd919e154f1291d31b4baa6f9e88d8f904c74fd54" + ], + [ + 81023, + "0x9846feb46081eafd35bd5f134cdf98f7a2a270b52a36779de7fb132250823a6b" + ], + [ + 81024, + "0x3716fe1bc70f07f53358e7936ff4b652a035eabe199265f99809dc4aac49d002" + ], + [ + 81025, + "0x841c89865fcece06f48c90ca283a1dbafb49e3fc422ff4cc5d3ec7fa68ac0f14" + ], + [ + 81026, + "0xcf2be45d7a88f2ffa7fa45366404b8b0af131758b7706f099ee8daa6b5d17b94" + ], + [ + 81027, + "0xf3c7dd48103bf10e13bea5800814475a543f4bf53958e1ad464b112d0c763d00" + ], + [ + 81028, + "0x6ebafde3e8fe7067903074cbeb854b9957bbc45859143f618ada972fccdc2c0d" + ], + [ + 81029, + "0x05146b0d5efdf34a927593169ecce2d2d112a3faa3b79cb2df1a0a2e6a4bd7c0" + ], + [ + 81030, + "0x33def97a4f3ca92c7af02bd19bb9703c516f6df8917928734953ec9fddae68ce" + ], + [ + 81031, + "0x18bbaeb286265fc8edb34abbfd3fade65196f09125dae4d31347401eb635b63f" + ], + [ + 81032, + "0x56aacb0d332a22fabb3bf8ac90b030fb82337b3b22b5be334d9a4f633847480c" + ], + [ + 81033, + "0x8d5d77464c966282353e2019598e0f1d77b5afb938475275f4a3b0d0c29d9c5e" + ], + [ + 81034, + "0xc8ccc7baba52e4fa56d488c78abf451919aeba4a029068fd364059633446c80d" + ], + [ + 81035, + "0xc295e160d3378e12e6b96de0eac43d903aa656f9255e58a3b0222bc229b18bbb" + ], + [ + 81036, + "0xd4cacd3eb6e8eebf312225c23284c15e347e73539816fe91954b482a9acb1753" + ], + [ + 81037, + "0x096f15eb2cfa8a6811e08ee533e3a9bfaf3447225c214157d33f27d532423598" + ], + [ + 81038, + "0xd2464a278a262d447f8d67feb035e8b9eaaaa849eb87fb330e1ac868e6b1d063" + ], + [ + 81039, + "0xfea1ffb0177d362f9e041268c395858091da08fb55de5c922493bf0f3dd19476" + ], + [ + 81040, + "0x5ab2d5c6f5789372abc97b4792b2c4e3476c70bd538764ba74d684ed102fdd0e" + ], + [ + 81041, + "0x2f5be38f279790c0e69e554fcaa28f82bd17ef645338539416db0c9252040f91" + ], + [ + 81042, + "0x844ebc844ba1c0c301a6a9189779e17e557f824d85d01247b4a0b1af49f50d56" + ], + [ + 81043, + "0x944855a63daa3b755a70f2ba5e6a956b037eeb210dd9e80cef73b7c814645adc" + ], + [ + 81044, + "0xf046f77959a6aa199f3d999b2c16dbeb1e9038c308cb1dcc9bba26141e0e5f35" + ], + [ + 81045, + "0x4e420f14ba78477c090eef99a098b62dfdadc3a61f7c0a5499797380763693db" + ], + [ + 81046, + "0xae10dd9883739a2cd9e95cafa4f9cc7b2e38e477207ec49fb3fe64e3fcf21613" + ], + [ + 81047, + "0xa9e686e36215cf1053ac16cb445e537ef8a633aac308ad755716f6374c3c1238" + ] + ], + "rewards": [ + [ + "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "0x32690d97c8236000" + ], + [ + "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "0x32690d97c8236000" + ], + [ + "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "0x32690d97c8236000" + ], + [ + "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "0x32690d97c8236000" + ] + ], + "proving_pool_credit": "0x32690d8e77f3d000", + "payouts": [], + "blocks": [ + { + "miner": "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "blue": true, + "txs": [] + }, + { + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "blue": true, + "txs": [] + }, + { + "miner": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "blue": true, + "txs": [] + }, + { + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "blue": true, + "txs": [] + } + ], + "pre_state": [ + { + "address": "0x0000000000000000000000000000000000000210", + "nonce": 1, + "balance": "0x0", + "code": "0x608060405234801561000f575f5ffd5b506004361061003f575f3560e01c8063aa67735414610043578063dea5c2e014610058578063fe7e05d51461009f575b5f5ffd5b6100566100513660046101e0565b6100ca565b005b610083610066366004610211565b6001600160a01b039081165f908152600160205260409020541690565b6040516001600160a01b03909116815260200160405180910390f35b6100836100ad366004610211565b6001600160a01b039081165f908152602081905260409020541690565b336001600160a01b03831614806100f957506001600160a01b038281165f908152600160205260409020541633145b6101635760405162461bcd60e51b815260206004820152603160248201527f446576656c6f70657252656769737472793a206e6f7420746865206163636f75604482015270373a1037b91034ba399031b932b0ba37b960791b606482015260840160405180910390fd5b6001600160a01b038281165f818152602081815260409182902080546001600160a01b031916948616948517905590513381527fa47563c41dab010f91a8ef9dc7ac2bcdfa0ef2af697e575048e71e6eec60dda3910160405180910390a35050565b80356001600160a01b03811681146101db575f5ffd5b919050565b5f5f604083850312156101f1575f5ffd5b6101fa836101c5565b9150610208602084016101c5565b90509250929050565b5f60208284031215610221575f5ffd5b61022a826101c5565b939250505056fea2646970667358221220cbf48f5aa6f911f83b3c2b09adf8c418f5da6fc530c3d5db9d4ba0be24c4bf5864736f6c63430008250033", + "storage": [] + }, + { + "address": "0x0000000000000000000000000000000000000220", + "nonce": 0, + "balance": "0x13cb62a624e65ec52000", + "code": "0x", + "storage": [] + }, + { + "address": "0x0f002c928c363c7041cafe788f099cf7d35452a1", + "nonce": 243, + "balance": "0x1bc36179476e9a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x18524811fa2e76dd0770d29308b51bcbdc0499f9", + "nonce": 242, + "balance": "0x1bb8aa483e652c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x1aca71f7872aebc85b4d4a9ad47a09abcc624d65", + "nonce": 0, + "balance": "0x2860cb14365c048800", + "code": "0x", + "storage": [] + }, + { + "address": "0x233d639f53225ea5012dc01ceb0a5a30891021cd", + "nonce": 264, + "balance": "0x1b8eea59c54b4200", + "code": "0x", + "storage": [] + }, + { + "address": "0x27fd8475c3db352fdddeaf29823e00c7d33ca39a", + "nonce": 248, + "balance": "0x1bccb45d83401000", + "code": "0x", + "storage": [] + }, + { + "address": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "nonce": 0, + "balance": "0x1e03108616f8d7d6800", + "code": "0x", + "storage": [] + }, + { + "address": "0x46681948060140945958068b373f94581e049d4f", + "nonce": 240, + "balance": "0x1baadca68fcb1a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4cb00bc539538d2ee7172474f2b06f021944ece4", + "nonce": 258, + "balance": "0x1b3e54c72c629c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4d7c0d6f3ad5466691649e98ebb59cd5b5194687", + "nonce": 0, + "balance": "0x1b598fbe8e08549c6000", + "code": "0x", + "storage": [] + }, + { + "address": "0x5a5e606bda0fe1b2b4298aadd55cc6ec57ff698a", + "nonce": 0, + "balance": "0x648cba8883ddb562c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x6613ca66c14338db771d01c238ae140c71310535", + "nonce": 243, + "balance": "0x1b98d3921da34800", + "code": "0x", + "storage": [] + }, + { + "address": "0x66577de387feeec96c10f7f8748f13912589c78c", + "nonce": 240, + "balance": "0x1bc37e3123b20a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x7e7b1db26094aa913933f65127b328c1861ff5b0", + "nonce": 256, + "balance": "0x1b89d2467bcd2600", + "code": "0x", + "storage": [] + }, + { + "address": "0x90acb15171deb958d3191d5c02d807afc367f0e8", + "nonce": 240, + "balance": "0x1b9e3c5e84fb6c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x9f296bf64eb7051912c9c8112da3e3ea6cf546e1", + "nonce": 266, + "balance": "0x1b69b11e00d88c00", + "code": "0x", + "storage": [] + }, + { + "address": "0xaa19223d63c82bf9be5543ec96e5504834d17bc8", + "nonce": 246, + "balance": "0x1b9abb1eb524dc00", + "code": "0x", + "storage": [] + }, + { + "address": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "nonce": 0, + "balance": "0x54f5c7db119fade000", + "code": "0x", + "storage": [] + }, + { + "address": "0xc2faa4a2866422a865c9b3b3cb5988bf4d07f0b4", + "nonce": 0, + "balance": "0xc05b9fc73002b32800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc45d23f49451faad9afadf877feead626ea8d083", + "nonce": 0, + "balance": "0x9ee8f806428adc9800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc8621de5921418701619ba715044575073d2ca62", + "nonce": 243, + "balance": "0x1bbb8989a888a800", + "code": "0x", + "storage": [] + }, + { + "address": "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "nonce": 0, + "balance": "0x1f299e7e965325154400", + "code": "0x", + "storage": [] + }, + { + "address": "0xd958300657e7931b06fc6f60dc4bfe7d8e8ea7b8", + "nonce": 252, + "balance": "0x1b79a326be219400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdadb11b01f6004eba6da3846a0e1bf6f4f71fb82", + "nonce": 244, + "balance": "0x1b900b01183e5600", + "code": "0x", + "storage": [] + }, + { + "address": "0xdd442fcbb964a3afdc90d49b408e8dd296fa86e8", + "nonce": 0, + "balance": "0xafd0c09109be7410400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdfaea67368f3e3753397d878f97efe6aa8020c2e", + "nonce": 16, + "balance": "0x4f20fde2ac0f98ec00", + "code": "0x", + "storage": [] + }, + { + "address": "0xfd4fca79b266a73a7defa776b04c7e15bdc70e86", + "nonce": 238, + "balance": "0x1bb5cf9909338400", + "code": "0x", + "storage": [] + } + ], + "fees": { + "base": { + "pgas": { + "version": 0, + "cycles_per_pgas": 1000, + "intrinsic_pgas_per_tx": 200, + "modexp_base": 1000, + "modexp_per_byte_numer": 10, + "modexp_per_byte_denom": 1 + }, + "block_proving_gas_limit": 30000000, + "shard_proving_gas_budget": 7500000, + "min_execution_base_fee_wei": 1000000000, + "min_proving_base_fee_wei": 1000000000, + "initial_execution_base_fee_wei": 1000000000, + "initial_proving_base_fee_wei": 1000000000, + "base_fee_change_denominator": 8 + }, + "v1_activation_daa": 18446744073709551615 + } + }, + "plan": { + "shard_budget": 7500000, + "consensus": true, + "shards": [ + { + "index": 0, + "tx_start": 0, + "tx_end": 0, + "over_budget": false, + "pre_root": "0xf607b0098f019577de100167e86771c9eec5cbc0193508abcf460385532aae09", + "post_root": "0xc6665c33a141e50cdb7a7194304afe5672016fa15f394ffa7f414d70912a7654", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "link_in": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "link_out": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "witness": [ + 4, + 0, + 6, + 14, + 14801 + ] + } + ] + }, + "expected": { + "pre_state_root": "0xf607b0098f019577de100167e86771c9eec5cbc0193508abcf460385532aae09", + "post_state_root": "0xc6665c33a141e50cdb7a7194304afe5672016fa15f394ffa7f414d70912a7654", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "tx_commitment": "0x0000000000000000000000000000000000000000000000000000000000000000", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "node_state_root": "0xc6665c33a141e50cdb7a7194304afe5672016fa15f394ffa7f414d70912a7654" + } +} \ No newline at end of file diff --git a/proving/fixtures/chain/block-81049.json b/proving/fixtures/chain/block-81049.json new file mode 100644 index 000000000..c945649da --- /dev/null +++ b/proving/fixtures/chain/block-81049.json @@ -0,0 +1,1316 @@ +{ + "format": "igneum-prove-fixture-v1", + "source": "live devnet export from node 1 at tip 81076, 5 October 2026 20:06 BST, proving v1 chain fixtures", + "block": { + "chain_id": 4463, + "env": { + "number": 81049, + "hash": "0xdba699a61f7e4b3b8356a7494afbb73237f00385d7ad279a77002bfd451b2643", + "parent_hash": "0xa7ab60b451ae4a7fa639faa0c8b1f2ad7bb6bcb853242797f4aca608b2e69193", + "timestamp": 1791227127, + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "prevrandao": "0xec8a60b550bb6da48cc2a76ad9cf98d512adc4e945819a62e876182175ad1eba", + "base_fee_exec": 1000000000, + "base_fee_proving": 1000000000, + "daa_score": 0 + }, + "block_hashes": [ + [ + 80793, + "0x1a75bb3615a32dad3eb935c2a1a3f4d2791cb266be47ef17bd4285962108ff15" + ], + [ + 80794, + "0xf657d95990ce07d72922d34f07dbc66bb681797076a3e9819b2b3349218efaec" + ], + [ + 80795, + "0x4bf9d010c0798a5340983b27eda6a9876e6fefb925104fdc46291a5789f1f70a" + ], + [ + 80796, + "0xb09af3831aa1d8f4945d50fa16a602848a2da22b14dbc1f7af2844cf19e90040" + ], + [ + 80797, + "0xe099536cfd64654690afc152b97866b32da6bf81f959a6c61f9e43bfd2ca2d9a" + ], + [ + 80798, + "0xd6af2cf6184f15ecca77ee9e7e8d2ca718d88f9cea39b475713ceb0e151ddc36" + ], + [ + 80799, + "0x9c76f9f6036c7035fd52a7a4e3fca99451715e729fcc9fc424487faee9a6134d" + ], + [ + 80800, + "0x9a046cd99eb89c551dd0c698f1969069c02ae36cde4140ff4b864d624fd9dff7" + ], + [ + 80801, + "0xc9bffa329d649c6540518cf249e25e845743dc01fcbcfdfb8c798f1455efeae6" + ], + [ + 80802, + "0xb124f2de3c17a985f10c0b58d2283631bcc12524de8d79b0955aecfd2d7f78aa" + ], + [ + 80803, + "0x97b21d6cd3f1b321d2c9729062c305a7955b97646d76910714f899c75d275568" + ], + [ + 80804, + "0xaa3f0e114e156f706e11a701fb940bf54e89b09d83953be7f6733478c7a33093" + ], + [ + 80805, + "0xe5a5868b02ec30304079652a8d5b8af4a9da0a5a676615c1a10d833feb4eb0f7" + ], + [ + 80806, + "0xd97c5c4288f56955e1a3e134eb63250e1becb907b0fdde9ce1383d8d2c8a3939" + ], + [ + 80807, + "0x8f7bf1128e5d03607cca246b86e63cfb4bfe73e74395c8660d027172376b5c36" + ], + [ + 80808, + "0x5d5fe4c3bee7d758986da5a433cfe5221be392419ac22e5acddd701cd12c8142" + ], + [ + 80809, + "0x9b0d0a871ce8eaab11a185a4344f9f48a58b7111f94a8f5221d557722cc5ca5a" + ], + [ + 80810, + "0x58cad4ab776ba7eca94e813323e28a6467a7fcf39d5aa2069e9dfef6f8058e5c" + ], + [ + 80811, + "0x706788247d6c05475eb3bdbc6f181e76429b57a97f168004832fe5ba4b20c7b7" + ], + [ + 80812, + "0xdf26fb97f999b1016865b3fee539e72e3f66d6d228b2aec7e2f59ab93e5a9ee6" + ], + [ + 80813, + "0x8e612c07fdb693a44353a08899e8beae8c93fb1b85845ad965f96e7d2f77fed4" + ], + [ + 80814, + "0x806f4c29ea28af1a3ce0c3ba775c776006ba8cf4b308ee617dc207dc3f740b29" + ], + [ + 80815, + "0xebd8e602049ddc2f8d11d5edc8ee837eefac77418e73e9bad97f6f1c7a85e8aa" + ], + [ + 80816, + "0xe2b069b77ac05cb0a95ae419966feeb20660e4ff3b9a68d044a5721e1525dd86" + ], + [ + 80817, + "0x1ca5b79e26dc425a94c3a6bfd339516caa0a7f8e5e2a62389bfa466186bca495" + ], + [ + 80818, + "0x7cf382f956597cbd36917036dd15b774500f97714e50587a35ec9617662052d6" + ], + [ + 80819, + "0xe70494836c5b5c68b6a167c868db166f9630140499f96e6660d3e3928cad6a14" + ], + [ + 80820, + "0x15f73eefc39e00f9d09390764f9ea0069908062c563686e73b49ec3919c24be3" + ], + [ + 80821, + "0x7f2bd2cf68fc14e6a5debaac39f368534657c41cc587402fac1156d24998e55b" + ], + [ + 80822, + "0x2641ec1c3397b06ab5855dcf6ec2e23de301f5831015a32ffa9f29e18cc2f084" + ], + [ + 80823, + "0xf643e963f5a62a49e935b4ef51aab9bf2689b0b9ec483924d7f42934c8ce5e07" + ], + [ + 80824, + "0x867ee0219bd30fc3117c79295cb1075ca21281fdd285467ae128aece1c1dd9a7" + ], + [ + 80825, + "0xab6bb19259a039d30c3022c9da3494af8c9a541dce30895df083252894c433df" + ], + [ + 80826, + "0x155511de8d34bef351ba0ae7af5bccf515e6e490135130aad5ce2eb02e103d75" + ], + [ + 80827, + "0x4e297a53901da9b364f4dce686f016f4039658ea8832d591dcb21aa44acd197b" + ], + [ + 80828, + "0x7d07b8bc9b06b20e7b56f96fff9051590caac01159d39b3f95bd592a4627daf4" + ], + [ + 80829, + "0x8392a1875388baec163e64854d6ee7664a31017402ee8e269ddc8aa42ca06647" + ], + [ + 80830, + "0x5ab4e5c586d2fb4272941bf26c15288cc129ca1092d3aaab82c478edd1942e29" + ], + [ + 80831, + "0xb0100f0a559daa435128cae74f47f9353ca2a93c2bdac8dd9d12a3130ce6a818" + ], + [ + 80832, + "0x326d1388d0911871c5087925500c4ee5091624331d07659b9d2fa09e83f5884a" + ], + [ + 80833, + "0xc01c186cfcebb8d53953589ce595afef213d46e481b8b5dee7e0a71abd773f6a" + ], + [ + 80834, + "0xe845934e79eaeb8a8ee0b9396b3a13f13c8ac45a0d4e81017c73042789013c38" + ], + [ + 80835, + "0xa45f8f03249624a2d35b54eb83f4edf54e8d1acd061fee3965d9164a132c72aa" + ], + [ + 80836, + "0xe72a59a64af9aa223b5341819a5872f8626908b5c1857d0617706f6999632e01" + ], + [ + 80837, + "0x94063f426663347bb4a3e79f4dedeb76fbce6fd69117a2c74ad1fb4714b856dd" + ], + [ + 80838, + "0x157ab5f9e7fbf98b48af049adcfb4a8e071253ec95b991d031a91668ff3b8d57" + ], + [ + 80839, + "0x5bff632aa54b1a71a7d9a2a57d9c42ef223eb6efe86c72f9aa8dfd67ab7b218a" + ], + [ + 80840, + "0x7f5d1952ad8ad27dd17bed3354395b646468985bae0341b4048376a425e43076" + ], + [ + 80841, + "0x8e69df81fd1b0f186813304260dda1d0c9dad5792126e69d021e679b837b01a5" + ], + [ + 80842, + "0x6733c5472e38830dfb273a0e0802924c1b60a2cc916d6efaf82785e1bf960724" + ], + [ + 80843, + "0x0838513dc8264efffb2c11f8c00ebbc9c57aeb67dd7be93f18e1c613872bc929" + ], + [ + 80844, + "0x8c2f26b17db95bb20983ac41df7a5fca3e3826d2e57831eba13bc3175daeee34" + ], + [ + 80845, + "0xe3edcab823e03598c66e16c5743f7a0917735c975ac946db4cb3ad862db40a21" + ], + [ + 80846, + "0xd075510e75159d941946e666e2c75604a3f3c959523cbe2c0f025427ae85d4a2" + ], + [ + 80847, + "0x57b4b6c0a6ca2532515653445bc872dd44e621f9774f3dd4e2000bf82a0f510a" + ], + [ + 80848, + "0x38a85c99bc3e489a4431ba9a125e74f4d807c630a239419287e039f05cc91ead" + ], + [ + 80849, + "0x3c95d08c99482d57bbcb9fd333cae4010628fc8b1c96b3ad4a68d8f3dd21695f" + ], + [ + 80850, + "0xe767527430f19da2c2b21bc648e33ed2b241f5b7c9413649d9c5a555491a9253" + ], + [ + 80851, + "0xe9941baa9f3378854aaff5563340d130700f0753197c7373229542c56ac91a8f" + ], + [ + 80852, + "0xb92e88ec7cc8c340959d8d0cd5434eec5c1033b389b953b02cda103d53292df0" + ], + [ + 80853, + "0x39f01f35ba2ea7b45e3eec1681c44bd85b066a6a8521a10fe6cec5d4d98aace8" + ], + [ + 80854, + "0x36456a5e36a9a66488656bc0a92705632326e59ca40bd09f6d5f99ccc976972a" + ], + [ + 80855, + "0x12e52067750279b8bc34b03cd21bdb2809ea18a3aa68739623780aaf1150129b" + ], + [ + 80856, + "0x64951c693d7f246647fac0df504f35ff9812416f966afaf4fecf7ca92c0e0b24" + ], + [ + 80857, + "0x3918b797fe9d8cd6c34a0b315aaef550de15f4d30dce932f3be8cdfdc27343a2" + ], + [ + 80858, + "0x86ec8713a68034acb9f34b015a48644c7562766c1760927a299e517ae8ef6f8a" + ], + [ + 80859, + "0xfe75a4c99fc17131b7aae61dd0a588b7d7019e6c662aab8c7d6a8c56370fe07f" + ], + [ + 80860, + "0xf4afa8f14a0f814639c88dea304a8ccf76450ee56f4fbb89caf29dbeef80399b" + ], + [ + 80861, + "0xd4c72bc192643e20658ca54b730ef865d1c29ebfb6d1f78263c796a23b6965b4" + ], + [ + 80862, + "0x7a73b7302ab47b7acc1c5359620df9026fdc08d78c4173bb4e45e73a66651767" + ], + [ + 80863, + "0x9033ce568146d2c69139d3674fa7e18cb8319abae5e6cd24f924f750d2ca3446" + ], + [ + 80864, + "0xe4d8da84ba631f233ea7dd1336a264f499083b08290c82b717f79bfb7ffa09a6" + ], + [ + 80865, + "0xc53ec9d6b11d592921f5cf850584ebb562119ea968a5fa7bb00af381d6eaef46" + ], + [ + 80866, + "0xedcbb7d2f9626af53341464fbfc54e838e6417b7ffcfb0c2f9758641c98d5f9e" + ], + [ + 80867, + "0xbe5f726816d43ca5bbcb70f13901e60ef9040debdfaf2a73a0692c3cbef578ec" + ], + [ + 80868, + "0xb7cf72378b14c086afb4094b9b0121fcf7c6929bf2c2134b5189862644e2e506" + ], + [ + 80869, + "0xa29e21de7e54defda356f088dea38a06cf1a1a75bf93c14e30afdba213df7ef3" + ], + [ + 80870, + "0x504d3f83a27c37c5a15cdbbdfc03285f53071ad01d82bf9be2d6f862548a9a8e" + ], + [ + 80871, + "0xe02107fefe75bc8d1abb0443bff2e7a658e3a9b2fccd5157c1a7c7ccc995fe17" + ], + [ + 80872, + "0xdb92e1b449e7670ea719a6d68f496965620656ff5d25a0521f1873eb8d7ae1ab" + ], + [ + 80873, + "0xf12689607f86ce8f762df5d67054a00b8b580300511a221ff5addcf732393988" + ], + [ + 80874, + "0x0e0fe961c78fe52f44f6742795bae3e96713f4618e4a95623045056eb9e046f6" + ], + [ + 80875, + "0x7fd5bb64e800686ae28c3b55613d27f181dbd91e35d342d33d7b3e5d4df17d00" + ], + [ + 80876, + "0x83c2feae4eacf2ada1c9a630a97614a6296df6b89bfb5f954c3e306b16f0ab2e" + ], + [ + 80877, + "0xd1f3af4d6464c5bc9416d73a291bcf3e1df4fa1a396bb3e9dff92dcd451f75d9" + ], + [ + 80878, + "0xa4cd4b49e1d475a5561a91a0de592ba0c00f63b39a61eab150c6486d1bf2b916" + ], + [ + 80879, + "0x714904f162c3b9013a9defd79d4552b58c5ca8747cc94187079f2ebefee2f0d9" + ], + [ + 80880, + "0xb2263d12431e74591e2843047d79933820d2dd352d5b5c977698b328cd59f46d" + ], + [ + 80881, + "0x67977d72a573f121d9a818dd035959cd502e5bec07a4a5c30a28fc66d4fdf964" + ], + [ + 80882, + "0xf1d5f1b271cd23a79f1ab32655418e4b664663c4a77ebaf0a213fc2456e59c5a" + ], + [ + 80883, + "0x91a9f94d339407747a77d79a9c2be608744d4ecc8b84699d0120814f55a5060b" + ], + [ + 80884, + "0x8f709bf4582d3a912e6006f6c417459bbbf182af2b500b4376cd8dff78201c25" + ], + [ + 80885, + "0xbc66848a66fc038200186f09665540b5905a350a3149bfdddf9575eabb2f0a98" + ], + [ + 80886, + "0x8827e43b9a02e7424513e6d0fb6db7e93cc5555fac9b832584570784708a8f90" + ], + [ + 80887, + "0x7c76b6619ea66b8e7fa2791ad1a3dcafe863388d14db50dd022f63748e2b6ccb" + ], + [ + 80888, + "0xace6e76c69db6530277c9749ff3e77a64c6e7123790fdf5f248cc67b761e10bd" + ], + [ + 80889, + "0x8f96a0d2f9524f0b2bf9f520f35516969fba795e005f003dac5721cc37c6b5ce" + ], + [ + 80890, + "0x68a635c0d7ba10b3fff9108deecc4e9af70d760d688c5614a000503b1f2f006c" + ], + [ + 80891, + "0x47dc91daaf1758699e90f8906c61bcb6dd020b3726a1e0d15a7fe5177bfab696" + ], + [ + 80892, + "0x79f1ff991dd34c0384b9ddb64a7f3791116891447f2c960de5c6d80f9a705b5f" + ], + [ + 80893, + "0xf41832c4af88cdd31eec42b233ccfae51a150cd41baf59e9dc6b64dba2b63c7a" + ], + [ + 80894, + "0xfaf2c58e9ba01047e4cbe2c147a4614d46ad39c1d64cbb43d1de60f9c75286b7" + ], + [ + 80895, + "0x65a671bcbc9b6f905811ce62594feda18a12b8c6ca1c72081c1dde3f6227d1a7" + ], + [ + 80896, + "0x815d16976077a9b599fd5d6a82eda30a0bd1d91c66718528aaea9d4dcd12274b" + ], + [ + 80897, + "0x710c0c1d4d44ed7925343b10c7ace3216741b31a91506fa0f1ad80c96c049b14" + ], + [ + 80898, + "0x5fff52604275c206678c800d95cd0569325ad67413b50ac15252c3440856cc78" + ], + [ + 80899, + "0xe67fe87df50a63d348d300de16353e560c1b86f1a79a0bd09e2a7771ae7e92ba" + ], + [ + 80900, + "0x5d854ceaef9de086a2361cfa1f829f843aac00181fa5ae27e69171c198853895" + ], + [ + 80901, + "0xca45572516a2d4a948cfe518d0bb5378a4d23b800522f515e59176c31c069160" + ], + [ + 80902, + "0xa724d3e05f16d3d8c97296f59802c7f9ebba4284808126d955d6fc96dc1c4729" + ], + [ + 80903, + "0x552061794bddc5185033e936ac8ad07d3a315c7a3d876989b2d61328ad0a3128" + ], + [ + 80904, + "0x23c1ee3ff43ae0d42b9b3f0d2c6c2fef6c8ab8fdec9772540b563cfebbb77a2c" + ], + [ + 80905, + "0xc7e33046cb4819a17a21a00e9c0efa55765a974d8a2364f4bed9a16034ec2828" + ], + [ + 80906, + "0xa7b5712f7d22449a6f6a8a55a33c2f76cb54be86faa2f0d892c8bf6aa9fe59d9" + ], + [ + 80907, + "0xbd8696f54d266714d694b41d3b6466ef242e999938b9effd2de035c3f2c5677b" + ], + [ + 80908, + "0x7ca324c51788d4cc173bd3a013702f80e68524b3b005299be9012621dc8e82c4" + ], + [ + 80909, + "0xaea0e3a560abbe0de857592d37bfd43bd986d1cd461f43139d951d123140d495" + ], + [ + 80910, + "0x640f0ed5ec76969f53d90d5e502f6fce0ec65ab28a728e92a81dd27c466dad09" + ], + [ + 80911, + "0xf4cb4beab37872930f84e37b38c15a0a5e5f5d95927c91284ae880337abcfd9a" + ], + [ + 80912, + "0x65121e870e1668a0a55c9a50d31166294cb8e57bc007897640a7414f0866ec3b" + ], + [ + 80913, + "0xb3ea0a195e00db811104db7362b555fd1ab8401814ee841ca7f4d3344843708c" + ], + [ + 80914, + "0xf92735fd64122d7f6f6df6f20f83fc0793675c81e25e54dc64b3ca027d8ef9a4" + ], + [ + 80915, + "0x78e07fd2fcedb51175785be4657b8d2a554c16c1d2cc396ca3738dbf6f9373f7" + ], + [ + 80916, + "0x0ce35f9c3a736ff2937e6b7addcb632855e347ab3c5c08b74d2dca5f82265426" + ], + [ + 80917, + "0x007b1e317c0d9e7565260e1fe7ebfe031de4291f6ddad75210f8563479e84178" + ], + [ + 80918, + "0xbdb4b040d7942b8bcd206e4e12632892d9cfb87c905f760f0c918dcb682d4eca" + ], + [ + 80919, + "0x1eb8ee38b4d246a0502cc56c8948ec6e17fa55b02002c9916e1a3d9867e3c852" + ], + [ + 80920, + "0x2bd9fea5702cc9c0b96fa3b34559c4a72c5ff4f1d572d9e702390a13b0b1c145" + ], + [ + 80921, + "0x12ae91b4df38f549004ab8c4482d9ed1a2841e43f649ee1a89156f9accea51ea" + ], + [ + 80922, + "0xb960711112f51fc11689061a66a9b4ed3c0bccb3a6f027bd2cceeb2d3786e780" + ], + [ + 80923, + "0x4d474df4920e9136e7564cf0e9ebd03bbcd8cb8bdb38c68b8a9c35120888bfea" + ], + [ + 80924, + "0x73d84f1013038f10bebf6606486c1bb6c6cf530a4d35f482be91df74fe15c93e" + ], + [ + 80925, + "0xf08dc8a601c46c43bb6f94edc1b0734a61e7f8efc82d915b33136fac63d5a407" + ], + [ + 80926, + "0x2b695a3c4d05e233cbe18213d50b10fb9d83914aca9a935aa6f6639b1f54c32e" + ], + [ + 80927, + "0x893a6295ba2f66f083feaa39e0570ccbd446166c2e15238d59a911f2826a47d8" + ], + [ + 80928, + "0xc2b2870e39e8344e18296012c4d846b78cd3bdc80145de6b595527c8845ac57b" + ], + [ + 80929, + "0xc32bf651e3ea72bd6808db5b4dbe4577078d418404221809ac12897082ab8bfa" + ], + [ + 80930, + "0xf1c2f5ed0a7d6b7f083c1a68f75004fbfc929ef3f6ab3b46acbba373feecfdde" + ], + [ + 80931, + "0x51638cd9a09c8ee630706d736c81ea0dde7d97234a515c6064be20d0b92b8418" + ], + [ + 80932, + "0xf3cb85e3f551c4e21d0c6c28e940c6bd965c2a20b5eb3375a9d8d32a65c58c4c" + ], + [ + 80933, + "0x65f05787c7f239dcd4dd9cac2e5c516e25f8a8b898eec284531c16a8da83ed2d" + ], + [ + 80934, + "0x58567dd84c5ce96fdfb6ea5794201199d259374e4879e76df9a6622694de95f3" + ], + [ + 80935, + "0x2031a54f25cefcf00c464f36897691a73616f777582bfcd9c03e5649d3339821" + ], + [ + 80936, + "0x0fc1b95d2579a6cb03f21083da53a989db2491f37bcc339195864baf21932823" + ], + [ + 80937, + "0xfc501a03f2666c7c7db743b6a7296692789824828c033f2254e3565de63bccec" + ], + [ + 80938, + "0xe752390723b297a6d326cc9d10f5a3219b5eede3159f9ea1c7fc7f26401b02f8" + ], + [ + 80939, + "0xfdba580d13c6964b46e7d2a66f2aeee150174facaf1cbff66016832b02d9f796" + ], + [ + 80940, + "0x2221d3b867caf0c1d4993845a60ba8389558ccd9d48bf97f51331e87ff51ce84" + ], + [ + 80941, + "0xb4ba8a208861ab1db5302570f7d589adc5f4826812b4759d625cabe2b68231a0" + ], + [ + 80942, + "0x0bc24048f15b64e0c836c387539686906198ed0b3a7adbe2aed37f41eaee5577" + ], + [ + 80943, + "0x5c3230d24716dfeb57c624bc44f7d80b65c95f4b810aa945f8e5309e45327104" + ], + [ + 80944, + "0x798ef63bf3ed33e8608b899ee956fc3a0a0b359a4466141728589ea914d42453" + ], + [ + 80945, + "0xb158a6b683a60e8d3debc803a89bd18af0f1307941039304a2e33a02f237325a" + ], + [ + 80946, + "0xd420374f3f533e0136b3fb1211357adbdde913776fbb851492ca80980e77ce2e" + ], + [ + 80947, + "0xbf7fea59019b30fec39a94c29b026285e0c2d6fe034da2555095999ca694914b" + ], + [ + 80948, + "0x3d232c12ca3a428474f8c8e21992efc614775842c365a29ef44d9b7f24caafc8" + ], + [ + 80949, + "0x87d17a79601cf126d2b5c95f1ac989f8a90f2b4e9f3b13d1d721776ee2a53c24" + ], + [ + 80950, + "0xd23290b4b63c19aafd1edefd1cfcf34829bdc6efa723df17cce5eb7612b17a0d" + ], + [ + 80951, + "0xa8d8db2850de878732f5a7b95bc08ddbb944e1e8e8884a2a2a6a98275bb7663d" + ], + [ + 80952, + "0xb2c005429b5d1f35d006693509651726e67fd19d34ac04a3473df90d8bed52ab" + ], + [ + 80953, + "0xfcdb4ee3a4fa1273ec429c49704ac1afe6c8c6f858640155276661b19a5c4ca0" + ], + [ + 80954, + "0x9b988ccb0403ebfb0c15486ed6b63ea070121d7754f88d1f14f85651c083f240" + ], + [ + 80955, + "0x39c3d647ae0dd09b1bbf1699803e3c1e76448787baf12cfced826b8d75dd4b7c" + ], + [ + 80956, + "0xb24d2ff081df7d34392e460ddf8fdb8c7cc5ee502e8cae2a279474f329fe7d05" + ], + [ + 80957, + "0xf1f1daa9a70d68fea94a204260b298285ca15469cdc1855be3ef731ae44707bb" + ], + [ + 80958, + "0xeadcfd836598031d7d1a6f5c14028edeffc5baed420b2fbd343a5e21cfdd1fd9" + ], + [ + 80959, + "0xe1391944a4bd9d2fe474c8b12de8b3d151bf3afc97addb223ebaa59592ddcd94" + ], + [ + 80960, + "0xdef9733fd939d6b260fe671867ee7cdcd802332b9afdadbba0ac24b964126d29" + ], + [ + 80961, + "0x641d150032ebcce0d0d7616b010eb29df976191edfad5c8751da8533f78f7946" + ], + [ + 80962, + "0x7480f68a44e03c35a1ece1cec7402fe9fb5b4d34d4f246c967c5178e1a7fa330" + ], + [ + 80963, + "0xefae600a6e868a4213011b106f295f79e5c86186b8382143c0a402761226a9ec" + ], + [ + 80964, + "0xc6ce05cbc32d8a1d786baa6473cc65c29892a5a47e7474f384b65c567d425748" + ], + [ + 80965, + "0xd3a3e3c3c964a77a0821984903c958079b294b195b7b13e109f26f3bafcfa092" + ], + [ + 80966, + "0xd42712815a77eaca169de5c74653b4c007159f47fb5d50d63f57cfd3f2c5868e" + ], + [ + 80967, + "0x9b309f8c13534ed43e399d45a5bcfdb38f49c5633b458f6db71d39571e9d8d41" + ], + [ + 80968, + "0x74729f1a9de838f15def5f5c183b7d3596b03afde1ed6302cabacf8a677f82a2" + ], + [ + 80969, + "0x4c07a8677c028eda23e1665416eb54d364959f3210e682978082d560ebf2bea4" + ], + [ + 80970, + "0xe3f992775c7dfbd2485ccecf516ca7ee9961c897296190c45fd7abb1431ee04b" + ], + [ + 80971, + "0xa8492953412a007f2f7696922b661a27fe2ec42e566504b3aacae2f97a24fab0" + ], + [ + 80972, + "0xe14d6f60ece95189dc8eaed8a2ea5c2da62cbdf988588222a4e99d242ea3beb8" + ], + [ + 80973, + "0x5c7ca844c8eec0cce4e4640ba6fc486c1739eb44aa50de86f55008a4d597357c" + ], + [ + 80974, + "0x2b33e852101ee35f2c61ee7b35152521eefb11364505bf6903a964c93d4001b6" + ], + [ + 80975, + "0x5516b628a8ec4328bcfdf90785fb5d96d22d60b4db638ea1b93bd837d7c739a3" + ], + [ + 80976, + "0x9959b993b66dc1285d0ecd8c95b74fe70949c50d5a97d822ce866291f94f2926" + ], + [ + 80977, + "0x04eb959364eae151a41d4a8d970795da0bf8a0ef2f64bde753cbca4ea3d745c4" + ], + [ + 80978, + "0x22907520934a83b056017bfaa72a0ce92fd8658c81dc8236f3b78f1f7e055aa6" + ], + [ + 80979, + "0x59fae978e8063799b4ab6335ca7ac097a25cbb4d60f9af04eb2f647a92ead4bc" + ], + [ + 80980, + "0x9ef7533f2e9c60f4acf9618efd1c47b147fa95bbb27caac54c1bb19540befeab" + ], + [ + 80981, + "0x9986f742ffa9d17f641293e788e53cd904b6edba3e971603470186f61f934e77" + ], + [ + 80982, + "0xbfb394920e4bd707342ddd41644a5733b7041e517d4e88863924046cd512ec19" + ], + [ + 80983, + "0xa3fef379a91df7a254bc4897cf2ef61ff25b380b2142374ff40d91d5140ed21c" + ], + [ + 80984, + "0x71626767d14a0735e3716b45b847611dc7fe04c6d7a4b22e3a6031d621b6d560" + ], + [ + 80985, + "0xb2c03912458086a6b421a8c83cee657c9276a1d38db128799a3170cba30e3266" + ], + [ + 80986, + "0x8f8000a6e85c09716e8766c9bbb9df1287d46f54625ff139d5e98142a58c9253" + ], + [ + 80987, + "0xfc88dba1647b7d43ef19a1a38780dd238228dd46beeb50ea0e00853f06c78471" + ], + [ + 80988, + "0x7e383c88765b3c499b05b60cdd4470cf9b23add0f6bd33990f207a43a51e431f" + ], + [ + 80989, + "0xa133fe8b34c9e839bcd4a75c2a5858874fa19d8bda43cd1c1af781cd2689753d" + ], + [ + 80990, + "0x466a4d175cb847a9899d98a6863900be93ecfec2f1681c8595db5a62b843478b" + ], + [ + 80991, + "0xb9b1bf393868e53ec5da954cb87fb65eef79ab9caf59cfb636f6cfd9c5c92f8e" + ], + [ + 80992, + "0x2f9eb9f04e16f420d5b98d6a28e25494207bc05affb6b914a1b98fb2244bba75" + ], + [ + 80993, + "0xc73829f98810dd017f45a0bd71ad7db3d07a1f5ff0a57c07994fa095989e6ef2" + ], + [ + 80994, + "0x9cdecf07b8f03979fdcb8f7d4d2c5ef4a13adc19edc0bb0c87a8e0beac6ff3be" + ], + [ + 80995, + "0x978c9e3471a4c37fb6aebab56d41d660720681073dfc68e5bef41d594d67ac8b" + ], + [ + 80996, + "0x739dad0614f0e852037aafbab92b4072f0f9dc9fc30c489465d67a18023f80a4" + ], + [ + 80997, + "0x1fbba34f589373e32f170a1cc3bf5c69083e8e2af5167e6ff38eab89ca2c0a72" + ], + [ + 80998, + "0xb58b192d29e3605fa4006c019698a7fbdc14639784ceeac087ae09148514e68e" + ], + [ + 80999, + "0x77f67acdb2afb4754c6e31005b27aea691cd3edce3094f87238df3f887e1bebe" + ], + [ + 81000, + "0x50e198025598c36879156ee586cdaa18bfa620f117c49baba1b2093d77b83831" + ], + [ + 81001, + "0x416e7d689843b2ddabbc9efd9b6f183b883dcb28e857977cd130f989ff7c72d6" + ], + [ + 81002, + "0xd639747fba2f48c62f7aeb91ead4b7ffa1292d55bd15ebec58879bc51b05863c" + ], + [ + 81003, + "0xc43806d1ab9e6d84be147ac3b74130fc2dca01254b4a847eb2713c3e7c1cd11c" + ], + [ + 81004, + "0xd854d522a33e98313a872fb4c47675f54eeed084d99794c5c72aab4464ee70b4" + ], + [ + 81005, + "0x752458c4e11fe4d564250bd4bb0122c1a1a4fabb124942ed9c3981887d612854" + ], + [ + 81006, + "0xe3eef7ded5487286de24e9802f5dfa1afe7c477666f2a26fca7606d3b9b8694c" + ], + [ + 81007, + "0xc525b44a5e888c3bdb0f66edc5bb4dbe9ce9116279b01fa510037760319ca15f" + ], + [ + 81008, + "0x9552445c57b6e2cd01d95bf4013bb913ce1611e5d707f71e528ad7522e830806" + ], + [ + 81009, + "0xd0eba42e3dd5cb37d42d6891eb47e5e2bdaeccbf9ddabeba06a09cd2b4b15f55" + ], + [ + 81010, + "0xcc6fd1b1758929e8a94d3bfc1b4c5444c55023f8939d4b4fb539b523219b00fd" + ], + [ + 81011, + "0x826ee09096704d3798aa4fccfc02cdedc59bb7472332b624d3716fa3aa14e612" + ], + [ + 81012, + "0x3f82182c620ce4d5ac0a776b0ec52b590f7d30c3bbea039d9a6ec3c6ec8a4aeb" + ], + [ + 81013, + "0xbdc7e4e07b4fc9245d8f5e7d7b207af203ce271b96912963aae2012de4d0dfd4" + ], + [ + 81014, + "0x797fd61f65fdc95f190d0dec1ad3eee5c98211c4529c83d013236838eda13d67" + ], + [ + 81015, + "0x7928f9107a62bd4d597894701cf4bca71365052e3560c23af7fdfa3374105b45" + ], + [ + 81016, + "0x78dc4dfd56d144d629278047ce28d8649e02b55768a137cf98260e40e6c35989" + ], + [ + 81017, + "0x3cd444ed1bd6f3012197839db12d566729ca5fb8c133af4353afbd756104c03c" + ], + [ + 81018, + "0x26350e310daef399afac9ac19b106d5f160d8dc6921fd289b196b4e979742bef" + ], + [ + 81019, + "0x99e1715d9ba9a3d125e77357595bd918766a4d46c20d9db1dfa9ebbd002a7bfa" + ], + [ + 81020, + "0xdf16ba27999f8dacfde10d403cf9982fd3b7545793e678219f09818e04ee64d6" + ], + [ + 81021, + "0xc4690b3c4b89b7282988cd2f71c9e43d3c17d081f0c735d04fbd01247a74bd6f" + ], + [ + 81022, + "0x94a1951db605f9a7dff70e2fd919e154f1291d31b4baa6f9e88d8f904c74fd54" + ], + [ + 81023, + "0x9846feb46081eafd35bd5f134cdf98f7a2a270b52a36779de7fb132250823a6b" + ], + [ + 81024, + "0x3716fe1bc70f07f53358e7936ff4b652a035eabe199265f99809dc4aac49d002" + ], + [ + 81025, + "0x841c89865fcece06f48c90ca283a1dbafb49e3fc422ff4cc5d3ec7fa68ac0f14" + ], + [ + 81026, + "0xcf2be45d7a88f2ffa7fa45366404b8b0af131758b7706f099ee8daa6b5d17b94" + ], + [ + 81027, + "0xf3c7dd48103bf10e13bea5800814475a543f4bf53958e1ad464b112d0c763d00" + ], + [ + 81028, + "0x6ebafde3e8fe7067903074cbeb854b9957bbc45859143f618ada972fccdc2c0d" + ], + [ + 81029, + "0x05146b0d5efdf34a927593169ecce2d2d112a3faa3b79cb2df1a0a2e6a4bd7c0" + ], + [ + 81030, + "0x33def97a4f3ca92c7af02bd19bb9703c516f6df8917928734953ec9fddae68ce" + ], + [ + 81031, + "0x18bbaeb286265fc8edb34abbfd3fade65196f09125dae4d31347401eb635b63f" + ], + [ + 81032, + "0x56aacb0d332a22fabb3bf8ac90b030fb82337b3b22b5be334d9a4f633847480c" + ], + [ + 81033, + "0x8d5d77464c966282353e2019598e0f1d77b5afb938475275f4a3b0d0c29d9c5e" + ], + [ + 81034, + "0xc8ccc7baba52e4fa56d488c78abf451919aeba4a029068fd364059633446c80d" + ], + [ + 81035, + "0xc295e160d3378e12e6b96de0eac43d903aa656f9255e58a3b0222bc229b18bbb" + ], + [ + 81036, + "0xd4cacd3eb6e8eebf312225c23284c15e347e73539816fe91954b482a9acb1753" + ], + [ + 81037, + "0x096f15eb2cfa8a6811e08ee533e3a9bfaf3447225c214157d33f27d532423598" + ], + [ + 81038, + "0xd2464a278a262d447f8d67feb035e8b9eaaaa849eb87fb330e1ac868e6b1d063" + ], + [ + 81039, + "0xfea1ffb0177d362f9e041268c395858091da08fb55de5c922493bf0f3dd19476" + ], + [ + 81040, + "0x5ab2d5c6f5789372abc97b4792b2c4e3476c70bd538764ba74d684ed102fdd0e" + ], + [ + 81041, + "0x2f5be38f279790c0e69e554fcaa28f82bd17ef645338539416db0c9252040f91" + ], + [ + 81042, + "0x844ebc844ba1c0c301a6a9189779e17e557f824d85d01247b4a0b1af49f50d56" + ], + [ + 81043, + "0x944855a63daa3b755a70f2ba5e6a956b037eeb210dd9e80cef73b7c814645adc" + ], + [ + 81044, + "0xf046f77959a6aa199f3d999b2c16dbeb1e9038c308cb1dcc9bba26141e0e5f35" + ], + [ + 81045, + "0x4e420f14ba78477c090eef99a098b62dfdadc3a61f7c0a5499797380763693db" + ], + [ + 81046, + "0xae10dd9883739a2cd9e95cafa4f9cc7b2e38e477207ec49fb3fe64e3fcf21613" + ], + [ + 81047, + "0xa9e686e36215cf1053ac16cb445e537ef8a633aac308ad755716f6374c3c1238" + ], + [ + 81048, + "0xa7ab60b451ae4a7fa639faa0c8b1f2ad7bb6bcb853242797f4aca608b2e69193" + ] + ], + "rewards": [ + [ + "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "0x32691598b1032000" + ] + ], + "proving_pool_credit": "0xc9a4563d834e400", + "payouts": [], + "blocks": [ + { + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "blue": true, + "txs": [] + } + ], + "pre_state": [ + { + "address": "0x0000000000000000000000000000000000000210", + "nonce": 1, + "balance": "0x0", + "code": "0x608060405234801561000f575f5ffd5b506004361061003f575f3560e01c8063aa67735414610043578063dea5c2e014610058578063fe7e05d51461009f575b5f5ffd5b6100566100513660046101e0565b6100ca565b005b610083610066366004610211565b6001600160a01b039081165f908152600160205260409020541690565b6040516001600160a01b03909116815260200160405180910390f35b6100836100ad366004610211565b6001600160a01b039081165f908152602081905260409020541690565b336001600160a01b03831614806100f957506001600160a01b038281165f908152600160205260409020541633145b6101635760405162461bcd60e51b815260206004820152603160248201527f446576656c6f70657252656769737472793a206e6f7420746865206163636f75604482015270373a1037b91034ba399031b932b0ba37b960791b606482015260840160405180910390fd5b6001600160a01b038281165f818152602081815260409182902080546001600160a01b031916948616948517905590513381527fa47563c41dab010f91a8ef9dc7ac2bcdfa0ef2af697e575048e71e6eec60dda3910160405180910390a35050565b80356001600160a01b03811681146101db575f5ffd5b919050565b5f5f604083850312156101f1575f5ffd5b6101fa836101c5565b9150610208602084016101c5565b90509250929050565b5f60208284031215610221575f5ffd5b61022a826101c5565b939250505056fea2646970667358221220cbf48f5aa6f911f83b3c2b09adf8c418f5da6fc530c3d5db9d4ba0be24c4bf5864736f6c63430008250033", + "storage": [] + }, + { + "address": "0x0000000000000000000000000000000000000220", + "nonce": 0, + "balance": "0x13cb950f3274d6b8f000", + "code": "0x", + "storage": [] + }, + { + "address": "0x0f002c928c363c7041cafe788f099cf7d35452a1", + "nonce": 243, + "balance": "0x1bc36179476e9a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x18524811fa2e76dd0770d29308b51bcbdc0499f9", + "nonce": 242, + "balance": "0x1bb8aa483e652c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x1aca71f7872aebc85b4d4a9ad47a09abcc624d65", + "nonce": 0, + "balance": "0x2860cb14365c048800", + "code": "0x", + "storage": [] + }, + { + "address": "0x233d639f53225ea5012dc01ceb0a5a30891021cd", + "nonce": 264, + "balance": "0x1b8eea59c54b4200", + "code": "0x", + "storage": [] + }, + { + "address": "0x27fd8475c3db352fdddeaf29823e00c7d33ca39a", + "nonce": 248, + "balance": "0x1bccb45d83401000", + "code": "0x", + "storage": [] + }, + { + "address": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "nonce": 0, + "balance": "0x1e063716f0755a0c800", + "code": "0x", + "storage": [] + }, + { + "address": "0x46681948060140945958068b373f94581e049d4f", + "nonce": 240, + "balance": "0x1baadca68fcb1a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4cb00bc539538d2ee7172474f2b06f021944ece4", + "nonce": 258, + "balance": "0x1b3e54c72c629c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4d7c0d6f3ad5466691649e98ebb59cd5b5194687", + "nonce": 0, + "balance": "0x1b598fbe8e08549c6000", + "code": "0x", + "storage": [] + }, + { + "address": "0x5a5e606bda0fe1b2b4298aadd55cc6ec57ff698a", + "nonce": 0, + "balance": "0x648cba8883ddb562c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x6613ca66c14338db771d01c238ae140c71310535", + "nonce": 243, + "balance": "0x1b98d3921da34800", + "code": "0x", + "storage": [] + }, + { + "address": "0x66577de387feeec96c10f7f8748f13912589c78c", + "nonce": 240, + "balance": "0x1bc37e3123b20a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x7e7b1db26094aa913933f65127b328c1861ff5b0", + "nonce": 256, + "balance": "0x1b89d2467bcd2600", + "code": "0x", + "storage": [] + }, + { + "address": "0x90acb15171deb958d3191d5c02d807afc367f0e8", + "nonce": 240, + "balance": "0x1b9e3c5e84fb6c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x9f296bf64eb7051912c9c8112da3e3ea6cf546e1", + "nonce": 266, + "balance": "0x1b69b11e00d88c00", + "code": "0x", + "storage": [] + }, + { + "address": "0xaa19223d63c82bf9be5543ec96e5504834d17bc8", + "nonce": 246, + "balance": "0x1b9abb1eb524dc00", + "code": "0x", + "storage": [] + }, + { + "address": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "nonce": 0, + "balance": "0x555a99f6412ff4a000", + "code": "0x", + "storage": [] + }, + { + "address": "0xc2faa4a2866422a865c9b3b3cb5988bf4d07f0b4", + "nonce": 0, + "balance": "0xc05b9fc73002b32800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc45d23f49451faad9afadf877feead626ea8d083", + "nonce": 0, + "balance": "0x9ee8f806428adc9800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc8621de5921418701619ba715044575073d2ca62", + "nonce": 243, + "balance": "0x1bbb8989a888a800", + "code": "0x", + "storage": [] + }, + { + "address": "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "nonce": 0, + "balance": "0x1f29d0e7a3eaed38a400", + "code": "0x", + "storage": [] + }, + { + "address": "0xd958300657e7931b06fc6f60dc4bfe7d8e8ea7b8", + "nonce": 252, + "balance": "0x1b79a326be219400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdadb11b01f6004eba6da3846a0e1bf6f4f71fb82", + "nonce": 244, + "balance": "0x1b900b01183e5600", + "code": "0x", + "storage": [] + }, + { + "address": "0xdd442fcbb964a3afdc90d49b408e8dd296fa86e8", + "nonce": 0, + "balance": "0xafd0c09109be7410400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdfaea67368f3e3753397d878f97efe6aa8020c2e", + "nonce": 16, + "balance": "0x4f20fde2ac0f98ec00", + "code": "0x", + "storage": [] + }, + { + "address": "0xfd4fca79b266a73a7defa776b04c7e15bdc70e86", + "nonce": 238, + "balance": "0x1bb5cf9909338400", + "code": "0x", + "storage": [] + } + ], + "fees": { + "base": { + "pgas": { + "version": 0, + "cycles_per_pgas": 1000, + "intrinsic_pgas_per_tx": 200, + "modexp_base": 1000, + "modexp_per_byte_numer": 10, + "modexp_per_byte_denom": 1 + }, + "block_proving_gas_limit": 30000000, + "shard_proving_gas_budget": 7500000, + "min_execution_base_fee_wei": 1000000000, + "min_proving_base_fee_wei": 1000000000, + "initial_execution_base_fee_wei": 1000000000, + "initial_proving_base_fee_wei": 1000000000, + "base_fee_change_denominator": 8 + }, + "v1_activation_daa": 18446744073709551615 + } + }, + "plan": { + "shard_budget": 7500000, + "consensus": true, + "shards": [ + { + "index": 0, + "tx_start": 0, + "tx_end": 0, + "over_budget": false, + "pre_root": "0xc6665c33a141e50cdb7a7194304afe5672016fa15f394ffa7f414d70912a7654", + "post_root": "0x57a18ca3df176250b8ec70274aa0d48e2b809f9751f2a0b17f8c67b71d00c5d1", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "link_in": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "link_out": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "witness": [ + 2, + 0, + 2, + 16, + 14064 + ] + } + ] + }, + "expected": { + "pre_state_root": "0xc6665c33a141e50cdb7a7194304afe5672016fa15f394ffa7f414d70912a7654", + "post_state_root": "0x57a18ca3df176250b8ec70274aa0d48e2b809f9751f2a0b17f8c67b71d00c5d1", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "tx_commitment": "0x0000000000000000000000000000000000000000000000000000000000000000", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "node_state_root": "0x57a18ca3df176250b8ec70274aa0d48e2b809f9751f2a0b17f8c67b71d00c5d1" + } +} \ No newline at end of file diff --git a/proving/fixtures/chain/block-81050.json b/proving/fixtures/chain/block-81050.json new file mode 100644 index 000000000..1edaa2335 --- /dev/null +++ b/proving/fixtures/chain/block-81050.json @@ -0,0 +1,1325 @@ +{ + "format": "igneum-prove-fixture-v1", + "source": "live devnet export from node 1 at tip 81076, 5 October 2026 20:06 BST, proving v1 chain fixtures", + "block": { + "chain_id": 4463, + "env": { + "number": 81050, + "hash": "0x42515b83488525c1c8e72417df7361aa61a62e3f78892446b06d6d6bccee6ada", + "parent_hash": "0xdba699a61f7e4b3b8356a7494afbb73237f00385d7ad279a77002bfd451b2643", + "timestamp": 1791227128, + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "prevrandao": "0x254aacca8137f722a752d25a5dddcb68c3ae710f7224cda7b59e8d307545806c", + "base_fee_exec": 1000000000, + "base_fee_proving": 1000000000, + "daa_score": 0 + }, + "block_hashes": [ + [ + 80794, + "0xf657d95990ce07d72922d34f07dbc66bb681797076a3e9819b2b3349218efaec" + ], + [ + 80795, + "0x4bf9d010c0798a5340983b27eda6a9876e6fefb925104fdc46291a5789f1f70a" + ], + [ + 80796, + "0xb09af3831aa1d8f4945d50fa16a602848a2da22b14dbc1f7af2844cf19e90040" + ], + [ + 80797, + "0xe099536cfd64654690afc152b97866b32da6bf81f959a6c61f9e43bfd2ca2d9a" + ], + [ + 80798, + "0xd6af2cf6184f15ecca77ee9e7e8d2ca718d88f9cea39b475713ceb0e151ddc36" + ], + [ + 80799, + "0x9c76f9f6036c7035fd52a7a4e3fca99451715e729fcc9fc424487faee9a6134d" + ], + [ + 80800, + "0x9a046cd99eb89c551dd0c698f1969069c02ae36cde4140ff4b864d624fd9dff7" + ], + [ + 80801, + "0xc9bffa329d649c6540518cf249e25e845743dc01fcbcfdfb8c798f1455efeae6" + ], + [ + 80802, + "0xb124f2de3c17a985f10c0b58d2283631bcc12524de8d79b0955aecfd2d7f78aa" + ], + [ + 80803, + "0x97b21d6cd3f1b321d2c9729062c305a7955b97646d76910714f899c75d275568" + ], + [ + 80804, + "0xaa3f0e114e156f706e11a701fb940bf54e89b09d83953be7f6733478c7a33093" + ], + [ + 80805, + "0xe5a5868b02ec30304079652a8d5b8af4a9da0a5a676615c1a10d833feb4eb0f7" + ], + [ + 80806, + "0xd97c5c4288f56955e1a3e134eb63250e1becb907b0fdde9ce1383d8d2c8a3939" + ], + [ + 80807, + "0x8f7bf1128e5d03607cca246b86e63cfb4bfe73e74395c8660d027172376b5c36" + ], + [ + 80808, + "0x5d5fe4c3bee7d758986da5a433cfe5221be392419ac22e5acddd701cd12c8142" + ], + [ + 80809, + "0x9b0d0a871ce8eaab11a185a4344f9f48a58b7111f94a8f5221d557722cc5ca5a" + ], + [ + 80810, + "0x58cad4ab776ba7eca94e813323e28a6467a7fcf39d5aa2069e9dfef6f8058e5c" + ], + [ + 80811, + "0x706788247d6c05475eb3bdbc6f181e76429b57a97f168004832fe5ba4b20c7b7" + ], + [ + 80812, + "0xdf26fb97f999b1016865b3fee539e72e3f66d6d228b2aec7e2f59ab93e5a9ee6" + ], + [ + 80813, + "0x8e612c07fdb693a44353a08899e8beae8c93fb1b85845ad965f96e7d2f77fed4" + ], + [ + 80814, + "0x806f4c29ea28af1a3ce0c3ba775c776006ba8cf4b308ee617dc207dc3f740b29" + ], + [ + 80815, + "0xebd8e602049ddc2f8d11d5edc8ee837eefac77418e73e9bad97f6f1c7a85e8aa" + ], + [ + 80816, + "0xe2b069b77ac05cb0a95ae419966feeb20660e4ff3b9a68d044a5721e1525dd86" + ], + [ + 80817, + "0x1ca5b79e26dc425a94c3a6bfd339516caa0a7f8e5e2a62389bfa466186bca495" + ], + [ + 80818, + "0x7cf382f956597cbd36917036dd15b774500f97714e50587a35ec9617662052d6" + ], + [ + 80819, + "0xe70494836c5b5c68b6a167c868db166f9630140499f96e6660d3e3928cad6a14" + ], + [ + 80820, + "0x15f73eefc39e00f9d09390764f9ea0069908062c563686e73b49ec3919c24be3" + ], + [ + 80821, + "0x7f2bd2cf68fc14e6a5debaac39f368534657c41cc587402fac1156d24998e55b" + ], + [ + 80822, + "0x2641ec1c3397b06ab5855dcf6ec2e23de301f5831015a32ffa9f29e18cc2f084" + ], + [ + 80823, + "0xf643e963f5a62a49e935b4ef51aab9bf2689b0b9ec483924d7f42934c8ce5e07" + ], + [ + 80824, + "0x867ee0219bd30fc3117c79295cb1075ca21281fdd285467ae128aece1c1dd9a7" + ], + [ + 80825, + "0xab6bb19259a039d30c3022c9da3494af8c9a541dce30895df083252894c433df" + ], + [ + 80826, + "0x155511de8d34bef351ba0ae7af5bccf515e6e490135130aad5ce2eb02e103d75" + ], + [ + 80827, + "0x4e297a53901da9b364f4dce686f016f4039658ea8832d591dcb21aa44acd197b" + ], + [ + 80828, + "0x7d07b8bc9b06b20e7b56f96fff9051590caac01159d39b3f95bd592a4627daf4" + ], + [ + 80829, + "0x8392a1875388baec163e64854d6ee7664a31017402ee8e269ddc8aa42ca06647" + ], + [ + 80830, + "0x5ab4e5c586d2fb4272941bf26c15288cc129ca1092d3aaab82c478edd1942e29" + ], + [ + 80831, + "0xb0100f0a559daa435128cae74f47f9353ca2a93c2bdac8dd9d12a3130ce6a818" + ], + [ + 80832, + "0x326d1388d0911871c5087925500c4ee5091624331d07659b9d2fa09e83f5884a" + ], + [ + 80833, + "0xc01c186cfcebb8d53953589ce595afef213d46e481b8b5dee7e0a71abd773f6a" + ], + [ + 80834, + "0xe845934e79eaeb8a8ee0b9396b3a13f13c8ac45a0d4e81017c73042789013c38" + ], + [ + 80835, + "0xa45f8f03249624a2d35b54eb83f4edf54e8d1acd061fee3965d9164a132c72aa" + ], + [ + 80836, + "0xe72a59a64af9aa223b5341819a5872f8626908b5c1857d0617706f6999632e01" + ], + [ + 80837, + "0x94063f426663347bb4a3e79f4dedeb76fbce6fd69117a2c74ad1fb4714b856dd" + ], + [ + 80838, + "0x157ab5f9e7fbf98b48af049adcfb4a8e071253ec95b991d031a91668ff3b8d57" + ], + [ + 80839, + "0x5bff632aa54b1a71a7d9a2a57d9c42ef223eb6efe86c72f9aa8dfd67ab7b218a" + ], + [ + 80840, + "0x7f5d1952ad8ad27dd17bed3354395b646468985bae0341b4048376a425e43076" + ], + [ + 80841, + "0x8e69df81fd1b0f186813304260dda1d0c9dad5792126e69d021e679b837b01a5" + ], + [ + 80842, + "0x6733c5472e38830dfb273a0e0802924c1b60a2cc916d6efaf82785e1bf960724" + ], + [ + 80843, + "0x0838513dc8264efffb2c11f8c00ebbc9c57aeb67dd7be93f18e1c613872bc929" + ], + [ + 80844, + "0x8c2f26b17db95bb20983ac41df7a5fca3e3826d2e57831eba13bc3175daeee34" + ], + [ + 80845, + "0xe3edcab823e03598c66e16c5743f7a0917735c975ac946db4cb3ad862db40a21" + ], + [ + 80846, + "0xd075510e75159d941946e666e2c75604a3f3c959523cbe2c0f025427ae85d4a2" + ], + [ + 80847, + "0x57b4b6c0a6ca2532515653445bc872dd44e621f9774f3dd4e2000bf82a0f510a" + ], + [ + 80848, + "0x38a85c99bc3e489a4431ba9a125e74f4d807c630a239419287e039f05cc91ead" + ], + [ + 80849, + "0x3c95d08c99482d57bbcb9fd333cae4010628fc8b1c96b3ad4a68d8f3dd21695f" + ], + [ + 80850, + "0xe767527430f19da2c2b21bc648e33ed2b241f5b7c9413649d9c5a555491a9253" + ], + [ + 80851, + "0xe9941baa9f3378854aaff5563340d130700f0753197c7373229542c56ac91a8f" + ], + [ + 80852, + "0xb92e88ec7cc8c340959d8d0cd5434eec5c1033b389b953b02cda103d53292df0" + ], + [ + 80853, + "0x39f01f35ba2ea7b45e3eec1681c44bd85b066a6a8521a10fe6cec5d4d98aace8" + ], + [ + 80854, + "0x36456a5e36a9a66488656bc0a92705632326e59ca40bd09f6d5f99ccc976972a" + ], + [ + 80855, + "0x12e52067750279b8bc34b03cd21bdb2809ea18a3aa68739623780aaf1150129b" + ], + [ + 80856, + "0x64951c693d7f246647fac0df504f35ff9812416f966afaf4fecf7ca92c0e0b24" + ], + [ + 80857, + "0x3918b797fe9d8cd6c34a0b315aaef550de15f4d30dce932f3be8cdfdc27343a2" + ], + [ + 80858, + "0x86ec8713a68034acb9f34b015a48644c7562766c1760927a299e517ae8ef6f8a" + ], + [ + 80859, + "0xfe75a4c99fc17131b7aae61dd0a588b7d7019e6c662aab8c7d6a8c56370fe07f" + ], + [ + 80860, + "0xf4afa8f14a0f814639c88dea304a8ccf76450ee56f4fbb89caf29dbeef80399b" + ], + [ + 80861, + "0xd4c72bc192643e20658ca54b730ef865d1c29ebfb6d1f78263c796a23b6965b4" + ], + [ + 80862, + "0x7a73b7302ab47b7acc1c5359620df9026fdc08d78c4173bb4e45e73a66651767" + ], + [ + 80863, + "0x9033ce568146d2c69139d3674fa7e18cb8319abae5e6cd24f924f750d2ca3446" + ], + [ + 80864, + "0xe4d8da84ba631f233ea7dd1336a264f499083b08290c82b717f79bfb7ffa09a6" + ], + [ + 80865, + "0xc53ec9d6b11d592921f5cf850584ebb562119ea968a5fa7bb00af381d6eaef46" + ], + [ + 80866, + "0xedcbb7d2f9626af53341464fbfc54e838e6417b7ffcfb0c2f9758641c98d5f9e" + ], + [ + 80867, + "0xbe5f726816d43ca5bbcb70f13901e60ef9040debdfaf2a73a0692c3cbef578ec" + ], + [ + 80868, + "0xb7cf72378b14c086afb4094b9b0121fcf7c6929bf2c2134b5189862644e2e506" + ], + [ + 80869, + "0xa29e21de7e54defda356f088dea38a06cf1a1a75bf93c14e30afdba213df7ef3" + ], + [ + 80870, + "0x504d3f83a27c37c5a15cdbbdfc03285f53071ad01d82bf9be2d6f862548a9a8e" + ], + [ + 80871, + "0xe02107fefe75bc8d1abb0443bff2e7a658e3a9b2fccd5157c1a7c7ccc995fe17" + ], + [ + 80872, + "0xdb92e1b449e7670ea719a6d68f496965620656ff5d25a0521f1873eb8d7ae1ab" + ], + [ + 80873, + "0xf12689607f86ce8f762df5d67054a00b8b580300511a221ff5addcf732393988" + ], + [ + 80874, + "0x0e0fe961c78fe52f44f6742795bae3e96713f4618e4a95623045056eb9e046f6" + ], + [ + 80875, + "0x7fd5bb64e800686ae28c3b55613d27f181dbd91e35d342d33d7b3e5d4df17d00" + ], + [ + 80876, + "0x83c2feae4eacf2ada1c9a630a97614a6296df6b89bfb5f954c3e306b16f0ab2e" + ], + [ + 80877, + "0xd1f3af4d6464c5bc9416d73a291bcf3e1df4fa1a396bb3e9dff92dcd451f75d9" + ], + [ + 80878, + "0xa4cd4b49e1d475a5561a91a0de592ba0c00f63b39a61eab150c6486d1bf2b916" + ], + [ + 80879, + "0x714904f162c3b9013a9defd79d4552b58c5ca8747cc94187079f2ebefee2f0d9" + ], + [ + 80880, + "0xb2263d12431e74591e2843047d79933820d2dd352d5b5c977698b328cd59f46d" + ], + [ + 80881, + "0x67977d72a573f121d9a818dd035959cd502e5bec07a4a5c30a28fc66d4fdf964" + ], + [ + 80882, + "0xf1d5f1b271cd23a79f1ab32655418e4b664663c4a77ebaf0a213fc2456e59c5a" + ], + [ + 80883, + "0x91a9f94d339407747a77d79a9c2be608744d4ecc8b84699d0120814f55a5060b" + ], + [ + 80884, + "0x8f709bf4582d3a912e6006f6c417459bbbf182af2b500b4376cd8dff78201c25" + ], + [ + 80885, + "0xbc66848a66fc038200186f09665540b5905a350a3149bfdddf9575eabb2f0a98" + ], + [ + 80886, + "0x8827e43b9a02e7424513e6d0fb6db7e93cc5555fac9b832584570784708a8f90" + ], + [ + 80887, + "0x7c76b6619ea66b8e7fa2791ad1a3dcafe863388d14db50dd022f63748e2b6ccb" + ], + [ + 80888, + "0xace6e76c69db6530277c9749ff3e77a64c6e7123790fdf5f248cc67b761e10bd" + ], + [ + 80889, + "0x8f96a0d2f9524f0b2bf9f520f35516969fba795e005f003dac5721cc37c6b5ce" + ], + [ + 80890, + "0x68a635c0d7ba10b3fff9108deecc4e9af70d760d688c5614a000503b1f2f006c" + ], + [ + 80891, + "0x47dc91daaf1758699e90f8906c61bcb6dd020b3726a1e0d15a7fe5177bfab696" + ], + [ + 80892, + "0x79f1ff991dd34c0384b9ddb64a7f3791116891447f2c960de5c6d80f9a705b5f" + ], + [ + 80893, + "0xf41832c4af88cdd31eec42b233ccfae51a150cd41baf59e9dc6b64dba2b63c7a" + ], + [ + 80894, + "0xfaf2c58e9ba01047e4cbe2c147a4614d46ad39c1d64cbb43d1de60f9c75286b7" + ], + [ + 80895, + "0x65a671bcbc9b6f905811ce62594feda18a12b8c6ca1c72081c1dde3f6227d1a7" + ], + [ + 80896, + "0x815d16976077a9b599fd5d6a82eda30a0bd1d91c66718528aaea9d4dcd12274b" + ], + [ + 80897, + "0x710c0c1d4d44ed7925343b10c7ace3216741b31a91506fa0f1ad80c96c049b14" + ], + [ + 80898, + "0x5fff52604275c206678c800d95cd0569325ad67413b50ac15252c3440856cc78" + ], + [ + 80899, + "0xe67fe87df50a63d348d300de16353e560c1b86f1a79a0bd09e2a7771ae7e92ba" + ], + [ + 80900, + "0x5d854ceaef9de086a2361cfa1f829f843aac00181fa5ae27e69171c198853895" + ], + [ + 80901, + "0xca45572516a2d4a948cfe518d0bb5378a4d23b800522f515e59176c31c069160" + ], + [ + 80902, + "0xa724d3e05f16d3d8c97296f59802c7f9ebba4284808126d955d6fc96dc1c4729" + ], + [ + 80903, + "0x552061794bddc5185033e936ac8ad07d3a315c7a3d876989b2d61328ad0a3128" + ], + [ + 80904, + "0x23c1ee3ff43ae0d42b9b3f0d2c6c2fef6c8ab8fdec9772540b563cfebbb77a2c" + ], + [ + 80905, + "0xc7e33046cb4819a17a21a00e9c0efa55765a974d8a2364f4bed9a16034ec2828" + ], + [ + 80906, + "0xa7b5712f7d22449a6f6a8a55a33c2f76cb54be86faa2f0d892c8bf6aa9fe59d9" + ], + [ + 80907, + "0xbd8696f54d266714d694b41d3b6466ef242e999938b9effd2de035c3f2c5677b" + ], + [ + 80908, + "0x7ca324c51788d4cc173bd3a013702f80e68524b3b005299be9012621dc8e82c4" + ], + [ + 80909, + "0xaea0e3a560abbe0de857592d37bfd43bd986d1cd461f43139d951d123140d495" + ], + [ + 80910, + "0x640f0ed5ec76969f53d90d5e502f6fce0ec65ab28a728e92a81dd27c466dad09" + ], + [ + 80911, + "0xf4cb4beab37872930f84e37b38c15a0a5e5f5d95927c91284ae880337abcfd9a" + ], + [ + 80912, + "0x65121e870e1668a0a55c9a50d31166294cb8e57bc007897640a7414f0866ec3b" + ], + [ + 80913, + "0xb3ea0a195e00db811104db7362b555fd1ab8401814ee841ca7f4d3344843708c" + ], + [ + 80914, + "0xf92735fd64122d7f6f6df6f20f83fc0793675c81e25e54dc64b3ca027d8ef9a4" + ], + [ + 80915, + "0x78e07fd2fcedb51175785be4657b8d2a554c16c1d2cc396ca3738dbf6f9373f7" + ], + [ + 80916, + "0x0ce35f9c3a736ff2937e6b7addcb632855e347ab3c5c08b74d2dca5f82265426" + ], + [ + 80917, + "0x007b1e317c0d9e7565260e1fe7ebfe031de4291f6ddad75210f8563479e84178" + ], + [ + 80918, + "0xbdb4b040d7942b8bcd206e4e12632892d9cfb87c905f760f0c918dcb682d4eca" + ], + [ + 80919, + "0x1eb8ee38b4d246a0502cc56c8948ec6e17fa55b02002c9916e1a3d9867e3c852" + ], + [ + 80920, + "0x2bd9fea5702cc9c0b96fa3b34559c4a72c5ff4f1d572d9e702390a13b0b1c145" + ], + [ + 80921, + "0x12ae91b4df38f549004ab8c4482d9ed1a2841e43f649ee1a89156f9accea51ea" + ], + [ + 80922, + "0xb960711112f51fc11689061a66a9b4ed3c0bccb3a6f027bd2cceeb2d3786e780" + ], + [ + 80923, + "0x4d474df4920e9136e7564cf0e9ebd03bbcd8cb8bdb38c68b8a9c35120888bfea" + ], + [ + 80924, + "0x73d84f1013038f10bebf6606486c1bb6c6cf530a4d35f482be91df74fe15c93e" + ], + [ + 80925, + "0xf08dc8a601c46c43bb6f94edc1b0734a61e7f8efc82d915b33136fac63d5a407" + ], + [ + 80926, + "0x2b695a3c4d05e233cbe18213d50b10fb9d83914aca9a935aa6f6639b1f54c32e" + ], + [ + 80927, + "0x893a6295ba2f66f083feaa39e0570ccbd446166c2e15238d59a911f2826a47d8" + ], + [ + 80928, + "0xc2b2870e39e8344e18296012c4d846b78cd3bdc80145de6b595527c8845ac57b" + ], + [ + 80929, + "0xc32bf651e3ea72bd6808db5b4dbe4577078d418404221809ac12897082ab8bfa" + ], + [ + 80930, + "0xf1c2f5ed0a7d6b7f083c1a68f75004fbfc929ef3f6ab3b46acbba373feecfdde" + ], + [ + 80931, + "0x51638cd9a09c8ee630706d736c81ea0dde7d97234a515c6064be20d0b92b8418" + ], + [ + 80932, + "0xf3cb85e3f551c4e21d0c6c28e940c6bd965c2a20b5eb3375a9d8d32a65c58c4c" + ], + [ + 80933, + "0x65f05787c7f239dcd4dd9cac2e5c516e25f8a8b898eec284531c16a8da83ed2d" + ], + [ + 80934, + "0x58567dd84c5ce96fdfb6ea5794201199d259374e4879e76df9a6622694de95f3" + ], + [ + 80935, + "0x2031a54f25cefcf00c464f36897691a73616f777582bfcd9c03e5649d3339821" + ], + [ + 80936, + "0x0fc1b95d2579a6cb03f21083da53a989db2491f37bcc339195864baf21932823" + ], + [ + 80937, + "0xfc501a03f2666c7c7db743b6a7296692789824828c033f2254e3565de63bccec" + ], + [ + 80938, + "0xe752390723b297a6d326cc9d10f5a3219b5eede3159f9ea1c7fc7f26401b02f8" + ], + [ + 80939, + "0xfdba580d13c6964b46e7d2a66f2aeee150174facaf1cbff66016832b02d9f796" + ], + [ + 80940, + "0x2221d3b867caf0c1d4993845a60ba8389558ccd9d48bf97f51331e87ff51ce84" + ], + [ + 80941, + "0xb4ba8a208861ab1db5302570f7d589adc5f4826812b4759d625cabe2b68231a0" + ], + [ + 80942, + "0x0bc24048f15b64e0c836c387539686906198ed0b3a7adbe2aed37f41eaee5577" + ], + [ + 80943, + "0x5c3230d24716dfeb57c624bc44f7d80b65c95f4b810aa945f8e5309e45327104" + ], + [ + 80944, + "0x798ef63bf3ed33e8608b899ee956fc3a0a0b359a4466141728589ea914d42453" + ], + [ + 80945, + "0xb158a6b683a60e8d3debc803a89bd18af0f1307941039304a2e33a02f237325a" + ], + [ + 80946, + "0xd420374f3f533e0136b3fb1211357adbdde913776fbb851492ca80980e77ce2e" + ], + [ + 80947, + "0xbf7fea59019b30fec39a94c29b026285e0c2d6fe034da2555095999ca694914b" + ], + [ + 80948, + "0x3d232c12ca3a428474f8c8e21992efc614775842c365a29ef44d9b7f24caafc8" + ], + [ + 80949, + "0x87d17a79601cf126d2b5c95f1ac989f8a90f2b4e9f3b13d1d721776ee2a53c24" + ], + [ + 80950, + "0xd23290b4b63c19aafd1edefd1cfcf34829bdc6efa723df17cce5eb7612b17a0d" + ], + [ + 80951, + "0xa8d8db2850de878732f5a7b95bc08ddbb944e1e8e8884a2a2a6a98275bb7663d" + ], + [ + 80952, + "0xb2c005429b5d1f35d006693509651726e67fd19d34ac04a3473df90d8bed52ab" + ], + [ + 80953, + "0xfcdb4ee3a4fa1273ec429c49704ac1afe6c8c6f858640155276661b19a5c4ca0" + ], + [ + 80954, + "0x9b988ccb0403ebfb0c15486ed6b63ea070121d7754f88d1f14f85651c083f240" + ], + [ + 80955, + "0x39c3d647ae0dd09b1bbf1699803e3c1e76448787baf12cfced826b8d75dd4b7c" + ], + [ + 80956, + "0xb24d2ff081df7d34392e460ddf8fdb8c7cc5ee502e8cae2a279474f329fe7d05" + ], + [ + 80957, + "0xf1f1daa9a70d68fea94a204260b298285ca15469cdc1855be3ef731ae44707bb" + ], + [ + 80958, + "0xeadcfd836598031d7d1a6f5c14028edeffc5baed420b2fbd343a5e21cfdd1fd9" + ], + [ + 80959, + "0xe1391944a4bd9d2fe474c8b12de8b3d151bf3afc97addb223ebaa59592ddcd94" + ], + [ + 80960, + "0xdef9733fd939d6b260fe671867ee7cdcd802332b9afdadbba0ac24b964126d29" + ], + [ + 80961, + "0x641d150032ebcce0d0d7616b010eb29df976191edfad5c8751da8533f78f7946" + ], + [ + 80962, + "0x7480f68a44e03c35a1ece1cec7402fe9fb5b4d34d4f246c967c5178e1a7fa330" + ], + [ + 80963, + "0xefae600a6e868a4213011b106f295f79e5c86186b8382143c0a402761226a9ec" + ], + [ + 80964, + "0xc6ce05cbc32d8a1d786baa6473cc65c29892a5a47e7474f384b65c567d425748" + ], + [ + 80965, + "0xd3a3e3c3c964a77a0821984903c958079b294b195b7b13e109f26f3bafcfa092" + ], + [ + 80966, + "0xd42712815a77eaca169de5c74653b4c007159f47fb5d50d63f57cfd3f2c5868e" + ], + [ + 80967, + "0x9b309f8c13534ed43e399d45a5bcfdb38f49c5633b458f6db71d39571e9d8d41" + ], + [ + 80968, + "0x74729f1a9de838f15def5f5c183b7d3596b03afde1ed6302cabacf8a677f82a2" + ], + [ + 80969, + "0x4c07a8677c028eda23e1665416eb54d364959f3210e682978082d560ebf2bea4" + ], + [ + 80970, + "0xe3f992775c7dfbd2485ccecf516ca7ee9961c897296190c45fd7abb1431ee04b" + ], + [ + 80971, + "0xa8492953412a007f2f7696922b661a27fe2ec42e566504b3aacae2f97a24fab0" + ], + [ + 80972, + "0xe14d6f60ece95189dc8eaed8a2ea5c2da62cbdf988588222a4e99d242ea3beb8" + ], + [ + 80973, + "0x5c7ca844c8eec0cce4e4640ba6fc486c1739eb44aa50de86f55008a4d597357c" + ], + [ + 80974, + "0x2b33e852101ee35f2c61ee7b35152521eefb11364505bf6903a964c93d4001b6" + ], + [ + 80975, + "0x5516b628a8ec4328bcfdf90785fb5d96d22d60b4db638ea1b93bd837d7c739a3" + ], + [ + 80976, + "0x9959b993b66dc1285d0ecd8c95b74fe70949c50d5a97d822ce866291f94f2926" + ], + [ + 80977, + "0x04eb959364eae151a41d4a8d970795da0bf8a0ef2f64bde753cbca4ea3d745c4" + ], + [ + 80978, + "0x22907520934a83b056017bfaa72a0ce92fd8658c81dc8236f3b78f1f7e055aa6" + ], + [ + 80979, + "0x59fae978e8063799b4ab6335ca7ac097a25cbb4d60f9af04eb2f647a92ead4bc" + ], + [ + 80980, + "0x9ef7533f2e9c60f4acf9618efd1c47b147fa95bbb27caac54c1bb19540befeab" + ], + [ + 80981, + "0x9986f742ffa9d17f641293e788e53cd904b6edba3e971603470186f61f934e77" + ], + [ + 80982, + "0xbfb394920e4bd707342ddd41644a5733b7041e517d4e88863924046cd512ec19" + ], + [ + 80983, + "0xa3fef379a91df7a254bc4897cf2ef61ff25b380b2142374ff40d91d5140ed21c" + ], + [ + 80984, + "0x71626767d14a0735e3716b45b847611dc7fe04c6d7a4b22e3a6031d621b6d560" + ], + [ + 80985, + "0xb2c03912458086a6b421a8c83cee657c9276a1d38db128799a3170cba30e3266" + ], + [ + 80986, + "0x8f8000a6e85c09716e8766c9bbb9df1287d46f54625ff139d5e98142a58c9253" + ], + [ + 80987, + "0xfc88dba1647b7d43ef19a1a38780dd238228dd46beeb50ea0e00853f06c78471" + ], + [ + 80988, + "0x7e383c88765b3c499b05b60cdd4470cf9b23add0f6bd33990f207a43a51e431f" + ], + [ + 80989, + "0xa133fe8b34c9e839bcd4a75c2a5858874fa19d8bda43cd1c1af781cd2689753d" + ], + [ + 80990, + "0x466a4d175cb847a9899d98a6863900be93ecfec2f1681c8595db5a62b843478b" + ], + [ + 80991, + "0xb9b1bf393868e53ec5da954cb87fb65eef79ab9caf59cfb636f6cfd9c5c92f8e" + ], + [ + 80992, + "0x2f9eb9f04e16f420d5b98d6a28e25494207bc05affb6b914a1b98fb2244bba75" + ], + [ + 80993, + "0xc73829f98810dd017f45a0bd71ad7db3d07a1f5ff0a57c07994fa095989e6ef2" + ], + [ + 80994, + "0x9cdecf07b8f03979fdcb8f7d4d2c5ef4a13adc19edc0bb0c87a8e0beac6ff3be" + ], + [ + 80995, + "0x978c9e3471a4c37fb6aebab56d41d660720681073dfc68e5bef41d594d67ac8b" + ], + [ + 80996, + "0x739dad0614f0e852037aafbab92b4072f0f9dc9fc30c489465d67a18023f80a4" + ], + [ + 80997, + "0x1fbba34f589373e32f170a1cc3bf5c69083e8e2af5167e6ff38eab89ca2c0a72" + ], + [ + 80998, + "0xb58b192d29e3605fa4006c019698a7fbdc14639784ceeac087ae09148514e68e" + ], + [ + 80999, + "0x77f67acdb2afb4754c6e31005b27aea691cd3edce3094f87238df3f887e1bebe" + ], + [ + 81000, + "0x50e198025598c36879156ee586cdaa18bfa620f117c49baba1b2093d77b83831" + ], + [ + 81001, + "0x416e7d689843b2ddabbc9efd9b6f183b883dcb28e857977cd130f989ff7c72d6" + ], + [ + 81002, + "0xd639747fba2f48c62f7aeb91ead4b7ffa1292d55bd15ebec58879bc51b05863c" + ], + [ + 81003, + "0xc43806d1ab9e6d84be147ac3b74130fc2dca01254b4a847eb2713c3e7c1cd11c" + ], + [ + 81004, + "0xd854d522a33e98313a872fb4c47675f54eeed084d99794c5c72aab4464ee70b4" + ], + [ + 81005, + "0x752458c4e11fe4d564250bd4bb0122c1a1a4fabb124942ed9c3981887d612854" + ], + [ + 81006, + "0xe3eef7ded5487286de24e9802f5dfa1afe7c477666f2a26fca7606d3b9b8694c" + ], + [ + 81007, + "0xc525b44a5e888c3bdb0f66edc5bb4dbe9ce9116279b01fa510037760319ca15f" + ], + [ + 81008, + "0x9552445c57b6e2cd01d95bf4013bb913ce1611e5d707f71e528ad7522e830806" + ], + [ + 81009, + "0xd0eba42e3dd5cb37d42d6891eb47e5e2bdaeccbf9ddabeba06a09cd2b4b15f55" + ], + [ + 81010, + "0xcc6fd1b1758929e8a94d3bfc1b4c5444c55023f8939d4b4fb539b523219b00fd" + ], + [ + 81011, + "0x826ee09096704d3798aa4fccfc02cdedc59bb7472332b624d3716fa3aa14e612" + ], + [ + 81012, + "0x3f82182c620ce4d5ac0a776b0ec52b590f7d30c3bbea039d9a6ec3c6ec8a4aeb" + ], + [ + 81013, + "0xbdc7e4e07b4fc9245d8f5e7d7b207af203ce271b96912963aae2012de4d0dfd4" + ], + [ + 81014, + "0x797fd61f65fdc95f190d0dec1ad3eee5c98211c4529c83d013236838eda13d67" + ], + [ + 81015, + "0x7928f9107a62bd4d597894701cf4bca71365052e3560c23af7fdfa3374105b45" + ], + [ + 81016, + "0x78dc4dfd56d144d629278047ce28d8649e02b55768a137cf98260e40e6c35989" + ], + [ + 81017, + "0x3cd444ed1bd6f3012197839db12d566729ca5fb8c133af4353afbd756104c03c" + ], + [ + 81018, + "0x26350e310daef399afac9ac19b106d5f160d8dc6921fd289b196b4e979742bef" + ], + [ + 81019, + "0x99e1715d9ba9a3d125e77357595bd918766a4d46c20d9db1dfa9ebbd002a7bfa" + ], + [ + 81020, + "0xdf16ba27999f8dacfde10d403cf9982fd3b7545793e678219f09818e04ee64d6" + ], + [ + 81021, + "0xc4690b3c4b89b7282988cd2f71c9e43d3c17d081f0c735d04fbd01247a74bd6f" + ], + [ + 81022, + "0x94a1951db605f9a7dff70e2fd919e154f1291d31b4baa6f9e88d8f904c74fd54" + ], + [ + 81023, + "0x9846feb46081eafd35bd5f134cdf98f7a2a270b52a36779de7fb132250823a6b" + ], + [ + 81024, + "0x3716fe1bc70f07f53358e7936ff4b652a035eabe199265f99809dc4aac49d002" + ], + [ + 81025, + "0x841c89865fcece06f48c90ca283a1dbafb49e3fc422ff4cc5d3ec7fa68ac0f14" + ], + [ + 81026, + "0xcf2be45d7a88f2ffa7fa45366404b8b0af131758b7706f099ee8daa6b5d17b94" + ], + [ + 81027, + "0xf3c7dd48103bf10e13bea5800814475a543f4bf53958e1ad464b112d0c763d00" + ], + [ + 81028, + "0x6ebafde3e8fe7067903074cbeb854b9957bbc45859143f618ada972fccdc2c0d" + ], + [ + 81029, + "0x05146b0d5efdf34a927593169ecce2d2d112a3faa3b79cb2df1a0a2e6a4bd7c0" + ], + [ + 81030, + "0x33def97a4f3ca92c7af02bd19bb9703c516f6df8917928734953ec9fddae68ce" + ], + [ + 81031, + "0x18bbaeb286265fc8edb34abbfd3fade65196f09125dae4d31347401eb635b63f" + ], + [ + 81032, + "0x56aacb0d332a22fabb3bf8ac90b030fb82337b3b22b5be334d9a4f633847480c" + ], + [ + 81033, + "0x8d5d77464c966282353e2019598e0f1d77b5afb938475275f4a3b0d0c29d9c5e" + ], + [ + 81034, + "0xc8ccc7baba52e4fa56d488c78abf451919aeba4a029068fd364059633446c80d" + ], + [ + 81035, + "0xc295e160d3378e12e6b96de0eac43d903aa656f9255e58a3b0222bc229b18bbb" + ], + [ + 81036, + "0xd4cacd3eb6e8eebf312225c23284c15e347e73539816fe91954b482a9acb1753" + ], + [ + 81037, + "0x096f15eb2cfa8a6811e08ee533e3a9bfaf3447225c214157d33f27d532423598" + ], + [ + 81038, + "0xd2464a278a262d447f8d67feb035e8b9eaaaa849eb87fb330e1ac868e6b1d063" + ], + [ + 81039, + "0xfea1ffb0177d362f9e041268c395858091da08fb55de5c922493bf0f3dd19476" + ], + [ + 81040, + "0x5ab2d5c6f5789372abc97b4792b2c4e3476c70bd538764ba74d684ed102fdd0e" + ], + [ + 81041, + "0x2f5be38f279790c0e69e554fcaa28f82bd17ef645338539416db0c9252040f91" + ], + [ + 81042, + "0x844ebc844ba1c0c301a6a9189779e17e557f824d85d01247b4a0b1af49f50d56" + ], + [ + 81043, + "0x944855a63daa3b755a70f2ba5e6a956b037eeb210dd9e80cef73b7c814645adc" + ], + [ + 81044, + "0xf046f77959a6aa199f3d999b2c16dbeb1e9038c308cb1dcc9bba26141e0e5f35" + ], + [ + 81045, + "0x4e420f14ba78477c090eef99a098b62dfdadc3a61f7c0a5499797380763693db" + ], + [ + 81046, + "0xae10dd9883739a2cd9e95cafa4f9cc7b2e38e477207ec49fb3fe64e3fcf21613" + ], + [ + 81047, + "0xa9e686e36215cf1053ac16cb445e537ef8a633aac308ad755716f6374c3c1238" + ], + [ + 81048, + "0xa7ab60b451ae4a7fa639faa0c8b1f2ad7bb6bcb853242797f4aca608b2e69193" + ], + [ + 81049, + "0xdba699a61f7e4b3b8356a7494afbb73237f00385d7ad279a77002bfd451b2643" + ] + ], + "rewards": [ + [ + "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "0x3269259a82c2a000" + ], + [ + "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "0x3269259a82c2a000" + ] + ], + "proving_pool_credit": "0x193492cd41615000", + "payouts": [], + "blocks": [ + { + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "blue": true, + "txs": [] + }, + { + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "blue": true, + "txs": [] + } + ], + "pre_state": [ + { + "address": "0x0000000000000000000000000000000000000210", + "nonce": 1, + "balance": "0x0", + "code": "0x608060405234801561000f575f5ffd5b506004361061003f575f3560e01c8063aa67735414610043578063dea5c2e014610058578063fe7e05d51461009f575b5f5ffd5b6100566100513660046101e0565b6100ca565b005b610083610066366004610211565b6001600160a01b039081165f908152600160205260409020541690565b6040516001600160a01b03909116815260200160405180910390f35b6100836100ad366004610211565b6001600160a01b039081165f908152602081905260409020541690565b336001600160a01b03831614806100f957506001600160a01b038281165f908152600160205260409020541633145b6101635760405162461bcd60e51b815260206004820152603160248201527f446576656c6f70657252656769737472793a206e6f7420746865206163636f75604482015270373a1037b91034ba399031b932b0ba37b960791b606482015260840160405180910390fd5b6001600160a01b038281165f818152602081815260409182902080546001600160a01b031916948616948517905590513381527fa47563c41dab010f91a8ef9dc7ac2bcdfa0ef2af697e575048e71e6eec60dda3910160405180910390a35050565b80356001600160a01b03811681146101db575f5ffd5b919050565b5f5f604083850312156101f1575f5ffd5b6101fa836101c5565b9150610208602084016101c5565b90509250929050565b5f60208284031215610221575f5ffd5b61022a826101c5565b939250505056fea2646970667358221220cbf48f5aa6f911f83b3c2b09adf8c418f5da6fc530c3d5db9d4ba0be24c4bf5864736f6c63430008250033", + "storage": [] + }, + { + "address": "0x0000000000000000000000000000000000000220", + "nonce": 0, + "balance": "0x13cba1a977d8aeedd400", + "code": "0x", + "storage": [] + }, + { + "address": "0x0f002c928c363c7041cafe788f099cf7d35452a1", + "nonce": 243, + "balance": "0x1bc36179476e9a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x18524811fa2e76dd0770d29308b51bcbdc0499f9", + "nonce": 242, + "balance": "0x1bb8aa483e652c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x1aca71f7872aebc85b4d4a9ad47a09abcc624d65", + "nonce": 0, + "balance": "0x2860cb14365c048800", + "code": "0x", + "storage": [] + }, + { + "address": "0x233d639f53225ea5012dc01ceb0a5a30891021cd", + "nonce": 264, + "balance": "0x1b8eea59c54b4200", + "code": "0x", + "storage": [] + }, + { + "address": "0x27fd8475c3db352fdddeaf29823e00c7d33ca39a", + "nonce": 248, + "balance": "0x1bccb45d83401000", + "code": "0x", + "storage": [] + }, + { + "address": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "nonce": 0, + "balance": "0x1e063716f0755a0c800", + "code": "0x", + "storage": [] + }, + { + "address": "0x46681948060140945958068b373f94581e049d4f", + "nonce": 240, + "balance": "0x1baadca68fcb1a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4cb00bc539538d2ee7172474f2b06f021944ece4", + "nonce": 258, + "balance": "0x1b3e54c72c629c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4d7c0d6f3ad5466691649e98ebb59cd5b5194687", + "nonce": 0, + "balance": "0x1b598fbe8e08549c6000", + "code": "0x", + "storage": [] + }, + { + "address": "0x5a5e606bda0fe1b2b4298aadd55cc6ec57ff698a", + "nonce": 0, + "balance": "0x648cba8883ddb562c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x6613ca66c14338db771d01c238ae140c71310535", + "nonce": 243, + "balance": "0x1b98d3921da34800", + "code": "0x", + "storage": [] + }, + { + "address": "0x66577de387feeec96c10f7f8748f13912589c78c", + "nonce": 240, + "balance": "0x1bc37e3123b20a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x7e7b1db26094aa913933f65127b328c1861ff5b0", + "nonce": 256, + "balance": "0x1b89d2467bcd2600", + "code": "0x", + "storage": [] + }, + { + "address": "0x90acb15171deb958d3191d5c02d807afc367f0e8", + "nonce": 240, + "balance": "0x1b9e3c5e84fb6c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x9f296bf64eb7051912c9c8112da3e3ea6cf546e1", + "nonce": 266, + "balance": "0x1b69b11e00d88c00", + "code": "0x", + "storage": [] + }, + { + "address": "0xaa19223d63c82bf9be5543ec96e5504834d17bc8", + "nonce": 246, + "balance": "0x1b9abb1eb524dc00", + "code": "0x", + "storage": [] + }, + { + "address": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "nonce": 0, + "balance": "0x558d030bd9e0f7c000", + "code": "0x", + "storage": [] + }, + { + "address": "0xc2faa4a2866422a865c9b3b3cb5988bf4d07f0b4", + "nonce": 0, + "balance": "0xc05b9fc73002b32800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc45d23f49451faad9afadf877feead626ea8d083", + "nonce": 0, + "balance": "0x9ee8f806428adc9800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc8621de5921418701619ba715044575073d2ca62", + "nonce": 243, + "balance": "0x1bbb8989a888a800", + "code": "0x", + "storage": [] + }, + { + "address": "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "nonce": 0, + "balance": "0x1f29d0e7a3eaed38a400", + "code": "0x", + "storage": [] + }, + { + "address": "0xd958300657e7931b06fc6f60dc4bfe7d8e8ea7b8", + "nonce": 252, + "balance": "0x1b79a326be219400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdadb11b01f6004eba6da3846a0e1bf6f4f71fb82", + "nonce": 244, + "balance": "0x1b900b01183e5600", + "code": "0x", + "storage": [] + }, + { + "address": "0xdd442fcbb964a3afdc90d49b408e8dd296fa86e8", + "nonce": 0, + "balance": "0xafd0c09109be7410400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdfaea67368f3e3753397d878f97efe6aa8020c2e", + "nonce": 16, + "balance": "0x4f20fde2ac0f98ec00", + "code": "0x", + "storage": [] + }, + { + "address": "0xfd4fca79b266a73a7defa776b04c7e15bdc70e86", + "nonce": 238, + "balance": "0x1bb5cf9909338400", + "code": "0x", + "storage": [] + } + ], + "fees": { + "base": { + "pgas": { + "version": 0, + "cycles_per_pgas": 1000, + "intrinsic_pgas_per_tx": 200, + "modexp_base": 1000, + "modexp_per_byte_numer": 10, + "modexp_per_byte_denom": 1 + }, + "block_proving_gas_limit": 30000000, + "shard_proving_gas_budget": 7500000, + "min_execution_base_fee_wei": 1000000000, + "min_proving_base_fee_wei": 1000000000, + "initial_execution_base_fee_wei": 1000000000, + "initial_proving_base_fee_wei": 1000000000, + "base_fee_change_denominator": 8 + }, + "v1_activation_daa": 18446744073709551615 + } + }, + "plan": { + "shard_budget": 7500000, + "consensus": true, + "shards": [ + { + "index": 0, + "tx_start": 0, + "tx_end": 0, + "over_budget": false, + "pre_root": "0x57a18ca3df176250b8ec70274aa0d48e2b809f9751f2a0b17f8c67b71d00c5d1", + "post_root": "0xb6dc9eb1ce5069a634647f05619d7b1bd477588aab401e5e814db9c3628e8831", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "link_in": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "link_out": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "witness": [ + 2, + 0, + 2, + 16, + 14132 + ] + } + ] + }, + "expected": { + "pre_state_root": "0x57a18ca3df176250b8ec70274aa0d48e2b809f9751f2a0b17f8c67b71d00c5d1", + "post_state_root": "0xb6dc9eb1ce5069a634647f05619d7b1bd477588aab401e5e814db9c3628e8831", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "tx_commitment": "0x0000000000000000000000000000000000000000000000000000000000000000", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "node_state_root": "0xb6dc9eb1ce5069a634647f05619d7b1bd477588aab401e5e814db9c3628e8831" + } +} \ No newline at end of file diff --git a/proving/fixtures/chain/block-81051.json b/proving/fixtures/chain/block-81051.json new file mode 100644 index 000000000..f4d6b9e18 --- /dev/null +++ b/proving/fixtures/chain/block-81051.json @@ -0,0 +1,1316 @@ +{ + "format": "igneum-prove-fixture-v1", + "source": "live devnet export from node 1 at tip 81076, 5 October 2026 20:06 BST, proving v1 chain fixtures", + "block": { + "chain_id": 4463, + "env": { + "number": 81051, + "hash": "0x9a3442c671ff6750549ac81e1a48d7ba70e8248b0f1122bf37a3767666cac4e6", + "parent_hash": "0x42515b83488525c1c8e72417df7361aa61a62e3f78892446b06d6d6bccee6ada", + "timestamp": 1791227132, + "miner": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "prevrandao": "0xe0cc2e4fadff0587f62bc19c9791534f94d2dabab1872b0abd8e27b6d9592cd2", + "base_fee_exec": 1000000000, + "base_fee_proving": 1000000000, + "daa_score": 0 + }, + "block_hashes": [ + [ + 80795, + "0x4bf9d010c0798a5340983b27eda6a9876e6fefb925104fdc46291a5789f1f70a" + ], + [ + 80796, + "0xb09af3831aa1d8f4945d50fa16a602848a2da22b14dbc1f7af2844cf19e90040" + ], + [ + 80797, + "0xe099536cfd64654690afc152b97866b32da6bf81f959a6c61f9e43bfd2ca2d9a" + ], + [ + 80798, + "0xd6af2cf6184f15ecca77ee9e7e8d2ca718d88f9cea39b475713ceb0e151ddc36" + ], + [ + 80799, + "0x9c76f9f6036c7035fd52a7a4e3fca99451715e729fcc9fc424487faee9a6134d" + ], + [ + 80800, + "0x9a046cd99eb89c551dd0c698f1969069c02ae36cde4140ff4b864d624fd9dff7" + ], + [ + 80801, + "0xc9bffa329d649c6540518cf249e25e845743dc01fcbcfdfb8c798f1455efeae6" + ], + [ + 80802, + "0xb124f2de3c17a985f10c0b58d2283631bcc12524de8d79b0955aecfd2d7f78aa" + ], + [ + 80803, + "0x97b21d6cd3f1b321d2c9729062c305a7955b97646d76910714f899c75d275568" + ], + [ + 80804, + "0xaa3f0e114e156f706e11a701fb940bf54e89b09d83953be7f6733478c7a33093" + ], + [ + 80805, + "0xe5a5868b02ec30304079652a8d5b8af4a9da0a5a676615c1a10d833feb4eb0f7" + ], + [ + 80806, + "0xd97c5c4288f56955e1a3e134eb63250e1becb907b0fdde9ce1383d8d2c8a3939" + ], + [ + 80807, + "0x8f7bf1128e5d03607cca246b86e63cfb4bfe73e74395c8660d027172376b5c36" + ], + [ + 80808, + "0x5d5fe4c3bee7d758986da5a433cfe5221be392419ac22e5acddd701cd12c8142" + ], + [ + 80809, + "0x9b0d0a871ce8eaab11a185a4344f9f48a58b7111f94a8f5221d557722cc5ca5a" + ], + [ + 80810, + "0x58cad4ab776ba7eca94e813323e28a6467a7fcf39d5aa2069e9dfef6f8058e5c" + ], + [ + 80811, + "0x706788247d6c05475eb3bdbc6f181e76429b57a97f168004832fe5ba4b20c7b7" + ], + [ + 80812, + "0xdf26fb97f999b1016865b3fee539e72e3f66d6d228b2aec7e2f59ab93e5a9ee6" + ], + [ + 80813, + "0x8e612c07fdb693a44353a08899e8beae8c93fb1b85845ad965f96e7d2f77fed4" + ], + [ + 80814, + "0x806f4c29ea28af1a3ce0c3ba775c776006ba8cf4b308ee617dc207dc3f740b29" + ], + [ + 80815, + "0xebd8e602049ddc2f8d11d5edc8ee837eefac77418e73e9bad97f6f1c7a85e8aa" + ], + [ + 80816, + "0xe2b069b77ac05cb0a95ae419966feeb20660e4ff3b9a68d044a5721e1525dd86" + ], + [ + 80817, + "0x1ca5b79e26dc425a94c3a6bfd339516caa0a7f8e5e2a62389bfa466186bca495" + ], + [ + 80818, + "0x7cf382f956597cbd36917036dd15b774500f97714e50587a35ec9617662052d6" + ], + [ + 80819, + "0xe70494836c5b5c68b6a167c868db166f9630140499f96e6660d3e3928cad6a14" + ], + [ + 80820, + "0x15f73eefc39e00f9d09390764f9ea0069908062c563686e73b49ec3919c24be3" + ], + [ + 80821, + "0x7f2bd2cf68fc14e6a5debaac39f368534657c41cc587402fac1156d24998e55b" + ], + [ + 80822, + "0x2641ec1c3397b06ab5855dcf6ec2e23de301f5831015a32ffa9f29e18cc2f084" + ], + [ + 80823, + "0xf643e963f5a62a49e935b4ef51aab9bf2689b0b9ec483924d7f42934c8ce5e07" + ], + [ + 80824, + "0x867ee0219bd30fc3117c79295cb1075ca21281fdd285467ae128aece1c1dd9a7" + ], + [ + 80825, + "0xab6bb19259a039d30c3022c9da3494af8c9a541dce30895df083252894c433df" + ], + [ + 80826, + "0x155511de8d34bef351ba0ae7af5bccf515e6e490135130aad5ce2eb02e103d75" + ], + [ + 80827, + "0x4e297a53901da9b364f4dce686f016f4039658ea8832d591dcb21aa44acd197b" + ], + [ + 80828, + "0x7d07b8bc9b06b20e7b56f96fff9051590caac01159d39b3f95bd592a4627daf4" + ], + [ + 80829, + "0x8392a1875388baec163e64854d6ee7664a31017402ee8e269ddc8aa42ca06647" + ], + [ + 80830, + "0x5ab4e5c586d2fb4272941bf26c15288cc129ca1092d3aaab82c478edd1942e29" + ], + [ + 80831, + "0xb0100f0a559daa435128cae74f47f9353ca2a93c2bdac8dd9d12a3130ce6a818" + ], + [ + 80832, + "0x326d1388d0911871c5087925500c4ee5091624331d07659b9d2fa09e83f5884a" + ], + [ + 80833, + "0xc01c186cfcebb8d53953589ce595afef213d46e481b8b5dee7e0a71abd773f6a" + ], + [ + 80834, + "0xe845934e79eaeb8a8ee0b9396b3a13f13c8ac45a0d4e81017c73042789013c38" + ], + [ + 80835, + "0xa45f8f03249624a2d35b54eb83f4edf54e8d1acd061fee3965d9164a132c72aa" + ], + [ + 80836, + "0xe72a59a64af9aa223b5341819a5872f8626908b5c1857d0617706f6999632e01" + ], + [ + 80837, + "0x94063f426663347bb4a3e79f4dedeb76fbce6fd69117a2c74ad1fb4714b856dd" + ], + [ + 80838, + "0x157ab5f9e7fbf98b48af049adcfb4a8e071253ec95b991d031a91668ff3b8d57" + ], + [ + 80839, + "0x5bff632aa54b1a71a7d9a2a57d9c42ef223eb6efe86c72f9aa8dfd67ab7b218a" + ], + [ + 80840, + "0x7f5d1952ad8ad27dd17bed3354395b646468985bae0341b4048376a425e43076" + ], + [ + 80841, + "0x8e69df81fd1b0f186813304260dda1d0c9dad5792126e69d021e679b837b01a5" + ], + [ + 80842, + "0x6733c5472e38830dfb273a0e0802924c1b60a2cc916d6efaf82785e1bf960724" + ], + [ + 80843, + "0x0838513dc8264efffb2c11f8c00ebbc9c57aeb67dd7be93f18e1c613872bc929" + ], + [ + 80844, + "0x8c2f26b17db95bb20983ac41df7a5fca3e3826d2e57831eba13bc3175daeee34" + ], + [ + 80845, + "0xe3edcab823e03598c66e16c5743f7a0917735c975ac946db4cb3ad862db40a21" + ], + [ + 80846, + "0xd075510e75159d941946e666e2c75604a3f3c959523cbe2c0f025427ae85d4a2" + ], + [ + 80847, + "0x57b4b6c0a6ca2532515653445bc872dd44e621f9774f3dd4e2000bf82a0f510a" + ], + [ + 80848, + "0x38a85c99bc3e489a4431ba9a125e74f4d807c630a239419287e039f05cc91ead" + ], + [ + 80849, + "0x3c95d08c99482d57bbcb9fd333cae4010628fc8b1c96b3ad4a68d8f3dd21695f" + ], + [ + 80850, + "0xe767527430f19da2c2b21bc648e33ed2b241f5b7c9413649d9c5a555491a9253" + ], + [ + 80851, + "0xe9941baa9f3378854aaff5563340d130700f0753197c7373229542c56ac91a8f" + ], + [ + 80852, + "0xb92e88ec7cc8c340959d8d0cd5434eec5c1033b389b953b02cda103d53292df0" + ], + [ + 80853, + "0x39f01f35ba2ea7b45e3eec1681c44bd85b066a6a8521a10fe6cec5d4d98aace8" + ], + [ + 80854, + "0x36456a5e36a9a66488656bc0a92705632326e59ca40bd09f6d5f99ccc976972a" + ], + [ + 80855, + "0x12e52067750279b8bc34b03cd21bdb2809ea18a3aa68739623780aaf1150129b" + ], + [ + 80856, + "0x64951c693d7f246647fac0df504f35ff9812416f966afaf4fecf7ca92c0e0b24" + ], + [ + 80857, + "0x3918b797fe9d8cd6c34a0b315aaef550de15f4d30dce932f3be8cdfdc27343a2" + ], + [ + 80858, + "0x86ec8713a68034acb9f34b015a48644c7562766c1760927a299e517ae8ef6f8a" + ], + [ + 80859, + "0xfe75a4c99fc17131b7aae61dd0a588b7d7019e6c662aab8c7d6a8c56370fe07f" + ], + [ + 80860, + "0xf4afa8f14a0f814639c88dea304a8ccf76450ee56f4fbb89caf29dbeef80399b" + ], + [ + 80861, + "0xd4c72bc192643e20658ca54b730ef865d1c29ebfb6d1f78263c796a23b6965b4" + ], + [ + 80862, + "0x7a73b7302ab47b7acc1c5359620df9026fdc08d78c4173bb4e45e73a66651767" + ], + [ + 80863, + "0x9033ce568146d2c69139d3674fa7e18cb8319abae5e6cd24f924f750d2ca3446" + ], + [ + 80864, + "0xe4d8da84ba631f233ea7dd1336a264f499083b08290c82b717f79bfb7ffa09a6" + ], + [ + 80865, + "0xc53ec9d6b11d592921f5cf850584ebb562119ea968a5fa7bb00af381d6eaef46" + ], + [ + 80866, + "0xedcbb7d2f9626af53341464fbfc54e838e6417b7ffcfb0c2f9758641c98d5f9e" + ], + [ + 80867, + "0xbe5f726816d43ca5bbcb70f13901e60ef9040debdfaf2a73a0692c3cbef578ec" + ], + [ + 80868, + "0xb7cf72378b14c086afb4094b9b0121fcf7c6929bf2c2134b5189862644e2e506" + ], + [ + 80869, + "0xa29e21de7e54defda356f088dea38a06cf1a1a75bf93c14e30afdba213df7ef3" + ], + [ + 80870, + "0x504d3f83a27c37c5a15cdbbdfc03285f53071ad01d82bf9be2d6f862548a9a8e" + ], + [ + 80871, + "0xe02107fefe75bc8d1abb0443bff2e7a658e3a9b2fccd5157c1a7c7ccc995fe17" + ], + [ + 80872, + "0xdb92e1b449e7670ea719a6d68f496965620656ff5d25a0521f1873eb8d7ae1ab" + ], + [ + 80873, + "0xf12689607f86ce8f762df5d67054a00b8b580300511a221ff5addcf732393988" + ], + [ + 80874, + "0x0e0fe961c78fe52f44f6742795bae3e96713f4618e4a95623045056eb9e046f6" + ], + [ + 80875, + "0x7fd5bb64e800686ae28c3b55613d27f181dbd91e35d342d33d7b3e5d4df17d00" + ], + [ + 80876, + "0x83c2feae4eacf2ada1c9a630a97614a6296df6b89bfb5f954c3e306b16f0ab2e" + ], + [ + 80877, + "0xd1f3af4d6464c5bc9416d73a291bcf3e1df4fa1a396bb3e9dff92dcd451f75d9" + ], + [ + 80878, + "0xa4cd4b49e1d475a5561a91a0de592ba0c00f63b39a61eab150c6486d1bf2b916" + ], + [ + 80879, + "0x714904f162c3b9013a9defd79d4552b58c5ca8747cc94187079f2ebefee2f0d9" + ], + [ + 80880, + "0xb2263d12431e74591e2843047d79933820d2dd352d5b5c977698b328cd59f46d" + ], + [ + 80881, + "0x67977d72a573f121d9a818dd035959cd502e5bec07a4a5c30a28fc66d4fdf964" + ], + [ + 80882, + "0xf1d5f1b271cd23a79f1ab32655418e4b664663c4a77ebaf0a213fc2456e59c5a" + ], + [ + 80883, + "0x91a9f94d339407747a77d79a9c2be608744d4ecc8b84699d0120814f55a5060b" + ], + [ + 80884, + "0x8f709bf4582d3a912e6006f6c417459bbbf182af2b500b4376cd8dff78201c25" + ], + [ + 80885, + "0xbc66848a66fc038200186f09665540b5905a350a3149bfdddf9575eabb2f0a98" + ], + [ + 80886, + "0x8827e43b9a02e7424513e6d0fb6db7e93cc5555fac9b832584570784708a8f90" + ], + [ + 80887, + "0x7c76b6619ea66b8e7fa2791ad1a3dcafe863388d14db50dd022f63748e2b6ccb" + ], + [ + 80888, + "0xace6e76c69db6530277c9749ff3e77a64c6e7123790fdf5f248cc67b761e10bd" + ], + [ + 80889, + "0x8f96a0d2f9524f0b2bf9f520f35516969fba795e005f003dac5721cc37c6b5ce" + ], + [ + 80890, + "0x68a635c0d7ba10b3fff9108deecc4e9af70d760d688c5614a000503b1f2f006c" + ], + [ + 80891, + "0x47dc91daaf1758699e90f8906c61bcb6dd020b3726a1e0d15a7fe5177bfab696" + ], + [ + 80892, + "0x79f1ff991dd34c0384b9ddb64a7f3791116891447f2c960de5c6d80f9a705b5f" + ], + [ + 80893, + "0xf41832c4af88cdd31eec42b233ccfae51a150cd41baf59e9dc6b64dba2b63c7a" + ], + [ + 80894, + "0xfaf2c58e9ba01047e4cbe2c147a4614d46ad39c1d64cbb43d1de60f9c75286b7" + ], + [ + 80895, + "0x65a671bcbc9b6f905811ce62594feda18a12b8c6ca1c72081c1dde3f6227d1a7" + ], + [ + 80896, + "0x815d16976077a9b599fd5d6a82eda30a0bd1d91c66718528aaea9d4dcd12274b" + ], + [ + 80897, + "0x710c0c1d4d44ed7925343b10c7ace3216741b31a91506fa0f1ad80c96c049b14" + ], + [ + 80898, + "0x5fff52604275c206678c800d95cd0569325ad67413b50ac15252c3440856cc78" + ], + [ + 80899, + "0xe67fe87df50a63d348d300de16353e560c1b86f1a79a0bd09e2a7771ae7e92ba" + ], + [ + 80900, + "0x5d854ceaef9de086a2361cfa1f829f843aac00181fa5ae27e69171c198853895" + ], + [ + 80901, + "0xca45572516a2d4a948cfe518d0bb5378a4d23b800522f515e59176c31c069160" + ], + [ + 80902, + "0xa724d3e05f16d3d8c97296f59802c7f9ebba4284808126d955d6fc96dc1c4729" + ], + [ + 80903, + "0x552061794bddc5185033e936ac8ad07d3a315c7a3d876989b2d61328ad0a3128" + ], + [ + 80904, + "0x23c1ee3ff43ae0d42b9b3f0d2c6c2fef6c8ab8fdec9772540b563cfebbb77a2c" + ], + [ + 80905, + "0xc7e33046cb4819a17a21a00e9c0efa55765a974d8a2364f4bed9a16034ec2828" + ], + [ + 80906, + "0xa7b5712f7d22449a6f6a8a55a33c2f76cb54be86faa2f0d892c8bf6aa9fe59d9" + ], + [ + 80907, + "0xbd8696f54d266714d694b41d3b6466ef242e999938b9effd2de035c3f2c5677b" + ], + [ + 80908, + "0x7ca324c51788d4cc173bd3a013702f80e68524b3b005299be9012621dc8e82c4" + ], + [ + 80909, + "0xaea0e3a560abbe0de857592d37bfd43bd986d1cd461f43139d951d123140d495" + ], + [ + 80910, + "0x640f0ed5ec76969f53d90d5e502f6fce0ec65ab28a728e92a81dd27c466dad09" + ], + [ + 80911, + "0xf4cb4beab37872930f84e37b38c15a0a5e5f5d95927c91284ae880337abcfd9a" + ], + [ + 80912, + "0x65121e870e1668a0a55c9a50d31166294cb8e57bc007897640a7414f0866ec3b" + ], + [ + 80913, + "0xb3ea0a195e00db811104db7362b555fd1ab8401814ee841ca7f4d3344843708c" + ], + [ + 80914, + "0xf92735fd64122d7f6f6df6f20f83fc0793675c81e25e54dc64b3ca027d8ef9a4" + ], + [ + 80915, + "0x78e07fd2fcedb51175785be4657b8d2a554c16c1d2cc396ca3738dbf6f9373f7" + ], + [ + 80916, + "0x0ce35f9c3a736ff2937e6b7addcb632855e347ab3c5c08b74d2dca5f82265426" + ], + [ + 80917, + "0x007b1e317c0d9e7565260e1fe7ebfe031de4291f6ddad75210f8563479e84178" + ], + [ + 80918, + "0xbdb4b040d7942b8bcd206e4e12632892d9cfb87c905f760f0c918dcb682d4eca" + ], + [ + 80919, + "0x1eb8ee38b4d246a0502cc56c8948ec6e17fa55b02002c9916e1a3d9867e3c852" + ], + [ + 80920, + "0x2bd9fea5702cc9c0b96fa3b34559c4a72c5ff4f1d572d9e702390a13b0b1c145" + ], + [ + 80921, + "0x12ae91b4df38f549004ab8c4482d9ed1a2841e43f649ee1a89156f9accea51ea" + ], + [ + 80922, + "0xb960711112f51fc11689061a66a9b4ed3c0bccb3a6f027bd2cceeb2d3786e780" + ], + [ + 80923, + "0x4d474df4920e9136e7564cf0e9ebd03bbcd8cb8bdb38c68b8a9c35120888bfea" + ], + [ + 80924, + "0x73d84f1013038f10bebf6606486c1bb6c6cf530a4d35f482be91df74fe15c93e" + ], + [ + 80925, + "0xf08dc8a601c46c43bb6f94edc1b0734a61e7f8efc82d915b33136fac63d5a407" + ], + [ + 80926, + "0x2b695a3c4d05e233cbe18213d50b10fb9d83914aca9a935aa6f6639b1f54c32e" + ], + [ + 80927, + "0x893a6295ba2f66f083feaa39e0570ccbd446166c2e15238d59a911f2826a47d8" + ], + [ + 80928, + "0xc2b2870e39e8344e18296012c4d846b78cd3bdc80145de6b595527c8845ac57b" + ], + [ + 80929, + "0xc32bf651e3ea72bd6808db5b4dbe4577078d418404221809ac12897082ab8bfa" + ], + [ + 80930, + "0xf1c2f5ed0a7d6b7f083c1a68f75004fbfc929ef3f6ab3b46acbba373feecfdde" + ], + [ + 80931, + "0x51638cd9a09c8ee630706d736c81ea0dde7d97234a515c6064be20d0b92b8418" + ], + [ + 80932, + "0xf3cb85e3f551c4e21d0c6c28e940c6bd965c2a20b5eb3375a9d8d32a65c58c4c" + ], + [ + 80933, + "0x65f05787c7f239dcd4dd9cac2e5c516e25f8a8b898eec284531c16a8da83ed2d" + ], + [ + 80934, + "0x58567dd84c5ce96fdfb6ea5794201199d259374e4879e76df9a6622694de95f3" + ], + [ + 80935, + "0x2031a54f25cefcf00c464f36897691a73616f777582bfcd9c03e5649d3339821" + ], + [ + 80936, + "0x0fc1b95d2579a6cb03f21083da53a989db2491f37bcc339195864baf21932823" + ], + [ + 80937, + "0xfc501a03f2666c7c7db743b6a7296692789824828c033f2254e3565de63bccec" + ], + [ + 80938, + "0xe752390723b297a6d326cc9d10f5a3219b5eede3159f9ea1c7fc7f26401b02f8" + ], + [ + 80939, + "0xfdba580d13c6964b46e7d2a66f2aeee150174facaf1cbff66016832b02d9f796" + ], + [ + 80940, + "0x2221d3b867caf0c1d4993845a60ba8389558ccd9d48bf97f51331e87ff51ce84" + ], + [ + 80941, + "0xb4ba8a208861ab1db5302570f7d589adc5f4826812b4759d625cabe2b68231a0" + ], + [ + 80942, + "0x0bc24048f15b64e0c836c387539686906198ed0b3a7adbe2aed37f41eaee5577" + ], + [ + 80943, + "0x5c3230d24716dfeb57c624bc44f7d80b65c95f4b810aa945f8e5309e45327104" + ], + [ + 80944, + "0x798ef63bf3ed33e8608b899ee956fc3a0a0b359a4466141728589ea914d42453" + ], + [ + 80945, + "0xb158a6b683a60e8d3debc803a89bd18af0f1307941039304a2e33a02f237325a" + ], + [ + 80946, + "0xd420374f3f533e0136b3fb1211357adbdde913776fbb851492ca80980e77ce2e" + ], + [ + 80947, + "0xbf7fea59019b30fec39a94c29b026285e0c2d6fe034da2555095999ca694914b" + ], + [ + 80948, + "0x3d232c12ca3a428474f8c8e21992efc614775842c365a29ef44d9b7f24caafc8" + ], + [ + 80949, + "0x87d17a79601cf126d2b5c95f1ac989f8a90f2b4e9f3b13d1d721776ee2a53c24" + ], + [ + 80950, + "0xd23290b4b63c19aafd1edefd1cfcf34829bdc6efa723df17cce5eb7612b17a0d" + ], + [ + 80951, + "0xa8d8db2850de878732f5a7b95bc08ddbb944e1e8e8884a2a2a6a98275bb7663d" + ], + [ + 80952, + "0xb2c005429b5d1f35d006693509651726e67fd19d34ac04a3473df90d8bed52ab" + ], + [ + 80953, + "0xfcdb4ee3a4fa1273ec429c49704ac1afe6c8c6f858640155276661b19a5c4ca0" + ], + [ + 80954, + "0x9b988ccb0403ebfb0c15486ed6b63ea070121d7754f88d1f14f85651c083f240" + ], + [ + 80955, + "0x39c3d647ae0dd09b1bbf1699803e3c1e76448787baf12cfced826b8d75dd4b7c" + ], + [ + 80956, + "0xb24d2ff081df7d34392e460ddf8fdb8c7cc5ee502e8cae2a279474f329fe7d05" + ], + [ + 80957, + "0xf1f1daa9a70d68fea94a204260b298285ca15469cdc1855be3ef731ae44707bb" + ], + [ + 80958, + "0xeadcfd836598031d7d1a6f5c14028edeffc5baed420b2fbd343a5e21cfdd1fd9" + ], + [ + 80959, + "0xe1391944a4bd9d2fe474c8b12de8b3d151bf3afc97addb223ebaa59592ddcd94" + ], + [ + 80960, + "0xdef9733fd939d6b260fe671867ee7cdcd802332b9afdadbba0ac24b964126d29" + ], + [ + 80961, + "0x641d150032ebcce0d0d7616b010eb29df976191edfad5c8751da8533f78f7946" + ], + [ + 80962, + "0x7480f68a44e03c35a1ece1cec7402fe9fb5b4d34d4f246c967c5178e1a7fa330" + ], + [ + 80963, + "0xefae600a6e868a4213011b106f295f79e5c86186b8382143c0a402761226a9ec" + ], + [ + 80964, + "0xc6ce05cbc32d8a1d786baa6473cc65c29892a5a47e7474f384b65c567d425748" + ], + [ + 80965, + "0xd3a3e3c3c964a77a0821984903c958079b294b195b7b13e109f26f3bafcfa092" + ], + [ + 80966, + "0xd42712815a77eaca169de5c74653b4c007159f47fb5d50d63f57cfd3f2c5868e" + ], + [ + 80967, + "0x9b309f8c13534ed43e399d45a5bcfdb38f49c5633b458f6db71d39571e9d8d41" + ], + [ + 80968, + "0x74729f1a9de838f15def5f5c183b7d3596b03afde1ed6302cabacf8a677f82a2" + ], + [ + 80969, + "0x4c07a8677c028eda23e1665416eb54d364959f3210e682978082d560ebf2bea4" + ], + [ + 80970, + "0xe3f992775c7dfbd2485ccecf516ca7ee9961c897296190c45fd7abb1431ee04b" + ], + [ + 80971, + "0xa8492953412a007f2f7696922b661a27fe2ec42e566504b3aacae2f97a24fab0" + ], + [ + 80972, + "0xe14d6f60ece95189dc8eaed8a2ea5c2da62cbdf988588222a4e99d242ea3beb8" + ], + [ + 80973, + "0x5c7ca844c8eec0cce4e4640ba6fc486c1739eb44aa50de86f55008a4d597357c" + ], + [ + 80974, + "0x2b33e852101ee35f2c61ee7b35152521eefb11364505bf6903a964c93d4001b6" + ], + [ + 80975, + "0x5516b628a8ec4328bcfdf90785fb5d96d22d60b4db638ea1b93bd837d7c739a3" + ], + [ + 80976, + "0x9959b993b66dc1285d0ecd8c95b74fe70949c50d5a97d822ce866291f94f2926" + ], + [ + 80977, + "0x04eb959364eae151a41d4a8d970795da0bf8a0ef2f64bde753cbca4ea3d745c4" + ], + [ + 80978, + "0x22907520934a83b056017bfaa72a0ce92fd8658c81dc8236f3b78f1f7e055aa6" + ], + [ + 80979, + "0x59fae978e8063799b4ab6335ca7ac097a25cbb4d60f9af04eb2f647a92ead4bc" + ], + [ + 80980, + "0x9ef7533f2e9c60f4acf9618efd1c47b147fa95bbb27caac54c1bb19540befeab" + ], + [ + 80981, + "0x9986f742ffa9d17f641293e788e53cd904b6edba3e971603470186f61f934e77" + ], + [ + 80982, + "0xbfb394920e4bd707342ddd41644a5733b7041e517d4e88863924046cd512ec19" + ], + [ + 80983, + "0xa3fef379a91df7a254bc4897cf2ef61ff25b380b2142374ff40d91d5140ed21c" + ], + [ + 80984, + "0x71626767d14a0735e3716b45b847611dc7fe04c6d7a4b22e3a6031d621b6d560" + ], + [ + 80985, + "0xb2c03912458086a6b421a8c83cee657c9276a1d38db128799a3170cba30e3266" + ], + [ + 80986, + "0x8f8000a6e85c09716e8766c9bbb9df1287d46f54625ff139d5e98142a58c9253" + ], + [ + 80987, + "0xfc88dba1647b7d43ef19a1a38780dd238228dd46beeb50ea0e00853f06c78471" + ], + [ + 80988, + "0x7e383c88765b3c499b05b60cdd4470cf9b23add0f6bd33990f207a43a51e431f" + ], + [ + 80989, + "0xa133fe8b34c9e839bcd4a75c2a5858874fa19d8bda43cd1c1af781cd2689753d" + ], + [ + 80990, + "0x466a4d175cb847a9899d98a6863900be93ecfec2f1681c8595db5a62b843478b" + ], + [ + 80991, + "0xb9b1bf393868e53ec5da954cb87fb65eef79ab9caf59cfb636f6cfd9c5c92f8e" + ], + [ + 80992, + "0x2f9eb9f04e16f420d5b98d6a28e25494207bc05affb6b914a1b98fb2244bba75" + ], + [ + 80993, + "0xc73829f98810dd017f45a0bd71ad7db3d07a1f5ff0a57c07994fa095989e6ef2" + ], + [ + 80994, + "0x9cdecf07b8f03979fdcb8f7d4d2c5ef4a13adc19edc0bb0c87a8e0beac6ff3be" + ], + [ + 80995, + "0x978c9e3471a4c37fb6aebab56d41d660720681073dfc68e5bef41d594d67ac8b" + ], + [ + 80996, + "0x739dad0614f0e852037aafbab92b4072f0f9dc9fc30c489465d67a18023f80a4" + ], + [ + 80997, + "0x1fbba34f589373e32f170a1cc3bf5c69083e8e2af5167e6ff38eab89ca2c0a72" + ], + [ + 80998, + "0xb58b192d29e3605fa4006c019698a7fbdc14639784ceeac087ae09148514e68e" + ], + [ + 80999, + "0x77f67acdb2afb4754c6e31005b27aea691cd3edce3094f87238df3f887e1bebe" + ], + [ + 81000, + "0x50e198025598c36879156ee586cdaa18bfa620f117c49baba1b2093d77b83831" + ], + [ + 81001, + "0x416e7d689843b2ddabbc9efd9b6f183b883dcb28e857977cd130f989ff7c72d6" + ], + [ + 81002, + "0xd639747fba2f48c62f7aeb91ead4b7ffa1292d55bd15ebec58879bc51b05863c" + ], + [ + 81003, + "0xc43806d1ab9e6d84be147ac3b74130fc2dca01254b4a847eb2713c3e7c1cd11c" + ], + [ + 81004, + "0xd854d522a33e98313a872fb4c47675f54eeed084d99794c5c72aab4464ee70b4" + ], + [ + 81005, + "0x752458c4e11fe4d564250bd4bb0122c1a1a4fabb124942ed9c3981887d612854" + ], + [ + 81006, + "0xe3eef7ded5487286de24e9802f5dfa1afe7c477666f2a26fca7606d3b9b8694c" + ], + [ + 81007, + "0xc525b44a5e888c3bdb0f66edc5bb4dbe9ce9116279b01fa510037760319ca15f" + ], + [ + 81008, + "0x9552445c57b6e2cd01d95bf4013bb913ce1611e5d707f71e528ad7522e830806" + ], + [ + 81009, + "0xd0eba42e3dd5cb37d42d6891eb47e5e2bdaeccbf9ddabeba06a09cd2b4b15f55" + ], + [ + 81010, + "0xcc6fd1b1758929e8a94d3bfc1b4c5444c55023f8939d4b4fb539b523219b00fd" + ], + [ + 81011, + "0x826ee09096704d3798aa4fccfc02cdedc59bb7472332b624d3716fa3aa14e612" + ], + [ + 81012, + "0x3f82182c620ce4d5ac0a776b0ec52b590f7d30c3bbea039d9a6ec3c6ec8a4aeb" + ], + [ + 81013, + "0xbdc7e4e07b4fc9245d8f5e7d7b207af203ce271b96912963aae2012de4d0dfd4" + ], + [ + 81014, + "0x797fd61f65fdc95f190d0dec1ad3eee5c98211c4529c83d013236838eda13d67" + ], + [ + 81015, + "0x7928f9107a62bd4d597894701cf4bca71365052e3560c23af7fdfa3374105b45" + ], + [ + 81016, + "0x78dc4dfd56d144d629278047ce28d8649e02b55768a137cf98260e40e6c35989" + ], + [ + 81017, + "0x3cd444ed1bd6f3012197839db12d566729ca5fb8c133af4353afbd756104c03c" + ], + [ + 81018, + "0x26350e310daef399afac9ac19b106d5f160d8dc6921fd289b196b4e979742bef" + ], + [ + 81019, + "0x99e1715d9ba9a3d125e77357595bd918766a4d46c20d9db1dfa9ebbd002a7bfa" + ], + [ + 81020, + "0xdf16ba27999f8dacfde10d403cf9982fd3b7545793e678219f09818e04ee64d6" + ], + [ + 81021, + "0xc4690b3c4b89b7282988cd2f71c9e43d3c17d081f0c735d04fbd01247a74bd6f" + ], + [ + 81022, + "0x94a1951db605f9a7dff70e2fd919e154f1291d31b4baa6f9e88d8f904c74fd54" + ], + [ + 81023, + "0x9846feb46081eafd35bd5f134cdf98f7a2a270b52a36779de7fb132250823a6b" + ], + [ + 81024, + "0x3716fe1bc70f07f53358e7936ff4b652a035eabe199265f99809dc4aac49d002" + ], + [ + 81025, + "0x841c89865fcece06f48c90ca283a1dbafb49e3fc422ff4cc5d3ec7fa68ac0f14" + ], + [ + 81026, + "0xcf2be45d7a88f2ffa7fa45366404b8b0af131758b7706f099ee8daa6b5d17b94" + ], + [ + 81027, + "0xf3c7dd48103bf10e13bea5800814475a543f4bf53958e1ad464b112d0c763d00" + ], + [ + 81028, + "0x6ebafde3e8fe7067903074cbeb854b9957bbc45859143f618ada972fccdc2c0d" + ], + [ + 81029, + "0x05146b0d5efdf34a927593169ecce2d2d112a3faa3b79cb2df1a0a2e6a4bd7c0" + ], + [ + 81030, + "0x33def97a4f3ca92c7af02bd19bb9703c516f6df8917928734953ec9fddae68ce" + ], + [ + 81031, + "0x18bbaeb286265fc8edb34abbfd3fade65196f09125dae4d31347401eb635b63f" + ], + [ + 81032, + "0x56aacb0d332a22fabb3bf8ac90b030fb82337b3b22b5be334d9a4f633847480c" + ], + [ + 81033, + "0x8d5d77464c966282353e2019598e0f1d77b5afb938475275f4a3b0d0c29d9c5e" + ], + [ + 81034, + "0xc8ccc7baba52e4fa56d488c78abf451919aeba4a029068fd364059633446c80d" + ], + [ + 81035, + "0xc295e160d3378e12e6b96de0eac43d903aa656f9255e58a3b0222bc229b18bbb" + ], + [ + 81036, + "0xd4cacd3eb6e8eebf312225c23284c15e347e73539816fe91954b482a9acb1753" + ], + [ + 81037, + "0x096f15eb2cfa8a6811e08ee533e3a9bfaf3447225c214157d33f27d532423598" + ], + [ + 81038, + "0xd2464a278a262d447f8d67feb035e8b9eaaaa849eb87fb330e1ac868e6b1d063" + ], + [ + 81039, + "0xfea1ffb0177d362f9e041268c395858091da08fb55de5c922493bf0f3dd19476" + ], + [ + 81040, + "0x5ab2d5c6f5789372abc97b4792b2c4e3476c70bd538764ba74d684ed102fdd0e" + ], + [ + 81041, + "0x2f5be38f279790c0e69e554fcaa28f82bd17ef645338539416db0c9252040f91" + ], + [ + 81042, + "0x844ebc844ba1c0c301a6a9189779e17e557f824d85d01247b4a0b1af49f50d56" + ], + [ + 81043, + "0x944855a63daa3b755a70f2ba5e6a956b037eeb210dd9e80cef73b7c814645adc" + ], + [ + 81044, + "0xf046f77959a6aa199f3d999b2c16dbeb1e9038c308cb1dcc9bba26141e0e5f35" + ], + [ + 81045, + "0x4e420f14ba78477c090eef99a098b62dfdadc3a61f7c0a5499797380763693db" + ], + [ + 81046, + "0xae10dd9883739a2cd9e95cafa4f9cc7b2e38e477207ec49fb3fe64e3fcf21613" + ], + [ + 81047, + "0xa9e686e36215cf1053ac16cb445e537ef8a633aac308ad755716f6374c3c1238" + ], + [ + 81048, + "0xa7ab60b451ae4a7fa639faa0c8b1f2ad7bb6bcb853242797f4aca608b2e69193" + ], + [ + 81049, + "0xdba699a61f7e4b3b8356a7494afbb73237f00385d7ad279a77002bfd451b2643" + ], + [ + 81050, + "0x42515b83488525c1c8e72417df7361aa61a62e3f78892446b06d6d6bccee6ada" + ] + ], + "rewards": [ + [ + "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "0x32692d9b6ba26000" + ] + ], + "proving_pool_credit": "0xc9a4b66dae89800", + "payouts": [], + "blocks": [ + { + "miner": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "blue": true, + "txs": [] + } + ], + "pre_state": [ + { + "address": "0x0000000000000000000000000000000000000210", + "nonce": 1, + "balance": "0x0", + "code": "0x608060405234801561000f575f5ffd5b506004361061003f575f3560e01c8063aa67735414610043578063dea5c2e014610058578063fe7e05d51461009f575b5f5ffd5b6100566100513660046101e0565b6100ca565b005b610083610066366004610211565b6001600160a01b039081165f908152600160205260409020541690565b6040516001600160a01b03909116815260200160405180910390f35b6100836100ad366004610211565b6001600160a01b039081165f908152602081905260409020541690565b336001600160a01b03831614806100f957506001600160a01b038281165f908152600160205260409020541633145b6101635760405162461bcd60e51b815260206004820152603160248201527f446576656c6f70657252656769737472793a206e6f7420746865206163636f75604482015270373a1037b91034ba399031b932b0ba37b960791b606482015260840160405180910390fd5b6001600160a01b038281165f818152602081815260409182902080546001600160a01b031916948616948517905590513381527fa47563c41dab010f91a8ef9dc7ac2bcdfa0ef2af697e575048e71e6eec60dda3910160405180910390a35050565b80356001600160a01b03811681146101db575f5ffd5b919050565b5f5f604083850312156101f1575f5ffd5b6101fa836101c5565b9150610208602084016101c5565b90509250929050565b5f60208284031215610221575f5ffd5b61022a826101c5565b939250505056fea2646970667358221220cbf48f5aa6f911f83b3c2b09adf8c418f5da6fc530c3d5db9d4ba0be24c4bf5864736f6c63430008250033", + "storage": [] + }, + { + "address": "0x0000000000000000000000000000000000000220", + "nonce": 0, + "balance": "0x13cbbade0aa5f04f2400", + "code": "0x", + "storage": [] + }, + { + "address": "0x0f002c928c363c7041cafe788f099cf7d35452a1", + "nonce": 243, + "balance": "0x1bc36179476e9a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x18524811fa2e76dd0770d29308b51bcbdc0499f9", + "nonce": 242, + "balance": "0x1bb8aa483e652c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x1aca71f7872aebc85b4d4a9ad47a09abcc624d65", + "nonce": 0, + "balance": "0x2860cb14365c048800", + "code": "0x", + "storage": [] + }, + { + "address": "0x233d639f53225ea5012dc01ceb0a5a30891021cd", + "nonce": 264, + "balance": "0x1b8eea59c54b4200", + "code": "0x", + "storage": [] + }, + { + "address": "0x27fd8475c3db352fdddeaf29823e00c7d33ca39a", + "nonce": 248, + "balance": "0x1bccb45d83401000", + "code": "0x", + "storage": [] + }, + { + "address": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "nonce": 0, + "balance": "0x1e063716f0755a0c800", + "code": "0x", + "storage": [] + }, + { + "address": "0x46681948060140945958068b373f94581e049d4f", + "nonce": 240, + "balance": "0x1baadca68fcb1a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4cb00bc539538d2ee7172474f2b06f021944ece4", + "nonce": 258, + "balance": "0x1b3e54c72c629c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4d7c0d6f3ad5466691649e98ebb59cd5b5194687", + "nonce": 0, + "balance": "0x1b598fbe8e08549c6000", + "code": "0x", + "storage": [] + }, + { + "address": "0x5a5e606bda0fe1b2b4298aadd55cc6ec57ff698a", + "nonce": 0, + "balance": "0x648cba8883ddb562c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x6613ca66c14338db771d01c238ae140c71310535", + "nonce": 243, + "balance": "0x1b98d3921da34800", + "code": "0x", + "storage": [] + }, + { + "address": "0x66577de387feeec96c10f7f8748f13912589c78c", + "nonce": 240, + "balance": "0x1bc37e3123b20a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x7e7b1db26094aa913933f65127b328c1861ff5b0", + "nonce": 256, + "balance": "0x1b89d2467bcd2600", + "code": "0x", + "storage": [] + }, + { + "address": "0x90acb15171deb958d3191d5c02d807afc367f0e8", + "nonce": 240, + "balance": "0x1b9e3c5e84fb6c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x9f296bf64eb7051912c9c8112da3e3ea6cf546e1", + "nonce": 266, + "balance": "0x1b69b11e00d88c00", + "code": "0x", + "storage": [] + }, + { + "address": "0xaa19223d63c82bf9be5543ec96e5504834d17bc8", + "nonce": 246, + "balance": "0x1b9abb1eb524dc00", + "code": "0x", + "storage": [] + }, + { + "address": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "nonce": 0, + "balance": "0x55f1d5570ee67d0000", + "code": "0x", + "storage": [] + }, + { + "address": "0xc2faa4a2866422a865c9b3b3cb5988bf4d07f0b4", + "nonce": 0, + "balance": "0xc05b9fc73002b32800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc45d23f49451faad9afadf877feead626ea8d083", + "nonce": 0, + "balance": "0x9ee8f806428adc9800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc8621de5921418701619ba715044575073d2ca62", + "nonce": 243, + "balance": "0x1bbb8989a888a800", + "code": "0x", + "storage": [] + }, + { + "address": "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "nonce": 0, + "balance": "0x1f29d0e7a3eaed38a400", + "code": "0x", + "storage": [] + }, + { + "address": "0xd958300657e7931b06fc6f60dc4bfe7d8e8ea7b8", + "nonce": 252, + "balance": "0x1b79a326be219400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdadb11b01f6004eba6da3846a0e1bf6f4f71fb82", + "nonce": 244, + "balance": "0x1b900b01183e5600", + "code": "0x", + "storage": [] + }, + { + "address": "0xdd442fcbb964a3afdc90d49b408e8dd296fa86e8", + "nonce": 0, + "balance": "0xafd0c09109be7410400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdfaea67368f3e3753397d878f97efe6aa8020c2e", + "nonce": 16, + "balance": "0x4f20fde2ac0f98ec00", + "code": "0x", + "storage": [] + }, + { + "address": "0xfd4fca79b266a73a7defa776b04c7e15bdc70e86", + "nonce": 238, + "balance": "0x1bb5cf9909338400", + "code": "0x", + "storage": [] + } + ], + "fees": { + "base": { + "pgas": { + "version": 0, + "cycles_per_pgas": 1000, + "intrinsic_pgas_per_tx": 200, + "modexp_base": 1000, + "modexp_per_byte_numer": 10, + "modexp_per_byte_denom": 1 + }, + "block_proving_gas_limit": 30000000, + "shard_proving_gas_budget": 7500000, + "min_execution_base_fee_wei": 1000000000, + "min_proving_base_fee_wei": 1000000000, + "initial_execution_base_fee_wei": 1000000000, + "initial_proving_base_fee_wei": 1000000000, + "base_fee_change_denominator": 8 + }, + "v1_activation_daa": 18446744073709551615 + } + }, + "plan": { + "shard_budget": 7500000, + "consensus": true, + "shards": [ + { + "index": 0, + "tx_start": 0, + "tx_end": 0, + "over_budget": false, + "pre_root": "0xb6dc9eb1ce5069a634647f05619d7b1bd477588aab401e5e814db9c3628e8831", + "post_root": "0xad123f2337b9d9df433d08e7f7ef6a19f885ab3eb100b19961aa051efa79f769", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "link_in": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "link_out": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "witness": [ + 2, + 0, + 3, + 14, + 14092 + ] + } + ] + }, + "expected": { + "pre_state_root": "0xb6dc9eb1ce5069a634647f05619d7b1bd477588aab401e5e814db9c3628e8831", + "post_state_root": "0xad123f2337b9d9df433d08e7f7ef6a19f885ab3eb100b19961aa051efa79f769", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "tx_commitment": "0x0000000000000000000000000000000000000000000000000000000000000000", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "node_state_root": "0xad123f2337b9d9df433d08e7f7ef6a19f885ab3eb100b19961aa051efa79f769" + } +} \ No newline at end of file diff --git a/proving/fixtures/chain/block-81052.json b/proving/fixtures/chain/block-81052.json new file mode 100644 index 000000000..49390e075 --- /dev/null +++ b/proving/fixtures/chain/block-81052.json @@ -0,0 +1,1334 @@ +{ + "format": "igneum-prove-fixture-v1", + "source": "live devnet export from node 1 at tip 81076, 5 October 2026 20:06 BST, proving v1 chain fixtures", + "block": { + "chain_id": 4463, + "env": { + "number": 81052, + "hash": "0x93319ca0399a5b00e77bbde1cf067c1f4f823cbeffaa1103c7953abd1b8db6c5", + "parent_hash": "0x9a3442c671ff6750549ac81e1a48d7ba70e8248b0f1122bf37a3767666cac4e6", + "timestamp": 1791227134, + "miner": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "prevrandao": "0x93582d11807624c877052a3cef7d658583d5e6595f72bb92412316de74fd4193", + "base_fee_exec": 1000000000, + "base_fee_proving": 1000000000, + "daa_score": 0 + }, + "block_hashes": [ + [ + 80796, + "0xb09af3831aa1d8f4945d50fa16a602848a2da22b14dbc1f7af2844cf19e90040" + ], + [ + 80797, + "0xe099536cfd64654690afc152b97866b32da6bf81f959a6c61f9e43bfd2ca2d9a" + ], + [ + 80798, + "0xd6af2cf6184f15ecca77ee9e7e8d2ca718d88f9cea39b475713ceb0e151ddc36" + ], + [ + 80799, + "0x9c76f9f6036c7035fd52a7a4e3fca99451715e729fcc9fc424487faee9a6134d" + ], + [ + 80800, + "0x9a046cd99eb89c551dd0c698f1969069c02ae36cde4140ff4b864d624fd9dff7" + ], + [ + 80801, + "0xc9bffa329d649c6540518cf249e25e845743dc01fcbcfdfb8c798f1455efeae6" + ], + [ + 80802, + "0xb124f2de3c17a985f10c0b58d2283631bcc12524de8d79b0955aecfd2d7f78aa" + ], + [ + 80803, + "0x97b21d6cd3f1b321d2c9729062c305a7955b97646d76910714f899c75d275568" + ], + [ + 80804, + "0xaa3f0e114e156f706e11a701fb940bf54e89b09d83953be7f6733478c7a33093" + ], + [ + 80805, + "0xe5a5868b02ec30304079652a8d5b8af4a9da0a5a676615c1a10d833feb4eb0f7" + ], + [ + 80806, + "0xd97c5c4288f56955e1a3e134eb63250e1becb907b0fdde9ce1383d8d2c8a3939" + ], + [ + 80807, + "0x8f7bf1128e5d03607cca246b86e63cfb4bfe73e74395c8660d027172376b5c36" + ], + [ + 80808, + "0x5d5fe4c3bee7d758986da5a433cfe5221be392419ac22e5acddd701cd12c8142" + ], + [ + 80809, + "0x9b0d0a871ce8eaab11a185a4344f9f48a58b7111f94a8f5221d557722cc5ca5a" + ], + [ + 80810, + "0x58cad4ab776ba7eca94e813323e28a6467a7fcf39d5aa2069e9dfef6f8058e5c" + ], + [ + 80811, + "0x706788247d6c05475eb3bdbc6f181e76429b57a97f168004832fe5ba4b20c7b7" + ], + [ + 80812, + "0xdf26fb97f999b1016865b3fee539e72e3f66d6d228b2aec7e2f59ab93e5a9ee6" + ], + [ + 80813, + "0x8e612c07fdb693a44353a08899e8beae8c93fb1b85845ad965f96e7d2f77fed4" + ], + [ + 80814, + "0x806f4c29ea28af1a3ce0c3ba775c776006ba8cf4b308ee617dc207dc3f740b29" + ], + [ + 80815, + "0xebd8e602049ddc2f8d11d5edc8ee837eefac77418e73e9bad97f6f1c7a85e8aa" + ], + [ + 80816, + "0xe2b069b77ac05cb0a95ae419966feeb20660e4ff3b9a68d044a5721e1525dd86" + ], + [ + 80817, + "0x1ca5b79e26dc425a94c3a6bfd339516caa0a7f8e5e2a62389bfa466186bca495" + ], + [ + 80818, + "0x7cf382f956597cbd36917036dd15b774500f97714e50587a35ec9617662052d6" + ], + [ + 80819, + "0xe70494836c5b5c68b6a167c868db166f9630140499f96e6660d3e3928cad6a14" + ], + [ + 80820, + "0x15f73eefc39e00f9d09390764f9ea0069908062c563686e73b49ec3919c24be3" + ], + [ + 80821, + "0x7f2bd2cf68fc14e6a5debaac39f368534657c41cc587402fac1156d24998e55b" + ], + [ + 80822, + "0x2641ec1c3397b06ab5855dcf6ec2e23de301f5831015a32ffa9f29e18cc2f084" + ], + [ + 80823, + "0xf643e963f5a62a49e935b4ef51aab9bf2689b0b9ec483924d7f42934c8ce5e07" + ], + [ + 80824, + "0x867ee0219bd30fc3117c79295cb1075ca21281fdd285467ae128aece1c1dd9a7" + ], + [ + 80825, + "0xab6bb19259a039d30c3022c9da3494af8c9a541dce30895df083252894c433df" + ], + [ + 80826, + "0x155511de8d34bef351ba0ae7af5bccf515e6e490135130aad5ce2eb02e103d75" + ], + [ + 80827, + "0x4e297a53901da9b364f4dce686f016f4039658ea8832d591dcb21aa44acd197b" + ], + [ + 80828, + "0x7d07b8bc9b06b20e7b56f96fff9051590caac01159d39b3f95bd592a4627daf4" + ], + [ + 80829, + "0x8392a1875388baec163e64854d6ee7664a31017402ee8e269ddc8aa42ca06647" + ], + [ + 80830, + "0x5ab4e5c586d2fb4272941bf26c15288cc129ca1092d3aaab82c478edd1942e29" + ], + [ + 80831, + "0xb0100f0a559daa435128cae74f47f9353ca2a93c2bdac8dd9d12a3130ce6a818" + ], + [ + 80832, + "0x326d1388d0911871c5087925500c4ee5091624331d07659b9d2fa09e83f5884a" + ], + [ + 80833, + "0xc01c186cfcebb8d53953589ce595afef213d46e481b8b5dee7e0a71abd773f6a" + ], + [ + 80834, + "0xe845934e79eaeb8a8ee0b9396b3a13f13c8ac45a0d4e81017c73042789013c38" + ], + [ + 80835, + "0xa45f8f03249624a2d35b54eb83f4edf54e8d1acd061fee3965d9164a132c72aa" + ], + [ + 80836, + "0xe72a59a64af9aa223b5341819a5872f8626908b5c1857d0617706f6999632e01" + ], + [ + 80837, + "0x94063f426663347bb4a3e79f4dedeb76fbce6fd69117a2c74ad1fb4714b856dd" + ], + [ + 80838, + "0x157ab5f9e7fbf98b48af049adcfb4a8e071253ec95b991d031a91668ff3b8d57" + ], + [ + 80839, + "0x5bff632aa54b1a71a7d9a2a57d9c42ef223eb6efe86c72f9aa8dfd67ab7b218a" + ], + [ + 80840, + "0x7f5d1952ad8ad27dd17bed3354395b646468985bae0341b4048376a425e43076" + ], + [ + 80841, + "0x8e69df81fd1b0f186813304260dda1d0c9dad5792126e69d021e679b837b01a5" + ], + [ + 80842, + "0x6733c5472e38830dfb273a0e0802924c1b60a2cc916d6efaf82785e1bf960724" + ], + [ + 80843, + "0x0838513dc8264efffb2c11f8c00ebbc9c57aeb67dd7be93f18e1c613872bc929" + ], + [ + 80844, + "0x8c2f26b17db95bb20983ac41df7a5fca3e3826d2e57831eba13bc3175daeee34" + ], + [ + 80845, + "0xe3edcab823e03598c66e16c5743f7a0917735c975ac946db4cb3ad862db40a21" + ], + [ + 80846, + "0xd075510e75159d941946e666e2c75604a3f3c959523cbe2c0f025427ae85d4a2" + ], + [ + 80847, + "0x57b4b6c0a6ca2532515653445bc872dd44e621f9774f3dd4e2000bf82a0f510a" + ], + [ + 80848, + "0x38a85c99bc3e489a4431ba9a125e74f4d807c630a239419287e039f05cc91ead" + ], + [ + 80849, + "0x3c95d08c99482d57bbcb9fd333cae4010628fc8b1c96b3ad4a68d8f3dd21695f" + ], + [ + 80850, + "0xe767527430f19da2c2b21bc648e33ed2b241f5b7c9413649d9c5a555491a9253" + ], + [ + 80851, + "0xe9941baa9f3378854aaff5563340d130700f0753197c7373229542c56ac91a8f" + ], + [ + 80852, + "0xb92e88ec7cc8c340959d8d0cd5434eec5c1033b389b953b02cda103d53292df0" + ], + [ + 80853, + "0x39f01f35ba2ea7b45e3eec1681c44bd85b066a6a8521a10fe6cec5d4d98aace8" + ], + [ + 80854, + "0x36456a5e36a9a66488656bc0a92705632326e59ca40bd09f6d5f99ccc976972a" + ], + [ + 80855, + "0x12e52067750279b8bc34b03cd21bdb2809ea18a3aa68739623780aaf1150129b" + ], + [ + 80856, + "0x64951c693d7f246647fac0df504f35ff9812416f966afaf4fecf7ca92c0e0b24" + ], + [ + 80857, + "0x3918b797fe9d8cd6c34a0b315aaef550de15f4d30dce932f3be8cdfdc27343a2" + ], + [ + 80858, + "0x86ec8713a68034acb9f34b015a48644c7562766c1760927a299e517ae8ef6f8a" + ], + [ + 80859, + "0xfe75a4c99fc17131b7aae61dd0a588b7d7019e6c662aab8c7d6a8c56370fe07f" + ], + [ + 80860, + "0xf4afa8f14a0f814639c88dea304a8ccf76450ee56f4fbb89caf29dbeef80399b" + ], + [ + 80861, + "0xd4c72bc192643e20658ca54b730ef865d1c29ebfb6d1f78263c796a23b6965b4" + ], + [ + 80862, + "0x7a73b7302ab47b7acc1c5359620df9026fdc08d78c4173bb4e45e73a66651767" + ], + [ + 80863, + "0x9033ce568146d2c69139d3674fa7e18cb8319abae5e6cd24f924f750d2ca3446" + ], + [ + 80864, + "0xe4d8da84ba631f233ea7dd1336a264f499083b08290c82b717f79bfb7ffa09a6" + ], + [ + 80865, + "0xc53ec9d6b11d592921f5cf850584ebb562119ea968a5fa7bb00af381d6eaef46" + ], + [ + 80866, + "0xedcbb7d2f9626af53341464fbfc54e838e6417b7ffcfb0c2f9758641c98d5f9e" + ], + [ + 80867, + "0xbe5f726816d43ca5bbcb70f13901e60ef9040debdfaf2a73a0692c3cbef578ec" + ], + [ + 80868, + "0xb7cf72378b14c086afb4094b9b0121fcf7c6929bf2c2134b5189862644e2e506" + ], + [ + 80869, + "0xa29e21de7e54defda356f088dea38a06cf1a1a75bf93c14e30afdba213df7ef3" + ], + [ + 80870, + "0x504d3f83a27c37c5a15cdbbdfc03285f53071ad01d82bf9be2d6f862548a9a8e" + ], + [ + 80871, + "0xe02107fefe75bc8d1abb0443bff2e7a658e3a9b2fccd5157c1a7c7ccc995fe17" + ], + [ + 80872, + "0xdb92e1b449e7670ea719a6d68f496965620656ff5d25a0521f1873eb8d7ae1ab" + ], + [ + 80873, + "0xf12689607f86ce8f762df5d67054a00b8b580300511a221ff5addcf732393988" + ], + [ + 80874, + "0x0e0fe961c78fe52f44f6742795bae3e96713f4618e4a95623045056eb9e046f6" + ], + [ + 80875, + "0x7fd5bb64e800686ae28c3b55613d27f181dbd91e35d342d33d7b3e5d4df17d00" + ], + [ + 80876, + "0x83c2feae4eacf2ada1c9a630a97614a6296df6b89bfb5f954c3e306b16f0ab2e" + ], + [ + 80877, + "0xd1f3af4d6464c5bc9416d73a291bcf3e1df4fa1a396bb3e9dff92dcd451f75d9" + ], + [ + 80878, + "0xa4cd4b49e1d475a5561a91a0de592ba0c00f63b39a61eab150c6486d1bf2b916" + ], + [ + 80879, + "0x714904f162c3b9013a9defd79d4552b58c5ca8747cc94187079f2ebefee2f0d9" + ], + [ + 80880, + "0xb2263d12431e74591e2843047d79933820d2dd352d5b5c977698b328cd59f46d" + ], + [ + 80881, + "0x67977d72a573f121d9a818dd035959cd502e5bec07a4a5c30a28fc66d4fdf964" + ], + [ + 80882, + "0xf1d5f1b271cd23a79f1ab32655418e4b664663c4a77ebaf0a213fc2456e59c5a" + ], + [ + 80883, + "0x91a9f94d339407747a77d79a9c2be608744d4ecc8b84699d0120814f55a5060b" + ], + [ + 80884, + "0x8f709bf4582d3a912e6006f6c417459bbbf182af2b500b4376cd8dff78201c25" + ], + [ + 80885, + "0xbc66848a66fc038200186f09665540b5905a350a3149bfdddf9575eabb2f0a98" + ], + [ + 80886, + "0x8827e43b9a02e7424513e6d0fb6db7e93cc5555fac9b832584570784708a8f90" + ], + [ + 80887, + "0x7c76b6619ea66b8e7fa2791ad1a3dcafe863388d14db50dd022f63748e2b6ccb" + ], + [ + 80888, + "0xace6e76c69db6530277c9749ff3e77a64c6e7123790fdf5f248cc67b761e10bd" + ], + [ + 80889, + "0x8f96a0d2f9524f0b2bf9f520f35516969fba795e005f003dac5721cc37c6b5ce" + ], + [ + 80890, + "0x68a635c0d7ba10b3fff9108deecc4e9af70d760d688c5614a000503b1f2f006c" + ], + [ + 80891, + "0x47dc91daaf1758699e90f8906c61bcb6dd020b3726a1e0d15a7fe5177bfab696" + ], + [ + 80892, + "0x79f1ff991dd34c0384b9ddb64a7f3791116891447f2c960de5c6d80f9a705b5f" + ], + [ + 80893, + "0xf41832c4af88cdd31eec42b233ccfae51a150cd41baf59e9dc6b64dba2b63c7a" + ], + [ + 80894, + "0xfaf2c58e9ba01047e4cbe2c147a4614d46ad39c1d64cbb43d1de60f9c75286b7" + ], + [ + 80895, + "0x65a671bcbc9b6f905811ce62594feda18a12b8c6ca1c72081c1dde3f6227d1a7" + ], + [ + 80896, + "0x815d16976077a9b599fd5d6a82eda30a0bd1d91c66718528aaea9d4dcd12274b" + ], + [ + 80897, + "0x710c0c1d4d44ed7925343b10c7ace3216741b31a91506fa0f1ad80c96c049b14" + ], + [ + 80898, + "0x5fff52604275c206678c800d95cd0569325ad67413b50ac15252c3440856cc78" + ], + [ + 80899, + "0xe67fe87df50a63d348d300de16353e560c1b86f1a79a0bd09e2a7771ae7e92ba" + ], + [ + 80900, + "0x5d854ceaef9de086a2361cfa1f829f843aac00181fa5ae27e69171c198853895" + ], + [ + 80901, + "0xca45572516a2d4a948cfe518d0bb5378a4d23b800522f515e59176c31c069160" + ], + [ + 80902, + "0xa724d3e05f16d3d8c97296f59802c7f9ebba4284808126d955d6fc96dc1c4729" + ], + [ + 80903, + "0x552061794bddc5185033e936ac8ad07d3a315c7a3d876989b2d61328ad0a3128" + ], + [ + 80904, + "0x23c1ee3ff43ae0d42b9b3f0d2c6c2fef6c8ab8fdec9772540b563cfebbb77a2c" + ], + [ + 80905, + "0xc7e33046cb4819a17a21a00e9c0efa55765a974d8a2364f4bed9a16034ec2828" + ], + [ + 80906, + "0xa7b5712f7d22449a6f6a8a55a33c2f76cb54be86faa2f0d892c8bf6aa9fe59d9" + ], + [ + 80907, + "0xbd8696f54d266714d694b41d3b6466ef242e999938b9effd2de035c3f2c5677b" + ], + [ + 80908, + "0x7ca324c51788d4cc173bd3a013702f80e68524b3b005299be9012621dc8e82c4" + ], + [ + 80909, + "0xaea0e3a560abbe0de857592d37bfd43bd986d1cd461f43139d951d123140d495" + ], + [ + 80910, + "0x640f0ed5ec76969f53d90d5e502f6fce0ec65ab28a728e92a81dd27c466dad09" + ], + [ + 80911, + "0xf4cb4beab37872930f84e37b38c15a0a5e5f5d95927c91284ae880337abcfd9a" + ], + [ + 80912, + "0x65121e870e1668a0a55c9a50d31166294cb8e57bc007897640a7414f0866ec3b" + ], + [ + 80913, + "0xb3ea0a195e00db811104db7362b555fd1ab8401814ee841ca7f4d3344843708c" + ], + [ + 80914, + "0xf92735fd64122d7f6f6df6f20f83fc0793675c81e25e54dc64b3ca027d8ef9a4" + ], + [ + 80915, + "0x78e07fd2fcedb51175785be4657b8d2a554c16c1d2cc396ca3738dbf6f9373f7" + ], + [ + 80916, + "0x0ce35f9c3a736ff2937e6b7addcb632855e347ab3c5c08b74d2dca5f82265426" + ], + [ + 80917, + "0x007b1e317c0d9e7565260e1fe7ebfe031de4291f6ddad75210f8563479e84178" + ], + [ + 80918, + "0xbdb4b040d7942b8bcd206e4e12632892d9cfb87c905f760f0c918dcb682d4eca" + ], + [ + 80919, + "0x1eb8ee38b4d246a0502cc56c8948ec6e17fa55b02002c9916e1a3d9867e3c852" + ], + [ + 80920, + "0x2bd9fea5702cc9c0b96fa3b34559c4a72c5ff4f1d572d9e702390a13b0b1c145" + ], + [ + 80921, + "0x12ae91b4df38f549004ab8c4482d9ed1a2841e43f649ee1a89156f9accea51ea" + ], + [ + 80922, + "0xb960711112f51fc11689061a66a9b4ed3c0bccb3a6f027bd2cceeb2d3786e780" + ], + [ + 80923, + "0x4d474df4920e9136e7564cf0e9ebd03bbcd8cb8bdb38c68b8a9c35120888bfea" + ], + [ + 80924, + "0x73d84f1013038f10bebf6606486c1bb6c6cf530a4d35f482be91df74fe15c93e" + ], + [ + 80925, + "0xf08dc8a601c46c43bb6f94edc1b0734a61e7f8efc82d915b33136fac63d5a407" + ], + [ + 80926, + "0x2b695a3c4d05e233cbe18213d50b10fb9d83914aca9a935aa6f6639b1f54c32e" + ], + [ + 80927, + "0x893a6295ba2f66f083feaa39e0570ccbd446166c2e15238d59a911f2826a47d8" + ], + [ + 80928, + "0xc2b2870e39e8344e18296012c4d846b78cd3bdc80145de6b595527c8845ac57b" + ], + [ + 80929, + "0xc32bf651e3ea72bd6808db5b4dbe4577078d418404221809ac12897082ab8bfa" + ], + [ + 80930, + "0xf1c2f5ed0a7d6b7f083c1a68f75004fbfc929ef3f6ab3b46acbba373feecfdde" + ], + [ + 80931, + "0x51638cd9a09c8ee630706d736c81ea0dde7d97234a515c6064be20d0b92b8418" + ], + [ + 80932, + "0xf3cb85e3f551c4e21d0c6c28e940c6bd965c2a20b5eb3375a9d8d32a65c58c4c" + ], + [ + 80933, + "0x65f05787c7f239dcd4dd9cac2e5c516e25f8a8b898eec284531c16a8da83ed2d" + ], + [ + 80934, + "0x58567dd84c5ce96fdfb6ea5794201199d259374e4879e76df9a6622694de95f3" + ], + [ + 80935, + "0x2031a54f25cefcf00c464f36897691a73616f777582bfcd9c03e5649d3339821" + ], + [ + 80936, + "0x0fc1b95d2579a6cb03f21083da53a989db2491f37bcc339195864baf21932823" + ], + [ + 80937, + "0xfc501a03f2666c7c7db743b6a7296692789824828c033f2254e3565de63bccec" + ], + [ + 80938, + "0xe752390723b297a6d326cc9d10f5a3219b5eede3159f9ea1c7fc7f26401b02f8" + ], + [ + 80939, + "0xfdba580d13c6964b46e7d2a66f2aeee150174facaf1cbff66016832b02d9f796" + ], + [ + 80940, + "0x2221d3b867caf0c1d4993845a60ba8389558ccd9d48bf97f51331e87ff51ce84" + ], + [ + 80941, + "0xb4ba8a208861ab1db5302570f7d589adc5f4826812b4759d625cabe2b68231a0" + ], + [ + 80942, + "0x0bc24048f15b64e0c836c387539686906198ed0b3a7adbe2aed37f41eaee5577" + ], + [ + 80943, + "0x5c3230d24716dfeb57c624bc44f7d80b65c95f4b810aa945f8e5309e45327104" + ], + [ + 80944, + "0x798ef63bf3ed33e8608b899ee956fc3a0a0b359a4466141728589ea914d42453" + ], + [ + 80945, + "0xb158a6b683a60e8d3debc803a89bd18af0f1307941039304a2e33a02f237325a" + ], + [ + 80946, + "0xd420374f3f533e0136b3fb1211357adbdde913776fbb851492ca80980e77ce2e" + ], + [ + 80947, + "0xbf7fea59019b30fec39a94c29b026285e0c2d6fe034da2555095999ca694914b" + ], + [ + 80948, + "0x3d232c12ca3a428474f8c8e21992efc614775842c365a29ef44d9b7f24caafc8" + ], + [ + 80949, + "0x87d17a79601cf126d2b5c95f1ac989f8a90f2b4e9f3b13d1d721776ee2a53c24" + ], + [ + 80950, + "0xd23290b4b63c19aafd1edefd1cfcf34829bdc6efa723df17cce5eb7612b17a0d" + ], + [ + 80951, + "0xa8d8db2850de878732f5a7b95bc08ddbb944e1e8e8884a2a2a6a98275bb7663d" + ], + [ + 80952, + "0xb2c005429b5d1f35d006693509651726e67fd19d34ac04a3473df90d8bed52ab" + ], + [ + 80953, + "0xfcdb4ee3a4fa1273ec429c49704ac1afe6c8c6f858640155276661b19a5c4ca0" + ], + [ + 80954, + "0x9b988ccb0403ebfb0c15486ed6b63ea070121d7754f88d1f14f85651c083f240" + ], + [ + 80955, + "0x39c3d647ae0dd09b1bbf1699803e3c1e76448787baf12cfced826b8d75dd4b7c" + ], + [ + 80956, + "0xb24d2ff081df7d34392e460ddf8fdb8c7cc5ee502e8cae2a279474f329fe7d05" + ], + [ + 80957, + "0xf1f1daa9a70d68fea94a204260b298285ca15469cdc1855be3ef731ae44707bb" + ], + [ + 80958, + "0xeadcfd836598031d7d1a6f5c14028edeffc5baed420b2fbd343a5e21cfdd1fd9" + ], + [ + 80959, + "0xe1391944a4bd9d2fe474c8b12de8b3d151bf3afc97addb223ebaa59592ddcd94" + ], + [ + 80960, + "0xdef9733fd939d6b260fe671867ee7cdcd802332b9afdadbba0ac24b964126d29" + ], + [ + 80961, + "0x641d150032ebcce0d0d7616b010eb29df976191edfad5c8751da8533f78f7946" + ], + [ + 80962, + "0x7480f68a44e03c35a1ece1cec7402fe9fb5b4d34d4f246c967c5178e1a7fa330" + ], + [ + 80963, + "0xefae600a6e868a4213011b106f295f79e5c86186b8382143c0a402761226a9ec" + ], + [ + 80964, + "0xc6ce05cbc32d8a1d786baa6473cc65c29892a5a47e7474f384b65c567d425748" + ], + [ + 80965, + "0xd3a3e3c3c964a77a0821984903c958079b294b195b7b13e109f26f3bafcfa092" + ], + [ + 80966, + "0xd42712815a77eaca169de5c74653b4c007159f47fb5d50d63f57cfd3f2c5868e" + ], + [ + 80967, + "0x9b309f8c13534ed43e399d45a5bcfdb38f49c5633b458f6db71d39571e9d8d41" + ], + [ + 80968, + "0x74729f1a9de838f15def5f5c183b7d3596b03afde1ed6302cabacf8a677f82a2" + ], + [ + 80969, + "0x4c07a8677c028eda23e1665416eb54d364959f3210e682978082d560ebf2bea4" + ], + [ + 80970, + "0xe3f992775c7dfbd2485ccecf516ca7ee9961c897296190c45fd7abb1431ee04b" + ], + [ + 80971, + "0xa8492953412a007f2f7696922b661a27fe2ec42e566504b3aacae2f97a24fab0" + ], + [ + 80972, + "0xe14d6f60ece95189dc8eaed8a2ea5c2da62cbdf988588222a4e99d242ea3beb8" + ], + [ + 80973, + "0x5c7ca844c8eec0cce4e4640ba6fc486c1739eb44aa50de86f55008a4d597357c" + ], + [ + 80974, + "0x2b33e852101ee35f2c61ee7b35152521eefb11364505bf6903a964c93d4001b6" + ], + [ + 80975, + "0x5516b628a8ec4328bcfdf90785fb5d96d22d60b4db638ea1b93bd837d7c739a3" + ], + [ + 80976, + "0x9959b993b66dc1285d0ecd8c95b74fe70949c50d5a97d822ce866291f94f2926" + ], + [ + 80977, + "0x04eb959364eae151a41d4a8d970795da0bf8a0ef2f64bde753cbca4ea3d745c4" + ], + [ + 80978, + "0x22907520934a83b056017bfaa72a0ce92fd8658c81dc8236f3b78f1f7e055aa6" + ], + [ + 80979, + "0x59fae978e8063799b4ab6335ca7ac097a25cbb4d60f9af04eb2f647a92ead4bc" + ], + [ + 80980, + "0x9ef7533f2e9c60f4acf9618efd1c47b147fa95bbb27caac54c1bb19540befeab" + ], + [ + 80981, + "0x9986f742ffa9d17f641293e788e53cd904b6edba3e971603470186f61f934e77" + ], + [ + 80982, + "0xbfb394920e4bd707342ddd41644a5733b7041e517d4e88863924046cd512ec19" + ], + [ + 80983, + "0xa3fef379a91df7a254bc4897cf2ef61ff25b380b2142374ff40d91d5140ed21c" + ], + [ + 80984, + "0x71626767d14a0735e3716b45b847611dc7fe04c6d7a4b22e3a6031d621b6d560" + ], + [ + 80985, + "0xb2c03912458086a6b421a8c83cee657c9276a1d38db128799a3170cba30e3266" + ], + [ + 80986, + "0x8f8000a6e85c09716e8766c9bbb9df1287d46f54625ff139d5e98142a58c9253" + ], + [ + 80987, + "0xfc88dba1647b7d43ef19a1a38780dd238228dd46beeb50ea0e00853f06c78471" + ], + [ + 80988, + "0x7e383c88765b3c499b05b60cdd4470cf9b23add0f6bd33990f207a43a51e431f" + ], + [ + 80989, + "0xa133fe8b34c9e839bcd4a75c2a5858874fa19d8bda43cd1c1af781cd2689753d" + ], + [ + 80990, + "0x466a4d175cb847a9899d98a6863900be93ecfec2f1681c8595db5a62b843478b" + ], + [ + 80991, + "0xb9b1bf393868e53ec5da954cb87fb65eef79ab9caf59cfb636f6cfd9c5c92f8e" + ], + [ + 80992, + "0x2f9eb9f04e16f420d5b98d6a28e25494207bc05affb6b914a1b98fb2244bba75" + ], + [ + 80993, + "0xc73829f98810dd017f45a0bd71ad7db3d07a1f5ff0a57c07994fa095989e6ef2" + ], + [ + 80994, + "0x9cdecf07b8f03979fdcb8f7d4d2c5ef4a13adc19edc0bb0c87a8e0beac6ff3be" + ], + [ + 80995, + "0x978c9e3471a4c37fb6aebab56d41d660720681073dfc68e5bef41d594d67ac8b" + ], + [ + 80996, + "0x739dad0614f0e852037aafbab92b4072f0f9dc9fc30c489465d67a18023f80a4" + ], + [ + 80997, + "0x1fbba34f589373e32f170a1cc3bf5c69083e8e2af5167e6ff38eab89ca2c0a72" + ], + [ + 80998, + "0xb58b192d29e3605fa4006c019698a7fbdc14639784ceeac087ae09148514e68e" + ], + [ + 80999, + "0x77f67acdb2afb4754c6e31005b27aea691cd3edce3094f87238df3f887e1bebe" + ], + [ + 81000, + "0x50e198025598c36879156ee586cdaa18bfa620f117c49baba1b2093d77b83831" + ], + [ + 81001, + "0x416e7d689843b2ddabbc9efd9b6f183b883dcb28e857977cd130f989ff7c72d6" + ], + [ + 81002, + "0xd639747fba2f48c62f7aeb91ead4b7ffa1292d55bd15ebec58879bc51b05863c" + ], + [ + 81003, + "0xc43806d1ab9e6d84be147ac3b74130fc2dca01254b4a847eb2713c3e7c1cd11c" + ], + [ + 81004, + "0xd854d522a33e98313a872fb4c47675f54eeed084d99794c5c72aab4464ee70b4" + ], + [ + 81005, + "0x752458c4e11fe4d564250bd4bb0122c1a1a4fabb124942ed9c3981887d612854" + ], + [ + 81006, + "0xe3eef7ded5487286de24e9802f5dfa1afe7c477666f2a26fca7606d3b9b8694c" + ], + [ + 81007, + "0xc525b44a5e888c3bdb0f66edc5bb4dbe9ce9116279b01fa510037760319ca15f" + ], + [ + 81008, + "0x9552445c57b6e2cd01d95bf4013bb913ce1611e5d707f71e528ad7522e830806" + ], + [ + 81009, + "0xd0eba42e3dd5cb37d42d6891eb47e5e2bdaeccbf9ddabeba06a09cd2b4b15f55" + ], + [ + 81010, + "0xcc6fd1b1758929e8a94d3bfc1b4c5444c55023f8939d4b4fb539b523219b00fd" + ], + [ + 81011, + "0x826ee09096704d3798aa4fccfc02cdedc59bb7472332b624d3716fa3aa14e612" + ], + [ + 81012, + "0x3f82182c620ce4d5ac0a776b0ec52b590f7d30c3bbea039d9a6ec3c6ec8a4aeb" + ], + [ + 81013, + "0xbdc7e4e07b4fc9245d8f5e7d7b207af203ce271b96912963aae2012de4d0dfd4" + ], + [ + 81014, + "0x797fd61f65fdc95f190d0dec1ad3eee5c98211c4529c83d013236838eda13d67" + ], + [ + 81015, + "0x7928f9107a62bd4d597894701cf4bca71365052e3560c23af7fdfa3374105b45" + ], + [ + 81016, + "0x78dc4dfd56d144d629278047ce28d8649e02b55768a137cf98260e40e6c35989" + ], + [ + 81017, + "0x3cd444ed1bd6f3012197839db12d566729ca5fb8c133af4353afbd756104c03c" + ], + [ + 81018, + "0x26350e310daef399afac9ac19b106d5f160d8dc6921fd289b196b4e979742bef" + ], + [ + 81019, + "0x99e1715d9ba9a3d125e77357595bd918766a4d46c20d9db1dfa9ebbd002a7bfa" + ], + [ + 81020, + "0xdf16ba27999f8dacfde10d403cf9982fd3b7545793e678219f09818e04ee64d6" + ], + [ + 81021, + "0xc4690b3c4b89b7282988cd2f71c9e43d3c17d081f0c735d04fbd01247a74bd6f" + ], + [ + 81022, + "0x94a1951db605f9a7dff70e2fd919e154f1291d31b4baa6f9e88d8f904c74fd54" + ], + [ + 81023, + "0x9846feb46081eafd35bd5f134cdf98f7a2a270b52a36779de7fb132250823a6b" + ], + [ + 81024, + "0x3716fe1bc70f07f53358e7936ff4b652a035eabe199265f99809dc4aac49d002" + ], + [ + 81025, + "0x841c89865fcece06f48c90ca283a1dbafb49e3fc422ff4cc5d3ec7fa68ac0f14" + ], + [ + 81026, + "0xcf2be45d7a88f2ffa7fa45366404b8b0af131758b7706f099ee8daa6b5d17b94" + ], + [ + 81027, + "0xf3c7dd48103bf10e13bea5800814475a543f4bf53958e1ad464b112d0c763d00" + ], + [ + 81028, + "0x6ebafde3e8fe7067903074cbeb854b9957bbc45859143f618ada972fccdc2c0d" + ], + [ + 81029, + "0x05146b0d5efdf34a927593169ecce2d2d112a3faa3b79cb2df1a0a2e6a4bd7c0" + ], + [ + 81030, + "0x33def97a4f3ca92c7af02bd19bb9703c516f6df8917928734953ec9fddae68ce" + ], + [ + 81031, + "0x18bbaeb286265fc8edb34abbfd3fade65196f09125dae4d31347401eb635b63f" + ], + [ + 81032, + "0x56aacb0d332a22fabb3bf8ac90b030fb82337b3b22b5be334d9a4f633847480c" + ], + [ + 81033, + "0x8d5d77464c966282353e2019598e0f1d77b5afb938475275f4a3b0d0c29d9c5e" + ], + [ + 81034, + "0xc8ccc7baba52e4fa56d488c78abf451919aeba4a029068fd364059633446c80d" + ], + [ + 81035, + "0xc295e160d3378e12e6b96de0eac43d903aa656f9255e58a3b0222bc229b18bbb" + ], + [ + 81036, + "0xd4cacd3eb6e8eebf312225c23284c15e347e73539816fe91954b482a9acb1753" + ], + [ + 81037, + "0x096f15eb2cfa8a6811e08ee533e3a9bfaf3447225c214157d33f27d532423598" + ], + [ + 81038, + "0xd2464a278a262d447f8d67feb035e8b9eaaaa849eb87fb330e1ac868e6b1d063" + ], + [ + 81039, + "0xfea1ffb0177d362f9e041268c395858091da08fb55de5c922493bf0f3dd19476" + ], + [ + 81040, + "0x5ab2d5c6f5789372abc97b4792b2c4e3476c70bd538764ba74d684ed102fdd0e" + ], + [ + 81041, + "0x2f5be38f279790c0e69e554fcaa28f82bd17ef645338539416db0c9252040f91" + ], + [ + 81042, + "0x844ebc844ba1c0c301a6a9189779e17e557f824d85d01247b4a0b1af49f50d56" + ], + [ + 81043, + "0x944855a63daa3b755a70f2ba5e6a956b037eeb210dd9e80cef73b7c814645adc" + ], + [ + 81044, + "0xf046f77959a6aa199f3d999b2c16dbeb1e9038c308cb1dcc9bba26141e0e5f35" + ], + [ + 81045, + "0x4e420f14ba78477c090eef99a098b62dfdadc3a61f7c0a5499797380763693db" + ], + [ + 81046, + "0xae10dd9883739a2cd9e95cafa4f9cc7b2e38e477207ec49fb3fe64e3fcf21613" + ], + [ + 81047, + "0xa9e686e36215cf1053ac16cb445e537ef8a633aac308ad755716f6374c3c1238" + ], + [ + 81048, + "0xa7ab60b451ae4a7fa639faa0c8b1f2ad7bb6bcb853242797f4aca608b2e69193" + ], + [ + 81049, + "0xdba699a61f7e4b3b8356a7494afbb73237f00385d7ad279a77002bfd451b2643" + ], + [ + 81050, + "0x42515b83488525c1c8e72417df7361aa61a62e3f78892446b06d6d6bccee6ada" + ], + [ + 81051, + "0x9a3442c671ff6750549ac81e1a48d7ba70e8248b0f1122bf37a3767666cac4e6" + ] + ], + "rewards": [ + [ + "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "0x326945a07a4d8400" + ], + [ + "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "0x326945a07a4d8400" + ], + [ + "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "0x326945a07a4d8400" + ] + ], + "proving_pool_credit": "0x25cef4369cb13800", + "payouts": [], + "blocks": [ + { + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "blue": true, + "txs": [] + }, + { + "miner": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "blue": true, + "txs": [] + }, + { + "miner": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "blue": true, + "txs": [] + } + ], + "pre_state": [ + { + "address": "0x0000000000000000000000000000000000000210", + "nonce": 1, + "balance": "0x0", + "code": "0x608060405234801561000f575f5ffd5b506004361061003f575f3560e01c8063aa67735414610043578063dea5c2e014610058578063fe7e05d51461009f575b5f5ffd5b6100566100513660046101e0565b6100ca565b005b610083610066366004610211565b6001600160a01b039081165f908152600160205260409020541690565b6040516001600160a01b03909116815260200160405180910390f35b6100836100ad366004610211565b6001600160a01b039081165f908152602081905260409020541690565b336001600160a01b03831614806100f957506001600160a01b038281165f908152600160205260409020541633145b6101635760405162461bcd60e51b815260206004820152603160248201527f446576656c6f70657252656769737472793a206e6f7420746865206163636f75604482015270373a1037b91034ba399031b932b0ba37b960791b606482015260840160405180910390fd5b6001600160a01b038281165f818152602081815260409182902080546001600160a01b031916948616948517905590513381527fa47563c41dab010f91a8ef9dc7ac2bcdfa0ef2af697e575048e71e6eec60dda3910160405180910390a35050565b80356001600160a01b03811681146101db575f5ffd5b919050565b5f5f604083850312156101f1575f5ffd5b6101fa836101c5565b9150610208602084016101c5565b90509250929050565b5f60208284031215610221575f5ffd5b61022a826101c5565b939250505056fea2646970667358221220cbf48f5aa6f911f83b3c2b09adf8c418f5da6fc530c3d5db9d4ba0be24c4bf5864736f6c63430008250033", + "storage": [] + }, + { + "address": "0x0000000000000000000000000000000000000220", + "nonce": 0, + "balance": "0x13cbc778560ccb37bc00", + "code": "0x", + "storage": [] + }, + { + "address": "0x0f002c928c363c7041cafe788f099cf7d35452a1", + "nonce": 243, + "balance": "0x1bc36179476e9a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x18524811fa2e76dd0770d29308b51bcbdc0499f9", + "nonce": 242, + "balance": "0x1bb8aa483e652c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x1aca71f7872aebc85b4d4a9ad47a09abcc624d65", + "nonce": 0, + "balance": "0x2860cb14365c048800", + "code": "0x", + "storage": [] + }, + { + "address": "0x233d639f53225ea5012dc01ceb0a5a30891021cd", + "nonce": 264, + "balance": "0x1b8eea59c54b4200", + "code": "0x", + "storage": [] + }, + { + "address": "0x27fd8475c3db352fdddeaf29823e00c7d33ca39a", + "nonce": 248, + "balance": "0x1bccb45d83401000", + "code": "0x", + "storage": [] + }, + { + "address": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "nonce": 0, + "balance": "0x1e095da9ca2c1432800", + "code": "0x", + "storage": [] + }, + { + "address": "0x46681948060140945958068b373f94581e049d4f", + "nonce": 240, + "balance": "0x1baadca68fcb1a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4cb00bc539538d2ee7172474f2b06f021944ece4", + "nonce": 258, + "balance": "0x1b3e54c72c629c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4d7c0d6f3ad5466691649e98ebb59cd5b5194687", + "nonce": 0, + "balance": "0x1b598fbe8e08549c6000", + "code": "0x", + "storage": [] + }, + { + "address": "0x5a5e606bda0fe1b2b4298aadd55cc6ec57ff698a", + "nonce": 0, + "balance": "0x648cba8883ddb562c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x6613ca66c14338db771d01c238ae140c71310535", + "nonce": 243, + "balance": "0x1b98d3921da34800", + "code": "0x", + "storage": [] + }, + { + "address": "0x66577de387feeec96c10f7f8748f13912589c78c", + "nonce": 240, + "balance": "0x1bc37e3123b20a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x7e7b1db26094aa913933f65127b328c1861ff5b0", + "nonce": 256, + "balance": "0x1b89d2467bcd2600", + "code": "0x", + "storage": [] + }, + { + "address": "0x90acb15171deb958d3191d5c02d807afc367f0e8", + "nonce": 240, + "balance": "0x1b9e3c5e84fb6c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x9f296bf64eb7051912c9c8112da3e3ea6cf546e1", + "nonce": 266, + "balance": "0x1b69b11e00d88c00", + "code": "0x", + "storage": [] + }, + { + "address": "0xaa19223d63c82bf9be5543ec96e5504834d17bc8", + "nonce": 246, + "balance": "0x1b9abb1eb524dc00", + "code": "0x", + "storage": [] + }, + { + "address": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "nonce": 0, + "balance": "0x55f1d5570ee67d0000", + "code": "0x", + "storage": [] + }, + { + "address": "0xc2faa4a2866422a865c9b3b3cb5988bf4d07f0b4", + "nonce": 0, + "balance": "0xc05b9fc73002b32800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc45d23f49451faad9afadf877feead626ea8d083", + "nonce": 0, + "balance": "0x9ee8f806428adc9800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc8621de5921418701619ba715044575073d2ca62", + "nonce": 243, + "balance": "0x1bbb8989a888a800", + "code": "0x", + "storage": [] + }, + { + "address": "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "nonce": 0, + "balance": "0x1f29d0e7a3eaed38a400", + "code": "0x", + "storage": [] + }, + { + "address": "0xd958300657e7931b06fc6f60dc4bfe7d8e8ea7b8", + "nonce": 252, + "balance": "0x1b79a326be219400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdadb11b01f6004eba6da3846a0e1bf6f4f71fb82", + "nonce": 244, + "balance": "0x1b900b01183e5600", + "code": "0x", + "storage": [] + }, + { + "address": "0xdd442fcbb964a3afdc90d49b408e8dd296fa86e8", + "nonce": 0, + "balance": "0xafd0c09109be7410400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdfaea67368f3e3753397d878f97efe6aa8020c2e", + "nonce": 16, + "balance": "0x4f20fde2ac0f98ec00", + "code": "0x", + "storage": [] + }, + { + "address": "0xfd4fca79b266a73a7defa776b04c7e15bdc70e86", + "nonce": 238, + "balance": "0x1bb5cf9909338400", + "code": "0x", + "storage": [] + } + ], + "fees": { + "base": { + "pgas": { + "version": 0, + "cycles_per_pgas": 1000, + "intrinsic_pgas_per_tx": 200, + "modexp_base": 1000, + "modexp_per_byte_numer": 10, + "modexp_per_byte_denom": 1 + }, + "block_proving_gas_limit": 30000000, + "shard_proving_gas_budget": 7500000, + "min_execution_base_fee_wei": 1000000000, + "min_proving_base_fee_wei": 1000000000, + "initial_execution_base_fee_wei": 1000000000, + "initial_proving_base_fee_wei": 1000000000, + "base_fee_change_denominator": 8 + }, + "v1_activation_daa": 18446744073709551615 + } + }, + "plan": { + "shard_budget": 7500000, + "consensus": true, + "shards": [ + { + "index": 0, + "tx_start": 0, + "tx_end": 0, + "over_budget": false, + "pre_root": "0xad123f2337b9d9df433d08e7f7ef6a19f885ab3eb100b19961aa051efa79f769", + "post_root": "0x7300a34f1521fc454efa0330508f4cf762541806c3d766a860e31900649d53bc", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "link_in": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "link_out": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "witness": [ + 3, + 0, + 4, + 15, + 14466 + ] + } + ] + }, + "expected": { + "pre_state_root": "0xad123f2337b9d9df433d08e7f7ef6a19f885ab3eb100b19961aa051efa79f769", + "post_state_root": "0x7300a34f1521fc454efa0330508f4cf762541806c3d766a860e31900649d53bc", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "tx_commitment": "0x0000000000000000000000000000000000000000000000000000000000000000", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "node_state_root": "0x7300a34f1521fc454efa0330508f4cf762541806c3d766a860e31900649d53bc" + } +} \ No newline at end of file diff --git a/proving/fixtures/chain/block-81053.json b/proving/fixtures/chain/block-81053.json new file mode 100644 index 000000000..f12e658b0 --- /dev/null +++ b/proving/fixtures/chain/block-81053.json @@ -0,0 +1,1316 @@ +{ + "format": "igneum-prove-fixture-v1", + "source": "live devnet export from node 1 at tip 81076, 5 October 2026 20:06 BST, proving v1 chain fixtures", + "block": { + "chain_id": 4463, + "env": { + "number": 81053, + "hash": "0x040acfe3dc3ca820cfe2565433d200cf9b2399ed3bcd1e04598fc05bcb1379c5", + "parent_hash": "0x93319ca0399a5b00e77bbde1cf067c1f4f823cbeffaa1103c7953abd1b8db6c5", + "timestamp": 1791227136, + "miner": "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "prevrandao": "0x03d007d8c2ab0d01d4b3189b6b2cb53957bf246da8b1ccd2af308be6b105b9c8", + "base_fee_exec": 1000000000, + "base_fee_proving": 1000000000, + "daa_score": 0 + }, + "block_hashes": [ + [ + 80797, + "0xe099536cfd64654690afc152b97866b32da6bf81f959a6c61f9e43bfd2ca2d9a" + ], + [ + 80798, + "0xd6af2cf6184f15ecca77ee9e7e8d2ca718d88f9cea39b475713ceb0e151ddc36" + ], + [ + 80799, + "0x9c76f9f6036c7035fd52a7a4e3fca99451715e729fcc9fc424487faee9a6134d" + ], + [ + 80800, + "0x9a046cd99eb89c551dd0c698f1969069c02ae36cde4140ff4b864d624fd9dff7" + ], + [ + 80801, + "0xc9bffa329d649c6540518cf249e25e845743dc01fcbcfdfb8c798f1455efeae6" + ], + [ + 80802, + "0xb124f2de3c17a985f10c0b58d2283631bcc12524de8d79b0955aecfd2d7f78aa" + ], + [ + 80803, + "0x97b21d6cd3f1b321d2c9729062c305a7955b97646d76910714f899c75d275568" + ], + [ + 80804, + "0xaa3f0e114e156f706e11a701fb940bf54e89b09d83953be7f6733478c7a33093" + ], + [ + 80805, + "0xe5a5868b02ec30304079652a8d5b8af4a9da0a5a676615c1a10d833feb4eb0f7" + ], + [ + 80806, + "0xd97c5c4288f56955e1a3e134eb63250e1becb907b0fdde9ce1383d8d2c8a3939" + ], + [ + 80807, + "0x8f7bf1128e5d03607cca246b86e63cfb4bfe73e74395c8660d027172376b5c36" + ], + [ + 80808, + "0x5d5fe4c3bee7d758986da5a433cfe5221be392419ac22e5acddd701cd12c8142" + ], + [ + 80809, + "0x9b0d0a871ce8eaab11a185a4344f9f48a58b7111f94a8f5221d557722cc5ca5a" + ], + [ + 80810, + "0x58cad4ab776ba7eca94e813323e28a6467a7fcf39d5aa2069e9dfef6f8058e5c" + ], + [ + 80811, + "0x706788247d6c05475eb3bdbc6f181e76429b57a97f168004832fe5ba4b20c7b7" + ], + [ + 80812, + "0xdf26fb97f999b1016865b3fee539e72e3f66d6d228b2aec7e2f59ab93e5a9ee6" + ], + [ + 80813, + "0x8e612c07fdb693a44353a08899e8beae8c93fb1b85845ad965f96e7d2f77fed4" + ], + [ + 80814, + "0x806f4c29ea28af1a3ce0c3ba775c776006ba8cf4b308ee617dc207dc3f740b29" + ], + [ + 80815, + "0xebd8e602049ddc2f8d11d5edc8ee837eefac77418e73e9bad97f6f1c7a85e8aa" + ], + [ + 80816, + "0xe2b069b77ac05cb0a95ae419966feeb20660e4ff3b9a68d044a5721e1525dd86" + ], + [ + 80817, + "0x1ca5b79e26dc425a94c3a6bfd339516caa0a7f8e5e2a62389bfa466186bca495" + ], + [ + 80818, + "0x7cf382f956597cbd36917036dd15b774500f97714e50587a35ec9617662052d6" + ], + [ + 80819, + "0xe70494836c5b5c68b6a167c868db166f9630140499f96e6660d3e3928cad6a14" + ], + [ + 80820, + "0x15f73eefc39e00f9d09390764f9ea0069908062c563686e73b49ec3919c24be3" + ], + [ + 80821, + "0x7f2bd2cf68fc14e6a5debaac39f368534657c41cc587402fac1156d24998e55b" + ], + [ + 80822, + "0x2641ec1c3397b06ab5855dcf6ec2e23de301f5831015a32ffa9f29e18cc2f084" + ], + [ + 80823, + "0xf643e963f5a62a49e935b4ef51aab9bf2689b0b9ec483924d7f42934c8ce5e07" + ], + [ + 80824, + "0x867ee0219bd30fc3117c79295cb1075ca21281fdd285467ae128aece1c1dd9a7" + ], + [ + 80825, + "0xab6bb19259a039d30c3022c9da3494af8c9a541dce30895df083252894c433df" + ], + [ + 80826, + "0x155511de8d34bef351ba0ae7af5bccf515e6e490135130aad5ce2eb02e103d75" + ], + [ + 80827, + "0x4e297a53901da9b364f4dce686f016f4039658ea8832d591dcb21aa44acd197b" + ], + [ + 80828, + "0x7d07b8bc9b06b20e7b56f96fff9051590caac01159d39b3f95bd592a4627daf4" + ], + [ + 80829, + "0x8392a1875388baec163e64854d6ee7664a31017402ee8e269ddc8aa42ca06647" + ], + [ + 80830, + "0x5ab4e5c586d2fb4272941bf26c15288cc129ca1092d3aaab82c478edd1942e29" + ], + [ + 80831, + "0xb0100f0a559daa435128cae74f47f9353ca2a93c2bdac8dd9d12a3130ce6a818" + ], + [ + 80832, + "0x326d1388d0911871c5087925500c4ee5091624331d07659b9d2fa09e83f5884a" + ], + [ + 80833, + "0xc01c186cfcebb8d53953589ce595afef213d46e481b8b5dee7e0a71abd773f6a" + ], + [ + 80834, + "0xe845934e79eaeb8a8ee0b9396b3a13f13c8ac45a0d4e81017c73042789013c38" + ], + [ + 80835, + "0xa45f8f03249624a2d35b54eb83f4edf54e8d1acd061fee3965d9164a132c72aa" + ], + [ + 80836, + "0xe72a59a64af9aa223b5341819a5872f8626908b5c1857d0617706f6999632e01" + ], + [ + 80837, + "0x94063f426663347bb4a3e79f4dedeb76fbce6fd69117a2c74ad1fb4714b856dd" + ], + [ + 80838, + "0x157ab5f9e7fbf98b48af049adcfb4a8e071253ec95b991d031a91668ff3b8d57" + ], + [ + 80839, + "0x5bff632aa54b1a71a7d9a2a57d9c42ef223eb6efe86c72f9aa8dfd67ab7b218a" + ], + [ + 80840, + "0x7f5d1952ad8ad27dd17bed3354395b646468985bae0341b4048376a425e43076" + ], + [ + 80841, + "0x8e69df81fd1b0f186813304260dda1d0c9dad5792126e69d021e679b837b01a5" + ], + [ + 80842, + "0x6733c5472e38830dfb273a0e0802924c1b60a2cc916d6efaf82785e1bf960724" + ], + [ + 80843, + "0x0838513dc8264efffb2c11f8c00ebbc9c57aeb67dd7be93f18e1c613872bc929" + ], + [ + 80844, + "0x8c2f26b17db95bb20983ac41df7a5fca3e3826d2e57831eba13bc3175daeee34" + ], + [ + 80845, + "0xe3edcab823e03598c66e16c5743f7a0917735c975ac946db4cb3ad862db40a21" + ], + [ + 80846, + "0xd075510e75159d941946e666e2c75604a3f3c959523cbe2c0f025427ae85d4a2" + ], + [ + 80847, + "0x57b4b6c0a6ca2532515653445bc872dd44e621f9774f3dd4e2000bf82a0f510a" + ], + [ + 80848, + "0x38a85c99bc3e489a4431ba9a125e74f4d807c630a239419287e039f05cc91ead" + ], + [ + 80849, + "0x3c95d08c99482d57bbcb9fd333cae4010628fc8b1c96b3ad4a68d8f3dd21695f" + ], + [ + 80850, + "0xe767527430f19da2c2b21bc648e33ed2b241f5b7c9413649d9c5a555491a9253" + ], + [ + 80851, + "0xe9941baa9f3378854aaff5563340d130700f0753197c7373229542c56ac91a8f" + ], + [ + 80852, + "0xb92e88ec7cc8c340959d8d0cd5434eec5c1033b389b953b02cda103d53292df0" + ], + [ + 80853, + "0x39f01f35ba2ea7b45e3eec1681c44bd85b066a6a8521a10fe6cec5d4d98aace8" + ], + [ + 80854, + "0x36456a5e36a9a66488656bc0a92705632326e59ca40bd09f6d5f99ccc976972a" + ], + [ + 80855, + "0x12e52067750279b8bc34b03cd21bdb2809ea18a3aa68739623780aaf1150129b" + ], + [ + 80856, + "0x64951c693d7f246647fac0df504f35ff9812416f966afaf4fecf7ca92c0e0b24" + ], + [ + 80857, + "0x3918b797fe9d8cd6c34a0b315aaef550de15f4d30dce932f3be8cdfdc27343a2" + ], + [ + 80858, + "0x86ec8713a68034acb9f34b015a48644c7562766c1760927a299e517ae8ef6f8a" + ], + [ + 80859, + "0xfe75a4c99fc17131b7aae61dd0a588b7d7019e6c662aab8c7d6a8c56370fe07f" + ], + [ + 80860, + "0xf4afa8f14a0f814639c88dea304a8ccf76450ee56f4fbb89caf29dbeef80399b" + ], + [ + 80861, + "0xd4c72bc192643e20658ca54b730ef865d1c29ebfb6d1f78263c796a23b6965b4" + ], + [ + 80862, + "0x7a73b7302ab47b7acc1c5359620df9026fdc08d78c4173bb4e45e73a66651767" + ], + [ + 80863, + "0x9033ce568146d2c69139d3674fa7e18cb8319abae5e6cd24f924f750d2ca3446" + ], + [ + 80864, + "0xe4d8da84ba631f233ea7dd1336a264f499083b08290c82b717f79bfb7ffa09a6" + ], + [ + 80865, + "0xc53ec9d6b11d592921f5cf850584ebb562119ea968a5fa7bb00af381d6eaef46" + ], + [ + 80866, + "0xedcbb7d2f9626af53341464fbfc54e838e6417b7ffcfb0c2f9758641c98d5f9e" + ], + [ + 80867, + "0xbe5f726816d43ca5bbcb70f13901e60ef9040debdfaf2a73a0692c3cbef578ec" + ], + [ + 80868, + "0xb7cf72378b14c086afb4094b9b0121fcf7c6929bf2c2134b5189862644e2e506" + ], + [ + 80869, + "0xa29e21de7e54defda356f088dea38a06cf1a1a75bf93c14e30afdba213df7ef3" + ], + [ + 80870, + "0x504d3f83a27c37c5a15cdbbdfc03285f53071ad01d82bf9be2d6f862548a9a8e" + ], + [ + 80871, + "0xe02107fefe75bc8d1abb0443bff2e7a658e3a9b2fccd5157c1a7c7ccc995fe17" + ], + [ + 80872, + "0xdb92e1b449e7670ea719a6d68f496965620656ff5d25a0521f1873eb8d7ae1ab" + ], + [ + 80873, + "0xf12689607f86ce8f762df5d67054a00b8b580300511a221ff5addcf732393988" + ], + [ + 80874, + "0x0e0fe961c78fe52f44f6742795bae3e96713f4618e4a95623045056eb9e046f6" + ], + [ + 80875, + "0x7fd5bb64e800686ae28c3b55613d27f181dbd91e35d342d33d7b3e5d4df17d00" + ], + [ + 80876, + "0x83c2feae4eacf2ada1c9a630a97614a6296df6b89bfb5f954c3e306b16f0ab2e" + ], + [ + 80877, + "0xd1f3af4d6464c5bc9416d73a291bcf3e1df4fa1a396bb3e9dff92dcd451f75d9" + ], + [ + 80878, + "0xa4cd4b49e1d475a5561a91a0de592ba0c00f63b39a61eab150c6486d1bf2b916" + ], + [ + 80879, + "0x714904f162c3b9013a9defd79d4552b58c5ca8747cc94187079f2ebefee2f0d9" + ], + [ + 80880, + "0xb2263d12431e74591e2843047d79933820d2dd352d5b5c977698b328cd59f46d" + ], + [ + 80881, + "0x67977d72a573f121d9a818dd035959cd502e5bec07a4a5c30a28fc66d4fdf964" + ], + [ + 80882, + "0xf1d5f1b271cd23a79f1ab32655418e4b664663c4a77ebaf0a213fc2456e59c5a" + ], + [ + 80883, + "0x91a9f94d339407747a77d79a9c2be608744d4ecc8b84699d0120814f55a5060b" + ], + [ + 80884, + "0x8f709bf4582d3a912e6006f6c417459bbbf182af2b500b4376cd8dff78201c25" + ], + [ + 80885, + "0xbc66848a66fc038200186f09665540b5905a350a3149bfdddf9575eabb2f0a98" + ], + [ + 80886, + "0x8827e43b9a02e7424513e6d0fb6db7e93cc5555fac9b832584570784708a8f90" + ], + [ + 80887, + "0x7c76b6619ea66b8e7fa2791ad1a3dcafe863388d14db50dd022f63748e2b6ccb" + ], + [ + 80888, + "0xace6e76c69db6530277c9749ff3e77a64c6e7123790fdf5f248cc67b761e10bd" + ], + [ + 80889, + "0x8f96a0d2f9524f0b2bf9f520f35516969fba795e005f003dac5721cc37c6b5ce" + ], + [ + 80890, + "0x68a635c0d7ba10b3fff9108deecc4e9af70d760d688c5614a000503b1f2f006c" + ], + [ + 80891, + "0x47dc91daaf1758699e90f8906c61bcb6dd020b3726a1e0d15a7fe5177bfab696" + ], + [ + 80892, + "0x79f1ff991dd34c0384b9ddb64a7f3791116891447f2c960de5c6d80f9a705b5f" + ], + [ + 80893, + "0xf41832c4af88cdd31eec42b233ccfae51a150cd41baf59e9dc6b64dba2b63c7a" + ], + [ + 80894, + "0xfaf2c58e9ba01047e4cbe2c147a4614d46ad39c1d64cbb43d1de60f9c75286b7" + ], + [ + 80895, + "0x65a671bcbc9b6f905811ce62594feda18a12b8c6ca1c72081c1dde3f6227d1a7" + ], + [ + 80896, + "0x815d16976077a9b599fd5d6a82eda30a0bd1d91c66718528aaea9d4dcd12274b" + ], + [ + 80897, + "0x710c0c1d4d44ed7925343b10c7ace3216741b31a91506fa0f1ad80c96c049b14" + ], + [ + 80898, + "0x5fff52604275c206678c800d95cd0569325ad67413b50ac15252c3440856cc78" + ], + [ + 80899, + "0xe67fe87df50a63d348d300de16353e560c1b86f1a79a0bd09e2a7771ae7e92ba" + ], + [ + 80900, + "0x5d854ceaef9de086a2361cfa1f829f843aac00181fa5ae27e69171c198853895" + ], + [ + 80901, + "0xca45572516a2d4a948cfe518d0bb5378a4d23b800522f515e59176c31c069160" + ], + [ + 80902, + "0xa724d3e05f16d3d8c97296f59802c7f9ebba4284808126d955d6fc96dc1c4729" + ], + [ + 80903, + "0x552061794bddc5185033e936ac8ad07d3a315c7a3d876989b2d61328ad0a3128" + ], + [ + 80904, + "0x23c1ee3ff43ae0d42b9b3f0d2c6c2fef6c8ab8fdec9772540b563cfebbb77a2c" + ], + [ + 80905, + "0xc7e33046cb4819a17a21a00e9c0efa55765a974d8a2364f4bed9a16034ec2828" + ], + [ + 80906, + "0xa7b5712f7d22449a6f6a8a55a33c2f76cb54be86faa2f0d892c8bf6aa9fe59d9" + ], + [ + 80907, + "0xbd8696f54d266714d694b41d3b6466ef242e999938b9effd2de035c3f2c5677b" + ], + [ + 80908, + "0x7ca324c51788d4cc173bd3a013702f80e68524b3b005299be9012621dc8e82c4" + ], + [ + 80909, + "0xaea0e3a560abbe0de857592d37bfd43bd986d1cd461f43139d951d123140d495" + ], + [ + 80910, + "0x640f0ed5ec76969f53d90d5e502f6fce0ec65ab28a728e92a81dd27c466dad09" + ], + [ + 80911, + "0xf4cb4beab37872930f84e37b38c15a0a5e5f5d95927c91284ae880337abcfd9a" + ], + [ + 80912, + "0x65121e870e1668a0a55c9a50d31166294cb8e57bc007897640a7414f0866ec3b" + ], + [ + 80913, + "0xb3ea0a195e00db811104db7362b555fd1ab8401814ee841ca7f4d3344843708c" + ], + [ + 80914, + "0xf92735fd64122d7f6f6df6f20f83fc0793675c81e25e54dc64b3ca027d8ef9a4" + ], + [ + 80915, + "0x78e07fd2fcedb51175785be4657b8d2a554c16c1d2cc396ca3738dbf6f9373f7" + ], + [ + 80916, + "0x0ce35f9c3a736ff2937e6b7addcb632855e347ab3c5c08b74d2dca5f82265426" + ], + [ + 80917, + "0x007b1e317c0d9e7565260e1fe7ebfe031de4291f6ddad75210f8563479e84178" + ], + [ + 80918, + "0xbdb4b040d7942b8bcd206e4e12632892d9cfb87c905f760f0c918dcb682d4eca" + ], + [ + 80919, + "0x1eb8ee38b4d246a0502cc56c8948ec6e17fa55b02002c9916e1a3d9867e3c852" + ], + [ + 80920, + "0x2bd9fea5702cc9c0b96fa3b34559c4a72c5ff4f1d572d9e702390a13b0b1c145" + ], + [ + 80921, + "0x12ae91b4df38f549004ab8c4482d9ed1a2841e43f649ee1a89156f9accea51ea" + ], + [ + 80922, + "0xb960711112f51fc11689061a66a9b4ed3c0bccb3a6f027bd2cceeb2d3786e780" + ], + [ + 80923, + "0x4d474df4920e9136e7564cf0e9ebd03bbcd8cb8bdb38c68b8a9c35120888bfea" + ], + [ + 80924, + "0x73d84f1013038f10bebf6606486c1bb6c6cf530a4d35f482be91df74fe15c93e" + ], + [ + 80925, + "0xf08dc8a601c46c43bb6f94edc1b0734a61e7f8efc82d915b33136fac63d5a407" + ], + [ + 80926, + "0x2b695a3c4d05e233cbe18213d50b10fb9d83914aca9a935aa6f6639b1f54c32e" + ], + [ + 80927, + "0x893a6295ba2f66f083feaa39e0570ccbd446166c2e15238d59a911f2826a47d8" + ], + [ + 80928, + "0xc2b2870e39e8344e18296012c4d846b78cd3bdc80145de6b595527c8845ac57b" + ], + [ + 80929, + "0xc32bf651e3ea72bd6808db5b4dbe4577078d418404221809ac12897082ab8bfa" + ], + [ + 80930, + "0xf1c2f5ed0a7d6b7f083c1a68f75004fbfc929ef3f6ab3b46acbba373feecfdde" + ], + [ + 80931, + "0x51638cd9a09c8ee630706d736c81ea0dde7d97234a515c6064be20d0b92b8418" + ], + [ + 80932, + "0xf3cb85e3f551c4e21d0c6c28e940c6bd965c2a20b5eb3375a9d8d32a65c58c4c" + ], + [ + 80933, + "0x65f05787c7f239dcd4dd9cac2e5c516e25f8a8b898eec284531c16a8da83ed2d" + ], + [ + 80934, + "0x58567dd84c5ce96fdfb6ea5794201199d259374e4879e76df9a6622694de95f3" + ], + [ + 80935, + "0x2031a54f25cefcf00c464f36897691a73616f777582bfcd9c03e5649d3339821" + ], + [ + 80936, + "0x0fc1b95d2579a6cb03f21083da53a989db2491f37bcc339195864baf21932823" + ], + [ + 80937, + "0xfc501a03f2666c7c7db743b6a7296692789824828c033f2254e3565de63bccec" + ], + [ + 80938, + "0xe752390723b297a6d326cc9d10f5a3219b5eede3159f9ea1c7fc7f26401b02f8" + ], + [ + 80939, + "0xfdba580d13c6964b46e7d2a66f2aeee150174facaf1cbff66016832b02d9f796" + ], + [ + 80940, + "0x2221d3b867caf0c1d4993845a60ba8389558ccd9d48bf97f51331e87ff51ce84" + ], + [ + 80941, + "0xb4ba8a208861ab1db5302570f7d589adc5f4826812b4759d625cabe2b68231a0" + ], + [ + 80942, + "0x0bc24048f15b64e0c836c387539686906198ed0b3a7adbe2aed37f41eaee5577" + ], + [ + 80943, + "0x5c3230d24716dfeb57c624bc44f7d80b65c95f4b810aa945f8e5309e45327104" + ], + [ + 80944, + "0x798ef63bf3ed33e8608b899ee956fc3a0a0b359a4466141728589ea914d42453" + ], + [ + 80945, + "0xb158a6b683a60e8d3debc803a89bd18af0f1307941039304a2e33a02f237325a" + ], + [ + 80946, + "0xd420374f3f533e0136b3fb1211357adbdde913776fbb851492ca80980e77ce2e" + ], + [ + 80947, + "0xbf7fea59019b30fec39a94c29b026285e0c2d6fe034da2555095999ca694914b" + ], + [ + 80948, + "0x3d232c12ca3a428474f8c8e21992efc614775842c365a29ef44d9b7f24caafc8" + ], + [ + 80949, + "0x87d17a79601cf126d2b5c95f1ac989f8a90f2b4e9f3b13d1d721776ee2a53c24" + ], + [ + 80950, + "0xd23290b4b63c19aafd1edefd1cfcf34829bdc6efa723df17cce5eb7612b17a0d" + ], + [ + 80951, + "0xa8d8db2850de878732f5a7b95bc08ddbb944e1e8e8884a2a2a6a98275bb7663d" + ], + [ + 80952, + "0xb2c005429b5d1f35d006693509651726e67fd19d34ac04a3473df90d8bed52ab" + ], + [ + 80953, + "0xfcdb4ee3a4fa1273ec429c49704ac1afe6c8c6f858640155276661b19a5c4ca0" + ], + [ + 80954, + "0x9b988ccb0403ebfb0c15486ed6b63ea070121d7754f88d1f14f85651c083f240" + ], + [ + 80955, + "0x39c3d647ae0dd09b1bbf1699803e3c1e76448787baf12cfced826b8d75dd4b7c" + ], + [ + 80956, + "0xb24d2ff081df7d34392e460ddf8fdb8c7cc5ee502e8cae2a279474f329fe7d05" + ], + [ + 80957, + "0xf1f1daa9a70d68fea94a204260b298285ca15469cdc1855be3ef731ae44707bb" + ], + [ + 80958, + "0xeadcfd836598031d7d1a6f5c14028edeffc5baed420b2fbd343a5e21cfdd1fd9" + ], + [ + 80959, + "0xe1391944a4bd9d2fe474c8b12de8b3d151bf3afc97addb223ebaa59592ddcd94" + ], + [ + 80960, + "0xdef9733fd939d6b260fe671867ee7cdcd802332b9afdadbba0ac24b964126d29" + ], + [ + 80961, + "0x641d150032ebcce0d0d7616b010eb29df976191edfad5c8751da8533f78f7946" + ], + [ + 80962, + "0x7480f68a44e03c35a1ece1cec7402fe9fb5b4d34d4f246c967c5178e1a7fa330" + ], + [ + 80963, + "0xefae600a6e868a4213011b106f295f79e5c86186b8382143c0a402761226a9ec" + ], + [ + 80964, + "0xc6ce05cbc32d8a1d786baa6473cc65c29892a5a47e7474f384b65c567d425748" + ], + [ + 80965, + "0xd3a3e3c3c964a77a0821984903c958079b294b195b7b13e109f26f3bafcfa092" + ], + [ + 80966, + "0xd42712815a77eaca169de5c74653b4c007159f47fb5d50d63f57cfd3f2c5868e" + ], + [ + 80967, + "0x9b309f8c13534ed43e399d45a5bcfdb38f49c5633b458f6db71d39571e9d8d41" + ], + [ + 80968, + "0x74729f1a9de838f15def5f5c183b7d3596b03afde1ed6302cabacf8a677f82a2" + ], + [ + 80969, + "0x4c07a8677c028eda23e1665416eb54d364959f3210e682978082d560ebf2bea4" + ], + [ + 80970, + "0xe3f992775c7dfbd2485ccecf516ca7ee9961c897296190c45fd7abb1431ee04b" + ], + [ + 80971, + "0xa8492953412a007f2f7696922b661a27fe2ec42e566504b3aacae2f97a24fab0" + ], + [ + 80972, + "0xe14d6f60ece95189dc8eaed8a2ea5c2da62cbdf988588222a4e99d242ea3beb8" + ], + [ + 80973, + "0x5c7ca844c8eec0cce4e4640ba6fc486c1739eb44aa50de86f55008a4d597357c" + ], + [ + 80974, + "0x2b33e852101ee35f2c61ee7b35152521eefb11364505bf6903a964c93d4001b6" + ], + [ + 80975, + "0x5516b628a8ec4328bcfdf90785fb5d96d22d60b4db638ea1b93bd837d7c739a3" + ], + [ + 80976, + "0x9959b993b66dc1285d0ecd8c95b74fe70949c50d5a97d822ce866291f94f2926" + ], + [ + 80977, + "0x04eb959364eae151a41d4a8d970795da0bf8a0ef2f64bde753cbca4ea3d745c4" + ], + [ + 80978, + "0x22907520934a83b056017bfaa72a0ce92fd8658c81dc8236f3b78f1f7e055aa6" + ], + [ + 80979, + "0x59fae978e8063799b4ab6335ca7ac097a25cbb4d60f9af04eb2f647a92ead4bc" + ], + [ + 80980, + "0x9ef7533f2e9c60f4acf9618efd1c47b147fa95bbb27caac54c1bb19540befeab" + ], + [ + 80981, + "0x9986f742ffa9d17f641293e788e53cd904b6edba3e971603470186f61f934e77" + ], + [ + 80982, + "0xbfb394920e4bd707342ddd41644a5733b7041e517d4e88863924046cd512ec19" + ], + [ + 80983, + "0xa3fef379a91df7a254bc4897cf2ef61ff25b380b2142374ff40d91d5140ed21c" + ], + [ + 80984, + "0x71626767d14a0735e3716b45b847611dc7fe04c6d7a4b22e3a6031d621b6d560" + ], + [ + 80985, + "0xb2c03912458086a6b421a8c83cee657c9276a1d38db128799a3170cba30e3266" + ], + [ + 80986, + "0x8f8000a6e85c09716e8766c9bbb9df1287d46f54625ff139d5e98142a58c9253" + ], + [ + 80987, + "0xfc88dba1647b7d43ef19a1a38780dd238228dd46beeb50ea0e00853f06c78471" + ], + [ + 80988, + "0x7e383c88765b3c499b05b60cdd4470cf9b23add0f6bd33990f207a43a51e431f" + ], + [ + 80989, + "0xa133fe8b34c9e839bcd4a75c2a5858874fa19d8bda43cd1c1af781cd2689753d" + ], + [ + 80990, + "0x466a4d175cb847a9899d98a6863900be93ecfec2f1681c8595db5a62b843478b" + ], + [ + 80991, + "0xb9b1bf393868e53ec5da954cb87fb65eef79ab9caf59cfb636f6cfd9c5c92f8e" + ], + [ + 80992, + "0x2f9eb9f04e16f420d5b98d6a28e25494207bc05affb6b914a1b98fb2244bba75" + ], + [ + 80993, + "0xc73829f98810dd017f45a0bd71ad7db3d07a1f5ff0a57c07994fa095989e6ef2" + ], + [ + 80994, + "0x9cdecf07b8f03979fdcb8f7d4d2c5ef4a13adc19edc0bb0c87a8e0beac6ff3be" + ], + [ + 80995, + "0x978c9e3471a4c37fb6aebab56d41d660720681073dfc68e5bef41d594d67ac8b" + ], + [ + 80996, + "0x739dad0614f0e852037aafbab92b4072f0f9dc9fc30c489465d67a18023f80a4" + ], + [ + 80997, + "0x1fbba34f589373e32f170a1cc3bf5c69083e8e2af5167e6ff38eab89ca2c0a72" + ], + [ + 80998, + "0xb58b192d29e3605fa4006c019698a7fbdc14639784ceeac087ae09148514e68e" + ], + [ + 80999, + "0x77f67acdb2afb4754c6e31005b27aea691cd3edce3094f87238df3f887e1bebe" + ], + [ + 81000, + "0x50e198025598c36879156ee586cdaa18bfa620f117c49baba1b2093d77b83831" + ], + [ + 81001, + "0x416e7d689843b2ddabbc9efd9b6f183b883dcb28e857977cd130f989ff7c72d6" + ], + [ + 81002, + "0xd639747fba2f48c62f7aeb91ead4b7ffa1292d55bd15ebec58879bc51b05863c" + ], + [ + 81003, + "0xc43806d1ab9e6d84be147ac3b74130fc2dca01254b4a847eb2713c3e7c1cd11c" + ], + [ + 81004, + "0xd854d522a33e98313a872fb4c47675f54eeed084d99794c5c72aab4464ee70b4" + ], + [ + 81005, + "0x752458c4e11fe4d564250bd4bb0122c1a1a4fabb124942ed9c3981887d612854" + ], + [ + 81006, + "0xe3eef7ded5487286de24e9802f5dfa1afe7c477666f2a26fca7606d3b9b8694c" + ], + [ + 81007, + "0xc525b44a5e888c3bdb0f66edc5bb4dbe9ce9116279b01fa510037760319ca15f" + ], + [ + 81008, + "0x9552445c57b6e2cd01d95bf4013bb913ce1611e5d707f71e528ad7522e830806" + ], + [ + 81009, + "0xd0eba42e3dd5cb37d42d6891eb47e5e2bdaeccbf9ddabeba06a09cd2b4b15f55" + ], + [ + 81010, + "0xcc6fd1b1758929e8a94d3bfc1b4c5444c55023f8939d4b4fb539b523219b00fd" + ], + [ + 81011, + "0x826ee09096704d3798aa4fccfc02cdedc59bb7472332b624d3716fa3aa14e612" + ], + [ + 81012, + "0x3f82182c620ce4d5ac0a776b0ec52b590f7d30c3bbea039d9a6ec3c6ec8a4aeb" + ], + [ + 81013, + "0xbdc7e4e07b4fc9245d8f5e7d7b207af203ce271b96912963aae2012de4d0dfd4" + ], + [ + 81014, + "0x797fd61f65fdc95f190d0dec1ad3eee5c98211c4529c83d013236838eda13d67" + ], + [ + 81015, + "0x7928f9107a62bd4d597894701cf4bca71365052e3560c23af7fdfa3374105b45" + ], + [ + 81016, + "0x78dc4dfd56d144d629278047ce28d8649e02b55768a137cf98260e40e6c35989" + ], + [ + 81017, + "0x3cd444ed1bd6f3012197839db12d566729ca5fb8c133af4353afbd756104c03c" + ], + [ + 81018, + "0x26350e310daef399afac9ac19b106d5f160d8dc6921fd289b196b4e979742bef" + ], + [ + 81019, + "0x99e1715d9ba9a3d125e77357595bd918766a4d46c20d9db1dfa9ebbd002a7bfa" + ], + [ + 81020, + "0xdf16ba27999f8dacfde10d403cf9982fd3b7545793e678219f09818e04ee64d6" + ], + [ + 81021, + "0xc4690b3c4b89b7282988cd2f71c9e43d3c17d081f0c735d04fbd01247a74bd6f" + ], + [ + 81022, + "0x94a1951db605f9a7dff70e2fd919e154f1291d31b4baa6f9e88d8f904c74fd54" + ], + [ + 81023, + "0x9846feb46081eafd35bd5f134cdf98f7a2a270b52a36779de7fb132250823a6b" + ], + [ + 81024, + "0x3716fe1bc70f07f53358e7936ff4b652a035eabe199265f99809dc4aac49d002" + ], + [ + 81025, + "0x841c89865fcece06f48c90ca283a1dbafb49e3fc422ff4cc5d3ec7fa68ac0f14" + ], + [ + 81026, + "0xcf2be45d7a88f2ffa7fa45366404b8b0af131758b7706f099ee8daa6b5d17b94" + ], + [ + 81027, + "0xf3c7dd48103bf10e13bea5800814475a543f4bf53958e1ad464b112d0c763d00" + ], + [ + 81028, + "0x6ebafde3e8fe7067903074cbeb854b9957bbc45859143f618ada972fccdc2c0d" + ], + [ + 81029, + "0x05146b0d5efdf34a927593169ecce2d2d112a3faa3b79cb2df1a0a2e6a4bd7c0" + ], + [ + 81030, + "0x33def97a4f3ca92c7af02bd19bb9703c516f6df8917928734953ec9fddae68ce" + ], + [ + 81031, + "0x18bbaeb286265fc8edb34abbfd3fade65196f09125dae4d31347401eb635b63f" + ], + [ + 81032, + "0x56aacb0d332a22fabb3bf8ac90b030fb82337b3b22b5be334d9a4f633847480c" + ], + [ + 81033, + "0x8d5d77464c966282353e2019598e0f1d77b5afb938475275f4a3b0d0c29d9c5e" + ], + [ + 81034, + "0xc8ccc7baba52e4fa56d488c78abf451919aeba4a029068fd364059633446c80d" + ], + [ + 81035, + "0xc295e160d3378e12e6b96de0eac43d903aa656f9255e58a3b0222bc229b18bbb" + ], + [ + 81036, + "0xd4cacd3eb6e8eebf312225c23284c15e347e73539816fe91954b482a9acb1753" + ], + [ + 81037, + "0x096f15eb2cfa8a6811e08ee533e3a9bfaf3447225c214157d33f27d532423598" + ], + [ + 81038, + "0xd2464a278a262d447f8d67feb035e8b9eaaaa849eb87fb330e1ac868e6b1d063" + ], + [ + 81039, + "0xfea1ffb0177d362f9e041268c395858091da08fb55de5c922493bf0f3dd19476" + ], + [ + 81040, + "0x5ab2d5c6f5789372abc97b4792b2c4e3476c70bd538764ba74d684ed102fdd0e" + ], + [ + 81041, + "0x2f5be38f279790c0e69e554fcaa28f82bd17ef645338539416db0c9252040f91" + ], + [ + 81042, + "0x844ebc844ba1c0c301a6a9189779e17e557f824d85d01247b4a0b1af49f50d56" + ], + [ + 81043, + "0x944855a63daa3b755a70f2ba5e6a956b037eeb210dd9e80cef73b7c814645adc" + ], + [ + 81044, + "0xf046f77959a6aa199f3d999b2c16dbeb1e9038c308cb1dcc9bba26141e0e5f35" + ], + [ + 81045, + "0x4e420f14ba78477c090eef99a098b62dfdadc3a61f7c0a5499797380763693db" + ], + [ + 81046, + "0xae10dd9883739a2cd9e95cafa4f9cc7b2e38e477207ec49fb3fe64e3fcf21613" + ], + [ + 81047, + "0xa9e686e36215cf1053ac16cb445e537ef8a633aac308ad755716f6374c3c1238" + ], + [ + 81048, + "0xa7ab60b451ae4a7fa639faa0c8b1f2ad7bb6bcb853242797f4aca608b2e69193" + ], + [ + 81049, + "0xdba699a61f7e4b3b8356a7494afbb73237f00385d7ad279a77002bfd451b2643" + ], + [ + 81050, + "0x42515b83488525c1c8e72417df7361aa61a62e3f78892446b06d6d6bccee6ada" + ], + [ + 81051, + "0x9a3442c671ff6750549ac81e1a48d7ba70e8248b0f1122bf37a3767666cac4e6" + ], + [ + 81052, + "0x93319ca0399a5b00e77bbde1cf067c1f4f823cbeffaa1103c7953abd1b8db6c5" + ] + ], + "rewards": [ + [ + "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "0x32694da1632d4400" + ] + ], + "proving_pool_credit": "0xc9a5367c3c85800", + "payouts": [], + "blocks": [ + { + "miner": "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "blue": true, + "txs": [] + } + ], + "pre_state": [ + { + "address": "0x0000000000000000000000000000000000000210", + "nonce": 1, + "balance": "0x0", + "code": "0x608060405234801561000f575f5ffd5b506004361061003f575f3560e01c8063aa67735414610043578063dea5c2e014610058578063fe7e05d51461009f575b5f5ffd5b6100566100513660046101e0565b6100ca565b005b610083610066366004610211565b6001600160a01b039081165f908152600160205260409020541690565b6040516001600160a01b03909116815260200160405180910390f35b6100836100ad366004610211565b6001600160a01b039081165f908152602081905260409020541690565b336001600160a01b03831614806100f957506001600160a01b038281165f908152600160205260409020541633145b6101635760405162461bcd60e51b815260206004820152603160248201527f446576656c6f70657252656769737472793a206e6f7420746865206163636f75604482015270373a1037b91034ba399031b932b0ba37b960791b606482015260840160405180910390fd5b6001600160a01b038281165f818152602081815260409182902080546001600160a01b031916948616948517905590513381527fa47563c41dab010f91a8ef9dc7ac2bcdfa0ef2af697e575048e71e6eec60dda3910160405180910390a35050565b80356001600160a01b03811681146101db575f5ffd5b919050565b5f5f604083850312156101f1575f5ffd5b6101fa836101c5565b9150610208602084016101c5565b90509250929050565b5f60208284031215610221575f5ffd5b61022a826101c5565b939250505056fea2646970667358221220cbf48f5aa6f911f83b3c2b09adf8c418f5da6fc530c3d5db9d4ba0be24c4bf5864736f6c63430008250033", + "storage": [] + }, + { + "address": "0x0000000000000000000000000000000000000220", + "nonce": 0, + "balance": "0x13cbed474a4367e8f400", + "code": "0x", + "storage": [] + }, + { + "address": "0x0f002c928c363c7041cafe788f099cf7d35452a1", + "nonce": 243, + "balance": "0x1bc36179476e9a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x18524811fa2e76dd0770d29308b51bcbdc0499f9", + "nonce": 242, + "balance": "0x1bb8aa483e652c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x1aca71f7872aebc85b4d4a9ad47a09abcc624d65", + "nonce": 0, + "balance": "0x2860cb14365c048800", + "code": "0x", + "storage": [] + }, + { + "address": "0x233d639f53225ea5012dc01ceb0a5a30891021cd", + "nonce": 264, + "balance": "0x1b8eea59c54b4200", + "code": "0x", + "storage": [] + }, + { + "address": "0x27fd8475c3db352fdddeaf29823e00c7d33ca39a", + "nonce": 248, + "balance": "0x1bccb45d83401000", + "code": "0x", + "storage": [] + }, + { + "address": "0x37b55c8531053a5cef7fe1f19ba67012719773f9", + "nonce": 0, + "balance": "0x1e0c843e2433b90ac00", + "code": "0x", + "storage": [] + }, + { + "address": "0x46681948060140945958068b373f94581e049d4f", + "nonce": 240, + "balance": "0x1baadca68fcb1a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4cb00bc539538d2ee7172474f2b06f021944ece4", + "nonce": 258, + "balance": "0x1b3e54c72c629c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x4d7c0d6f3ad5466691649e98ebb59cd5b5194687", + "nonce": 0, + "balance": "0x1b598fbe8e08549c6000", + "code": "0x", + "storage": [] + }, + { + "address": "0x5a5e606bda0fe1b2b4298aadd55cc6ec57ff698a", + "nonce": 0, + "balance": "0x648cba8883ddb562c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x6613ca66c14338db771d01c238ae140c71310535", + "nonce": 243, + "balance": "0x1b98d3921da34800", + "code": "0x", + "storage": [] + }, + { + "address": "0x66577de387feeec96c10f7f8748f13912589c78c", + "nonce": 240, + "balance": "0x1bc37e3123b20a00", + "code": "0x", + "storage": [] + }, + { + "address": "0x7e7b1db26094aa913933f65127b328c1861ff5b0", + "nonce": 256, + "balance": "0x1b89d2467bcd2600", + "code": "0x", + "storage": [] + }, + { + "address": "0x90acb15171deb958d3191d5c02d807afc367f0e8", + "nonce": 240, + "balance": "0x1b9e3c5e84fb6c00", + "code": "0x", + "storage": [] + }, + { + "address": "0x9f296bf64eb7051912c9c8112da3e3ea6cf546e1", + "nonce": 266, + "balance": "0x1b69b11e00d88c00", + "code": "0x", + "storage": [] + }, + { + "address": "0xaa19223d63c82bf9be5543ec96e5504834d17bc8", + "nonce": 246, + "balance": "0x1b9abb1eb524dc00", + "code": "0x", + "storage": [] + }, + { + "address": "0xb03779ce9fc5d0236a5a73dbbbf455ecfe2839fd", + "nonce": 0, + "balance": "0x5656a7e24fdb180800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc2faa4a2866422a865c9b3b3cb5988bf4d07f0b4", + "nonce": 0, + "balance": "0xc05b9fc73002b32800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc45d23f49451faad9afadf877feead626ea8d083", + "nonce": 0, + "balance": "0x9ee8f806428adc9800", + "code": "0x", + "storage": [] + }, + { + "address": "0xc8621de5921418701619ba715044575073d2ca62", + "nonce": 243, + "balance": "0x1bbb8989a888a800", + "code": "0x", + "storage": [] + }, + { + "address": "0xcafc6e743c6848d637ab6283db5ba9fa3023516a", + "nonce": 0, + "balance": "0x1f29d0e7a3eaed38a400", + "code": "0x", + "storage": [] + }, + { + "address": "0xd958300657e7931b06fc6f60dc4bfe7d8e8ea7b8", + "nonce": 252, + "balance": "0x1b79a326be219400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdadb11b01f6004eba6da3846a0e1bf6f4f71fb82", + "nonce": 244, + "balance": "0x1b900b01183e5600", + "code": "0x", + "storage": [] + }, + { + "address": "0xdd442fcbb964a3afdc90d49b408e8dd296fa86e8", + "nonce": 0, + "balance": "0xafd0c09109be7410400", + "code": "0x", + "storage": [] + }, + { + "address": "0xdfaea67368f3e3753397d878f97efe6aa8020c2e", + "nonce": 16, + "balance": "0x4f20fde2ac0f98ec00", + "code": "0x", + "storage": [] + }, + { + "address": "0xfd4fca79b266a73a7defa776b04c7e15bdc70e86", + "nonce": 238, + "balance": "0x1bb5cf9909338400", + "code": "0x", + "storage": [] + } + ], + "fees": { + "base": { + "pgas": { + "version": 0, + "cycles_per_pgas": 1000, + "intrinsic_pgas_per_tx": 200, + "modexp_base": 1000, + "modexp_per_byte_numer": 10, + "modexp_per_byte_denom": 1 + }, + "block_proving_gas_limit": 30000000, + "shard_proving_gas_budget": 7500000, + "min_execution_base_fee_wei": 1000000000, + "min_proving_base_fee_wei": 1000000000, + "initial_execution_base_fee_wei": 1000000000, + "initial_proving_base_fee_wei": 1000000000, + "base_fee_change_denominator": 8 + }, + "v1_activation_daa": 18446744073709551615 + } + }, + "plan": { + "shard_budget": 7500000, + "consensus": true, + "shards": [ + { + "index": 0, + "tx_start": 0, + "tx_end": 0, + "over_budget": false, + "pre_root": "0x7300a34f1521fc454efa0330508f4cf762541806c3d766a860e31900649d53bc", + "post_root": "0xbe1ac38932436b515a6ef939e28dc50b4299c26094563fffb3e3fc710439c874", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "link_in": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "link_out": "0x38046159e1bf364d16df7645adc5a18f254dd70ffab279ed30785b727b52a3d5", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "witness": [ + 2, + 0, + 3, + 14, + 14093 + ] + } + ] + }, + "expected": { + "pre_state_root": "0x7300a34f1521fc454efa0330508f4cf762541806c3d766a860e31900649d53bc", + "post_state_root": "0xbe1ac38932436b515a6ef939e28dc50b4299c26094563fffb3e3fc710439c874", + "receipts_root": "0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421", + "tx_commitment": "0x0000000000000000000000000000000000000000000000000000000000000000", + "gas_used": 0, + "pgas_used": 0, + "executed": 0, + "skipped": 0, + "node_state_root": "0xbe1ac38932436b515a6ef939e28dc50b4299c26094563fffb3e3fc710439c874" + } +} \ No newline at end of file diff --git a/proving/igneum-prove/host/src/main.rs b/proving/igneum-prove/host/src/main.rs index d74a48af7..6891a9658 100644 --- a/proving/igneum-prove/host/src/main.rs +++ b/proving/igneum-prove/host/src/main.rs @@ -78,10 +78,25 @@ fn run() -> Result<()> { // proving v0 (spec 7.7): the node's proof pool verifies a submitted shard proof off the consensus path return run_verify(&pinned, &arg("--proof").context("--proof ")?, &arg("--statement").context("--statement 0x")?); } - let path = args.get(1).filter(|a| !a.starts_with("--")).context("usage: igneum-prove-host [--mode native|execute|shard|compressed|block|all] [--shard N] [--prover 0x..] [--out results.json]; igneum-prove-host --mode verify --proof --statement 0x..; igneum-prove-host --mode id")?; - let shard_index: usize = arg("--shard").map(|s| s.parse()).transpose()?.unwrap_or(0); + if mode == "verify-segment" { + // proving v1 (spec 7.8): the node's pool verifies an aggregated segment proof against the pinned aggregator key + return run_verify_segment(&pinned, &arg("--proof").context("--proof ")?, &arg("--statement").context("--statement 0x")?); + } let prover: Address = arg("--prover").map(|s| s.parse()).transpose()?.unwrap_or_else(|| Address::from_slice(&[0x19; 20])); let out_path = arg("--out"); + if mode == "aggregate" { + // proving v1: the live aggregator, from shard proof files (the node's pool) and the previous segment proof + return run_aggregate(&pinned, &arg("--proofs").context("--proofs (the segment's shard proofs in shard order)")?, &arg("--parent").context("--parent 0x")?, arg("--prev").as_deref(), out_path.as_deref()); + } + if mode == "chain" { + // proving v1: N consecutive fixtures proven shard by shard, each block aggregated with the previous block's + // proof (the chain rule of design 5.3), the measurement of docs/plans/proving-v1.md step 2 + let list = arg("--chain").or_else(|| args.get(1).filter(|a| !a.starts_with("--")).cloned()).context("--chain (consecutive fixtures)")?; + let fixtures: Vec = list.split(',').map(|s| s.trim().to_string()).filter(|s| !s.is_empty()).collect(); + return run_chain(&pinned, &fixtures, prover, out_path.as_deref()); + } + let path = args.get(1).filter(|a| !a.starts_with("--")).context("usage: igneum-prove-host [--mode native|execute|shard|compressed|block|all] [--shard N] [--prover 0x..] [--out results.json]; --mode chain --chain [--prover 0x..] [--out results.json]; --mode aggregate --proofs --parent 0x.. [--prev prev.bin] [--out results.json]; --mode verify --proof --statement 0x..; --mode verify-segment --proof --statement 0x..; --mode id")?; + let shard_index: usize = arg("--shard").map(|s| s.parse()).transpose()?.unwrap_or(0); let fixture: Fixture = serde_json::from_str(&std::fs::read_to_string(path).with_context(|| format!("read {path}"))?)?; if fixture.format != igneum_prove_core::fixture::FORMAT { bail!("fixture format {} is not {} (regenerate with igneum-prove-export)", fixture.format, igneum_prove_core::fixture::FORMAT); @@ -531,6 +546,314 @@ fn run_block(sp1: &Sp1ProofSystem, shards: &[BuiltShard], claim: &SegmentClaim, Ok(()) } +/// Loads a fixture with the checks `run` makes (format, the plan's budget against the fee set at its DAA score). +fn load_fixture(path: &str) -> Result { + let fixture: Fixture = serde_json::from_str(&std::fs::read_to_string(path).with_context(|| format!("read {path}"))?)?; + if fixture.format != igneum_prove_core::fixture::FORMAT { + bail!("fixture format {} is not {} (regenerate with igneum-prove-export)", fixture.format, igneum_prove_core::fixture::FORMAT); + } + let fee_set = fixture.block.fees.at(fixture.block.env.daa_score); + if fixture.plan.consensus && fixture.plan.shard_budget != fee_set.shard_proving_gas_budget { + bail!("{path}: the fixture's plan says consensus budget {} but the fee set at DAA score {} ({}) has S_p {}; regenerate with igneum-prove-export", fixture.plan.shard_budget, fixture.block.env.daa_score, fee_set.name(), fee_set.shard_proving_gas_budget); + } + Ok(fixture) +} + +fn setup_sp1(pinned: &pinned::Pinned, results: &mut serde_json::Map) -> Result { + stage("setup"); + let t = Instant::now(); + let sp1 = Sp1ProofSystem::from_env(pinned.shard_elf(), pinned.agg_elf())?; + let setup_s = t.elapsed().as_secs_f64(); + println!("RESULT setup: {:.2} s, ProofSystem v{} shard program id {} aggregator id {} at {}", setup_s, Sp1ProofSystem::VERSION, sp1.program_id(), sp1.aggregator_id(), now()); + if sp1.program_id() != pinned.shard_id || sp1.aggregator_id() != pinned.agg_id { + bail!("SP1's key setup derived shard program id {} and aggregator id {} from the embedded guests, the pinned manifest says {} and {}: this host would make proofs no other node accepts (re-pin with proving/igneum-prove/pin-guests.sh)", sp1.program_id(), sp1.aggregator_id(), pinned.shard_id, pinned.agg_id); + } + results.insert("setup_seconds".into(), setup_s.into()); + results.insert("shard_program_id".into(), sp1.program_id().to_string().into()); + results.insert("aggregator_id".into(), sp1.aggregator_id().to_string().into()); + results.insert("prover".into(), std::env::var("SP1_PROVER").unwrap_or_else(|_| "cpu".into()).into()); + Ok(sp1) +} + +/// What a segment record carries about an aggregated proof (spec 7.8): the public values, their keccak (the +/// statement), the proof's SHA-256 and where the proof bytes went. +fn segment_results(seg: &proof_system::Sp1SegmentProof, out_dir: &std::path::Path, results: &mut serde_json::Map) -> Result<(B256, usize)> { + let bytes = bincode::serialize(&seg.proof)?; + let pv = seg.proof.public_values.as_slice().to_vec(); + let statement = alloy_primitives::keccak256(&pv); + let proof_hash: [u8; 32] = sha2::Sha256::digest(&bytes).into(); + let file = out_dir.join(format!("segment-{}-aggregated.bin", seg.output.number)); + std::fs::write(&file, &bytes)?; + results.insert("segment_number".into(), seg.output.number.into()); + results.insert("segment_block_hash".into(), seg.output.block_hash.to_string().into()); + results.insert("segment_chain_len".into(), seg.output.chain_len.into()); + results.insert("segment_public_values".into(), format!("0x{}", hex::encode(&pv)).into()); + results.insert("segment_statement".into(), statement.to_string().into()); + results.insert("segment_proof_sha256".into(), format!("0x{}", hex::encode(proof_hash)).into()); + results.insert("segment_proof_bytes".into(), bytes.len().into()); + results.insert("segment_proof_file".into(), file.display().to_string().into()); + results.insert("segment_provers".into(), seg.output.provers.to_string().into()); + println!("segment proof written to {} ({} bytes); statement {statement} proof sha256 0x{}", file.display(), bytes.len(), hex::encode(proof_hash)); + Ok((statement, bytes.len())) +} + +fn out_dir_of(out_path: Option<&str>) -> std::path::PathBuf { + out_path.and_then(|p| std::path::Path::new(p).parent().map(|d| d.to_path_buf())).filter(|d| !d.as_os_str().is_empty()).unwrap_or_else(|| std::path::PathBuf::from(".")) +} + +/// `--mode chain`: every fixture in order, consecutive on the chain (number and parent hash), each block's shards +/// proven compressed and aggregated with the previous block's aggregated proof (`AggInput.prev`, the chain rule), +/// every proof verified. One RESULT line per shard, per block (with the running totals) and for the chain. +fn run_chain(pinned: &pinned::Pinned, fixtures: &[String], prover: Address, out_path: Option<&str>) -> Result<()> { + if fixtures.is_empty() { + bail!("--chain needs at least one fixture"); + } + let mut results = serde_json::Map::new(); + results.insert("mode".into(), "chain".into()); + results.insert("fixtures".into(), fixtures.iter().map(|f| serde_json::Value::from(f.as_str())).collect::>().into()); + let mut loaded = Vec::with_capacity(fixtures.len()); + for path in fixtures { + let f = load_fixture(path)?; + if let Some(prev) = loaded.last().map(|(_, f): &(String, Fixture)| &f.block.env) { + if f.block.env.number != prev.number + 1 || f.block.env.parent_hash != prev.hash { + bail!("{path}: block {} (parent {}) does not follow block {} ({}); the chain needs consecutive fixtures", f.block.env.number, f.block.env.parent_hash, prev.number, prev.hash); + } + } + loaded.push((path.clone(), f)); + } + let first = loaded[0].1.block.env.number; + let last = loaded[loaded.len() - 1].1.block.env.number; + println!("igneum-prove-host sources {}: chain of {} consecutive blocks {first}..={last}, prover payout {prover}, SP1_PROVER={}; {}", env!("IGNEUM_PROVE_SOURCES"), loaded.len(), std::env::var("SP1_PROVER").unwrap_or_else(|_| "cpu".into()), now()); + // native first: every block's cut and every shard statement must reproduce the fixture before a proof is made + stage("native"); + let mut built: Vec<(u64, B256, Vec, B256)> = Vec::with_capacity(loaded.len()); + for (path, f) in &loaded { + let (outcome, pre_root, shards) = build_shards(&f.block, f.plan.shard_budget, prover); + let e = &f.expected; + if pre_root != e.pre_state_root || outcome.state_root != e.post_state_root || outcome.receipts_root != e.receipts_root || outcome.gas_used != e.gas_used || outcome.pgas_used != e.pgas_used || shards.len() != f.plan.shards.len() { + bail!("{path}: native execution differs from the fixture; regenerate it with igneum-prove-export"); + } + if let Some((_, _, prev_shards, _)) = built.last() { + let prev_post = prev_shards.last().map(|s| s.output.post_root).unwrap_or_default(); + if pre_root != prev_post { + bail!("{path}: block {} starts from state root {pre_root}, the previous block ended at {prev_post}: the state does not chain", f.block.env.number); + } + } + println!("RESULT chain native block {}: {} shard(s), pgas {}, gas {}, pre {} post {}", f.block.env.number, shards.len(), outcome.pgas_used, outcome.gas_used, pre_root, outcome.state_root); + built.push((f.block.env.number, f.block.env.hash, shards, pre_root)); + } + let sp1 = setup_sp1(pinned, &mut results)?; + let chain_t = Instant::now(); + let mut prev: Option = None; + let mut blocks_json = Vec::with_capacity(built.len()); + let (mut shard_total, mut agg_total, mut shards_total) = (0.0f64, 0.0f64, 0usize); + let out_dir = out_dir_of(out_path); + for (number, _hash, shards, _) in &built { + let block_t = Instant::now(); + let mut proofs: Vec = Vec::with_capacity(shards.len()); + let mut shard_secs = Vec::new(); + for s in shards { + let i = s.output.shard_index; + stage(&format!("chain block {number} compressed shard {i}")); + let p = sp1.prove_shard(&ShardWitness { input: s.input.clone() })?; + let dt = sp1.last_timing("compressed").unwrap_or_default().as_secs_f64(); + let (ok, vdt) = sp1.verify_shard(&p.proof, &s.output); + println!("RESULT chain block {number} shard {i}: compressed prove {dt:.1} s, proof {} bytes, verify {:.3} s, {}, pgas {} at {}", bincode::serialize(&p.proof)?.len(), vdt.as_secs_f64(), if ok { "VERIFIED" } else { "VERIFY FAILED" }, p.output.pgas_used, now()); + if !ok { + bail!("compressed proof of block {number} shard {i} did not verify"); + } + shard_secs.push(dt); + shard_total += dt; + proofs.push(p); + } + shards_total += proofs.len(); + stage(&format!("chain block {number} aggregate {} shards{}", proofs.len(), if prev.is_some() { " with the previous block proof" } else { "" })); + let seg = sp1.aggregate(prev.as_ref(), &proofs)?; + let adt = sp1.last_timing("aggregate").unwrap_or_default().as_secs_f64(); + agg_total += adt; + let claim = SegmentClaim::from_block(&seg.output); + let ok = sp1.verify_segment(&seg, &claim); + let vdt = sp1.last_timing("verify-block").unwrap_or_default().as_secs_f64(); + let bytes = bincode::serialize(&seg.proof)?.len(); + let block_s = block_t.elapsed().as_secs_f64(); + let cumulative = chain_t.elapsed().as_secs_f64(); + println!( + "RESULT chain block {number}: {} shards ({:.1} s of shard proofs), aggregate prove {adt:.1} s, proof {bytes} bytes, verify {vdt:.3} s, {}; chain_len {}, agg_vk {}; this block {block_s:.1} s, cumulative {cumulative:.1} s over {} block(s) at {}", + seg.output.shard_count, + shard_secs.iter().sum::(), + if ok { "VERIFIED (shard program id, aggregator id and claim checked)" } else { "VERIFY FAILED" }, + seg.output.chain_len, + seg.output.agg_vk, + blocks_json.len() + 1, + now() + ); + if !ok { + bail!("the aggregated proof of block {number} did not verify"); + } + let expected_len = blocks_json.len() as u64 + 1; + if seg.output.chain_len != expected_len { + bail!("block {number}: chain_len {} is not {expected_len}", seg.output.chain_len); + } + blocks_json.push(serde_json::json!({ + "number": number, "shards": seg.output.shard_count, "shard_prove_seconds": shard_secs, "aggregate_prove_seconds": adt, + "aggregate_verify_seconds": vdt, "proof_bytes": bytes, "chain_len": seg.output.chain_len, "block_seconds": block_s, "cumulative_seconds": cumulative, + "post_root": seg.output.post_root.to_string(), "statement": alloy_primitives::keccak256(seg.output.to_bytes()).to_string(), + })); + prev = Some(seg); + } + let seg = prev.unwrap(); + let total = chain_t.elapsed().as_secs_f64(); + let (statement, bytes) = segment_results(&seg, &out_dir, &mut results)?; + println!( + "RESULT chain: {} blocks {first}..={last}, {shards_total} shards, shard proofs {shard_total:.1} s, aggregation {agg_total:.1} s, end to end {total:.1} s; final proof {bytes} bytes attests chain_len {} (statement {statement}), pre {} post {} provers {} at {}", + built.len(), + seg.output.chain_len, + built[0].3, + seg.output.post_root, + seg.output.provers, + now() + ); + results.insert("blocks".into(), blocks_json.into()); + results.insert("first".into(), first.into()); + results.insert("last".into(), last.into()); + results.insert("shards".into(), shards_total.into()); + results.insert("shard_prove_seconds_total".into(), shard_total.into()); + results.insert("aggregate_prove_seconds_total".into(), agg_total.into()); + results.insert("chain_seconds".into(), total.into()); + drop(sp1); + finish(results, out_path.map(|s| s.to_string())) +} + +/// `--mode aggregate`: the live aggregator (the app's segment step, spec 7.8). `--proofs` names the shard proof +/// files of one or more consecutive chain blocks: blocks separated by `;`, a block's shards by `,` (what the node's +/// pool holds, `igneum_getProofBytes`); `--parent` is the first block's parent chain block hash; `--prev` the +/// previous segment's aggregated proof when the chain continues. Each block is aggregated with the previous +/// block's proof in one process (one key setup); the output is the last block's aggregated proof, its public +/// values, the statement and the proof hash the segment record carries. +fn run_aggregate(pinned: &pinned::Pinned, proofs: &str, parent: &str, prev_path: Option<&str>, out_path: Option<&str>) -> Result<()> { + let first_parent: B256 = parent.parse().context("--parent is not 32 bytes of hex")?; + let mut results = serde_json::Map::new(); + results.insert("mode".into(), "aggregate".into()); + let mut blocks: Vec> = Vec::new(); + let mut parent_hash = first_parent; + for group in proofs.split(';').map(str::trim).filter(|s| !s.is_empty()) { + let mut shards: Vec = Vec::new(); + for path in group.split(',').map(str::trim).filter(|s| !s.is_empty()) { + let bytes = std::fs::read(path).with_context(|| format!("read {path}"))?; + let proof: sp1_sdk::SP1ProofWithPublicValues = bincode::deserialize(&bytes).with_context(|| format!("{path} is not a bincode SP1 proof"))?; + let output = ShardOutput::from_bytes(proof.public_values.as_slice()).with_context(|| format!("{path}: public values are not a shard statement"))?; + if let Some(c) = pinned::claimed_program_id(&proof) { + if c != pinned.shard_id { + bail!("{path}: shard proof made with program id {c}, ours is {}", pinned.shard_id); + } + } + shards.push(Sp1ShardProof { proof, output, parent_hash }); + } + if shards.is_empty() { + bail!("a block in --proofs names no file"); + } + shards.sort_by_key(|s| s.output.shard_index); + if let Some(prev) = blocks.last().and_then(|b: &Vec| b.first()) { + if shards[0].output.number != prev.output.number + 1 { + bail!("block {} does not follow block {}: the blocks of --proofs must be consecutive", shards[0].output.number, prev.output.number); + } + } + parent_hash = shards[0].output.block_hash; + blocks.push(shards); + } + if blocks.is_empty() { + bail!("--proofs names no file"); + } + let mut prev = match prev_path { + None => None, + Some(p) => { + let bytes = std::fs::read(p).with_context(|| format!("read {p}"))?; + let proof: sp1_sdk::SP1ProofWithPublicValues = bincode::deserialize(&bytes).with_context(|| format!("{p} is not a bincode SP1 proof"))?; + let output = BlockOutput::from_bytes(proof.public_values.as_slice()).with_context(|| format!("{p}: public values are not a block statement"))?; + Some(proof_system::Sp1SegmentProof { proof, output }) + } + }; + let first = blocks[0][0].output.number; + let last = blocks[blocks.len() - 1][0].output.number; + println!("igneum-prove-host sources {}: aggregate blocks {first}..={last} ({} shard proofs){}; {}", env!("IGNEUM_PROVE_SOURCES"), blocks.iter().map(|b| b.len()).sum::(), prev.as_ref().map(|p| format!(", chaining to block {} (chain_len {})", p.output.number, p.output.chain_len)).unwrap_or_default(), now()); + let sp1 = setup_sp1(pinned, &mut results)?; + let t_all = Instant::now(); + let mut per_block = Vec::new(); + for shards in &blocks { + let number = shards[0].output.number; + stage(&format!("aggregate block {number}, {} shards", shards.len())); + let seg = sp1.aggregate(prev.as_ref(), shards)?; + let adt = sp1.last_timing("aggregate").unwrap_or_default().as_secs_f64(); + let claim = SegmentClaim::from_block(&seg.output); + let ok = sp1.verify_segment(&seg, &claim); + let vdt = sp1.last_timing("verify-block").unwrap_or_default().as_secs_f64(); + println!("RESULT aggregate block {number}: {} shards, prove {adt:.1} s, proof {} bytes, verify {vdt:.3} s, {}; chain_len {}, post {} at {}", seg.output.shard_count, bincode::serialize(&seg.proof)?.len(), if ok { "VERIFIED" } else { "VERIFY FAILED" }, seg.output.chain_len, seg.output.post_root, now()); + if !ok { + bail!("the aggregated proof of block {number} did not verify"); + } + per_block.push(serde_json::json!({ "number": number, "shards": seg.output.shard_count, "aggregate_prove_seconds": adt, "aggregate_verify_seconds": vdt, "chain_len": seg.output.chain_len })); + prev = Some(seg); + } + let seg = prev.unwrap(); + let (statement, bytes) = segment_results(&seg, &out_dir_of(out_path), &mut results)?; + println!("RESULT aggregate: blocks {first}..={last} in {:.1} s, final proof {bytes} bytes, chain_len {}, statement {statement} at {}", t_all.elapsed().as_secs_f64(), seg.output.chain_len, now()); + results.insert("blocks".into(), per_block.into()); + results.insert("first".into(), first.into()); + results.insert("last".into(), last.into()); + results.insert("aggregate_seconds".into(), t_all.elapsed().as_secs_f64().into()); + drop(sp1); + finish(results, out_path.map(|s| s.to_string())) +} + +/// `--mode verify-segment --proof --statement 0x..`: the node's verifier for an aggregated segment record +/// (spec 7.8): SP1's light verifier against the PINNED aggregator key, the public values' keccak against the +/// statement, the shard program id inside the statement against ours, the aggregator id inside it against ours +/// (or zero when the proof chains to nothing). Exit 0 = verified, 3 = not verified. +fn run_verify_segment(pinned: &pinned::Pinned, proof_path: &str, statement: &str) -> Result<()> { + use sp1_sdk::blocking::{LightProver, Prover}; + let bytes = std::fs::read(proof_path).with_context(|| format!("read {proof_path}"))?; + let want: B256 = statement.parse().context("statement is not 32 bytes of hex")?; + stage("setup"); + let t = Instant::now(); + let verifier = LightProver::new(); + println!("RESULT setup: {:.3} s (light verifier, pinned aggregator key), aggregator id {} shard program id {} at {}", t.elapsed().as_secs_f64(), pinned.agg_id, pinned.shard_id, now()); + stage("verify-segment"); + let t = Instant::now(); + let proof: sp1_sdk::SP1ProofWithPublicValues = bincode::deserialize(&bytes).context("the file is not a bincode SP1 proof")?; + let got = alloy_primitives::keccak256(proof.public_values.as_slice()); + let output = BlockOutput::from_bytes(proof.public_values.as_slice()); + let claimed = pinned::claimed_program_id(&proof); + let same_program = claimed == Some(pinned.agg_id); + let crypto_ok = same_program && verifier.verify(&proof, &pinned.agg_vk, None).is_ok(); + let ids_ok = output.as_ref().map(|o| o.shard_vk == pinned.shard_id && (o.agg_vk == pinned.agg_id || (o.chain_len == 1 && o.agg_vk == B256::ZERO))).unwrap_or(false); + let ok = crypto_ok && ids_ok && got == want; + let dt = t.elapsed().as_secs_f64(); + let program = match claimed { + Some(c) if same_program => format!("aggregator id {c} (ours)"), + Some(c) => format!("aggregator id {c} IS NOT OURS {} (the aggregator runs another guest build)", pinned.agg_id), + None => "not a compressed proof".to_string(), + }; + match &output { + Some(o) => println!( + "RESULT verify-segment: {} in {dt:.3} s; block {} ({}) chain_len {} shards {} statement {got} (want {want}) {program}; inner ids {}; proof {} bytes at {}", + if ok { "VERIFIED" } else { "NOT VERIFIED" }, + o.number, + o.block_hash, + o.chain_len, + o.shard_count, + if ids_ok { "ours".to_string() } else { format!("NOT OURS (shard {} agg {} chain_len {})", o.shard_vk, o.agg_vk, o.chain_len) }, + bytes.len(), + now() + ), + None => println!("RESULT verify-segment: NOT VERIFIED in {dt:.3} s; public values are not a block statement; {program} at {}", now()), + } + if ok { + Ok(()) + } else { + std::process::exit(3) + } +} + fn check_shard_output(out: &ShardOutput, native: &ShardOutput) -> Result<()> { if out != native { bail!("the guest's public values differ from the native run:\n guest {out:?}\n native {native:?}"); diff --git a/proving/igneum-prove/host/src/pinned.rs b/proving/igneum-prove/host/src/pinned.rs index c55cbd6b6..c87256d1e 100644 --- a/proving/igneum-prove/host/src/pinned.rs +++ b/proving/igneum-prove/host/src/pinned.rs @@ -74,8 +74,7 @@ pub fn program_id_of(vk: &SP1VerifyingKey) -> B256 { pub struct Pinned { pub manifest: Manifest, pub shard_vk: SP1VerifyingKey, - /// For the aggregated record's verifier (proving v0 next step); checked against the manifest today. - #[allow(dead_code)] + /// The aggregated segment record's verifier (`--mode verify-segment`, proving v1). pub agg_vk: SP1VerifyingKey, pub shard_id: B256, pub agg_id: B256, diff --git a/tools/prove-fixtures/node_modules b/tools/prove-fixtures/node_modules new file mode 120000 index 000000000..12da4c7fc --- /dev/null +++ b/tools/prove-fixtures/node_modules @@ -0,0 +1 @@ +/Users/joshm/Projects/igneum/tools/prove-fixtures/node_modules \ No newline at end of file diff --git a/tools/proving-v1/coverage.mjs b/tools/proving-v1/coverage.mjs new file mode 100644 index 000000000..701f06d92 --- /dev/null +++ b/tools/proving-v1/coverage.mjs @@ -0,0 +1,89 @@ +#!/usr/bin/env node +// Proving v1 step 3: the live devnet's proven-block share and the proof latency distribution over a window, read +// from one node's execution-layer RPC (what the chain carried: every node agrees on it) and, beside it, the live +// page's 10-minute proving object. Read-only: no job, no switch. +// +// node tools/proving-v1/coverage.mjs [--rpc http://127.0.0.1:26790] [--minutes 30] [--live https://igneum.network/api/live] +// [--out ] [--watch] +// +// Without --watch: one pass over the last --minutes of chain blocks (by block timestamp) at the tip: per block the +// shard plan (shards), the paid rows (carrierNumber) and the carrier's timestamp; prints the table for the bench log. +// With --watch: samples the live page every 60 s for --minutes, then the pass. +// +// Latency here is on-chain: the carrying chain block's timestamp minus the proven block's timestamp (the record was +// verified by its producer before that, so this bounds "block time to verified record" from above). + +import { writeFileSync } from 'node:fs'; +const args = process.argv.slice(2); +const opt = (n, d) => { const i = args.indexOf(n); return i >= 0 && args[i + 1] !== undefined ? args[i + 1] : d; }; +const RPC = opt('--rpc', 'http://127.0.0.1:26790'); +const LIVE = opt('--live', 'https://igneum.network/api/live'); +const MINUTES = +opt('--minutes', 30); +const OUT = opt('--out', null); +const WATCH = args.includes('--watch'); +const log = (...a) => console.log(new Date().toISOString().slice(11, 19), ...a); +const sleep = (ms) => new Promise((r) => setTimeout(r, ms)); +let id = 0; +async function rpc(method, params = []) { + const r = await fetch(RPC, { method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify({ jsonrpc: '2.0', id: ++id, method, params }), signal: AbortSignal.timeout(30000) }); + const j = await r.json(); if (j.error) throw new Error(j.error.message); return j.result; +} +const hexn = (v) => (v == null ? null : Number(BigInt(v))); +const pct = (a, b) => (b ? (100 * a / b).toFixed(1) + '%' : 'n/a'); +function quantiles(xs) { + if (!xs.length) return { n: 0 }; + const s = [...xs].sort((a, b) => a - b); + const q = (p) => s[Math.min(s.length - 1, Math.floor(p * (s.length - 1)))]; + return { n: s.length, min: s[0], p50: q(0.5), p90: q(0.9), p99: q(0.99), max: s[s.length - 1], mean: +(s.reduce((a, b) => a + b, 0) / s.length).toFixed(1) }; +} + +async function pass() { + const tip = hexn(await rpc('eth_blockNumber')); + const tipBlock = await rpc('eth_getBlockByNumber', ['0x' + tip.toString(16), false]); + const tipTs = hexn(tipBlock.timestamp); + const from = tipTs - MINUTES * 60; + // walk back by timestamp + const ts = new Map(); + const tsOf = async (n) => { if (!ts.has(n)) { const b = await rpc('eth_getBlockByNumber', ['0x' + n.toString(16), false]); ts.set(n, hexn(b.timestamp)); } return ts.get(n); }; + let first = tip; + while (first > 1 && (await tsOf(first - 1)) >= from) first--; + log(`window: chain blocks ${first}..${tip} (${tip - first + 1} blocks, ${MINUTES} min by block timestamp), tip DAA via status`); + const status = await rpc('igneum_getProvingStatus'); + const rows = []; + for (let n = first; n <= tip; n++) { + const r = await rpc('igneum_getProofRecords', ['0x' + n.toString(16)]); + const plan = await rpc('igneum_getShardPlan', ['0x' + n.toString(16)]).catch(() => null); + const shards = plan ? plan.shards.length : 0; + const paid = (r.paid || []).filter(Boolean); + const bt = await tsOf(n); + const lat = []; + for (const p of paid) { const c = hexn(p.carrierNumber); if (c != null) lat.push((await tsOf(c)) - bt); } + rows.push({ n, shards, pgas: plan ? plan.shards.reduce((a, s) => a + hexn(s.pgas), 0) : 0, paid: paid.length, latencies: lat, carried: (r.carried || []).length }); + } + const blocks = rows.length; + const any = rows.filter((x) => x.paid > 0).length; + const full = rows.filter((x) => x.shards > 0 && x.paid === x.shards).length; + const shardsTotal = rows.reduce((a, x) => a + x.shards, 0), shardsPaid = rows.reduce((a, x) => a + x.paid, 0); + const lat = quantiles(rows.flatMap((x) => x.latencies)); + const content = rows.filter((x) => x.pgas > 0); + const contentPaid = content.filter((x) => x.paid === x.shards).length; + const summary = { rpc: RPC, window: { first, last: tip, blocks, minutes: MINUTES, from_ts: from, to_ts: tipTs }, blocks_with_any_paid_shard: any, blocks_fully_proven: full, share_any: pct(any, blocks), share_full: pct(full, blocks), shards_total: shardsTotal, shards_paid: shardsPaid, share_shards: pct(shardsPaid, shardsTotal), content_blocks: content.length, content_blocks_fully_proven: contentPaid, latency_s: lat, status: { paidShards: status.paidShards, pool: status.pool, verifier: status.verifier, tipDaa: hexn(status.tipDaa), v1: status.v1 || null } }; + log(`RESULT coverage: ${blocks} blocks, ${any} with a paid shard (${summary.share_any}), ${full} fully proven (${summary.share_full}); shards ${shardsPaid}/${shardsTotal} (${summary.share_shards}); content blocks ${content.length}, fully proven ${contentPaid}; on-chain latency s: ${JSON.stringify(lat)}`); + return { summary, rows }; +} + +const samples = []; +if (WATCH) { + const end = Date.now() + MINUTES * 60 * 1000; + while (Date.now() < end) { + const lv = await fetch(LIVE, { signal: AbortSignal.timeout(20000) }).then((r) => r.json()).catch((e) => ({ error: e.message })); + const p = lv.proving || {}; + const s = { t: new Date().toISOString(), block_count: lv.state?.block_count, blocks_per_second_60s: lv.state?.blocks_per_second_60s, shards_proven_10m: p.shards_proven_10m, blocks_fully_proven_10m: p.blocks_fully_proven_10m, median_proof_lag_s: p.median_proof_lag_s, provers_10m: p.provers_10m, error: lv.error }; + samples.push(s); + log(`live: ${JSON.stringify(s)}`); + await sleep(60000); + } +} +const result = await pass(); +result.live_samples = samples; +if (OUT) { writeFileSync(OUT, JSON.stringify(result, null, 2)); log(`written ${OUT}`); } diff --git a/tools/proving-v1/net.mjs b/tools/proving-v1/net.mjs new file mode 100644 index 000000000..fd7d7271f --- /dev/null +++ b/tools/proving-v1/net.mjs @@ -0,0 +1,246 @@ +#!/usr/bin/env node +// Proving v1 (spec 7.8, docs/plans/proving-v1.md step 4) on a private 3-node fast-time test network: ports 29950 and +// up, network igneum-devnet-956, data under /tmp/igneum-proving-v1, infra/fast-time's 60x profile with +// skip_proof_of_work, proving v0 at DAA 60 and proving v1 at DAA 120 (4 blocks a segment, unproven after 60 DAA, a +// tenth of the pool credit to the aggregator). Every node runs in trust mode (IGNEUM_PROOF_VERIFY=trust): the rule is +// what is under test, not SP1, so the proof bytes are placeholders and the statement is the node's own native block +// statement (igneum_getSegmentStatement), signed with igneum-miner sign-segment-record. +// +// Cases, each timed and asserted (the CLAUDE.md rule: a gate is trusted only after one known-finished and one +// known-failed case): +// 1. known-finished: the first v1 segment's record (a fresh chain, chain_len 4) relays, verifies, is carried and +// pays the aggregator's share (4 x 10% of the credits) to the payout address on every node +// 2. the chain rule: the second segment refuses a fresh-chain record while the first is proven, and accepts the +// continuing one (chain_len 8) +// 3. known-failed: the third segment gets no record; after its deadline it is unproven, a late record for it is +// refused, and the fourth segment's fresh-chain record is accepted and paid (the chain restarts) +// 4. the v0 side after the switch: a shard's shardWei is 90% of its block's share +// Never touches the live devnet (26610/26611, 26640/28640), the proving-v0 harness (29800+) or the C4 runs (29900+). +// +// node tools/proving-v1/net.mjs [--secs 1200] [--v1 120] [--segment 4] [--unproven 60] +// IGNEUM_PV1_BIN= (default vendor/igneum-node/target-pv1/release) + +import { spawn, spawnSync } from 'node:child_process'; +import { mkdirSync, rmSync, writeFileSync, readFileSync, openSync, existsSync } from 'node:fs'; +import { createHash } from 'node:crypto'; +import { connectRpc } from '../finality-attacks/lib/rpc.mjs'; +import { privateKeyToAccount } from '../prove-fixtures/node_modules/viem/_esm/accounts/index.js'; + +const ROOT = new URL('../../', import.meta.url).pathname; +const FILE = `${ROOT}infra/fast-time/override-60x.json`; +const REL = process.env.IGNEUM_PV1_BIN || `${ROOT}vendor/igneum-node/target-pv1/release`; +const IGNEUMD = `${REL}/igneumd`; +const MINER = `${REL}/igneum-miner`; +const TMP = '/tmp/igneum-proving-v1'; +const BASE = 29950, SUFFIX = 956, CHAIN_NAME = `igneum-devnet-${SUFFIX}`; +const args = process.argv.slice(2); +const flag = (name, dflt) => { const i = args.indexOf(name); return i >= 0 && args[i + 1] !== undefined ? +args[i + 1] : dflt; }; +const SECS = flag('--secs', 1200); +const V0 = 60, V1 = flag('--v1', 120), SEG = flag('--segment', 4), UNPROVEN = flag('--unproven', 60), SHARE_BPS = 1000; +const started = []; +const t0 = Date.now(); +const since = () => ((Date.now() - t0) / 1000).toFixed(1); +const log = (...a) => console.log(new Date().toISOString().slice(11, 23), `t=${since()}s`, ...a); +const sleep = (ms) => new Promise(r => setTimeout(r, ms)); +const hexn = (v) => Number(BigInt(v)); +for (const b of [IGNEUMD, MINER]) if (!existsSync(b)) { console.error(`missing ${b}`); process.exit(2); } + +const funded = privateKeyToAccount('0x59c6995e998f97a5a0044966f0945389dc9e86dae88c7a8412f4603b6b78690d'); +const PAYOUT = '0x4343434343434343434343434343434343434343'; + +rmSync(TMP, { recursive: true, force: true }); mkdirSync(TMP, { recursive: true }); +const override = `${TMP}/override.json`; +let overrideText = readFileSync(FILE, 'utf8') + .replace(/"proving_v0_activation_daa":\s*\d+/, `"proving_v0_activation_daa": ${V0}`) + .replace(/"proving_v1_activation_daa":\s*\d+/, `"proving_v1_activation_daa": ${V1}`) + .replace(/"proving_v1_segment_blocks":\s*\d+/, `"proving_v1_segment_blocks": ${SEG}`) + .replace(/"proving_v1_unproven_daa":\s*\d+/, `"proving_v1_unproven_daa": ${UNPROVEN}`) + .replace(/"proving_v1_aggregator_share_bps":\s*\d+/, `"proving_v1_aggregator_share_bps": ${SHARE_BPS}`) + .replace(/"skip_proof_of_work":\s*(true|false)/, '"skip_proof_of_work": true'); +for (const re of [/"skip_proof_of_work": true/, new RegExp(`"proving_v1_activation_daa": ${V1}`), new RegExp(`"proving_v1_segment_blocks": ${SEG}`), new RegExp(`"proving_v1_unproven_daa": ${UNPROVEN}`)]) if (!re.test(overrideText)) throw new Error(`override edit failed: ${re}`); +writeFileSync(override, overrideText); + +class Node { + constructor(i, connect = [], env = {}) { + this.i = i; this.grpcPort = BASE + i * 10; this.p2pPort = BASE + i * 10 + 1; this.jsonPort = BASE + i * 10 + 2; this.evmPort = BASE + i * 10 + 3; + this.connect = connect; this.env = env; this.dir = `${TMP}/n${i}`; this.logFile = `${this.dir}/node.log`; + } + get grpc() { return `grpc://127.0.0.1:${this.grpcPort}`; } + get evm() { return `http://127.0.0.1:${this.evmPort}`; } + async start() { + mkdirSync(this.dir, { recursive: true }); + const a = ['--devnet', `--devnet-suffix=${SUFFIX}`, '--nodnsseed', '--disable-upnp', '--nologfiles', '--enable-unsynced-mining', '--utxoindex', '--unsaferpc', + `--appdir=${this.dir}`, `--rpclisten=127.0.0.1:${this.grpcPort}`, `--rpclisten-json=127.0.0.1:${this.jsonPort}`, `--evm-rpclisten=127.0.0.1:${this.evmPort}`, + `--listen=127.0.0.1:${this.p2pPort}`, `--override-params-file=${override}`, '--loglevel=info', '--yes']; + if (this.connect.length) a.push(...this.connect.map(c => `--connect=${c}`)); else a.push('--outpeers=0'); + const out = openSync(this.logFile, 'a'); + this.proc = spawn(IGNEUMD, a, { stdio: ['ignore', out, out], env: { ...process.env, ...this.env } }); + started.push(this.proc); + await sleep(800); + this.rpc = await connectRpc(`ws://127.0.0.1:${this.jsonPort}`); + log(`n${this.i} up pid ${this.proc.pid} json ${this.jsonPort} evm ${this.evmPort} p2p ${this.p2pPort} env ${JSON.stringify(this.env)}`); + return this; + } + async eth(method, params = []) { + const r = await fetch(this.evm, { method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify({ jsonrpc: '2.0', id: 1, method, params }) }); + const j = await r.json(); + if (j.error) throw new Error(`${method}: ${j.error.message}`); + return j.result; + } + grepLog(re) { try { return readFileSync(this.logFile, 'utf8').split('\n').filter(l => re.test(l)); } catch { return []; } } +} +function miner(argv, name) { + const out = openSync(`${TMP}/${name}.log`, 'a'); + const p = spawn(MINER, argv, { stdio: ['ignore', out, out] }); + started.push(p); + return p; +} +async function stopAll() { + for (const p of started.reverse()) { try { p.kill('SIGINT'); } catch { } } + await sleep(1500); + for (const p of started) { try { p.kill('SIGKILL'); } catch { } } +} +process.on('SIGINT', async () => { await stopAll(); process.exit(130); }); +const report = { v0: V0, v1: V1, segment: SEG, unproven: UNPROVEN, share_bps: SHARE_BPS, steps: {}, checks: [] }; +const step = (k, v) => { report.steps[k] = v; log(`STEP ${k}: ${JSON.stringify(v)}`); }; +const check = (name, ok, detail) => { report.checks.push({ name, ok, detail }); log(`${ok ? 'PASS' : 'FAIL'} ${name}${detail ? ': ' + JSON.stringify(detail) : ''}`); if (!ok) throw new Error(`check failed: ${name} ${JSON.stringify(detail)}`); }; +function run(cmd, argv, env = {}) { + const r = spawnSync(cmd, argv, { encoding: 'utf8', maxBuffer: 1 << 28, env: { ...process.env, ...env } }); + return { code: r.status, out: (r.stdout || '') + (r.stderr || '') }; +} +const sha256 = (buf) => '0x' + createHash('sha256').update(buf).digest('hex'); +async function waitDaa(n, want, label) { + let daa = 0; + while (daa < want) { await sleep(1000); const s = await n.eth('igneum_getProvingStatus'); daa = hexn(s.tipDaa); } + log(`${label}: daa ${daa}`); + return daa; +} +async function waitExecuted(n, number) { + for (let k = 0; k < 600; k++) { const tip = hexn(await n.eth('eth_blockNumber')); if (tip >= number) return tip; await sleep(500); } + throw new Error(`block ${number} not executed in 300 s`); +} +// signs a segment record over the given public values with v0's key; the proof bytes are a placeholder (trust mode) +function signSegment(first, last, hash, pv, tag) { + const proof = Buffer.from(`igneum-proving-v1-harness-${tag}-${first}-${last}`); + const sg = run(MINER, ['sign-segment-record', 'v0', CHAIN_NAME, hash, String(first), String(last), PAYOUT, pv, sha256(proof)]); + if (sg.code !== 0) throw new Error(`sign-segment-record failed: ${sg.out}`); + const signed = JSON.parse(sg.out.trim().split('\n').pop()); + return { record: signed.record, proof: '0x' + proof.toString('hex'), keyHash: signed.keyHash, statement: signed.statement }; +} +async function waitSegmentPaid(n, first, secs = 240) { + const deadline = Date.now() + secs * 1000; + while (Date.now() < deadline) { + const r = await n.eth('igneum_getSegmentRecords', ['0x' + first.toString(16)]); + if (r.paid) return r; + await sleep(1000); + } + throw new Error(`segment ${first} not paid within ${secs} s on n${n.i}`); +} + +try { + const n0 = await new Node(0, [], { IGNEUM_PROOF_VERIFY: 'trust' }).start(); + const n1 = await new Node(1, [`127.0.0.1:${n0.p2pPort}`], { IGNEUM_PROOF_VERIFY: 'trust' }).start(); + const n2 = await new Node(2, [`127.0.0.1:${n0.p2pPort}`, `127.0.0.1:${n1.p2pPort}`], { IGNEUM_PROOF_VERIFY: 'trust' }).start(); + const nodes = [n0, n1, n2]; + log(`node 0 says: ${n0.grepLog(/proving v1/).join(' | ') || '(no proving v1 line)'}`); + nodes.forEach((n, i) => miner(['vmine', n.grpc, String(SECS), '--label', `v${i}`, '--share', String(1 / 3), '--bps', '1', ...(i === 0 ? ['--evm-address', funded.address] : [])], `vmine-v${i}`)); + const keyHashes = []; + for (let i = 0; i < 3; i++) { + for (let k = 0; k < 50 && !keyHashes[i]; k++) { const m = (() => { try { return readFileSync(`${TMP}/vmine-v${i}.log`, 'utf8').match(/key=([0-9a-f]{64})/); } catch { return null; } })(); if (m) keyHashes[i] = '0x' + m[1]; else await sleep(200); } + } + log(`vote keys: ${keyHashes.join(' ')}`); + const st0 = await n0.eth('igneum_getProvingStatus'); + step('status', { v0: st0.activationDaa, v1: st0.v1 }); + check('the node carries the v1 parameters', hexn(st0.v1.activationDaa) === V1 && hexn(st0.v1.segmentBlocks) === SEG && hexn(st0.v1.unprovenDaa) === UNPROVEN && hexn(st0.v1.aggregatorShareBps) === SHARE_BPS, st0.v1); + + // the v1 start: the first chain block at DAA >= V1 + await waitDaa(n0, V1 + 8, 'past the v1 activation'); + let st = await n0.eth('igneum_getProvingStatus'); + const S = hexn(st.v1.start); + step('v1_start', { start: S, tipDaa: hexn(st.tipDaa) }); + check('every node agrees on the v1 start', (await Promise.all(nodes.map(n => n.eth('igneum_getProvingStatus')))).every(x => hexn(x.v1.start) === S), S); + + // ---- case 1: the first segment, a fresh chain, carried and paid + const seg0 = [S, S + SEG - 1]; + await waitExecuted(n0, seg0[1] + 1); + const stmt0 = await n0.eth('igneum_getSegmentStatement', ['0x' + seg0[0].toString(16)]); + check('segment 0 is aligned at the start and pending', hexn(stmt0.first) === seg0[0] && hexn(stmt0.last) === seg0[1] && stmt0.status.status === 'pending', { first: stmt0.first, last: stmt0.last, status: stmt0.status }); + check('the native statement is the same on every node', (await Promise.all(nodes.map(n => n.eth('igneum_getSegmentStatement', ['0x' + seg0[0].toString(16)])))).every(x => x.publicValuesFresh === stmt0.publicValuesFresh), stmt0.publicValuesFresh.slice(0, 24)); + const expectedWei0 = stmt0.blocks.reduce((a, b) => a + BigInt(b.aggregatorWei), 0n); + const creditSum0 = stmt0.blocks.reduce((a, b) => a + BigInt(b.poolCreditWei), 0n); + check('the aggregator share is a tenth of the segment credits', expectedWei0 === creditSum0 / 10n, { expectedWei0: expectedWei0.toString(), creditSum0: creditSum0.toString() }); + const lastHash0 = stmt0.blocks[stmt0.blocks.length - 1].hash; + const rec0 = signSegment(seg0[0], seg0[1], lastHash0, stmt0.publicValuesFresh, 'seg0'); + const tSubmit0 = Date.now(); + const sub0 = await n1.eth('igneum_submitSegmentRecord', [{ record: rec0.record, proof: rec0.proof }]); + check('known-finished: segment 0 record accepted by n1', sub0.accepted === true && sub0.new === true, sub0); + const paid0 = await waitSegmentPaid(n0, seg0[0]); + const paidAt0 = (Date.now() - tSubmit0) / 1000; + check('known-finished: segment 0 paid on n0 (relayed over p2p, verified in trust mode, carried)', BigInt(paid0.paid.wei) === expectedWei0 && paid0.paid.payout.toLowerCase() === PAYOUT, { paid: paid0.paid, secs: paidAt0 }); + await sleep(3000); + const agree0 = await Promise.all(nodes.map(n => n.eth('igneum_getSegmentRecords', ['0x' + seg0[0].toString(16)]).then(r => r.paid && r.paid.wei))); + check('every node paid segment 0 the same', agree0.every(w => w && BigInt(w) === expectedWei0), agree0); + const bal0 = BigInt(await n0.eth('eth_getBalance', [PAYOUT, 'latest'])); + check('the payout address holds the aggregator share', bal0 === expectedWei0, { balance: bal0.toString() }); + step('case1', { segment: seg0, wei: expectedWei0.toString(), paidSecs: paidAt0, carrier: paid0.paid.carrierNumber }); + + // ---- case 2: the chain rule on segment 1 + const seg1 = [S + SEG, S + 2 * SEG - 1]; + await waitExecuted(n0, seg1[1] + 1); + const stmt1 = await n0.eth('igneum_getSegmentStatement', ['0x' + seg1[0].toString(16)]); + check('segment 1 names segment 0 as its proven previous', stmt1.previous && hexn(stmt1.previous.first) === seg0[0] && hexn(stmt1.previous.chainLen) === SEG, stmt1.previous); + const lastHash1 = stmt1.blocks[stmt1.blocks.length - 1].hash; + const fresh1 = signSegment(seg1[0], seg1[1], lastHash1, stmt1.publicValuesFresh, 'seg1-fresh'); + const subFresh = await n0.eth('igneum_submitSegmentRecord', [{ record: fresh1.record, proof: fresh1.proof }]); + check('chain rule: a fresh-chain record for segment 1 is refused while segment 0 is proven', subFresh.accepted === false && /does not chain to segment/.test(subFresh.reason), subFresh); + const cont1 = signSegment(seg1[0], seg1[1], lastHash1, stmt1.publicValuesContinuing, 'seg1-cont'); + const subCont = await n2.eth('igneum_submitSegmentRecord', [{ record: cont1.record, proof: cont1.proof }]); + check('chain rule: the continuing record (chain_len 8) is accepted', subCont.accepted === true, subCont); + const paid1 = await waitSegmentPaid(n0, seg1[0]); + check('segment 1 paid with chain_len 8', hexn(paid1.paid.chainLen) === 2 * SEG, paid1.paid); + step('case2', { segment: seg1, refused: subFresh.reason, paid: paid1.paid }); + + // ---- case 3: segment 2 left unproven; segment 3 restarts the chain + const seg2 = [S + 2 * SEG, S + 3 * SEG - 1]; + const seg3 = [S + 3 * SEG, S + 4 * SEG - 1]; + await waitExecuted(n0, seg3[1] + 1); + const stmt2 = await n0.eth('igneum_getSegmentStatement', ['0x' + seg2[0].toString(16)]); + const deadline2 = hexn(stmt2.status.deadline_daa); + check('segment 2 is pending with a deadline', stmt2.status.status === 'pending' && deadline2 > 0, stmt2.status); + const stmt3early = await n0.eth('igneum_getSegmentStatement', ['0x' + seg3[0].toString(16)]); + const lastHash3 = stmt3early.blocks[stmt3early.blocks.length - 1].hash; + const fresh3 = signSegment(seg3[0], seg3[1], lastHash3, stmt3early.publicValuesFresh, 'seg3-fresh'); + const subEarly = await n0.eth('igneum_submitSegmentRecord', [{ record: fresh3.record, proof: fresh3.proof }]); + check('known-failed: segment 3 cannot start a fresh chain while segment 2 is pending', subEarly.accepted === false && /pending until DAA/.test(subEarly.reason), subEarly); + await waitDaa(n0, deadline2 + 2, 'past segment 2 deadline'); + const stmt2late = await n0.eth('igneum_getSegmentStatement', ['0x' + seg2[0].toString(16)]); + check('known-failed: segment 2 is unproven after its deadline', stmt2late.status.status === 'unproven', stmt2late.status); + const lastHash2 = stmt2late.blocks[stmt2late.blocks.length - 1].hash; + const late2 = signSegment(seg2[0], seg2[1], lastHash2, stmt2late.publicValuesContinuing || stmt2late.publicValuesFresh, 'seg2-late'); + const subLate = await n0.eth('igneum_submitSegmentRecord', [{ record: late2.record, proof: late2.proof }]); + check('known-failed: a late record for segment 2 pays nothing (refused as unproven)', subLate.accepted === false && /unproven/.test(subLate.reason), subLate); + const subRestart = await n1.eth('igneum_submitSegmentRecord', [{ record: fresh3.record, proof: fresh3.proof }]); + check('segment 3 restarts the chain with a fresh-chain record after the unproven segment', subRestart.accepted === true, subRestart); + const paid3 = await waitSegmentPaid(n0, seg3[0]); + check('segment 3 paid with chain_len 4', hexn(paid3.paid.chainLen) === SEG, paid3.paid); + st = await n0.eth('igneum_getProvingStatus'); + check('the status counts proven and unproven segments', st.v1.segmentsInWindow.proven >= 3 && st.v1.segmentsInWindow.unproven >= 1, st.v1.segmentsInWindow); + step('case3', { unproven: seg2, restarted: seg3, status: st.v1 }); + + // ---- case 4: the shard side after the switch + const work = await n0.eth('igneum_getAssignedShards', [[keyHashes[0]], 40]); + const w = work.find(x => hexn(x.number) >= S); + check('a v1 shard is paid 90% of its block share (one shard a block here)', w && BigInt(w.shardWei) === BigInt(w.poolCreditWei) - BigInt(w.poolCreditWei) / 10n, w && { number: w.number, shardWei: w.shardWei, credit: w.poolCreditWei }); + const before = work.find(x => hexn(x.number) < S); + if (before) check('a pre-v1 shard keeps the whole share', BigInt(before.shardWei) === BigInt(before.poolCreditWei), { number: before.number }); + report.ok = report.checks.every(c => c.ok); + log(`RESULT proving v1 harness: ${report.ok ? 'PASSED' : 'FAILED'} (${report.checks.length} checks) in ${since()} s`); +} catch (e) { + report.error = e.message; + log(`FAILED: ${e.message}`); +} finally { + writeFileSync(`${TMP}/report.json`, JSON.stringify(report, null, 2)); + console.log(JSON.stringify(report, null, 2)); + await stopAll(); + process.exit(report.ok ? 0 : 1); +} diff --git a/tools/proving-v1/pc2-chain.ps1 b/tools/proving-v1/pc2-chain.ps1 new file mode 100644 index 000000000..a56be84e2 --- /dev/null +++ b/tools/proving-v1/pc2-chain.ps1 @@ -0,0 +1,88 @@ +# Proving v1, step 2 (5 October 2026): aggregated chains on PC 2's RTX 5090. A signed `run` job (shell powershell, +# not elevated; the miners keep mining; the live prover in /opt/igneum is untouched: this build lands in /opt/igneum-pv1). +# 1. exports the chain from PC 2's own node (igneum_exportSegments 0..tip) to the job folder +# 2. inside WSL2 (root, Ubuntu-24.04): the fetched package igneum-prove-wsl2-pv1 -> ~/igneum-prove-pv1, every file +# re-stamped, built with the cuda feature against the live build's warm target dir, installed to /opt/igneum-pv1 +# 3. cuts 8 consecutive live fixtures (tip-30 .. tip-23) with the new exporter, runs --mode native on each +# 4. --mode chain over the 8 (shards compressed, each block aggregated with the previous block's proof, verified), +# SP1_PROVER=cuda, a 1-s nvidia-smi sampler underneath for the GPU memory peak +# 5. --mode verify-segment on the final proof (the node's light-verifier path) for its timing +# Every number is a RESULT line. The results JSON is printed at the end (RESULTS-JSON ... END). +$ErrorActionPreference = 'Continue' +function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') } +function Rpc($port, $method, $params) { + try { (Invoke-RestMethod -Method Post -Uri "http://127.0.0.1:$port" -ContentType 'application/json' -Body (@{jsonrpc='2.0'; id=1; method=$method; params=$params} | ConvertTo-Json -Compress -Depth 6) -TimeoutSec 180).result } catch { "rpc error: $_" | Out-Host; $null } +} +$evmPort = $null +foreach ($p in 26790, 26800, 26810) { if (Rpc $p 'igneum_getProvingStatus' @()) { $evmPort = $p; break } } +$job = $env:IGNEUM_JOB_DIR +if (-not $job) { $job = Join-Path $env:TEMP 'igneum-pv1-chain' } +New-Item -ItemType Directory -Force -Path $job | Out-Null +$tipHex = Rpc $evmPort 'eth_blockNumber' @() +$tip = [Convert]::ToInt64($tipHex, 16) +$first = $tip - 30; $last = $first + 7 +"RESULT start $(Stamp) node_evm_port=$evmPort tip=$tip chain blocks $first..$last" +# 1. the export (the whole chain: the exporter replays from genesis) +$t = Get-Date +$body = (@{jsonrpc='2.0'; id=1; method='igneum_exportSegments'; params=@('0x0', ('0x{0:x}' -f $last))} | ConvertTo-Json -Compress) +$seqFile = Join-Path $job 'seq.json' +try { + $resp = Invoke-WebRequest -Method Post -Uri "http://127.0.0.1:$evmPort" -ContentType 'application/json' -Body $body -TimeoutSec 600 -UseBasicParsing + [IO.File]::WriteAllBytes($seqFile, $resp.Content) +} catch { "RESULT export FAILED: $_"; exit 1 } +$len = (Get-Item $seqFile).Length +"RESULT export $(Stamp) $len bytes in $([math]::Round(((Get-Date) - $t).TotalSeconds,1)) s to $seqFile" +# the JSON-RPC envelope: the exporter wants the result object; unwrap with python inside WSL (below) +$pkg = Join-Path $env:LOCALAPPDATA 'igneum\prove\igneum-prove-wsl2-pv1\igneum-prove-wsl2' +if (-not (Test-Path $pkg)) { $pkg = Join-Path $env:LOCALAPPDATA 'igneum\prove\igneum-prove-wsl2-pv1' } +"RESULT package $(Stamp) $pkg exists=$(Test-Path (Join-Path $pkg 'package'))" +function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } } +$pkgW = WslPath $pkg; $jobW = WslPath $job +$bash = @" +set -uo pipefail +export PATH="`$HOME/.cargo/bin:`$HOME/.sp1/bin:`$PATH" +CUDA_DIR="`$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "`$CUDA_DIR" ] && export PATH="`$CUDA_DIR/bin:`$PATH" && export LD_LIBRARY_PATH="`$CUDA_DIR/lib64:/usr/lib/wsl/lib:`${LD_LIBRARY_PATH:-}" +stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; } +PKG='$pkgW'; JOB='$jobW'; DEST="`$HOME/igneum-prove-pv1"; LIVE_TARGET="`$HOME/igneum-prove/proving/igneum-prove/target" +mkdir -p "`$DEST" +rsync -a --delete --exclude target "`$PKG/package/" "`$DEST/" +find "`$DEST" -name target -prune -o -type f -exec touch {} + 2>/dev/null +ls -la "`$DEST/proving/igneum-prove/elf/" | sed 's/^/elf: /' +grep -o '"program_id": "0x[0-9a-f]*"' "`$DEST/proving/igneum-prove/elf/manifest.json" | sed 's/^/RESULT manifest /' +cd "`$DEST/proving/igneum-prove" +echo "RESULT build start `$(stamp) target `$LIVE_TARGET (warm from the live build; the three workspace crates recompile)" +t0=`$(date +%s) +if ! CARGO_TARGET_DIR="`$LIVE_TARGET" cargo build --release -p igneum-prove-export -p igneum-prove-host --features igneum-prove-host/cuda 2>&1 | tail -3; then echo "RESULT build FAILED"; exit 1; fi +echo "RESULT build `$(stamp) exit 0 in `$(( `$(date +%s) - t0 )) s" +mkdir -p /opt/igneum-pv1 && cp "`$LIVE_TARGET/release/igneum-prove-host" "`$LIVE_TARGET/release/igneum-prove-export" /opt/igneum-pv1/ +H=/opt/igneum-pv1/igneum-prove-host; X=/opt/igneum-pv1/igneum-prove-export +echo "RESULT installed `$(sha256sum `$H | cut -c1-16) host, `$(sha256sum `$X | cut -c1-16) export; live /opt/igneum untouched: `$(sha256sum /opt/igneum/igneum-prove-host | cut -c1-16)" +`$H --mode id | sed 's/^/RESULT pv1-host /' +# 3. the fixtures: unwrap the JSON-RPC envelope, cut eight consecutive blocks +python3 -c "import json,sys; d=json.load(open('`$JOB/seq.json')); json.dump(d['result'], open('`$JOB/export.json','w'))" +LIST="" +for n in `$(seq $first $last); do + if ! `$X "`$JOB/export.json" `$n "`$JOB/block-`$n.json" --source "PC 2 live devnet export, proving v1 chain run" 2>&1 | tail -1 | sed "s/^/export `$n: /"; then echo "RESULT cut `$n FAILED"; exit 1; fi + `$H "`$JOB/block-`$n.json" --mode native 2>&1 | grep -E "^RESULT (native|plan)" | sed "s/^/block `$n /" + LIST="`$LIST`${LIST:+,}`$JOB/block-`$n.json" +done +echo "RESULT fixtures `$(stamp) `$LIST" +# 4. the chain on the GPU, with the memory sampler +nvidia-smi --query-gpu=timestamp,index,memory.used,utilization.gpu,power.draw --format=csv,noheader,nounits -l 1 > "`$JOB/smi-chain.csv" 2>/dev/null & +SMI=`$! +t0=`$(date +%s) +SP1_PROVER=cuda RUST_LOG=off `$H --mode chain --chain "`$LIST" --prover 0xCAfc6e74000000000000000000000000000000c2 --out "`$JOB/chain-results.json" 2>&1 | grep -E "^(RESULT|STAGE|igneum-prove-host sources|segment proof written)" | sed 's/^/chain: /' +echo "RESULT chain wall `$(( `$(date +%s) - t0 )) s at `$(stamp)" +kill `$SMI 2>/dev/null; sleep 1 +awk -F', *' '{ if (`$3+0 > max[`$2]) max[`$2]=`$3+0; n[`$2]++ } END { for (i in max) print "RESULT gpu_memory_peak_chain index=" i " samples=" n[i] " memory_used_max_mib=" max[i] }' "`$JOB/smi-chain.csv" +free -m | awk '/Mem:/ {print "RESULT wsl_ram_now total_mb=" `$2 " used_mb=" `$3}' +# 5. the node's verifier path on the final proof +STMT=`$(python3 -c "import json; print(json.load(open('`$JOB/chain-results.json'))['segment_statement'])") +PROOF=`$(python3 -c "import json; print(json.load(open('`$JOB/chain-results.json'))['segment_proof_file'])") +for i in 1 2 3; do `$H --mode verify-segment --proof "`$PROOF" --statement "`$STMT" 2>&1 | grep -E "^RESULT" | sed "s/^/verify-segment run `$i: /"; done +echo "RESULTS-JSON"; cat "`$JOB/chain-results.json"; echo; echo "END" +"@ +$bashFile = Join-Path $job 'chain.sh' +[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false)) +& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') } +"RESULT end $(Stamp)" diff --git a/tools/proving-v1/pc2-memory-sweep.ps1 b/tools/proving-v1/pc2-memory-sweep.ps1 new file mode 100644 index 000000000..42cb96f49 --- /dev/null +++ b/tools/proving-v1/pc2-memory-sweep.ps1 @@ -0,0 +1,76 @@ +# Proving v1, the 12 GB requirement (5 October 2026, Josh: "make sure we can prove on 12gb cards"): the GPU memory peak +# of one shard proof on PC 2's RTX 5090 under SP1 6.8.1's knobs, the miners STOPPED by the job (stop_miners_first) and +# the live prover switched off for the run (its sp1-gpu-server would otherwise be the one the client connects to, with +# the live environment, not this one). Every config: the server killed, a 1-s nvidia-smi sampler, one compressed shard +# proof, the peak and the time as a RESULT line. Fixtures: the empty live shard (block 83616, the chain job's cut), +# the full shard at S_p (block-338-shard1, 6.75 M pgas) and block 344 (27 M pgas) re-cut at S_p/2 and S_p/4. +# Leaves the prover ON. The runner restores the miners. +$ErrorActionPreference = 'Continue' +$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' } +if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' } +$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/') +function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') } +function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } } +"RESULT start $(Stamp) prover off for the sweep: $(Prove $false)" +Start-Sleep -Seconds 45 +"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,power.draw --format=csv,noheader,nounits 2>$null) -join ' | ')" +$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-pv1-mem' }; New-Item -ItemType Directory -Force -Path $job | Out-Null +function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } } +$jobW = WslPath $job +$emptyW = WslPath (Join-Path $env:LOCALAPPDATA 'igneum\app\jobs\chain-pc2-pv1c\block-83616.json') +$bash = @" +set -uo pipefail +export PATH="`$HOME/.cargo/bin:`$HOME/.sp1/bin:`$PATH" +CUDA_DIR="`$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "`$CUDA_DIR" ] && export PATH="`$CUDA_DIR/bin:`$PATH" && export LD_LIBRARY_PATH="`$CUDA_DIR/lib64:/usr/lib/wsl/lib:`${LD_LIBRARY_PATH:-}" +stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; } +JOB='$jobW'; H=/opt/igneum-pv1/igneum-prove-host; X=/opt/igneum-pv1/igneum-prove-export; FX="`$HOME/igneum-prove-pv1/proving/fixtures" +EMPTY='$emptyW'; FULL="`$FX/block-338-shard1.json" +pkill -f sp1-gpu-server 2>/dev/null; sleep 2 +echo "RESULT servers_before `$(pgrep -a sp1-gpu-server | tr '\n' ' ' || echo none)" +# the S_p/2 and S_p/4 cuts of block 344 (27 M pgas: 8 and 16 shards), from the package's own export +SEQ="`$HOME/igneum-prove-pv1/tools/prove-fixtures/seq.json" +if [ -f "`$SEQ" ]; then + `$X "`$SEQ" 344 "`$JOB/block-344-half.json" --budget 3750000 --source "block 344 cut at S_p/2 for the 12 GB sweep" 2>&1 | tail -1 + `$X "`$SEQ" 344 "`$JOB/block-344-quarter.json" --budget 1875000 --source "block 344 cut at S_p/4 for the 12 GB sweep" 2>&1 | tail -1 +fi +run() { # name fixture env... + local name="`$1" fx="`$2"; shift 2 + pkill -f sp1-gpu-server 2>/dev/null; sleep 2 + local csv="`$JOB/smi-`$name-`$(basename `$fx .json).csv" + nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv,noheader,nounits -l 1 > "`$csv" 2>/dev/null & + local SMI=`$! + local t0=`$(date +%s) + env SP1_PROVER=cuda RUST_LOG=off "`$@" `$H "`$fx" --mode compressed --shard 0 --out "`$JOB/res-`$name-`$(basename `$fx .json).json" > "`$JOB/log-`$name-`$(basename `$fx .json).txt" 2>&1 + local rc=`$? + local wall=`$(( `$(date +%s) - t0 )) + kill `$SMI 2>/dev/null; sleep 1 + local peak=`$(awk -F', *' '{ if (`$2+0 > m) m=`$2+0 } END { print m+0 }' "`$csv") + local n=`$(wc -l < "`$csv") + local line=`$(grep -E "^RESULT compressed shard" "`$JOB/log-`$name-`$(basename `$fx .json).txt" | tail -1 | sed -E 's/.*prove ([0-9.]+) s, proof ([0-9]+) bytes, verify ([0-9.]+) s, ([A-Z ]+);.*/prove_s=\1 bytes=\2 verify_s=\3 \4/') + local cyc=`$(grep -E "^RESULT execute shard" "`$JOB/log-`$name-`$(basename `$fx .json).txt" | tail -1 | sed -E 's/.*: ([0-9]+) cycles.*/\1/') + local err=`$(grep -iE "error|panick|out of memory|OOM" "`$JOB/log-`$name-`$(basename `$fx .json).txt" | head -1 | cut -c1-160) + echo "RESULT sweep cfg=`$name fixture=`$(basename `$fx .json) peak_mib=`$peak samples=`$n wall_s=`$wall cycles=`${cyc:-na} `${line:-no_result} exit=`$rc env='`$*' `${err:+err=`$err}" +} +W1="SP1_WORKER_NUM_CORE_WORKERS=1 SP1_WORKER_CORE_BUFFER_SIZE=1 SP1_WORKER_NUM_RECURSION_PROVER_WORKERS=1 SP1_WORKER_RECURSION_PROVER_BUFFER_SIZE=1 SP1_WORKER_NUM_RECURSION_EXECUTOR_WORKERS=1 SP1_WORKER_RECURSION_EXECUTOR_BUFFER_SIZE=1 SP1_WORKER_NUM_PREPARE_REDUCE_WORKERS=1 SP1_WORKER_PREPARE_REDUCE_BUFFER_SIZE=1 SP1_WORKER_NUM_SETUP_WORKERS=1 SP1_WORKER_SETUP_BUFFER_SIZE=1 SP1_WORKER_NUM_DEFERRED_WORKERS=1 SP1_WORKER_DEFERRED_BUFFER_SIZE=1 SP1_WORKER_NUM_SPLICING_WORKERS=1 SP1_WORKER_SPLICING_BUFFER_SIZE=1" +W2="SP1_WORKER_NUM_CORE_WORKERS=2 SP1_WORKER_CORE_BUFFER_SIZE=2 SP1_WORKER_NUM_RECURSION_PROVER_WORKERS=2 SP1_WORKER_RECURSION_PROVER_BUFFER_SIZE=2 SP1_WORKER_NUM_RECURSION_EXECUTOR_WORKERS=2 SP1_WORKER_RECURSION_EXECUTOR_BUFFER_SIZE=2 SP1_WORKER_NUM_PREPARE_REDUCE_WORKERS=2 SP1_WORKER_PREPARE_REDUCE_BUFFER_SIZE=2" +run baseline "`$EMPTY" +run baseline "`$FULL" +run elem27 "`$FULL" ELEMENT_THRESHOLD=134217728 +run elem26 "`$FULL" ELEMENT_THRESHOLD=67108864 HEIGHT_THRESHOLD=2097152 +run workers1 "`$FULL" `$W1 +run workers2 "`$FULL" `$W2 +run w1elem27 "`$FULL" ELEMENT_THRESHOLD=134217728 `$W1 +run w1elem26 "`$FULL" ELEMENT_THRESHOLD=67108864 HEIGHT_THRESHOLD=2097152 `$W1 +run w1elem26chunk "`$FULL" ELEMENT_THRESHOLD=67108864 HEIGHT_THRESHOLD=2097152 MINIMAL_TRACE_CHUNK_THRESHOLD=4194304 TRACE_CHUNK_SLOTS=2 `$W1 +run w1elem25 "`$FULL" ELEMENT_THRESHOLD=33554432 HEIGHT_THRESHOLD=1048576 `$W1 +run w1elem26 "`$EMPTY" ELEMENT_THRESHOLD=67108864 HEIGHT_THRESHOLD=2097152 `$W1 +[ -f "`$JOB/block-344-half.json" ] && run w1elem26 "`$JOB/block-344-half.json" ELEMENT_THRESHOLD=67108864 HEIGHT_THRESHOLD=2097152 `$W1 +[ -f "`$JOB/block-344-quarter.json" ] && run w1elem26 "`$JOB/block-344-quarter.json" ELEMENT_THRESHOLD=67108864 HEIGHT_THRESHOLD=2097152 `$W1 +[ -f "`$JOB/block-344-quarter.json" ] && run baseline "`$JOB/block-344-quarter.json" +pkill -f sp1-gpu-server 2>/dev/null +echo "RESULT sweep_end `$(stamp)" +"@ +$bashFile = Join-Path $job 'sweep.sh' +[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false)) +& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') } +"RESULT end $(Stamp) prover back on: $(Prove $true)" diff --git a/tools/proving-v1/pc2-prover-cost.ps1 b/tools/proving-v1/pc2-prover-cost.ps1 new file mode 100644 index 000000000..a21565eab --- /dev/null +++ b/tools/proving-v1/pc2-prover-cost.ps1 @@ -0,0 +1,111 @@ +# Proving v1, step 1 (5 October 2026): what the prover costs a mining machine. A signed `run` job for PC 2 +# (app/igneum-app/src/jobrun.rs, shell powershell, not elevated). Never stops the miners. +# 1. waits (up to 20 min) for the NVIDIA worker to report a hash rate, so both phases see the same miner +# 2. prover OFF (POST /api/prove {"on":false}): 5 min of samples every 15 s +# 3. prover ON: 5 min of samples; a 1-s nvidia-smi sampler runs underneath for the GPU memory peak +# 4. inside WSL2: the sp1-gpu-server's version and its compiled SM targets (cuobjdump, strings) +# 5. leaves the prover ON +# Every number is a RESULT line; the per-sample lines are SAMPLE lines. Read with `node tools/jobs.mjs `. +$ErrorActionPreference = 'Continue' +$base = (Get-Content (Join-Path $env:LOCALAPPDATA 'igneum\app\app.url') -Raw).Trim().TrimEnd('/') +function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') } +function State { try { Invoke-RestMethod -Uri "$base/api/state" -TimeoutSec 10 } catch { $null } } +function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } } +function Rpc($port, $method, $params) { + try { (Invoke-RestMethod -Method Post -Uri "http://127.0.0.1:$port" -ContentType 'application/json' -Body (@{jsonrpc='2.0'; id=1; method=$method; params=$params} | ConvertTo-Json -Compress -Depth 6) -TimeoutSec 20).result } catch { $null } +} +function Smi { try { (& nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,power.draw --format=csv,noheader,nounits 2>$null) -join ' | ' } catch { 'nvidia-smi failed' } } +function Ram { + $os = Get-CimInstance Win32_OperatingSystem + $usedMb = [math]::Round(($os.TotalVisibleMemorySize - $os.FreePhysicalMemory) / 1024) + $vm = (Get-Process -ErrorAction SilentlyContinue | Where-Object { $_.ProcessName -like 'vmmem*' } | Measure-Object WorkingSet64 -Sum).Sum + $vmMb = if ($vm) { [math]::Round($vm / 1MB) } else { 0 } + "host_used_mb=$usedMb total_mb=$([math]::Round($os.TotalVisibleMemorySize / 1024)) vmmem_mb=$vmMb" +} +function Card($st) { if ($null -eq $st) { return $null }; $st.mining.cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1 } +function Sample($phase, $i) { + $st = State; $c = Card $st; $pv = if ($st) { $st.proving } else { $null } + $cards = if ($st) { ($st.mining.cards | ForEach-Object { "$($_.name):$($_.state):$([math]::Round($_.hash_now,1))" }) -join ',' } else { 'no state' } + "SAMPLE $phase $i $(Stamp) nvidia_hash_now=$(if ($c) { [math]::Round($c.hash_now,2) } else { 'na' }) nvidia_hash_avg=$(if ($c) { [math]::Round($c.hash_avg,2) } else { 'na' }) nvidia_state=$(if ($c) { $c.state } else { 'na' }) restarts=$(if ($c) { $c.restarts } else { 'na' }) cards=$cards proving=$(if ($pv) { "$($pv.status)/proved=$($pv.proved)/submitted=$($pv.submitted)/last_prove_s=$([math]::Round($pv.last_prove_s,1))/current='$($pv.current)'" } else { 'na' }) smi=[$(Smi)] $(Ram)" +} +$evmPort = $null +foreach ($p in 26790, 26800, 26810) { if (Rpc $p 'igneum_getProvingStatus' @()) { $evmPort = $p; break } } +$st0 = State +"RESULT start $(Stamp) app $($st0.version) machine $($st0.machine_id) node_evm_port=$evmPort prove_setting=$($st0.settings.prove) proving_status=$($st0.proving.status)/$($st0.proving.backend)" +"RESULT gpus $(Stamp) $(Smi)" +# 1. the NVIDIA worker must be hashing before anything is measured (5 October 2026, 18:35Z: the 5090 worker was +# exiting on a pack seed mismatch; a baseline of 0 MH/s is not a baseline) +$waited = 0 +while ($waited -lt 1200) { + $c = Card (State) + if ($c -and $c.hash_now -gt 10) { break } + if ($waited % 120 -eq 0) { "WAIT $(Stamp) nvidia worker: state=$(if ($c) { $c.state } else { 'na' }) hash_now=$(if ($c) { $c.hash_now } else { 'na' }) restarts=$(if ($c) { $c.restarts } else { 'na' }) (waiting up to 20 min)" } + Start-Sleep -Seconds 20; $waited += 20 +} +$c = Card (State) +$minerUp = ($c -and $c.hash_now -gt 10) +"RESULT miner $(Stamp) nvidia worker $(if ($minerUp) { 'hashing' } else { 'NOT hashing after 20 min: the mining-cost rows below are void' }) hash_now=$(if ($c) { $c.hash_now } else { 'na' }) after waiting $waited s" +$paid0 = Rpc $evmPort 'igneum_getProvingStatus' @() +# 2. prover off +"RESULT prove_off $(Stamp) $(Prove $false)" +Start-Sleep -Seconds 30 +$off = @(); for ($i = 1; $i -le 20; $i++) { $line = Sample 'off' $i; $line; $off += (State); Start-Sleep -Seconds 15 } +$pvOffEnd = (State).proving +# 3. prover on, with a 1-s GPU memory sampler underneath +$smiFile = Join-Path $env:TEMP "igneum-pv1-smi-$(Get-Date -Format yyyyMMdd-HHmmss).csv" +$smiProc = Start-Process -FilePath 'nvidia-smi' -ArgumentList '--query-gpu=timestamp,index,memory.used,utilization.gpu,power.draw --format=csv,noheader,nounits -l 1' -RedirectStandardOutput $smiFile -NoNewWindow -PassThru +"RESULT prove_on $(Stamp) $(Prove $true) (gpu sampler pid $($smiProc.Id) -> $smiFile)" +$on = @(); for ($i = 1; $i -le 20; $i++) { $line = Sample 'on' $i; $line; $on += (State); Start-Sleep -Seconds 15 } +try { Stop-Process -Id $smiProc.Id -Force -ErrorAction SilentlyContinue } catch {} +Start-Sleep -Seconds 2 +$pvOnEnd = (State).proving +$paid1 = Rpc $evmPort 'igneum_getProvingStatus' @() +function Stats($states) { + $h = @(); foreach ($s in $states) { $c = Card $s; if ($c) { $h += [double]$c.hash_now } } + if ($h.Count -eq 0) { return 'n=0' } + $m = ($h | Measure-Object -Average -Minimum -Maximum) + $sorted = $h | Sort-Object; $p50 = $sorted[[math]::Floor(($sorted.Count - 1) / 2)] + "n=$($h.Count) mean=$([math]::Round($m.Average,2)) p50=$([math]::Round($p50,2)) min=$([math]::Round($m.Minimum,2)) max=$([math]::Round($m.Maximum,2)) MH/s" +} +"RESULT mining_alone $(Stamp) nvidia hash_now over 5 min: $(Stats $off); proved during the phase: $($pvOffEnd.proved - $off[0].proving.proved)" +"RESULT mining_and_proving $(Stamp) nvidia hash_now over 5 min: $(Stats $on); proved during the phase: $($pvOnEnd.proved - $on[0].proving.proved) shards (submitted $($pvOnEnd.submitted - $on[0].proving.submitted)), last_prove_s=$($pvOnEnd.last_prove_s); node paidShards $($paid0.paidShards) -> $($paid1.paidShards)" +# the GPU memory peak from the 1-s sampler: per card index, max memory.used (MiB) and the sample count +try { + $rows = Get-Content $smiFile | Where-Object { $_ -match ',' } | ForEach-Object { $f = $_ -split ',\s*'; [pscustomobject]@{ t = $f[0]; idx = $f[1]; mem = [int]$f[2]; util = [int]$f[3]; power = [double]$f[4] } } + foreach ($g in ($rows | Group-Object idx)) { + $mx = ($g.Group | Measure-Object mem -Maximum).Maximum; $mn = ($g.Group | Measure-Object mem -Minimum).Minimum; $ut = ($g.Group | Measure-Object util -Average).Average; $pw = ($g.Group | Measure-Object power -Maximum).Maximum + "RESULT gpu_memory_peak $(Stamp) index=$($g.Name) samples=$($g.Count) memory_used_min_mib=$mn memory_used_max_mib=$mx util_mean_pct=$([math]::Round($ut,1)) power_max_w=$pw (1-s nvidia-smi samples while mining with the prover on)" + } +} catch { "RESULT gpu_memory_peak error: $_" } +# host RAM peak over both phases, from the 15-s samples (host used and the WSL2 VM's working set) +$ramPeakOff = 0; $vmPeakOff = 0; $ramPeakOn = 0; $vmPeakOn = 0 +# (re-sampled here from the SAMPLE lines' fields was simpler; the per-sample Ram() strings are above; compute again from a fresh sample for the record) +"RESULT ram_now $(Stamp) $(Ram)" +# 4. the SP1 GPU server inside WSL2: version and compiled SM targets +$bash = @' +export PATH="$HOME/.cargo/bin:$HOME/.sp1/bin:$PATH" +CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH" +f="$HOME/.sp1/bin/sp1-gpu-server" +echo "RESULT gpu_server_file $(ls -la "$f" 2>&1)" +echo "RESULT gpu_server_version $("$f" --version 2>&1 | head -2 | tr '\n' ' ')" +echo "RESULT gpu_server_sha256 $(sha256sum "$f" 2>/dev/null | cut -c1-64)" +echo "RESULT nvcc $(nvcc --version 2>&1 | tail -1)" +echo "RESULT wsl_nvidia_smi $(nvidia-smi --query-gpu=name,memory.total,driver_version --format=csv,noheader 2>&1 | head -1)" +if command -v cuobjdump >/dev/null; then + echo "RESULT cuobjdump_elf $(cuobjdump --list-elf "$f" 2>&1 | grep -oE 'sm_[0-9]+' | sort | uniq -c | tr '\n' ' ')" + echo "RESULT cuobjdump_ptx $(cuobjdump --list-ptx "$f" 2>&1 | grep -oE 'sm_[0-9]+|compute_[0-9]+' | sort | uniq -c | tr '\n' ' ')" + echo "RESULT cuobjdump_head $(cuobjdump --list-elf "$f" 2>&1 | head -5 | tr '\n' ' ')" +else + echo "RESULT cuobjdump missing on PATH ($PATH)" +fi +echo "RESULT strings_sm $(strings "$f" 2>/dev/null | grep -oE '\b(sm|compute)_[0-9]{2,3}\b' | sort | uniq -c | tr '\n' ' ')" +echo "RESULT strings_arch_hints $(strings "$f" 2>/dev/null | grep -iE 'cuda_arch|gencode|arch=|--generate-code|nvcc' | sort -u | head -8 | tr '\n' ' ' | cut -c1-600)" +echo "RESULT host_elf_ids $(/opt/igneum/igneum-prove-host --mode id 2>&1 | tail -1)" +echo "RESULT free $(free -m | awk '/Mem:/ {print "wsl_total_mb="$2" used_mb="$3}')" +'@ +$bashFile = Join-Path $env:TEMP 'igneum-pv1-arch.sh' +[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false)) +$wslPath = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($bashFile -replace '\\', '/') 2>$null) +if (-not $wslPath) { $wslPath = '/mnt/c' + ($bashFile.Substring(2) -replace '\\', '/') } +& wsl.exe -d Ubuntu-24.04 -u root -- bash $wslPath 2>&1 | ForEach-Object { ($_ -replace "`0", '') } +"RESULT end $(Stamp) prover left ON: $(Prove $true); app proving status $((State).proving.status)" From 13dc9ce24ea06de8bc6247c8d5e5a244fb5c66a9 Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:16:47 +0100 Subject: [PATCH 011/311] Proving v1: the harness passes (21 checks), the bench-log entry with the step 1 GPU memory, RAM and SM-target numbers, the CPU chain of 2 and verify-segment Co-Authored-By: Claude Fable 5.1 --- docs/bench-log.md | 38 +++ tools/proving-v1/net.mjs | 2 +- tools/proving-v1/pc2-prover-cost.ps1 | 10 +- tools/proving-v1/report-2026-10-05.json | 332 ++++++++++++++++++++++++ 4 files changed, 379 insertions(+), 3 deletions(-) create mode 100644 tools/proving-v1/report-2026-10-05.json diff --git a/docs/bench-log.md b/docs/bench-log.md index 0ecb53182..834986c21 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1586,3 +1586,41 @@ The 3-node run (`tools/txgen/relay-net.mjs`, new; this Mac, load 7 to 8 at the e Reading. Every transaction given to A was mined by B or C within 6 s, two thirds of them within 3 s, with no skipped copy: the hold on block-added kept B's and C's parallel blocks from carrying the same transfer. The afternoon run on the live devnet, through one node with the cooldown, had p50 40.7 s and p90 110.8 s with 50-s quiet stretches; here the 1.5 s p50 is one fast-time block plus the relay and the executor's lag. The pool depth matching on all three nodes at every sample is the convergence. Not measured here: a transaction flood above the per-peer rate (the bucket is unit-tested only), a 13 peer in the fleet (the digest check and the version gate are the evidence), and the hold's 30-s expiry on a block that never reaches the chain (not seen in 149 chain blocks). Commands: `IGNEUMD=vendor/igneum-node/target-txgossip/release/igneumd IGNEUM_MINER=vendor/igneum-node/target-txgossip/release/igneum-miner tools/lock/with-lock.sh run node tools/txgen/relay-net.mjs --rate 2 --duration 120 --wallets 16 --fund 2`; the Mac binaries from the fork worktree with `CARGO_TARGET_DIR=vendor/igneum-node/target-txgossip cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow` under the build lock (an APFS clone of `target-036`, 2 min 15 s to clone, 5 min 06 s to build); the suites with `node tools/build-job.mjs run --target 1ccfe586 --node vendor/igneum-node-txgossip --targets linux --node-tests "igneum-exec kaspa-p2p-flows" --no-app`. +## 5 October 2026 (evening), proving v1: segment records, the chain rule, the unproven rule; what was measured tonight (proving engineer) + +Branches `proving-v1` (main repository, worktree `igneum-wt-proving-v1`; fork `vendor/igneum-node-pv1` from a24ab01a). Rules: spec 7.8; plan `docs/plans/proving-v1.md`. Every row names its command. The live devnet was in a degraded state the whole evening: from 18:35Z the RTX 5090 workers on PC 1 and PC 2 exited at start on a pack seed mismatch (`the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT`, restart 60+ on PC 2 by 19:05Z, another agent's branch `pack-loop`), the Mac app node was down from 17:45Z, so PC 2 mined 3.4 MH/s from its iGPU and PC 2's prover was the only prover; the coordinator held every PC 2 measurement at 19:00Z until the fleet mines again. + +### Step 1, the prover default and its cost + +| What | Measured | +|---|---| +| The default rule (`app/igneum-app/src/provedefault.rs`) | `cargo test --release -p igneum-app provedefault` on this Mac (the app crate, build lock, 19:05Z): 5 passed (a 5090 with WSL2 on Windows is on; Windows without WSL2 off with the Set up hint; Linux needs no WSL2 and the 12 GB gate holds, a 10 GB 3080 and a 16 GB AMD card stay off; Apple silicon off; the biggest qualifying card is named) | +| Mining alone against mining with the prover (PC 2 job `prover-cost-pc2-pv1`, `tools/proving-v1/pc2-prover-cost.ps1`, 5 + 5 min, published 18:40:48Z, ran 18:41:13Z) | VOID: the job waited 20 min for the 5090 worker to hash and it never did (the pack fault above); "mining alone" was 0 MH/s. Re-run after the coordinator's go. The job's `/api/state` read also came back empty on PC 2 (`cards=::0`, every sample) while `igneum_getProvingStatus` on 127.0.0.1:26790 answered; the re-run script prints the URL file and the error | +| GPU memory during proving (the same job, phase B: the prover on for 5 min, 298 one-second `nvidia-smi --query-gpu=memory.used` samples, the 5090 worker dead so the card held nothing else) | memory.used min 1,654 MiB, max 13,816 MiB, utilisation mean 2.7%, power max 190.6 W. The shards were the devnet's empty shards (0 pgas); the peak on a full shard at `S_p` is the chain job's row (held). Read: a 16 GB card clears tonight's peak, a 12 GB card does not, and the 24 GB of a 4090 has 10 GB of headroom | +| Shards per minute with the mining worker dead | the node's `paidShards` 510 -> 518 over the 5-min phase: 1.6 shards a minute from one 5090 through the app's loop (export, cut, prove, sign, submit) | +| Host RAM (Windows `Win32_OperatingSystem` and the `vmmem` working set, sampled every 15 s) | host used 25,550 MB of 63,132 MB at the end; the WSL2 VM's working set 7,915 MB (2,334 MB used of 30,914 MB inside the distribution) | +| The SP1 GPU server's compiled targets (`cuobjdump --list-elf /root/.sp1/bin/sp1-gpu-server` inside PC 2's Ubuntu-24.04, CUDA 12.8, driver 610.47) | `sp1-gpu-server` 6.8.1 (251,306,680 bytes, sha256 c2642ad1c42e85d8525159cf0c7cd5200d8766c9be1283f452a1f9bf9fea725c, the asset `sp1_gpu_server_v6.8.1_x86_64.tar.gz` the SDK downloads, `sp1-cuda-6.8.1/src/server.rs`): one ELF each for sm_80, sm_86, sm_89, sm_90, sm_100 and sm_120; `strings` finds compute_120 PTX as well. So sm_89 (Ada: RTX 4090, 4080) is compiled in natively, no JIT; so are Ampere (3090, 3060), Hopper, Blackwell datacentre (sm_100) and consumer (sm_120, the 5090). Nothing for AMD (no HIP path in SP1) | + +### Step 2, aggregated chains + +| What | Measured | +|---|---| +| The new host (`--mode chain`, `aggregate`, `verify-segment`) against every fixture natively | `igneum-prove-host --mode native` on the Mac for the 12 fixtures of `proving/fixtures/` (9 block, 3 fee-switch), host built from this branch 19:06Z: every one MATCHES (the package gate's native half); `--mode id`: shard `0x2b1a81cb...`, aggregator `0x474678f3...`, the 0.3.9 pin, unchanged | +| Eight consecutive live fixtures | `igneum_exportSegments 0x0..0x13cb4` on node 1's exec RPC (127.0.0.1:26790, read-only, 20:06 BST, tip 81,076): 71,042,616 bytes, 81,077 segments, 28 accounts, 0.5 s; `igneum-prove-export export.json block-.json` for 81046..81053: replayed 81,077 segments from genesis in 1.8 s each, every state root equal to the node's; one shard a block, 0 pgas (no transactions on the devnet tonight), `proving/fixtures/chain/` | +| Chain of 2 on the Mac CPU (the known-finished case of `--mode chain` before the GPU; M5 Max under the live nodes, the harness and two builds) | `SP1_PROVER=cpu igneum-prove-host --mode chain --chain block-81046.json,block-81047.json --out results.json` under the run lock, 19:07:48Z to 19:11:28Z: setup 12.2 s; block 81046: shard 0 compressed 55.4 s (1,272,897 bytes, verify 0.036 s), aggregate 52.0 s (1,272,909 bytes, verify 0.031 s), chain_len 1, agg_vk zero; block 81047: shard 41.3 s, aggregate WITH the previous block proof 59.1 s, chain_len 2, agg_vk = the pinned aggregator id; end to end 207.9 s; final proof 1,272,909 bytes, statement 0x232276f4... The recursion over the previous proof cost 7 s more than the first aggregation on this CPU | +| `--mode verify-segment` on that proof (the node's path: SP1 light verifier, pinned aggregator key) | VERIFIED in 0.032 s (0.27 s wall, three runs: 0.033, 0.032, 0.032); known-failed: a wrong statement NOT VERIFIED (0.032 s); the shard verifier (`--mode verify`) on the segment proof NOT VERIFIED, "program id 0x474678f3... IS NOT OURS 0x2b1a81cb..." | +| Chain of 8 on the RTX 5090 (N = 2, 4, 8) | HELD with step 1 (`tools/proving-v1/pc2-chain.ps1`, package `igneum-prove-wsl2-pv1.zip` eb6dccf8..., 1.5 MB, built from this branch with the gate's native half run above) | + +### Step 3, coverage + +| What | Measured | +|---|---| +| A 3-minute window at 18:57Z on node 1 (`node tools/proving-v1/coverage.mjs --minutes 3`, chain blocks 80754..80839, 86 blocks) | 4 blocks with a paid shard (4.7%), 4 fully proven, 4 of 86 shards; on-chain latency (carrier timestamp minus block timestamp) n 4: min 36 s, p50 39 s, max 44 s; 0 content blocks. One prover (PC 2), the Mac verifier node down, PC 2 producing few blocks (3.4 MH/s): the degraded state above, not the fleet's number | +| The 30-minute window | waits for the fleet (the coordinator's go) | + +### Step 4, the rule + +| What | Measured | +|---|---| +| Unit tests | `cargo test --release -p kaspa-consensus-core -p igneum-exec --lib -- proving config::params::tests::override_params_carry_the_proving_v1 config::params::tests::consensus_digest` on this Mac (target `vendor/igneum-node/target-pv1`, 19:09Z): consensus core 13 passed (the segment record round trip, signature and the three nested sections; the credit split; the params switch and the digest that moves only once the switch is set), exec 8 passed (the segment grid and the split; the record checks: alignment, block, chain length, the veto naming the field, the deadline, the window; the chain rule both ways; the unproven restart; the shard side at 90%; the pool offering the segment section). The six full node suites go to PC 2 as a build job when the fleet is back | +| The fast-time 3-node harness (`tools/proving-v1/net.mjs`, 29950+, suffix 956, every node in trust mode, three vmine voters, v0 at DAA 60, v1 at DAA 120, 4 blocks a segment, unproven after 60 DAA, a tenth to the aggregator; fork b177718e built on this Mac) | run 2, 19:13:01Z to 19:16:19Z, under the run lock: PASSED, 21 checks in 197.3 s (`tools/proving-v1/report-2026-10-05.json`). v1 start = chain block 119 on all three nodes; the native statement identical on all three. Known-finished: segment 119..122's fresh-chain record submitted to n1 at t=131.1 s, relayed, verified (trust) and PAID on n0 1.0 s later at chain block 129, 253,611,648,000,000,000 wei = a tenth of the four credits, the same on every node, the payout address holding it. Chain rule: segment 123..126's fresh-chain record refused ("does not chain to segment 119..122 ... proven (record paid at chain block 129)"), the continuing one (chain_len 8) accepted and paid. Known-failed: segment 127..130 left without a record: a fresh-chain record for 131..134 refused while 127..130 was pending ("pending until DAA 191"); at DAA 192 the status read unproven, a late record for 127..130 refused ("unproven: carried after the deadline"), the fresh-chain record for 131..134 accepted and paid with chain_len 4; `segmentsInWindow` proven 3, unproven 1. The shard side: a v1 shard's `shardWei` = 90% of its block's credit. Run 1 (19:10Z) failed in its own tooling (the signer's argument order), fixed | diff --git a/tools/proving-v1/net.mjs b/tools/proving-v1/net.mjs index fd7d7271f..6ba403324 100644 --- a/tools/proving-v1/net.mjs +++ b/tools/proving-v1/net.mjs @@ -122,7 +122,7 @@ async function waitExecuted(n, number) { // signs a segment record over the given public values with v0's key; the proof bytes are a placeholder (trust mode) function signSegment(first, last, hash, pv, tag) { const proof = Buffer.from(`igneum-proving-v1-harness-${tag}-${first}-${last}`); - const sg = run(MINER, ['sign-segment-record', 'v0', CHAIN_NAME, hash, String(first), String(last), PAYOUT, pv, sha256(proof)]); + const sg = run(MINER, ['sign-segment-record', 'v0', CHAIN_NAME, String(first), String(last), hash, PAYOUT, pv, sha256(proof)]); if (sg.code !== 0) throw new Error(`sign-segment-record failed: ${sg.out}`); const signed = JSON.parse(sg.out.trim().split('\n').pop()); return { record: signed.record, proof: '0x' + proof.toString('hex'), keyHash: signed.keyHash, statement: signed.statement }; diff --git a/tools/proving-v1/pc2-prover-cost.ps1 b/tools/proving-v1/pc2-prover-cost.ps1 index a21565eab..82aa21d9f 100644 --- a/tools/proving-v1/pc2-prover-cost.ps1 +++ b/tools/proving-v1/pc2-prover-cost.ps1 @@ -7,9 +7,15 @@ # 5. leaves the prover ON # Every number is a RESULT line; the per-sample lines are SAMPLE lines. Read with `node tools/jobs.mjs `. $ErrorActionPreference = 'Continue' -$base = (Get-Content (Join-Path $env:LOCALAPPDATA 'igneum\app\app.url') -Raw).Trim().TrimEnd('/') +$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' } +if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' } +$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/') function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') } -function State { try { Invoke-RestMethod -Uri "$base/api/state" -TimeoutSec 10 } catch { $null } } +$script:stateErr = '' +function State { try { Invoke-RestMethod -Uri "$base/api/state" -TimeoutSec 20 } catch { $script:stateErr = "$_"; $null } } +"RESULT app_url_file $(Stamp) $urlFile exists=$(Test-Path $urlFile) base_len=$($base.Length)" +$probe = State +"RESULT state_probe $(Stamp) ok=$($null -ne $probe) version=$($probe.version) cards=$(($probe.mining.cards | Measure-Object).Count) err=$script:stateErr" function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } } function Rpc($port, $method, $params) { try { (Invoke-RestMethod -Method Post -Uri "http://127.0.0.1:$port" -ContentType 'application/json' -Body (@{jsonrpc='2.0'; id=1; method=$method; params=$params} | ConvertTo-Json -Compress -Depth 6) -TimeoutSec 20).result } catch { $null } diff --git a/tools/proving-v1/report-2026-10-05.json b/tools/proving-v1/report-2026-10-05.json new file mode 100644 index 000000000..440a30412 --- /dev/null +++ b/tools/proving-v1/report-2026-10-05.json @@ -0,0 +1,332 @@ +{ + "v0": 60, + "v1": 120, + "segment": 4, + "unproven": 60, + "share_bps": 1000, + "steps": { + "status": { + "v0": "0x3c", + "v1": { + "activationDaa": "0x78", + "active": false, + "aggregatorId": null, + "aggregatorShareBps": "0x3e8", + "paidSegmentWei": "0x0", + "paidSegments": 0, + "pool": { + "entries": 0, + "failed": 0, + "pending": 0, + "verified": 0 + }, + "segmentBlocks": "0x4", + "segmentsInWindow": { + "pending": 0, + "proven": 0, + "unproven": 0 + }, + "shardProgramId": null, + "start": null, + "unprovenDaa": "0x3c" + } + }, + "v1_start": { + "start": 119, + "tipDaa": 128 + }, + "case1": { + "segment": [ + 119, + 122 + ], + "wei": "253611648000000000", + "paidSecs": 1.008, + "carrier": "0x81" + }, + "case2": { + "segment": [ + 123, + 126 + ], + "refused": "segment 123..126 does not chain to segment 119..122 (chain_len 4), which is proven (record paid at chain block 129)", + "paid": { + "carrierNumber": "0x87", + "chainLen": "0x8", + "keyHash": "0xc379ba2cdb239d2884d6bdb391d243ddfa869ab95ec4e72372e0e31fc9025fa0", + "payout": "0x4343434343434343434343434343434343434343", + "wei": "0x38505a6ce4a8000" + } + }, + "case3": { + "unproven": [ + 127, + 130 + ], + "restarted": [ + 131, + 134 + ], + "status": { + "activationDaa": "0x78", + "active": true, + "aggregatorId": null, + "aggregatorShareBps": "0x3e8", + "paidSegmentWei": "0xa8f1428e9a42800", + "paidSegments": 3, + "pool": { + "entries": 3, + "failed": 0, + "pending": 0, + "verified": 0 + }, + "segmentBlocks": "0x4", + "segmentsInWindow": { + "pending": 15, + "proven": 3, + "unproven": 1 + }, + "shardProgramId": null, + "start": "0x77", + "unprovenDaa": "0x3c" + } + } + }, + "checks": [ + { + "name": "the node carries the v1 parameters", + "ok": true, + "detail": { + "activationDaa": "0x78", + "active": false, + "aggregatorId": null, + "aggregatorShareBps": "0x3e8", + "paidSegmentWei": "0x0", + "paidSegments": 0, + "pool": { + "entries": 0, + "failed": 0, + "pending": 0, + "verified": 0 + }, + "segmentBlocks": "0x4", + "segmentsInWindow": { + "pending": 0, + "proven": 0, + "unproven": 0 + }, + "shardProgramId": null, + "start": null, + "unprovenDaa": "0x3c" + } + }, + { + "name": "every node agrees on the v1 start", + "ok": true, + "detail": 119 + }, + { + "name": "segment 0 is aligned at the start and pending", + "ok": true, + "detail": { + "first": "0x77", + "last": "0x7a", + "status": { + "deadline_daa": 183, + "status": "pending" + } + } + }, + { + "name": "the native statement is the same on every node", + "ok": true, + "detail": "0x000000000000116f000000" + }, + { + "name": "the aggregator share is a tenth of the segment credits", + "ok": true, + "detail": { + "expectedWei0": "253611648000000000", + "creditSum0": "2536116480000000000" + } + }, + { + "name": "known-finished: segment 0 record accepted by n1", + "ok": true, + "detail": { + "accepted": true, + "block": "0xe03d6896eff5476fffb90d066efd77feb0e2eb9c527b967423b74bb5948f36d2", + "first": "0x77", + "keyHash": "0xc379ba2cdb239d2884d6bdb391d243ddfa869ab95ec4e72372e0e31fc9025fa0", + "last": "0x7a", + "new": true, + "reason": "accepted" + } + }, + { + "name": "known-finished: segment 0 paid on n0 (relayed over p2p, verified in trust mode, carried)", + "ok": true, + "detail": { + "paid": { + "carrierNumber": "0x81", + "chainLen": "0x4", + "keyHash": "0xc379ba2cdb239d2884d6bdb391d243ddfa869ab95ec4e72372e0e31fc9025fa0", + "payout": "0x4343434343434343434343434343434343434343", + "wei": "0x38502733df10000" + }, + "secs": 1.008 + } + }, + { + "name": "every node paid segment 0 the same", + "ok": true, + "detail": [ + "0x38502733df10000", + "0x38502733df10000", + "0x38502733df10000" + ] + }, + { + "name": "the payout address holds the aggregator share", + "ok": true, + "detail": { + "balance": "253611648000000000" + } + }, + { + "name": "segment 1 names segment 0 as its proven previous", + "ok": true, + "detail": { + "carrierNumber": "0x81", + "chainLen": "0x4", + "first": "0x77", + "keyHash": "0xc379ba2cdb239d2884d6bdb391d243ddfa869ab95ec4e72372e0e31fc9025fa0", + "last": "0x7a", + "proofHash": "0xed42fb45a37c81f151b2fde1b2b9b35ce0c76f7b59145eaa428f6760b739cae4", + "proofInPool": true + } + }, + { + "name": "chain rule: a fresh-chain record for segment 1 is refused while segment 0 is proven", + "ok": true, + "detail": { + "accepted": false, + "block": "0x43b605649b1d40550fc280a520950a328b25e733b4efe9cc0a204e31e9351022", + "first": "0x7b", + "keyHash": "0xc379ba2cdb239d2884d6bdb391d243ddfa869ab95ec4e72372e0e31fc9025fa0", + "last": "0x7e", + "new": false, + "reason": "segment 123..126 does not chain to segment 119..122 (chain_len 4), which is proven (record paid at chain block 129)" + } + }, + { + "name": "chain rule: the continuing record (chain_len 8) is accepted", + "ok": true, + "detail": { + "accepted": true, + "block": "0x43b605649b1d40550fc280a520950a328b25e733b4efe9cc0a204e31e9351022", + "first": "0x7b", + "keyHash": "0xc379ba2cdb239d2884d6bdb391d243ddfa869ab95ec4e72372e0e31fc9025fa0", + "last": "0x7e", + "new": true, + "reason": "accepted" + } + }, + { + "name": "segment 1 paid with chain_len 8", + "ok": true, + "detail": { + "carrierNumber": "0x87", + "chainLen": "0x8", + "keyHash": "0xc379ba2cdb239d2884d6bdb391d243ddfa869ab95ec4e72372e0e31fc9025fa0", + "payout": "0x4343434343434343434343434343434343434343", + "wei": "0x38505a6ce4a8000" + } + }, + { + "name": "segment 2 is pending with a deadline", + "ok": true, + "detail": { + "deadline_daa": 191, + "status": "pending" + } + }, + { + "name": "known-failed: segment 3 cannot start a fresh chain while segment 2 is pending", + "ok": true, + "detail": { + "accepted": false, + "block": "0x4f502ffcbe86530352096707a93074ee8e1d43e3d2e94b06b8c71735c9eb0b46", + "first": "0x83", + "keyHash": "0xc379ba2cdb239d2884d6bdb391d243ddfa869ab95ec4e72372e0e31fc9025fa0", + "last": "0x86", + "new": false, + "reason": "segment 131..134 does not chain to segment 127..130 (chain_len 4), which is pending until DAA 191" + } + }, + { + "name": "known-failed: segment 2 is unproven after its deadline", + "ok": true, + "detail": { + "deadline_daa": 191, + "status": "unproven" + } + }, + { + "name": "known-failed: a late record for segment 2 pays nothing (refused as unproven)", + "ok": true, + "detail": { + "accepted": false, + "block": "0xb46575c11a6e42efd2cdbc8a450c0f169c069409a1f6e107c7418cea9c594623", + "first": "0x7f", + "keyHash": "0xc379ba2cdb239d2884d6bdb391d243ddfa869ab95ec4e72372e0e31fc9025fa0", + "last": "0x82", + "new": false, + "reason": "unproven: the record is carried at DAA 194, after the segment's deadline DAA 191" + } + }, + { + "name": "segment 3 restarts the chain with a fresh-chain record after the unproven segment", + "ok": true, + "detail": { + "accepted": true, + "block": "0x4f502ffcbe86530352096707a93074ee8e1d43e3d2e94b06b8c71735c9eb0b46", + "first": "0x83", + "keyHash": "0xc379ba2cdb239d2884d6bdb391d243ddfa869ab95ec4e72372e0e31fc9025fa0", + "last": "0x86", + "new": true, + "reason": "accepted" + } + }, + { + "name": "segment 3 paid with chain_len 4", + "ok": true, + "detail": { + "carrierNumber": "0xc1", + "chainLen": "0x4", + "keyHash": "0xc379ba2cdb239d2884d6bdb391d243ddfa869ab95ec4e72372e0e31fc9025fa0", + "payout": "0x4343434343434343434343434343434343434343", + "wei": "0x3850c0edd68a800" + } + }, + { + "name": "the status counts proven and unproven segments", + "ok": true, + "detail": { + "pending": 15, + "proven": 3, + "unproven": 1 + } + }, + { + "name": "a v1 shard is paid 90% of its block share (one shard a block here)", + "ok": true, + "detail": { + "number": "0xc3", + "shardWei": "0x7ebcd81877b9c00", + "credit": "0x8cd1d3a96895800" + } + } + ], + "ok": true +} \ No newline at end of file From b5e27037f1ed64854d50fb5645ea4f6be4808f3b Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:18:21 +0100 Subject: [PATCH 012/311] Proving v1: the fleet-size table from the measured inputs, the numbers table in the plan Co-Authored-By: Claude Fable 5.1 --- docs/bench-log.md | 15 +++++++++++++++ docs/plans/proving-v1.md | 15 +++++++++++++-- 2 files changed, 28 insertions(+), 2 deletions(-) diff --git a/docs/bench-log.md b/docs/bench-log.md index 834986c21..bafc934bf 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1618,6 +1618,21 @@ Branches `proving-v1` (main repository, worktree `igneum-wt-proving-v1`; fork `v | A 3-minute window at 18:57Z on node 1 (`node tools/proving-v1/coverage.mjs --minutes 3`, chain blocks 80754..80839, 86 blocks) | 4 blocks with a paid shard (4.7%), 4 fully proven, 4 of 86 shards; on-chain latency (carrier timestamp minus block timestamp) n 4: min 36 s, p50 39 s, max 44 s; 0 content blocks. One prover (PC 2), the Mac verifier node down, PC 2 producing few blocks (3.4 MH/s): the degraded state above, not the fleet's number | | The 30-minute window | waits for the fleet (the coordinator's go) | + +### Step 3, the fleet size (arithmetic from measured inputs; every input names its entry) + +Inputs, all RTX 5090 (PC 2), SP1 6.8.1 cuda: a full shard at the provisional `S_p` (6.75 M pgas) compressed in 10.9 s and the four shards of a near-`B_p` block in 10.2 to 10.7 s each (bench-log 4 October 2026, "shard proving on the RTX 5090", runs run-20261004-173115 and run-20261004-r3-shards); one aggregation 2.2 s (two shards) to 2.5 s (four shards), the same entry; tonight's chain of 2 on the Mac CPU shows the recursion over the previous block proof costs the same order as a first aggregation (52.0 s against 59.1 s), so the GPU figure for a chained aggregation is taken as 2.5 s, approximate, until the held PC 2 chain job measures it; the app's live loop tonight: 1.6 shards a minute per card on empty shards (export, cut, key setup, prove, sign, submit: about 37 s a shard, of which the proof is a few seconds), bench-log step 1 above. A 5090 proves one thing at a time. + +| Block content at 1 block/s | Shard proofs a second (fleet) | Card-seconds a second for shards | Aggregations a second | Card-seconds a second for aggregation | 5090-class cards for 100% | Rule | +|---|---|---|---|---|---|---| +| empty blocks (tonight's devnet), the app's loop as it is | 1 | 37 | 1 | 2.5 (approximate) | 40 | one shard per block, the loop's 37 s each, one card does 1.6 a minute | +| empty blocks, the loop with one key setup per process and the export cached (the chain mode's shape: 12 s setup once, then proofs back to back) | 1 | about 5 (approximate: an empty shard's compressed proof on the 5090 is not measured; the 200-pgas shard took 2.7 s on 4 October) | 1 | 2.5 | 8 (approximate) | the held PC 2 chain job gives the empty-shard proof time | +| one full shard a block (`S_p`, 6.75 M pgas) | 1 | 10.9 | 1 | 2.5 | 14 | 10.9 + 2.5 card-seconds per block-second | +| blocks at `B_p` (four full shards) | 4 | 42.5 | 1 | 2.5 | 45 | 4 x 10.6 + 2.5 | +| at the adopted v1 budgets (`B_p` 120,000 pgas, `S_p` 30,000, from DAA 210,000 on the devnet): a v1 shard of transfers ran at 213 to 236 cycles per pgas (bench-log 5 October, "the prover carries both fee tables"), 7 M cycles a shard against 60 M for the prototype shard | 4 | under 42.5 (the 5090 time for a 7 M-cycle shard is not measured; scaling 10.9 s by cycles gives about 1.3 s, approximate) | 1 | 2.5 | about 8 (approximate) | measure before the switch lands | + +Reading. The card count is the sum of card-seconds of work per block-second, rounded up, with no slack for the exclusive window, the relay or a card's idle gaps; the devnet's own numbers tonight (one card, 1.6 shards a minute, 4.7% of blocks) are the first row. The lever is the loop, not the proof: a shard's proof on the 5090 is 3 to 11 s and its carriage through export, cut and a 12-s key setup is 25 s more. The host's `--mode aggregate` and `--mode chain` already hold one key setup per process; the prover loop should do the same (one host process per segment, the 0.3.11 item in the plan). + ### Step 4, the rule | What | Measured | diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md index 2f4ef493d..4bff98a95 100644 --- a/docs/plans/proving-v1.md +++ b/docs/plans/proving-v1.md @@ -26,9 +26,20 @@ that say how many cards cover the chain. | 4 | The chain rule and the unproven rule in consensus behind `proving_v1_activation_daa` (spec 7.8 items 2, 6, 7); unit tests; the fast-time 3-node harness `tools/proving-v1/net.mjs` (ports 29950+, suffix 956, trust mode) with the known-finished and known-failed cases | Implemented; the harness run waits for the Mac build of the fork (`vendor/igneum-node/target-pv1`) | | 5 | This plan: the rollout for 0.3.11 and Josh's decisions | Written below | -## Numbers (every one from `docs/bench-log.md`, "proving v1") +## Numbers (every one from `docs/bench-log.md`, "proving v1: segment records ...", 5 October 2026 evening) -(filled as the measurements land; see the bench log entries of 5 October 2026 named "proving v1 ...") +| What | Number | +|---|---| +| GPU memory during proving on the 5090 (empty shards, 298 one-second samples) | min 1,654 MiB, max 13,816 MiB; so a 16 GB card clears it and a 12 GB card does not; the full-shard peak is the held chain job's row | +| Host RAM | host used 25.6 GB of 63 GB; the WSL2 VM 7.9 GB working set | +| sp1-gpu-server 6.8.1 compiled targets (cuobjdump) | sm_80, sm_86, sm_89, sm_90, sm_100, sm_120 and compute_120 PTX: Ada (4090) is native, no JIT; nothing for AMD | +| Shards a minute, one 5090 through the app's loop (empty shards) | 1.6 | +| Chain of 2 live blocks on the Mac CPU (`--mode chain`) | shard 55.4 and 41.3 s, aggregate 52.0 s then 59.1 s with the previous proof, chain_len 2, final proof 1,272,909 bytes, `verify-segment` 0.032 s | +| Unit tests | consensus core 13, exec 8, app 5, all passing on the Mac | +| The harness (3 nodes, fast time, trust mode) | PASSED, 21 checks in 197 s: paid 1.0 s after submit, every node agreeing; the fresh chain refused after a proven segment; the unproven segment skipped after its deadline; shards at 90% | +| Coverage, 3-min window, the degraded fleet (one card, the Mac verifier down) | 4.7% of blocks proven, on-chain latency p50 39 s | +| 5090-class cards for 100% at 1 block/s | 40 with the loop as it is (empty blocks), 14 at one full shard a block, 45 at `B_p` (four full shards), about 8 at the adopted v1 budgets (approximate): the table in the bench log | +| HELD (coordinator, 19:00Z): mining alone against mining with the prover; the chain of 8 on the 5090 (N = 2, 4, 8); the 30-min coverage with the fleet mining | re-run after the go: jobs `prover-cost-pc2-pv1` (script fixed to print its state error) and `pc2-chain.ps1` with the package `igneum-prove-wsl2-pv1.zip` | ## The rule, in one paragraph (spec 7.8) From e5bb296c2e279d50211507a0447e4ebff224106b Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:53:31 +0100 Subject: [PATCH 013/311] Proving v1: step 1 measured with the fleet mining (the prover costs the 5090 4.0%, GPU peak 15.6 GB with the miner), the fleet coverage window (2.4%, p50 44 s) Co-Authored-By: Claude Fable 5.1 --- docs/bench-log.md | 9 ++++++--- docs/plans/proving-v1.md | 6 ++++-- 2 files changed, 10 insertions(+), 5 deletions(-) diff --git a/docs/bench-log.md b/docs/bench-log.md index bafc934bf..6e29548f9 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1595,8 +1595,10 @@ Branches `proving-v1` (main repository, worktree `igneum-wt-proving-v1`; fork `v | What | Measured | |---|---| | The default rule (`app/igneum-app/src/provedefault.rs`) | `cargo test --release -p igneum-app provedefault` on this Mac (the app crate, build lock, 19:05Z): 5 passed (a 5090 with WSL2 on Windows is on; Windows without WSL2 off with the Set up hint; Linux needs no WSL2 and the 12 GB gate holds, a 10 GB 3080 and a 16 GB AMD card stay off; Apple silicon off; the biggest qualifying card is named) | -| Mining alone against mining with the prover (PC 2 job `prover-cost-pc2-pv1`, `tools/proving-v1/pc2-prover-cost.ps1`, 5 + 5 min, published 18:40:48Z, ran 18:41:13Z) | VOID: the job waited 20 min for the 5090 worker to hash and it never did (the pack fault above); "mining alone" was 0 MH/s. Re-run after the coordinator's go. The job's `/api/state` read also came back empty on PC 2 (`cards=::0`, every sample) while `igneum_getProvingStatus` on 127.0.0.1:26790 answered; the re-run script prints the URL file and the error | -| GPU memory during proving (the same job, phase B: the prover on for 5 min, 298 one-second `nvidia-smi --query-gpu=memory.used` samples, the 5090 worker dead so the card held nothing else) | memory.used min 1,654 MiB, max 13,816 MiB, utilisation mean 2.7%, power max 190.6 W. The shards were the devnet's empty shards (0 pgas); the peak on a full shard at `S_p` is the chain job's row (held). Read: a 16 GB card clears tonight's peak, a 12 GB card does not, and the 24 GB of a 4090 has 10 GB of headroom | +| Mining alone against mining with the prover, first try (PC 2 job `prover-cost-pc2-pv1`, `tools/proving-v1/pc2-prover-cost.ps1`, published 18:40:48Z, ran 18:41:13Z) | VOID: the job waited 20 min for the 5090 worker to hash and it never did (the pack fault above); "mining alone" was 0 MH/s | +| Mining alone against mining with the prover, the re-run after the coordinator's go (job `prover-cost-pc2-pv1b`, ran 19:20:36Z to 19:51:17Z; the 5090 worker restored at 19:16Z and hashing throughout; prover OFF by `POST /api/prove {"on":false}` 19:40:39Z, back ON 19:46:09Z, left on). The job's own `/api/state` samples stayed empty on PC 2 (`Invoke-RestMethod` returns an object PowerShell 5.1 cannot walk, `cards=0`, the fix is for the next run), so the hash rate is read from the miner's own STATUS lines (`miner-nvidia-1ccfe586-1` uploads, `now=... MH/s wall`, one every 30 s, the intake table `miner_logs`) | prover OFF, 19:41:09 to 19:46:09Z: n 10, mean 124.72 MH/s, p50 124.81, min 124.10, max 125.38. Prover ON, 19:46:39 to 19:51:13Z: n 9, mean 119.74, p50 118.87, min 118.08, max 123.42. The 15 min before the job with the prover on (19:25 to 19:40Z): n 30, mean 119.88, p50 118.79. So the prover costs the 5090 5.0 MH/s, 4.0% of its hash rate, while it proves the devnet's empty shards one after another (1.4 a minute here: the node's paidShards 559 -> 566 over the 5-min phase). A full shard at `S_p` keeps the card busier (the 4 October run proved one in 10.9 s); the cost at that load is the chain job's row | +| GPU memory during proving, first try (phase B of the first job: the prover on for 5 min, 298 one-second `nvidia-smi --query-gpu=memory.used` samples, the 5090 worker dead so the card held nothing else) | memory.used min 1,654 MiB, max 13,816 MiB, utilisation mean 2.7%, power max 190.6 W: the prover alone on empty shards | +| GPU memory with the miner AND the prover on the card (the re-run's phase B, 298 one-second samples, 19:46 to 19:51Z) | memory.used min 3,396 MiB (the miner's dataset and program resident), max 15,590 MiB, utilisation mean 92.9%, power max 328.6 W. So the prover's own peak is about 12.2 GB on an empty shard (15,590 minus the miner's 3,396), and the two together need 15.6 GB: a 16 GB card (5080, 9070 XT class, if it had a CUDA path) sits 0.4 GB under tonight's peak with no room for a full shard, a 24 GB 4090 has 8.4 GB of headroom, a 12 GB card cannot mine and prove at once on this build. The full-shard peak is the chain job's row | | Shards per minute with the mining worker dead | the node's `paidShards` 510 -> 518 over the 5-min phase: 1.6 shards a minute from one 5090 through the app's loop (export, cut, prove, sign, submit) | | Host RAM (Windows `Win32_OperatingSystem` and the `vmmem` working set, sampled every 15 s) | host used 25,550 MB of 63,132 MB at the end; the WSL2 VM's working set 7,915 MB (2,334 MB used of 30,914 MB inside the distribution) | | The SP1 GPU server's compiled targets (`cuobjdump --list-elf /root/.sp1/bin/sp1-gpu-server` inside PC 2's Ubuntu-24.04, CUDA 12.8, driver 610.47) | `sp1-gpu-server` 6.8.1 (251,306,680 bytes, sha256 c2642ad1c42e85d8525159cf0c7cd5200d8766c9be1283f452a1f9bf9fea725c, the asset `sp1_gpu_server_v6.8.1_x86_64.tar.gz` the SDK downloads, `sp1-cuda-6.8.1/src/server.rs`): one ELF each for sm_80, sm_86, sm_89, sm_90, sm_100 and sm_120; `strings` finds compute_120 PTX as well. So sm_89 (Ada: RTX 4090, 4080) is compiled in natively, no JIT; so are Ampere (3090, 3060), Hopper, Blackwell datacentre (sm_100) and consumer (sm_120, the 5090). Nothing for AMD (no HIP path in SP1) | @@ -1616,7 +1618,8 @@ Branches `proving-v1` (main repository, worktree `igneum-wt-proving-v1`; fork `v | What | Measured | |---|---| | A 3-minute window at 18:57Z on node 1 (`node tools/proving-v1/coverage.mjs --minutes 3`, chain blocks 80754..80839, 86 blocks) | 4 blocks with a paid shard (4.7%), 4 fully proven, 4 of 86 shards; on-chain latency (carrier timestamp minus block timestamp) n 4: min 36 s, p50 39 s, max 44 s; 0 content blocks. One prover (PC 2), the Mac verifier node down, PC 2 producing few blocks (3.4 MH/s): the degraded state above, not the fleet's number | -| The 30-minute window | waits for the fleet (the coordinator's go) | +| A 30-minute window, 19:13 to 19:43Z, the degraded fleet (PC 2 the only prover, its 5090 worker restored at 19:16Z, the Mac app node down by decision: the Mac app is attached to node 1) | `node tools/proving-v1/coverage.mjs --minutes 30 --watch` on node 1: chain blocks 81236..82668, 1,433 blocks; 38 with a paid shard (2.7%), all 38 fully proven (one shard a block, 0 content blocks); on-chain latency n 38: min 36, p50 44, p90 52, p99 62, max 65 s. The live page's 10-minute proving object read 0 shards and 0 provers at 19:42Z (it counts what its own node verified; that node is the Mac app node, down), so the chain's own count is the number | +| A 30-minute window with the fleet mining (PC 2 at 119 MH/s from 19:16Z, PC 1 at 128.8 from 19:18Z; PC 2 still the only prover, its prover OFF for the 5 min of the cost job's phase A inside this window; the Mac app node down by decision) | `coverage.mjs --minutes 30 --watch`, 19:21 to 19:51Z on node 1: chain blocks 81644..83069, 1,426 blocks; 34 with a paid shard (2.4%), all fully proven (one shard a block, no content); on-chain latency n 34: min 38, p50 44, p90 51, p99 52, max 53 s. One 5090 through the app's loop as it is covers 2.4 to 2.7% of the blocks; the latency from block to carried record is 44 s at the median, under the litepaper's minute, and would be the same for every block if the fleet were 40 cards (the table below) | ### Step 3, the fleet size (arithmetic from measured inputs; every input names its entry) diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md index 4bff98a95..890a7649b 100644 --- a/docs/plans/proving-v1.md +++ b/docs/plans/proving-v1.md @@ -30,7 +30,9 @@ that say how many cards cover the chain. | What | Number | |---|---| -| GPU memory during proving on the 5090 (empty shards, 298 one-second samples) | min 1,654 MiB, max 13,816 MiB; so a 16 GB card clears it and a 12 GB card does not; the full-shard peak is the held chain job's row | +| The prover's cost to a mining 5090 (the re-run with the fleet mining, hash rate from the miner's own STATUS lines) | 124.7 MH/s alone, 119.7 MH/s with the prover on: 5.0 MH/s, 4.0%, on empty shards at 1.4 a minute | +| GPU memory on the 5090: the prover alone (empty shards) / the miner and the prover together | max 13,816 MiB / max 15,590 MiB (the miner holds 3,396 MiB); a 24 GB 4090 has 8.4 GB of headroom, a 16 GB card 0.4 GB, a 12 GB card cannot do both on this build; the full-shard peak is the chain job's row | +| Coverage, 30-min window with the fleet mining, one prover | 2.4% of blocks, latency p50 44 s, p99 52 s | | Host RAM | host used 25.6 GB of 63 GB; the WSL2 VM 7.9 GB working set | | sp1-gpu-server 6.8.1 compiled targets (cuobjdump) | sm_80, sm_86, sm_89, sm_90, sm_100, sm_120 and compute_120 PTX: Ada (4090) is native, no JIT; nothing for AMD | | Shards a minute, one 5090 through the app's loop (empty shards) | 1.6 | @@ -39,7 +41,7 @@ that say how many cards cover the chain. | The harness (3 nodes, fast time, trust mode) | PASSED, 21 checks in 197 s: paid 1.0 s after submit, every node agreeing; the fresh chain refused after a proven segment; the unproven segment skipped after its deadline; shards at 90% | | Coverage, 3-min window, the degraded fleet (one card, the Mac verifier down) | 4.7% of blocks proven, on-chain latency p50 39 s | | 5090-class cards for 100% at 1 block/s | 40 with the loop as it is (empty blocks), 14 at one full shard a block, 45 at `B_p` (four full shards), about 8 at the adopted v1 budgets (approximate): the table in the bench log | -| HELD (coordinator, 19:00Z): mining alone against mining with the prover; the chain of 8 on the 5090 (N = 2, 4, 8); the 30-min coverage with the fleet mining | re-run after the go: jobs `prover-cost-pc2-pv1` (script fixed to print its state error) and `pc2-chain.ps1` with the package `igneum-prove-wsl2-pv1.zip` | +| The chain of 8 on the 5090 (N = 2, 4, 8), job `chain-pc2-pv1` (published 19:52Z after the go) | (the bench log row when it ends) | ## The rule, in one paragraph (spec 7.8) From 353a6e813b027f3c14479de5cb436eaaf5a2e48a Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:57:05 +0100 Subject: [PATCH 014/311] pc2-chain.ps1: the export goes through curl.exe to a file (Invoke-WebRequest's Content is a string; the first run failed on WriteAllBytes) Co-Authored-By: Claude Fable 5.1 --- tools/proving-v1/pc2-chain.ps1 | 10 ++++++---- 1 file changed, 6 insertions(+), 4 deletions(-) diff --git a/tools/proving-v1/pc2-chain.ps1 b/tools/proving-v1/pc2-chain.ps1 index a56be84e2..cbb8992e6 100644 --- a/tools/proving-v1/pc2-chain.ps1 +++ b/tools/proving-v1/pc2-chain.ps1 @@ -26,10 +26,12 @@ $first = $tip - 30; $last = $first + 7 $t = Get-Date $body = (@{jsonrpc='2.0'; id=1; method='igneum_exportSegments'; params=@('0x0', ('0x{0:x}' -f $last))} | ConvertTo-Json -Compress) $seqFile = Join-Path $job 'seq.json' -try { - $resp = Invoke-WebRequest -Method Post -Uri "http://127.0.0.1:$evmPort" -ContentType 'application/json' -Body $body -TimeoutSec 600 -UseBasicParsing - [IO.File]::WriteAllBytes($seqFile, $resp.Content) -} catch { "RESULT export FAILED: $_"; exit 1 } +$bodyFile = Join-Path $job 'export-request.json' +[IO.File]::WriteAllText($bodyFile, $body, (New-Object System.Text.UTF8Encoding $false)) +# curl.exe (Windows 10+ ships it) streams the 70 MB reply to a file; Invoke-WebRequest's Content is a string there +# and WriteAllBytes refused it (the first run, 19:52Z) +& curl.exe -s -S -m 600 -X POST "http://127.0.0.1:$evmPort" -H 'Content-Type: application/json' --data-binary "@$bodyFile" -o $seqFile 2>&1 | ForEach-Object { "curl: $_" } +if (-not (Test-Path $seqFile) -or (Get-Item $seqFile).Length -lt 1000) { "RESULT export FAILED: no reply file"; exit 1 } $len = (Get-Item $seqFile).Length "RESULT export $(Stamp) $len bytes in $([math]::Round(((Get-Date) - $t).TotalSeconds,1)) s to $seqFile" # the JSON-RPC envelope: the exporter wants the result object; unwrap with python inside WSL (below) From 5d07095d3a052d3e9403e0f0b33d7091fa127456 Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:03:15 +0100 Subject: [PATCH 015/311] Proving v1: the chain of 2 on the 5090 (31.8 s, the chained aggregation 9.5 s), the job aborted by the 0.3.10 restart at block 3 Co-Authored-By: Claude Fable 5.1 --- docs/bench-log.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/bench-log.md b/docs/bench-log.md index 6e29548f9..161692642 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1611,7 +1611,7 @@ Branches `proving-v1` (main repository, worktree `igneum-wt-proving-v1`; fork `v | Eight consecutive live fixtures | `igneum_exportSegments 0x0..0x13cb4` on node 1's exec RPC (127.0.0.1:26790, read-only, 20:06 BST, tip 81,076): 71,042,616 bytes, 81,077 segments, 28 accounts, 0.5 s; `igneum-prove-export export.json block-.json` for 81046..81053: replayed 81,077 segments from genesis in 1.8 s each, every state root equal to the node's; one shard a block, 0 pgas (no transactions on the devnet tonight), `proving/fixtures/chain/` | | Chain of 2 on the Mac CPU (the known-finished case of `--mode chain` before the GPU; M5 Max under the live nodes, the harness and two builds) | `SP1_PROVER=cpu igneum-prove-host --mode chain --chain block-81046.json,block-81047.json --out results.json` under the run lock, 19:07:48Z to 19:11:28Z: setup 12.2 s; block 81046: shard 0 compressed 55.4 s (1,272,897 bytes, verify 0.036 s), aggregate 52.0 s (1,272,909 bytes, verify 0.031 s), chain_len 1, agg_vk zero; block 81047: shard 41.3 s, aggregate WITH the previous block proof 59.1 s, chain_len 2, agg_vk = the pinned aggregator id; end to end 207.9 s; final proof 1,272,909 bytes, statement 0x232276f4... The recursion over the previous proof cost 7 s more than the first aggregation on this CPU | | `--mode verify-segment` on that proof (the node's path: SP1 light verifier, pinned aggregator key) | VERIFIED in 0.032 s (0.27 s wall, three runs: 0.033, 0.032, 0.032); known-failed: a wrong statement NOT VERIFIED (0.032 s); the shard verifier (`--mode verify`) on the segment proof NOT VERIFIED, "program id 0x474678f3... IS NOT OURS 0x2b1a81cb..." | -| Chain of 8 on the RTX 5090 (N = 2, 4, 8) | HELD with step 1 (`tools/proving-v1/pc2-chain.ps1`, package `igneum-prove-wsl2-pv1.zip` eb6dccf8..., 1.5 MB, built from this branch with the gate's native half run above) | +| Chain of 8 on the RTX 5090 (N = 2, 4, 8), job `chain-pc2-pv1b` (`tools/proving-v1/pc2-chain.ps1`; the package `igneum-prove-wsl2-pv1.zip` eb6dccf8..., 1.5 MB, fetched by `fetch-prove-pv1` 19:51Z; the first try `chain-pc2-pv1` died in its own export step, fixed) | Ran 19:58:37Z: the export from PC 2's node (72,901,414 bytes, 1.4 s), the host built in WSL2 against the live build's warm target dir in 6 s and installed to `/opt/igneum-pv1` (the live `/opt/igneum` host untouched, sha 29cc4768...), `--mode id` the pinned pair; eight consecutive fixtures 83346..83353 cut, every one MATCHES natively. The chain on the GPU (SP1_PROVER=cuda, the miner mining on the same card at 119 MH/s): setup 12.7 s; block 83346: shard 7.4 s, aggregate 7.6 s (chain_len 1), 15.1 s; block 83347: shard 7.2 s, aggregate WITH the previous proof 9.5 s (chain_len 2, agg_vk the pinned aggregator id), 16.8 s, cumulative 31.8 s over 2 blocks; block 83348: shard 7.0 s, then at 20:01:09Z the app quit for the 0.3.10 update and aborted the job ("aborted (the app is quitting)"). So N = 2 measured: 31.8 s of GPU time for two empty blocks, the chained aggregation 1.9 s dearer than the first; N = 4 and 8 are the re-run `chain-pc2-pv1c` after the restart. An empty shard's compressed proof on the 5090 is 7.0 to 7.4 s (the 200-pgas shard of 4 October took 2.7 s with the card to itself; tonight the miner held it at 92% utilisation) | ### Step 3, coverage From b3997089ff20d4e608c99665d062008815d193f4 Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 21:10:50 +0100 Subject: [PATCH 016/311] Proving v1: the chain of 8 on the 5090 measured (N = 2, 4, 8: 32.6, 66.8, 135.6 s; chained aggregation 9.7 s a block on a mining card), the fleet table re-cut on the measured rows Co-Authored-By: Claude Fable 5.1 --- docs/bench-log.md | 18 ++- docs/plans/proving-v1.md | 8 +- tools/proving-v1/chain-pc2-2026-10-05.json | 154 +++++++++++++++++++++ 3 files changed, 170 insertions(+), 10 deletions(-) create mode 100644 tools/proving-v1/chain-pc2-2026-10-05.json diff --git a/docs/bench-log.md b/docs/bench-log.md index 161692642..934e49d5c 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1612,6 +1612,10 @@ Branches `proving-v1` (main repository, worktree `igneum-wt-proving-v1`; fork `v | Chain of 2 on the Mac CPU (the known-finished case of `--mode chain` before the GPU; M5 Max under the live nodes, the harness and two builds) | `SP1_PROVER=cpu igneum-prove-host --mode chain --chain block-81046.json,block-81047.json --out results.json` under the run lock, 19:07:48Z to 19:11:28Z: setup 12.2 s; block 81046: shard 0 compressed 55.4 s (1,272,897 bytes, verify 0.036 s), aggregate 52.0 s (1,272,909 bytes, verify 0.031 s), chain_len 1, agg_vk zero; block 81047: shard 41.3 s, aggregate WITH the previous block proof 59.1 s, chain_len 2, agg_vk = the pinned aggregator id; end to end 207.9 s; final proof 1,272,909 bytes, statement 0x232276f4... The recursion over the previous proof cost 7 s more than the first aggregation on this CPU | | `--mode verify-segment` on that proof (the node's path: SP1 light verifier, pinned aggregator key) | VERIFIED in 0.032 s (0.27 s wall, three runs: 0.033, 0.032, 0.032); known-failed: a wrong statement NOT VERIFIED (0.032 s); the shard verifier (`--mode verify`) on the segment proof NOT VERIFIED, "program id 0x474678f3... IS NOT OURS 0x2b1a81cb..." | | Chain of 8 on the RTX 5090 (N = 2, 4, 8), job `chain-pc2-pv1b` (`tools/proving-v1/pc2-chain.ps1`; the package `igneum-prove-wsl2-pv1.zip` eb6dccf8..., 1.5 MB, fetched by `fetch-prove-pv1` 19:51Z; the first try `chain-pc2-pv1` died in its own export step, fixed) | Ran 19:58:37Z: the export from PC 2's node (72,901,414 bytes, 1.4 s), the host built in WSL2 against the live build's warm target dir in 6 s and installed to `/opt/igneum-pv1` (the live `/opt/igneum` host untouched, sha 29cc4768...), `--mode id` the pinned pair; eight consecutive fixtures 83346..83353 cut, every one MATCHES natively. The chain on the GPU (SP1_PROVER=cuda, the miner mining on the same card at 119 MH/s): setup 12.7 s; block 83346: shard 7.4 s, aggregate 7.6 s (chain_len 1), 15.1 s; block 83347: shard 7.2 s, aggregate WITH the previous proof 9.5 s (chain_len 2, agg_vk the pinned aggregator id), 16.8 s, cumulative 31.8 s over 2 blocks; block 83348: shard 7.0 s, then at 20:01:09Z the app quit for the 0.3.10 update and aborted the job ("aborted (the app is quitting)"). So N = 2 measured: 31.8 s of GPU time for two empty blocks, the chained aggregation 1.9 s dearer than the first; N = 4 and 8 are the re-run `chain-pc2-pv1c` after the restart. An empty shard's compressed proof on the 5090 is 7.0 to 7.4 s (the 200-pgas shard of 4 October took 2.7 s with the card to itself; tonight the miner held it at 92% utilisation) | +| The chain of 8, the third run `chain-pc2-pv1c` (20:05:21Z to 20:08:33Z, after the app restart; blocks 83616..83623 from PC 2's node at tip 83646, the same script; results `tools/proving-v1/chain-pc2-2026-10-05.json`) | Build 5 s (warm), eight fixtures cut and MATCHING natively, setup 15.7 s, then on the GPU with the miner mining on the same card: shard proofs 7.3 to 7.7 s each (8 x, 59.5 s), aggregations 7.9 s for the first block and 9.6 to 9.7 s for every chained one (75.5 s), every proof VERIFIED, end to end 135.6 s for 8 blocks (17.0 s a block from the second on). Cumulative: N = 2 at 32.6 s, N = 4 at 66.8 s, N = 8 at 135.6 s. The final proof is 1,272,909 bytes whatever N (chain_len 8, agg_vk the pinned aggregator id), the record 586 bytes; `--mode verify-segment` on it: VERIFIED in 0.039, 0.037, 0.040 s after a 0.26-s light-verifier setup, the same three runs each time. GPU memory over the chain (152 one-second samples): max 16,751 MiB with the miner's 3.4 GB resident, so the chained aggregation holds about 13.4 GB, 1.2 GB over the shard-only peak; WSL used 2,456 MB | + +Reading the chain numbers. Aggregation is a fixed cost per block (9.7 s here), not per segment: the recursion verifies one more proof whatever `chain_len`, so the record for N blocks costs N aggregations and the verifier one. Against 4 October with the miner stopped (aggregate 2.2 to 2.5 s, a 200-pgas shard 2.7 s), tonight's 9.7 s and 7.3 s say the miner's 92% utilisation slows the prover about 3 to 4x while the prover slows the miner 4%: the card is shared, and the lottery wins the arbitration. A machine that mines and proves at once delivers one empty block's proof and aggregation in 17 s; one that only proves, about 5 s (approximate, from the 4 October stages). + ### Step 3, coverage @@ -1628,13 +1632,15 @@ Inputs, all RTX 5090 (PC 2), SP1 6.8.1 cuda: a full shard at the provisional `S_ | Block content at 1 block/s | Shard proofs a second (fleet) | Card-seconds a second for shards | Aggregations a second | Card-seconds a second for aggregation | 5090-class cards for 100% | Rule | |---|---|---|---|---|---|---| -| empty blocks (tonight's devnet), the app's loop as it is | 1 | 37 | 1 | 2.5 (approximate) | 40 | one shard per block, the loop's 37 s each, one card does 1.6 a minute | -| empty blocks, the loop with one key setup per process and the export cached (the chain mode's shape: 12 s setup once, then proofs back to back) | 1 | about 5 (approximate: an empty shard's compressed proof on the 5090 is not measured; the 200-pgas shard took 2.7 s on 4 October) | 1 | 2.5 | 8 (approximate) | the held PC 2 chain job gives the empty-shard proof time | -| one full shard a block (`S_p`, 6.75 M pgas) | 1 | 10.9 | 1 | 2.5 | 14 | 10.9 + 2.5 card-seconds per block-second | -| blocks at `B_p` (four full shards) | 4 | 42.5 | 1 | 2.5 | 45 | 4 x 10.6 + 2.5 | -| at the adopted v1 budgets (`B_p` 120,000 pgas, `S_p` 30,000, from DAA 210,000 on the devnet): a v1 shard of transfers ran at 213 to 236 cycles per pgas (bench-log 5 October, "the prover carries both fee tables"), 7 M cycles a shard against 60 M for the prototype shard | 4 | under 42.5 (the 5090 time for a 7 M-cycle shard is not measured; scaling 10.9 s by cycles gives about 1.3 s, approximate) | 1 | 2.5 | about 8 (approximate) | measure before the switch lands | +| empty blocks (tonight's devnet), the app's loop as it is, the card also mining | 1 | 37 | 1 | 9.7 (measured, `chain-pc2-pv1c`) | 47 | one shard per block, the loop's 37 s each plus a chained aggregation | +| empty blocks, the chain mode's shape (one key setup per process, proofs back to back), the card also mining | 1 | 7.4 (measured) | 1 | 9.7 (measured) | 18 | 17.1 card-seconds a block, `chain-pc2-pv1c` | +| empty blocks, cards that only prove | 1 | 2.7 (4 October, a 200-pgas shard) | 1 | 2.5 (4 October) | 6 (approximate) | the miner's 92% utilisation costs the prover 3 to 4x | +| one full shard a block (`S_p`, 6.75 M pgas), cards that only prove | 1 | 10.9 | 1 | 2.5 | 14 | 4 October's stages | +| one full shard a block, the card also mining | 1 | about 35 (approximate: 10.9 x 3.2, tonight's ratio) | 1 | 9.7 | about 45 (approximate) | the full-shard proof with the miner on the card is not measured | +| blocks at `B_p` (four full shards), cards that only prove | 4 | 42.5 | 1 | 2.5 | 45 | 4 x 10.6 + 2.5 | +| at the adopted v1 budgets (`B_p` 120,000 pgas, `S_p` 30,000, from DAA 210,000 on the devnet): a v1 shard of transfers ran at 213 to 236 cycles per pgas (bench-log 5 October, "the prover carries both fee tables"), 7 M cycles a shard against 60 M for the prototype shard | 4 | under 42.5 (the 5090 time for a 7 M-cycle shard is not measured; scaling 10.9 s by cycles gives about 1.3 s, approximate) | 1 | 2.5 to 9.7 | 8 to 15 (approximate) | measure before the switch lands | -Reading. The card count is the sum of card-seconds of work per block-second, rounded up, with no slack for the exclusive window, the relay or a card's idle gaps; the devnet's own numbers tonight (one card, 1.6 shards a minute, 4.7% of blocks) are the first row. The lever is the loop, not the proof: a shard's proof on the 5090 is 3 to 11 s and its carriage through export, cut and a 12-s key setup is 25 s more. The host's `--mode aggregate` and `--mode chain` already hold one key setup per process; the prover loop should do the same (one host process per segment, the 0.3.11 item in the plan). +Reading. The card count is the sum of card-seconds of work per block-second, rounded up, with no slack for the exclusive window, the relay or a card's idle gaps; the devnet's own numbers tonight (one card, 1.4 to 1.6 shards a minute, 2.4 to 4.7% of blocks) are the first row. Two levers, both measured tonight: the loop (a shard's carriage through export, cut and a 12-s key setup is 25 s on top of a 7-s proof; the host's `--mode aggregate` and `--mode chain` hold one key setup per process and the prover loop should do the same, the 0.3.11 item in the plan) and the card's other job (a mining card proves 3 to 4x slower than an idle one, `chain-pc2-pv1c` against 4 October; the prover's cost to mining is 4%). A fleet of 18 mining 5090s, or 6 proving-only ones, covers an empty-block chain at 1 block/s through the chain mode; the mandatory rule waits for the measured share to reach one, not for these rows. ### Step 4, the rule diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md index 890a7649b..9f1ff1ea5 100644 --- a/docs/plans/proving-v1.md +++ b/docs/plans/proving-v1.md @@ -40,8 +40,8 @@ that say how many cards cover the chain. | Unit tests | consensus core 13, exec 8, app 5, all passing on the Mac | | The harness (3 nodes, fast time, trust mode) | PASSED, 21 checks in 197 s: paid 1.0 s after submit, every node agreeing; the fresh chain refused after a proven segment; the unproven segment skipped after its deadline; shards at 90% | | Coverage, 3-min window, the degraded fleet (one card, the Mac verifier down) | 4.7% of blocks proven, on-chain latency p50 39 s | -| 5090-class cards for 100% at 1 block/s | 40 with the loop as it is (empty blocks), 14 at one full shard a block, 45 at `B_p` (four full shards), about 8 at the adopted v1 budgets (approximate): the table in the bench log | -| The chain of 8 on the 5090 (N = 2, 4, 8), job `chain-pc2-pv1` (published 19:52Z after the go) | (the bench log row when it ends) | +| The chain of 8 live blocks on the 5090, the card also mining (`chain-pc2-pv1c`) | shard 7.3 to 7.7 s, first aggregation 7.9 s, every chained one 9.6 to 9.7 s; N = 2 in 32.6 s, N = 4 in 66.8 s, N = 8 in 135.6 s (17.0 s a block); the final proof 1,272,909 bytes whatever N, the record 586 bytes, `verify-segment` 0.037 to 0.040 s; GPU peak 16,751 MiB with the miner resident. Against 4 October with the miner stopped (aggregate 2.2 s): the miner slows the prover 3 to 4x | +| 5090-class cards for 100% at 1 block/s, measured rows | 47 with the loop as it is, 18 through the chain mode on mining cards, 6 (approximate) on proving-only cards, at empty blocks; 14 proving-only at one full shard a block; 45 at `B_p`: the table in the bench log | ## The rule, in one paragraph (spec 7.8) @@ -88,8 +88,8 @@ the rolling upgrade does not partition the network. The order, each step with it | Decision | Proposed | Why | |---|---|---| -| `proving_v1_segment_blocks` (N) | 4 | the N = 2, 4, 8 measurement on the 5090 decides; 4 keeps a record every 4 s at 1 block/s and the recursion cost per block constant | -| `proving_v1_unproven_daa` (T) | 600 | equals the 600-block record window of v0: nothing is payable for a segment after it either way; the fleet's measured latency (p99) must sit well inside it | +| `proving_v1_segment_blocks` (N) | 4 | measured: aggregation is a fixed 9.7 s per block on a mining 5090 whatever N, so N only sets how often a record is carried (every 4 s at 1 block/s) and how much a missed deadline forfeits (4 blocks' aggregator share); 8 halves the record traffic for the same card time | +| `proving_v1_unproven_daa` (T) | 600 | equals the 600-block record window of v0: nothing is payable for a segment after it either way; tonight's measured latency from block to carried shard record is p99 52 to 62 s, so 600 leaves 10x | | `proving_v1_aggregator_share_bps` | 1,000 (a tenth) | the aggregation is one recursion per block, far cheaper than the shards; a tenth pays a second role without starving the shard provers; it is a consensus parameter in the digest | | `proving_v1_activation_daa` (H) | 24 h after the 0.3.11 publish | the fee-switch rule | | Aggregator sortition | none in v1 (first valid record wins) | design 5.3's VRF draw is O-7.3; with one or two aggregators on the devnet a draw changes nothing yet | diff --git a/tools/proving-v1/chain-pc2-2026-10-05.json b/tools/proving-v1/chain-pc2-2026-10-05.json new file mode 100644 index 000000000..d07901661 --- /dev/null +++ b/tools/proving-v1/chain-pc2-2026-10-05.json @@ -0,0 +1,154 @@ +{ + "aggregate_prove_seconds_total": 75.492162475, + "aggregator_id": "0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896", + "blocks": [ + { + "aggregate_prove_seconds": 7.903165861, + "aggregate_verify_seconds": 0.037313475, + "block_seconds": 15.53678578, + "chain_len": 1, + "cumulative_seconds": 15.53678733, + "number": 83616, + "post_root": "0x85e67dcea1453afbc17cc0c83ddee0ae12e43d000a4821b4f1c6b2351d677fbc", + "proof_bytes": 1272909, + "shard_prove_seconds": [ + 7.557706469 + ], + "shards": 1, + "statement": "0x1f3c209362cc0e6a8b0f4c988728e073331f00d4a2ebdc58cc5cfe078cdf42df" + }, + { + "aggregate_prove_seconds": 9.625608758, + "aggregate_verify_seconds": 0.037767114, + "block_seconds": 17.043759121, + "chain_len": 2, + "cumulative_seconds": 32.580617039, + "number": 83617, + "post_root": "0x6e85cac1c9c7973c2cb4d0ff38f96edc054c092fa414f26bf53b237e08e81b77", + "proof_bytes": 1272909, + "shard_prove_seconds": [ + 7.344317298 + ], + "shards": 1, + "statement": "0x903b6bf10915403c9ed8a7bc87951c6bc975ba992f947c3b02af48c7e9331306" + }, + { + "aggregate_prove_seconds": 9.554146543, + "aggregate_verify_seconds": 0.037548738, + "block_seconds": 16.93764002, + "chain_len": 3, + "cumulative_seconds": 49.518392437, + "number": 83618, + "post_root": "0x4a386002c4df5c9a0ea0494442d252c02860bf2c44bc626321c8b4bb11c5b7d5", + "proof_bytes": 1272909, + "shard_prove_seconds": [ + 7.307433102 + ], + "shards": 1, + "statement": "0x0913be6ca9c2671c53e202a144e9a57ecb9b45224ad764031b5e01487dccfa94" + }, + { + "aggregate_prove_seconds": 9.725408301, + "aggregate_verify_seconds": 0.038074529, + "block_seconds": 17.316632646, + "chain_len": 4, + "cumulative_seconds": 66.835133749, + "number": 83619, + "post_root": "0x2d46b11d9f257e3e6a6d6e3ebd947a3498c3c5d9a80ef2651caf72f5d2af47ce", + "proof_bytes": 1272909, + "shard_prove_seconds": [ + 7.516456343 + ], + "shards": 1, + "statement": "0x00e0a2c9f1e16e06491ce8d26f457b35c4ac60cd3b609e8a87134a31a06a1fec" + }, + { + "aggregate_prove_seconds": 9.694312635, + "aggregate_verify_seconds": 0.038877759, + "block_seconds": 17.159775091, + "chain_len": 5, + "cumulative_seconds": 83.995023181, + "number": 83620, + "post_root": "0x209d61c1fa5d410a252a829f186b22b01840729619b69177647ad6d73c87af5d", + "proof_bytes": 1272909, + "shard_prove_seconds": [ + 7.389889021 + ], + "shards": 1, + "statement": "0xf26202b09c544d042d78d4844e6589d467f8621d9c4f13d61685fa36d58444d9" + }, + { + "aggregate_prove_seconds": 9.664895146, + "aggregate_verify_seconds": 0.035052656, + "block_seconds": 17.453177945, + "chain_len": 6, + "cumulative_seconds": 101.448329353, + "number": 83621, + "post_root": "0x78034f6b455e956da74f6e685b89ead732a45b665ff4706be6deea7cbb402c53", + "proof_bytes": 1272909, + "shard_prove_seconds": [ + 7.71706568 + ], + "shards": 1, + "statement": "0x53b07bfc69d17ca7504590bfe52f1d0915e39b2e3e497c57117d871ab275d2c9" + }, + { + "aggregate_prove_seconds": 9.634209172, + "aggregate_verify_seconds": 0.037378761, + "block_seconds": 17.047352194, + "chain_len": 7, + "cumulative_seconds": 118.49578211, + "number": 83622, + "post_root": "0x258fbe0b700833451b4bde8139cc6eb60da4fd959dafdd2784572ead6e39d908", + "proof_bytes": 1272909, + "shard_prove_seconds": [ + 7.338077294 + ], + "shards": 1, + "statement": "0xc425c86b6569c80481691e6c89abfbd1481cc555030b791178ef536b1b3264a7" + }, + { + "aggregate_prove_seconds": 9.690416059, + "aggregate_verify_seconds": 0.036976436, + "block_seconds": 17.115775928, + "chain_len": 8, + "cumulative_seconds": 135.611679366, + "number": 83623, + "post_root": "0x8840c082c4406252bcd65a31ed0102a2eba1b59130582bbba0f87d64ebbc80ed", + "proof_bytes": 1272909, + "shard_prove_seconds": [ + 7.347491181 + ], + "shards": 1, + "statement": "0xa99aba5397bae3fd8323dcbc4a93105ae531356fff92feaccca0c3284168ca49" + } + ], + "chain_seconds": 135.611788663, + "first": 83616, + "fixtures": [ + "/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83616.json", + "/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83617.json", + "/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83618.json", + "/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83619.json", + "/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83620.json", + "/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83621.json", + "/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83622.json", + "/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83623.json" + ], + "last": 83623, + "mode": "chain", + "prover": "cuda", + "segment_block_hash": "0x6a3295e7481cb3d5db67758d95280debb72b000a3a0c8a99fe30bf4f3d66aed1", + "segment_chain_len": 8, + "segment_number": 83623, + "segment_proof_bytes": 1272909, + "segment_proof_file": "/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/segment-83623-aggregated.bin", + "segment_proof_sha256": "0xc28c2c1e8054523c1cff5e38cef8d86498bc004b9beeff010a1d2acc7ef1435c", + "segment_provers": "0x1805afcd50a68f5237f8a5e2e41a4270457be031d1158f5247af6f68808b0d11", + "segment_public_values": "0x000000000000116f00000000000146a76a3295e7481cb3d5db67758d95280debb72b000a3a0c8a99fe30bf4f3d66aed16cad584f6a86085d410c13728d11fa5c4daf583ebbc443512999174b9f7d3b68000000010000000000000000000000000000000000000000000000000000000000000000258fbe0b700833451b4bde8139cc6eb60da4fd959dafdd2784572ead6e39d9088840c082c4406252bcd65a31ed0102a2eba1b59130582bbba0f87d64ebbc80ed6fcd3e8e97da273711ccefb79abdd246c5663c7d61057f82bae745ceac5dcc750000000000000000000000000000000000000000000000001805afcd50a68f5237f8a5e2e41a4270457be031d1158f5247af6f68808b0d112b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b091238960000000000000008", + "segment_statement": "0xa99aba5397bae3fd8323dcbc4a93105ae531356fff92feaccca0c3284168ca49", + "setup_seconds": 15.734044172, + "shard_program_id": "0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a", + "shard_prove_seconds_total": 59.518436388000005, + "shards": 8 +} \ No newline at end of file From a4658816e03502099d005e7d4f0f1aa323eeb21e Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:21:49 +0000 Subject: [PATCH 017/311] scratch-soundness.md: the five findings, the recompute-versus-store arithmetic at 64, 256 and 2,048 slots, the named on-die-cache chip row per variant (2.4x at every share under the cap) beside the M16 mixer lever, the host tag contract, the vector requirements, the Metal results (228 of 228 pass, 3 of 3 built-in failures caught); bench-log entry with the commands and counts Co-Authored-By: Claude Fable 5.1 --- docs/analysis/scratch-soundness.md | 408 +++++++++++++++++++++++++++++ docs/bench-log.md | 23 ++ 2 files changed, 431 insertions(+) create mode 100644 docs/analysis/scratch-soundness.md diff --git a/docs/analysis/scratch-soundness.md b/docs/analysis/scratch-soundness.md new file mode 100644 index 000000000..8093f0bb1 --- /dev/null +++ b/docs/analysis/scratch-soundness.md @@ -0,0 +1,408 @@ +# Layer 3 soundness: the per-warp scratch with read-modify-writes + +5 October 2026 (night), cryptographer role, Counter ASIC 2.0 plan step 4 (`docs/plans/counter-asic-2.md`). Branch +`ca2-soundness` on top of `readwidth` b970dda (the scratch as a class parameter, 32 or 128 KiB per warp). Tests: +`igneum-pow/tests/scratch.rs`; Metal runs through `proto-metal/packbench` on the M5 Max; commands and counts in +`docs/bench-log.md` (entry of the same date). Nothing here touches the lottery hash as shipped: variant 5 is behind +`LoadClass::scratch(k, kb)` and is never emitted by generator version 2. + +Every figure below is measured (machine, date, command named) or cited; "approximate" marks a figure from memory. + +## 0. The five findings + +| # | Question | Finding | Status | +|---|---|---|---| +| 1 | Is what is written uniform and beyond a chip's precomputation? | The fill is a bijection of the lane nonce, the rewrite a bijection of the fold value in each word; written words show no bit bias over 3 to 12 million rewrites per class (worst 3.63 sigma of 6). The fill IS precomputable, by design, and at 64 slots 78.5 percent of reads are fill reads. | sound as a function; see 2 for what that means | +| 2 | Does any short cut avoid the writes? | No short cut inside a unit: a slot after d read-modify-writes needs all d fold values (replay test). But the live state is bounded by the read-modify-write count, not by the scratch size, because CPU verification resets the scratch per unit: 64 to 320 bytes per lane at scr2 to scr8, whatever the nominal 32 KiB, 128 KiB or 1 MiB. The named chip (cache mirror plus recompute) keeps that in SRAM at under 5 percent of its mirror and its gain does not move at any share under the 6 GB cap. | NOT sound as an anti-chip layer | +| 3 | Is the verifier's one-warp simulation exact? | Exact when the GPU's lazy per-unit tag is unique over the arena's life and the arena holds no stale tag. The kernels rely on this and neither host guarantees it (no clear at allocation, no clear at the 32-bit wrap of the tag counter, 16.4 minutes on a 5090). With the host contract of section 4.3 the simulation is exact: 14 edge packs twice, 200 fuzz packs, consecutive units on one warp and the wrap inside a launch all match the CPU on Metal (228 of 228); a broken tag and a broken fill are caught (3 of 3). | sound with a host contract; today it is luck | +| 4 | The attack surface of the writes | Out of bounds: impossible by the mask, 42 of 42 emitted kernels pass the static check, which catches six deliberate breaks. Aliasing: none, lane-major arenas disjoint by (warp, lane), two logical units of a wave64 get two arenas. Ordering: one lane, one slot, program order; no cross-lane sharing, no atomics needed. Alignment: 16-byte slots at 16-byte offsets from a 256-byte-aligned base. Wrap: identical to the CPU, tested at the launch level. | sound | +| 5 | What a conformance vector must carry | The class and geometry, the fill and rewrite, the host contract (tags, clearing, groups a multiple of warps), two consecutive units on one warp with a forced slot collision, a unit in the top 256 nonces with the wrap inside the launch, and the fingerprint declared independent of the warp count. The standard three-unit vectors catch a broken tag only through base 1,000,000 and would miss it at a 1 MiB scratch. | defined in section 6 | + +Recommendation (section 10): do not adopt layer 3 as the plan states it (read-modify-writes taken from the 16 +dataset loads). It replaces latency-bound dataset reads with cache-bound ones for the GPU, costs the named chip +nothing it cannot keep in a few megabytes of SRAM, and leaves that chip's gain at 2.4x at every share. The lever +that moves that chip is the mixer multiplier of the M16 analysis (x2 brings it to 1.2x, x4 to 0.6x, under the +verifier's 10 ms gate). If a scratch is kept for another reason, add the read-modify-writes beside the 128 loads, +never in their place, and ship the host contract and the vector of section 6 with it. + +## 1. What the branch implements + +| Piece | Where | What | +|---|---|---| +| Class | `igneum-pow/src/generator.rs:170-230` | `LoadClass { scratch: Some(k), scratch_kb }`: `k` of the 16 memory slots are `Op::Scratch`; `scratch_kb` KiB per warp of 16-byte slots, lane-major, `slots = kb x 2` per lane (32 KiB: 64, 128 KiB: 256); `scratch_slot_mask() = slots - 1` | +| Draw | `generator.rs:488-491` | the first `k` of the 16 drawn load slots become scratch ops (a uniform k-subset); the source register follows the fresh-source rule like a load | +| Fill | `igneum-pow/src/verify.rs:30` | `scratch_fill(seed, base, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot x 0x9e3779b1 + (j + 1) x 0x85ebca77)`, j in 0..2 | +| Fold | `verify.rs:18` | `x = dst ^ w0; x = (rotl(x, 11) x 0x9e3779b1) ^ w1; x = (rotl(x, 11) x 0x9e3779b1) ^ w2; dst = x` (the read-width fold over the three data words) | +| Rewrite | `verify.rs:41` | the slot becomes `(x ^ w1, rotl(x, 7) ^ w2, x + w0)` | +| CPU model | `verify.rs:48-100`, `:305-312` | `ScratchModel`: per (lane, slot) a written bit and three words; an unwritten slot reads as its fill; one model per unit, so a unit starts from the fill | +| Acceptance | `igneum-pow/src/accept.rs:202-215, 374` | a scratch site that reads one slot in all 32 lanes rejects the program (lane-constant site); scratch slots carry bit 31 in the address list and are left out of the distinct-address bound, which now covers the dataset loads only | +| GPU statement | `igneum-pow/src/emit.rs:143-157` | `s_ = rN & mask; v_ = 16-byte load of slot s_; m_ = (v_.x == tag) ? ~0 : 0; w = (v_.yzw & m_) \| (fill & ~m_); fold; dst = x_; 16-byte store of (tag, x_ ^ w1_, rotl(x_, 7) ^ w2_, x_ + w0_)` in Metal, CUDA and OpenCL | +| Persistent prologue | `emit.rs:159-175` | `lane = tid & 31; warp_ = tid >> 5; arena = scratch + (warp_ x 32 + lane) x words_per_lane; for (g_ = warp_; g_ < groups; g_ += nwarps_) { gbase = baseNonce + g_ x 32; tag = salt + g_; ... }` | +| Hosts | `proto-metal/packbench.swift:144-164`, `proto-opencl/host.c:1025-1033, 1268` | the arena is allocated and never written by the host; `salt` starts at 1 and advances by the launch's unit count; no clear at allocation, none at the wrap | + +The constraint of the night (coordinator, 5 October 2026): the whole working set on an 8 GB card stays under 6 GB +(1 GiB table, the layer 5 hot table, the scratch of every resident warp, buffers), which caps the scratch at tens of +KiB per warp. On an RTX 5090 at full occupancy (170 SMs x 64 warps = 10,880 warps, approximate hardware maximum; +the measured version 2 kernel ran 24 warps per SM, 4,080 warps, `docs/bench-log.md` M11, 4 October 2026): + +| Scratch per warp | 10,880 warps | 4,080 warps (measured occupancy) | Table + scratch at 10,880 | Under 6 GB with a 1 GiB table | +|---|---|---|---|---| +| 32 KiB | 340 MiB | 128 MiB | 1,364 MiB | yes | +| 128 KiB | 1,360 MiB | 510 MiB | 2,384 MiB | yes | +| 1 MiB (the first experiment) | 10,880 MiB | 4,080 MiB | 11,904 MiB | no | + +## 2. Question 1: uniformity of what is written + +### 2.1 As functions + +The fill of word j of slot s for lane nonce n is `splitmix32(((n ^ seed[j]) + s x 0x9e3779b1 + (j + 1) x 0x85ebca77))`. +`splitmix32` is a bijection of its 32-bit input; for fixed (seed, s, j) the input is a bijection of n. So over any +2^32 consecutive nonces every 32-bit value appears once as the fill of (s, j): uniform. Test +`fill_is_a_bijection_of_the_nonce`: 2^16 consecutive nonces give 2^16 distinct words for 7 slots x 3 word +positions; the fill of lane l at base b equals the fill of lane 0 at base b + l; it wraps with the nonce +(base 0xffffffe0, lane 32 equals nonce 0). + +The rewrite `(x ^ w1, rotl(x, 7) ^ w2, x + w0)` is, for fixed old content w, a bijection of the fold value x in +EACH word. Test `rewrite_is_a_bijection_of_the_fold_value`: 2^16 consecutive x give 2^16 distinct words in each +position for 16 random w. Consequence: a uniform x gives a uniform word in every position, and the three words +are three images of the same x, so a rewritten slot carries exactly 32 bits of new state behind 96 bits of +storage (from w and any one written word, x is recovered; the test checks all three inversions). + +The fold value x is `fold(dst, w)`, a bijection of `dst` for fixed w (xor, then rotate-multiply-xor twice; the +multiplier is odd). So the written words are uniform whenever `dst` is, and `dst` is a register of the running +program. + +### 2.2 The attack: what a chip can precompute + +The fill is a pure function of (seed, nonce, slot): precomputable, and meant to be (the verifier computes it too). +A chip never stores a fill; it computes it in about 10 integer operations when a slot is first touched. The written +words depend on `dst`, the register state at that instruction, which depends on every earlier instruction of the +hash, including the dataset loads. Nothing about them is precomputable before the hash runs. This is the whole of +what question 1 can give: the writes are as unpredictable as the registers. What that is worth is question 2. + +### 2.3 The stats run (the `TESTS.md` section 3 shape) + +Test `written_words_unbiased_and_rehit_rates`, M5 Max, 5 October 2026, `cargo test --test scratch`: for each +class, programs of `igneum-genesis`, `igneum-genesis/stats1`, `igneum-genesis/stats2`, 2^11 units each (196,608 +hashes per class), closed-form dataset, every read-modify-write traced (`verify::interpret_warp_scratch`). Ones +count per bit of every written word and of the change each rewrite makes (written XOR read), sigma = sqrt(N)/2, +limit 6 sigma like the acceptance rule's output check. + +| Class | Slots per lane | RMW per hash per lane | Rewrites traced | Max bias, written words (sigma) | Max bias, written XOR read (sigma) | +|---|---|---|---|---|---| +| scr2k32 | 64 | 16 | 3,145,728 | 2.61 | 3.40 | +| scr4k32 | 64 | 32 | 6,291,456 | 2.18 | 3.81 | +| scr8k32 | 64 | 64 | 12,582,912 | 3.63 | 2.25 | +| scr2k128 | 256 | 16 | 3,145,728 | 3.36 | 2.19 | +| scr4k128 | 256 | 32 | 6,291,456 | 2.71 | 3.68 | +| scr8k128 | 256 | 64 | 12,582,912 | 2.73 | 2.60 | + +576 bit positions (6 classes x 3 words x 32 bits) at under 4 sigma is what fair coins give. Verdict: no structural +bias in what is written. Like `TESTS.md` section 3 this is a sanity check, not a proof of strength. + +## 3. Question 2: no short cut avoids the writes + +### 3.1 Inside a unit: the chain is dependent + +Slot s of lane l, touched d times in a unit, holds `w_d = rewrite(x_d, w_{d-1})`, `w_0 = fill`, with +`x_i = fold(dst_i, w_{i-1})`. `x_i` depends on the slot content before it, which depends on every earlier fold +value of that slot; and `dst_i` is the register state, which the earlier fold values entered. Test +`slot_is_replayable_from_its_fold_values`: a slot after 64 read-modify-writes is reproduced from the fill and the +64 fold values; dropping one diverges. So a chip cannot skip a write and still read the slot later. It has three +ways to hold a slot, all exact: + +| Store | Bytes per lane | Cost on a re-hit | +|---|---|---| +| Dense: every slot, 12 data bytes plus a valid bit | 12 x slots: 776 (64 slots), 3,104 (256), 24,832 (2,048) | one SRAM read | +| Sparse: only touched slots, 12 bytes plus a slot index | about 13 x distinct: 185 to 820 (table below) | one lookup | +| Implicit: only the fold values, 4 bytes plus a slot index per read-modify-write, replay on a re-hit | 5 x 8k: 80 (scr2), 160 (scr4), 320 (scr8) | d rewrites of 5 integer ops | + +The implicit store is smaller than the dense one whenever `slots > 8k / 3`: at scr4 above 10.7 slots, at scr8 +above 21.3. So "the smallest scratch at which keeping it implicitly is dearer than storing it" is 8k/3 slots per +lane, 2.7 to 5.3 KiB per warp at scr4 to scr8. Every size on the table, 32 KiB and above, is past it: a chip +keeps the scratch implicitly in 80 to 320 bytes per lane at any nominal size, and the replay cost is bounded by +the re-hit depth, which the next table measures. + +### 3.2 The re-hit rate at 64 and 256 slots (and at 2,048) + +Measured in the same test run (every read-modify-write of 196,608 hashes per class traced; a re-hit is a read of a +slot the same unit wrote earlier). Birthday: `distinct = S (1 - (1 - 1/S)^n)` for n uniform draws from S slots. + +| Class | S | n = RMW per hash | Distinct slots, birthday | Re-hits, birthday | Re-hit %, birthday | Re-hit %, measured | Max chain depth seen | Slot histogram against uniform | +|---|---|---|---|---|---|---|---|---| +| scr2k32 | 64 | 16 | 14.26 | 1.74 | 10.9 | 12.58 | 7 | chi2 z 22,023; hottest slot 2.74x, coldest 0.83x | +| scr4k32 | 64 | 32 | 25.33 | 6.67 | 20.8 | 21.47 | 8 | z 10,880; 1.87x, 0.92x | +| scr8k32 | 64 | 64 | 40.64 | 23.36 | 36.5 | 36.99 | 9 | z 7,587; 1.39x, 0.91x | +| scr2k128 | 256 | 16 | 15.54 | 0.46 | 2.9 | 3.84 | 5 | z 19,146; 5.10x, 0.82x | +| scr4k128 | 256 | 32 | 30.14 | 1.86 | 5.8 | 6.25 | 6 | z 9,632; 3.06x, 0.90x | +| scr8k128 | 256 | 64 | 56.72 | 7.28 | 11.4 | 11.89 | 6 | z 5,873; 1.98x, 0.90x | +| 1 MiB (not run) | 2,048 | 32 | 31.76 | 0.24 | 0.8 | | | | + +Two readings. First, the slot a read-modify-write addresses is the low 6 or 8 bits of a program register, and +those bits are not uniform: `or` sets them, `mul` clears them, so one slot of 256 is addressed 5.1 times as often +as the mean and the re-hit rate runs 2 to 33 percent above the birthday rate. For the dataset the same bias on the +low bits of a 28-bit address is harmless (it moves the read inside an item); for a 64-slot scratch it concentrates +the chain. Second, the chain depth is small: at scr4k32 the deepest slot in 196,608 hashes saw 8 earlier +read-modify-writes; a replay costs at most 8 x 5 integer operations, against about 1,170 for one dataset item. + +### 3.3 The live state is bounded by the read-modify-write count, not by the size + +The verifier evaluates one unit from nothing but (program, day, nonce group): `ScratchModel::new` per unit, +`verify.rs:296`. Every conforming GPU must therefore start every unit from the fill, which the tag does +(section 4). So no state crosses a unit boundary, and the state a unit can ever read back is what it wrote itself: +at most 8k slots per lane. The nominal size only sets how often those 8k writes land on the same slot (the table +above). The scratch's "memory" is 8k x 16 bytes per lane of touched slots, 256 bytes to 1 KiB at scr2 to scr8, +and a chip holds it implicitly in 80 to 320 bytes. + +The attack of rolling back or sharing scratch between units has nothing to take: a unit starts from the fill +whatever ran before it, so a chip that clears 64 valid bits per unit has rolled back, and nothing one unit wrote +is readable by another. The CPU verifier is that chip. + +### 3.4 The named chip, and what the scratch costs it + +The strongest chip the plan has priced (coordinator, 5 October 2026): the whole 256 MiB cache on the die, computing +every dataset item on the fly. Its cache SRAM, from `docs/analysis/sram-mirror.md` revision 2 (`ca2-analysis` +e6085c6), headline at shipped-product density / bit-cell lower bound, dollars per good die approximate: 164 / 83 +mm^2 and $30 / $13 at N7 (shipped density from AMD 3D V-Cache, 64 MB on 41 mm^2, Hot Chips 2021); 128 / 64 mm^2 and +$46 / $21 at N5, N3E and Intel 18A (TSMC N5 HD macro 31.8 Mib/mm^2 after assist overhead, SemiAnalysis, December +2022); 106 / 54 mm^2 and $56 / $26 at N2; with a 96 MB hot table 226 / 114 at N7, 175 / 89 at N5, 146 / 74 at N2. +The chip's cache cost in the table below is the N5 headline, 128 mm^2 and $46 per good die. It computes every item +through the mixer (`docs/analysis/m16-recompute-attacker-2026-10-05.md`: 128 items per hash, about 1,170 integer +operations per item, 150,000 per hash; at a 50 T op/s integer budget equal to a 5090's, approximate, 0.33 Ghash/s). +Against the measured version 2 rate of the RTX 5090, 139.7 MH/s (`docs/bench-log.md` M11, 4 October 2026), that is +2.4x before any fixed-function factor, 7x with the 3x the M16 analysis allows (approximate). + +Units in flight on that chip. It has no DRAM latency to cover: every one of its 1,024 cache reads per hash is an +on-die SRAM read. Its hash latency is the dependent chain: 128 items x (8 dependent SRAM reads plus 9 mixer +applications). At about 10 ns per on-die read and about 40 ns per 130-operation mixer on a 16-wide integer +pipeline at 2 GHz (both approximate), an item is about 0.4 us and a hash about 50 us; at 0.33 Ghash/s that is +about 17,000 hashes in flight, 530 units of 32 lanes. A tighter pipeline halves it. The GPU covers DRAM latency (40 to 48 ns row +cycle, MEMSYS 2018, more under load) with 130,560 lanes in flight at the measured occupancy (4,080 warps x 32), 348,160 at +full occupancy, that is 8 to 20 times more lanes than the chip needs. + +What the scratch costs that chip, per variant, with the arithmetic: + +Chip cache mirror: 128 mm^2, $46 per good die (N5 headline; 64 mm^2, $21 bit-cell lower bound). Chip scratch SRAM at +the same two densities (2.1 MB/mm^2 headline, 4.2 MB/mm^2 lower bound at N5): + +| Variant | Dataset loads per hash | Chip ops per hash | Chip rate at 50 T op/s | 5090 rate | Chip gain | Chip scratch SRAM at 17,000 lanes, implicit store | Same, dense 64-slot store | Dense store as mm^2, headline / lower bound (N5) | Share of the 256 MiB mirror (any density) | +|---|---|---|---|---|---|---|---|---|---| +| scr0 (control), 128 loads | 128 | 150,000 | 333 MH/s | 139.7 measured | 2.4x | 0 | 0 | 0 | 0 | +| 12.5% replaced (scr2) | 112 | 131,400 | 381 | 160 projected (128/112 x 139.7) | 2.4x | 1.4 MB | 13 MB | 6.2 / 3.1 mm^2 | 4.9% | +| 25% replaced (scr4) | 96 | 112,800 | 443 | 186 projected | 2.4x | 2.7 MB | 13 MB | 6.2 / 3.1 | 4.9% | +| 50% replaced (scr8) | 64 | 75,600 | 661 | 279 projected | 2.4x | 5.4 MB | 13 MB | 6.2 / 3.1 | 4.9% | +| 12.5% added (16 RMW beside 128 loads) | 128 | 150,200 | 333 | 139.7 or below | 2.4x or more | 1.4 MB | 13 MB | 6.2 / 3.1 | 4.9% | +| 25% added | 128 | 150,400 | 332 | 139.7 or below | 2.4x or more | 2.7 MB | 13 MB | 6.2 / 3.1 | 4.9% | +| 50% added | 128 | 150,800 | 332 | 139.7 or below | 2.4x or more | 5.4 MB | 13 MB | 6.2 / 3.1 | 4.9% | +| 256-slot dense store (128 KiB class), any share | | | | | | | 53 MB | 25 / 12.6 | 20% | + +How the rows are computed: a read-modify-write costs the chip about 12 integer operations (fold and rewrite) and +one SRAM access; replacing a load removes an item derivation (1,170 operations); the 5090's rate for a replaced +load is projected from the measured distinct-load bound (the card's rate tracks distinct dataset loads per hash, +`docs/bench-log.md` 3 October, 23.7 G loads/s at 1 GiB; the readwidth agent's M5 Max measurement of the night, +relayed by the coordinator, shows the same: 27.7 MH/s at v2 to 29.4-31.7 at 25 percent replaced and 44.4-49.1 at 50 +percent, 32 KiB per warp). The scratch SRAM is 17,000 lanes x 80 to 320 bytes (implicit) or x 776 bytes (dense at +64 slots) or x 3,104 bytes (dense at 256 slots); its share of the mirror is a ratio of bytes, 4.9 or 20 percent, +whichever density is used for both; the implicit store (the chip's cheaper choice at every size, section 3.1) is +0.5 to 2 percent. The chip's gain is set by operations per dataset item and the GPU's distinct-load bound, and the +scratch touches neither. + +Plain answer to the coordinator's question: no read-modify-write share under the 6 GB cap, replaced or added, +brings the named chip under 2x. The share would be chosen as the smallest at which the chip falls under 1.5x, and +there is none: the gain is 2.4x at 0, 12.5, 25 and 50 percent, 32 or 128 KiB. This changes nothing about the public +claim that layer 3 would have changed: the claim must rest on the mixer, not on the scratch. + +The lever that does move that chip, from the M16 table, beside it: + +| Mixer cost multiplier | Chip ops per hash | Chip rate | Gain against 139.7 MH/s, no fixed-function factor | With a 3x factor (approximate) | CPU verify per warp (M16 table, scaled from 0.41 to 1.2 ms) | 5090 daily dataset build | +|---|---|---|---|---|---|---| +| x1 (today) | 150,000 | 333 MH/s | 2.4x | 7.2x | 0.4 to 1.2 ms | 13.4 ms | +| x2 | 300,000 | 167 | 1.2x | 3.6x | 0.8 to 2.4 ms | 27 ms | +| x4 | 600,000 | 83 | 0.6x | 1.8x | 1.6 to 4.8 ms | 54 ms | +| x8 | 1,200,000 | 42 | 0.3x | 0.9x | 3.3 to 9.6 ms | 107 ms | + +The mixer multiplier leaves the honest hash rate untouched (the miner pays the mixer once a day), costs the chip +linearly, and is bounded by the 10 ms verification gate (x8 is at the gate's edge on this core, and the 2019-class +core of O-1.14 is unmeasured). The scratch costs the honest GPU a measured share of its rate when it spills the +cache and nothing when it does not, and costs the chip a few megabytes. The comparison is not close. + +### 3.5 Where the GPU's writes would cost DRAM latency, and why that does not help + +The GPU's hot scratch footprint is not the nominal size either: it is the slots in-flight units have touched, +about `warps x 32 lanes x distinct slots x 16 bytes` (x 2 at a 32-byte sector, approximate): on the 5090 at 4,080 +resident warps and scr4, 25.3 slots at 64 or 30.1 at 256, 53 to 63 MB of slots, 100 to 125 MB in sectors, around +the card's 96 MiB L2 (`docs/bench-log.md`, 3 October). The readwidth agent's M5 Max rows (coordinator's message: +the rate rises with the share at 32 and 128 KiB) show the scratch sitting in that chip's caches at 4,096 warps. +To push the writes to DRAM latency the hot footprint must pass the last-level cache at the resident count: +`96 MiB / 4,080 warps = 24 KiB per warp`, which at 512 bytes of touched slots per lane per read-modify-write slot +means `8k x 512 B > 24 KiB`, k above 6 (above 48 read-modify-writes per hash) at ANY nominal size on the table, or +a higher resident count. That fits the 6 GB cap (it is the hot set, not the arena, that matters), and it costs +the honest miner a DRAM-latency read-modify-write per slot (a DRAM row cycle is 40 to 48 ns across DDR4, GDDR5 and +HBM2, Li, Reddy and Jacob, MEMSYS 2018; the loaded latency a GPU kernel sees is higher, approximate; DRAM latency +improved 1.3x in two decades while bandwidth improved 20x, Chang 2017, so no memory technology an attacker could +buy removes it, and no shipped mining chip has used HBM or stacked memory) while the named chip still keeps the same +hot set in a few megabytes of SRAM at 8 to 20 times fewer lanes in flight. The write path cannot be made to cost the +chip more than the GPU, because the GPU must keep 8 to 20 times more of it live. + +## 4. Question 3: the verifier's one-warp simulation is exact + +### 4.1 Lazy fill on both sides + +The CPU initialises lazily with a written bit per (lane, slot), one model per unit. The GPU initialises lazily with +a 32-bit tag in word 0 of each 16-byte slot: a slot whose tag equals the unit's tag reads as written, any other +reads as the fill (`emit.rs:143-157`). There is no explicit fill and no reset between units of a persistent warp +(`emit.rs:159-175`: the loop over `g_` keeps the arena). The two agree if and only if, when a unit first touches a +slot, that slot does not already carry the unit's tag. That is: + +1. Tags are unique over the life of the arena's contents (`tag = salt + g_`, `salt` the host's running counter). +2. The arena holds no word equal to a live tag in a slot's tag position before the unit writes it. + +### 4.2 The attacks (the bug classes) + +| Case | What happens | Today | +|---|---|---| +| Recycled allocation | A fresh process starts `salt` at 1 (`packbench.swift:144`, `host.c:1027`). If the driver hands back the previous process's arena with its contents (Metal, CUDA and OpenCL do not promise zeroed memory, approximate), slots tagged 1..N from the old run match the new run's first units exactly, and those units read stale words instead of the fill: a CPU mismatch on every colliding slot. | not guarded; passes on this Mac because fresh allocations read as zero in practice and tag 0 is never issued (luck, not contract) | +| Tag counter wrap | `salt` is 32 bits and advances by units per launch. A 5090 at 139.7 MH/s runs 4.37 M units/s, 2^32 units in 984 s: the counter wraps every 16.4 minutes on one card (81.8 minutes on the M5 Max at 28 MH/s). After the wrap a slot whose LAST writer carried the repeated tag reads as written. With 10,880 arenas each slot is rewritten about 395,000 times between two uses of one tag (at 64 slots a unit leaves a slot untouched with probability 0.60; 0.60^395,000 is 0), so on a full card the wrap is harmless in practice; on a one-warp launch repeated 2^32 times it is not. | not guarded | +| Tag 0 on zeroed memory | A host that starts `salt` at 0 gives unit 0 the tag 0, which a zeroed arena carries in every slot: unit 0 reads zeros for every first touch. | both hosts start at 1; nothing in the pack says they must | +| `groups` not a multiple of the warp count | Warps run different trip counts; the OpenCL local-memory exchange path carries a barrier inside the loop (spec 1.9), so a short warp hangs or desynchronises. | `packbench` refuses it; `host.c` rounds the batch | + +### 4.3 The host contract that makes the simulation exact + +A host of a scratch class MUST: allocate the arena as `warps x 32 x words_per_lane` words and zero it; issue tags +from a 32-bit counter that starts at 1 and advances by the unit count of every launch; zero the arena again before +any launch whose tags would pass 2^32 - 1 (tag 0 is never issued); launch `groups` as a multiple of the warp count. +The zeroing costs one memset of the arena (340 MiB at 32 KiB x 10,880 warps) every 2^32 units, 16 minutes on a +5090. This is the class fix for all four rows: with it the GPU's tag test and the CPU's written bit are the same +predicate. + +### 4.4 The tests (Metal, M5 Max, 5 October 2026) + +Two consecutive units on one persistent warp and the wrap inside a launch (`packbench --warps 1`, +`--batch-base 4294967040`, the option added on this branch); the hand-built edge programs that force every +read-modify-write of a hash onto one slot (so two consecutive units on one arena collide on every slot); the +deliberate breaks. Results in section 7.2. On the CPU, the same edge programs against an independent hand model +(a second interpreter with its own slot store, `tests/scratch.rs`): 56 of 56 cases match, and the hand model with +its rewrite words swapped mismatches on every case (the comparison has teeth). + +## 5. Question 4: the attack surface of the writes + +| Surface | Argument | Test | +|---|---|---| +| Out of bounds | `s_ = rN & (slots - 1)`, so `s_ < slots`; the lane's arena is `(warp_ x 32 + lane) x 4 x slots` words from the base, the access is `arena + 4 x s_ + 0..3`, the largest index is `warps x 32 x 4 x slots - 1`, the host's allocation. The emitter has one scratch template (`emit.rs:143`) and it masks. | `scr_packs_regenerate_and_pass_the_static_scratch_check`: 42 of 42 emitted kernels (7 scr packs x 6 files, the OpenCL bound file carrying two kernels) regenerate byte for byte from program.json and pass the text check: k masked slot definitions with the class mask, k tagged stores, 3k fill calls, one arena definition with the class stride, one tag definition, no `scratch[`; six deliberate breaks caught (section 8) | +| Aliasing between lanes | Lane-major: lane l of warp w owns words `[(32w + l) x 4S, (32w + l + 1) x 4S)`; two (w, l) pairs give disjoint ranges. Inside the range a slot is 4 words at `4 x s_`, so two slots of one lane are disjoint too. | the `lanevar` edge program: one init-dependent slot per lane, 32 lanes at 64 slots share slots in pairs by the birthday bound; any cross-lane aliasing would change the fold; 128 of 128 lanes on Metal (section 7.2) | +| Wave64 (two logical units in one hardware wave) | `warp_ = tid >> 5`, so the two halves get `warp_ = 2w` and `2w + 1`, two arenas; `gbase` and `tag` are per `g_`, per half. | not run on wave64 hardware (the OpenCL emulator's persistent launch is on the readwidth commit; unverified here) | +| Determinism: alignment | A slot is 16 bytes at byte offset `16 x (lane_base + s_)`; the arena base is the buffer base: Metal, CUDA and OpenCL allocations are at least 128-byte aligned (CUDA 256, OpenCL `CL_DEVICE_MEM_BASE_ADDR_ALIGN` at least the largest built-in type, approximate from memory), so every 16-byte vector access is aligned. | Metal: every run of section 7 | +| Determinism: ordering | A lane's two read-modify-writes of the same slot in one hash are a load and a store, then a load and a store, from one thread to one address: program order within a thread holds in every model. No other thread touches the slot (aliasing row), so no atomics, fences or barriers are needed and none are emitted. | `slot0` and `sixteen` edge programs: 64 and 128 dependent read-modify-writes on one slot per lane per hash, standalone and as the second unit on a warp | +| Determinism: vendors | The statement is integer only: xor, rotate by immediate, multiply, add, a 16-byte load and store. Bit-exact across Metal, CUDA and OpenCL by construction; measured only on Metal here. | Metal; CUDA and OpenCL runs are PC jobs (not mine tonight) | +| 32-bit nonce wrap | `gbase = baseNonce + g_ x 32` and `nonce = baseNonce + gid` wrap in 32-bit arithmetic; `scr_fill(gbase + lane)` wraps like the CPU's `base.wrapping_add(lane)`; `out[gid]` indexes by launch position, not by nonce. An aligned unit never straddles 2^32 (spec 1.9), so the wrap case is a launch whose unit SEQUENCE crosses it. | `packbench --batch-base 4294967040 --batch-log2 9`: 16 units from 0xffffff00, the ninth at gbase 0; fingerprint identical at 1 and 4 warps (section 7.2); every fuzz pack runs that launch | + +## 6. Question 5: what a vector for the scratch class must carry + +Before a scratch pack can be a conformance vector (plan step 4, "only then a vector"), it must carry, beyond what +`igneum-program-pack-3` carries today: + +1. The class in the program id and the pack (`scrk`: it is, `program_id_class`, `generator.rs:400-412`) + and the geometry (slots per lane, words per lane, bytes per warp: it is, `program.h`). +2. The fill and the rewrite as text (it is, `program.json` "scratch"). +3. The host contract of section 4.3 as text in `program.h` and `program.json`: tag counter from 1, zero at + allocation and at the wrap, `groups` a multiple of the warp count. Not there today. +4. Vectors that exercise the tag path, which the three standard units do not reliably: two consecutive units on + one warp (bases 0 and 32 in one one-warp launch) for a program whose consecutive units collide on a slot. At + 64 slots any generated program collides (25 touched of 64 per unit; the broken-tag run of section 8 was caught by + base 1,000,000, a warp's 16th unit, and NOT by a two-unit launch whose vectors lack base 32). At 2,048 slots two + consecutive units share a touched slot with probability about 0.4 (32 x 32 / 2,048 expected overlaps = 0.5), so + the standard vectors would miss a broken tag at the 1 MiB size with probability about 0.6 per unit pair. The + edge programs `slot0` and `sixteen` collide on every slot at every size: a vector set should carry one. +5. A unit in the top 256 nonces with the launch crossing 2^32 (`--batch-base` near the top, at least two warps). +6. The batch fingerprint declared independent of the warp count (`8c07620f4d9adefd` for scr4k32 at 2^12 nonces + from base 0 at 1, 2 and 128 warps, section 7.2): unit independence is the property the per-unit reset gives, and + a fingerprint that moved with the warp count would mean a unit read another unit's slot. + +## 7. Tests and results + +### 7.1 CPU (`igneum-pow/tests/scratch.rs`, `cargo test -j4 --test scratch`, M5 Max, 5 October 2026, 3.6 s) + +| Test | What | Result | +|---|---|---| +| `rewrite_is_a_bijection_of_the_fold_value` | 16 random slot contents x 2^16 consecutive fold values, each written word distinct; the three inversions | pass | +| `fill_is_a_bijection_of_the_nonce` | 7 slots x 3 words x 2^16 nonces distinct; lane and base interchange; wrap | pass | +| `written_words_unbiased_and_rehit_rates` | 6 classes x 3 seeds x 2^11 units, every rewrite traced: bias within 6 sigma (worst 3.63), re-hit rate within 0.9x to 2x of birthday, slot histogram, depth histogram | pass (tables of sections 2.3 and 3.2) | +| `edge_programs_match_the_hand_model` | 7 edge programs x 2 geometries x 4 bases (0, 32, 0x7ffffff0, 0xffffffe0) against an independent hand model; the slots driven and the re-hit counts as built; the mutated hand model mismatches | 56 of 56 pass, 56 of 56 teeth | +| `scr_packs_regenerate_and_pass_the_static_scratch_check` | 7 scr packs: program and program id from program.json, 6 kernel texts byte for byte, static scratch check on all 42, the pack's vectors from the CPU; six deliberate breaks caught | pass | +| `fuzz_scr_programs_cpu` | 200 generated programs over the six classes, generator contract and acceptance on every one, 4 units each (one in 0..224, one around 2^31, one in the top 256 nonces, one uniform), traced run equal to the untraced run, every slot inside the lane; writes the 214 packs for Metal with `IGNEUM_SCRATCH_PACKS_OUT` | pass; 200 of 200 have a unit in the top 256 | +| `slot_is_replayable_from_its_fold_values` | 64 dependent read-modify-writes replayed from the fill and the fold values; one dropped diverges | pass | + +The rest of the crate: 33 of 34 lib tests and all pack tests pass; `verify::tests::fold_and_wide_fetch` fails on the +readwidth tip itself (`verify.rs:508`, `k as u32 * 0x9E37_79B1` overflows under the test profile's overflow +checks; the readwidth agent's test, reported to its owner, not touched here). + +### 7.2 Metal (`proto-metal/packbench` built from this branch, M5 Max, 5 October 2026, under `with-lock.sh run`) + +| Run | Launch | Expected | Result | +|---|---|---|---| +| scr4k32, standard pack | 2,048 warps, 2^24 nonces, 1 batch | 3 of 3 standalone, 3 of 3 in batch | PASS, fingerprint `3d1af881bd978fb9`; 1.8 s wall for the whole run (compile, cache, 1 GiB build, vectors, batch) | +| scr4k32, warp-count independence | 2^12 nonces (128 units) at 1, 2 and 128 warps | one fingerprint | `8c07620f4d9adefd` at all three, PASS | +| scr4k32, wrap inside the launch | 512 nonces from 0xffffff00 at 1 and 4 warps | one fingerprint, the base-0 vector inside the window after the wrap | `8e9e233234d3a297` at both, in-batch 1 of 1, PASS | +| scr4k32, broken tag (`tag = salt`), standard vectors | 2,048 warps, 2^24 | the base-1,000,000 vector (warp 530's 16th unit) fails | standalone 3 of 3, in batch 2 of 3, overall FAIL (caught) | +| scr4k32, broken tag, two units on one warp | 1 warp, 2^6 | nothing to catch it: the standard vectors have no base 32 | standalone 3 of 3, in batch 1 of 1, PASS (missed: the point of section 6 item 4) | +| 14 edge packs (7 programs x 32 and 128 KiB), run A | 1 warp, 2^6 (units at bases 0 and 32 on one arena) | 4 of 4 standalone, 2 of 2 in batch each | 14 of 14 PASS (56 of 56 standalone units, 28 of 28 in batch) | +| 14 edge packs, run B | 1 warp, 2^9 from 0xffffff00 (16 units on one arena, the wrap inside) | 4 of 4 standalone, 3 of 3 in batch each | 14 of 14 PASS (56 of 56, 42 of 42) | +| edge `slot0` at 32 and 128 KiB, broken tag (`tag = salt`) | 1 warp, 2^6 | standalone 4 of 4, in batch 1 of 2, FAIL | as expected at both geometries: the second unit read the first's slot 0 and FAILED; the standalone units passed | +| edge `slot0` at 32 KiB, broken lazy fill (`m_` forced to all ones: a first touch reads the stale words) | 1 warp, 2^6 | standalone fails | 0 of 4 standalone, 0 of 2 in batch, FAIL (lane 0 of base 0: GPU `64b49aeb987dae69`, expected `9ff3a2021f66b5be`) | +| 200 fuzz packs (scr2k32 29, scr4k32 26, scr8k32 42, scr2k128 32, scr4k128 26, scr8k128 45; datasets 64 MiB, 256 MiB, 1 GiB) | 2 warps, 2^9 from 0xffffff00 (8 units per warp, the wrap inside) | 4 of 4 standalone, 2 of 2 in the window, 200 of 200 PASS | 200 of 200 PASS: 800 of 800 standalone units (25,600 hashes), 400 of 400 in batch; 91 s for the 200 runs | + +Totals on Metal: 228 of 228 runs PASS where a pass was expected, 3 of 3 FAIL where a failure was built in. + +## 8. Deliberate breaks (the watcher rule) + +| Break | Where | Caught by | Evidence | +|---|---|---|---| +| One slot mask dropped (Metal) | copy of scr4k32 `program.metal` | static check: "masked slot followed by the load: 3, expected 4" | test output | +| Mask 63 changed to 127 on every RMW (Metal) | same | "masked slot followed by the load: 0, expected 4" | test output | +| Arena stride 256 changed to 128 words (Metal) | same | "arena definition: 0, expected 1" | test output | +| A stray `arena[0]` and `scratch[1]` access (Metal) | same | "arena mentions: 10, expected 9; direct scratch indexing: 1, expected 0" | test output | +| One slot mask dropped (OpenCL, CUDA) | copies of scr4k32 `kernel.cl`, `kernel.cu` | "masked slot followed by the load: 3, expected 4" | test output | +| Wrong class geometry or RMW count or kernel count passed against a right text | the same text | the check fails | test output | +| `tag = salt` (every unit of a launch shares the tag) | copy of scr4k32 `program.metal`, on the GPU | the base-1,000,000 vector in a 2,048-warp batch | `vectors standalone 3/3, in batch 2/3`, overall FAIL | +| the same on the `slot0` edge pack, two units on one warp | on the GPU | in-batch 1 of 2 | bench-log entry | +| lazy fill broken (`m_` all ones) | copy of the `slot0` edge pack, on the GPU | standalone vectors | bench-log entry | +| The hand model's rewrite words swapped | `tests/scratch.rs` | every edge case mismatches | 56 of 56 | + +The out-of-bounds break (mask dropped) was not run on the GPU on purpose: Metal does not bounds-check device +buffers (`TESTS.md` section 5), so a run would read another lane's or another buffer's words and "did not crash" would +prove nothing. The static check is the guard, as it is for the dataset mask. + +## 9. What is unverified + +1. CUDA and OpenCL runs of the scratch packs on NVIDIA and AMD (PC jobs, reserved for the readwidth agent tonight); + the 5090's rate per variant, so the "projected" column of section 3.4 is the distinct-load bound, not a + measurement. Wave64 hardware for the two-arena argument. +2. The chip-side latency figures of section 3.4 (10 ns SRAM read, 40 ns mixer) are approximate; the conclusion + does not depend on them: at ten times the in-flight count the scratch is still under a sixth of the mirror. +3. The recycled-allocation case was not reproduced (it needs a driver that hands back live contents); the argument + is that nothing forbids it and the contract of 4.3 removes it. +4. The slot-bias finding (section 3.2) was measured on three seeds per class; the hottest-slot ratio will vary by + program. +5. `verify::tests::fold_and_wide_fetch` on the readwidth tip (section 7.1). + +## 10. Recommendation + +1. Layer 3 is sound as a construct: the written words are uniform, the chain inside a unit has no short cut, the + kernels cannot write out of bounds, and with the host contract of section 4.3 the CPU's one-warp simulation is + exact (14 edge packs, 200 fuzz packs, the wrap, consecutive units on one arena, on Metal). +2. Layer 3 is not sound as a chip-resistance layer, at the capped size or at any size: CPU verification resets the + scratch per unit, so its live state is 8k slots per lane whatever the arena, a chip keeps it implicitly in 80 to + 320 bytes per lane, and the named chip (on-die cache mirror plus recompute) keeps its whole scratch in 1.4 to + 13 MB of SRAM at 530 units in flight, 3 to 5 percent of its mirror. Its gain stays at 2.4x (7x with a 3x + fixed-function factor, approximate) at 0, 12.5, 25 and 50 percent, replaced or added, 32 or 128 KiB. No share + under the 6 GB cap brings it under 2x. +3. Taking the read-modify-writes from the 16 dataset loads makes the hash less memory-hard for everyone: the GPU + measured faster at every share on the M5 Max (readwidth rows), and the chip's operations per hash fall with the + loads. If a scratch is kept at all, add it beside the 128 loads. There is no reason found here to keep one. +4. The lever that moves the named chip is the M16 mixer multiplier: x2 to 1.2x, x4 to 0.6x against the measured + 5090 rate, at 0.8 to 4.8 ms of verification per warp against the 10 ms gate. Decision 2 should price that + against the gate on the 2019-class core (O-1.14) rather than layer 3. +5. If Josh keeps layer 3 for a reason outside this analysis: ship the host contract in the pack, add the four + vector items of section 6 (consecutive units with a forced collision, the wrap launch, the warp-count-independent + fingerprint, the contract text), and run the CUDA and OpenCL twins of section 7.2 on the PCs before the class + becomes a genesis rule. diff --git a/docs/bench-log.md b/docs/bench-log.md index 30b299cbc..a4170c675 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1524,3 +1524,26 @@ What is measured: one BLS12-381 aggregate signature over 16 summed G1 keys plus | on, split 90 s | v3 | 0 / 2 | none / 3 | 278 / 265 | apart | none | 3 on n0 | 2 (n0 reconnected 6 s after the heal, A's chain at about 58 DAA, inside the table) | Reading (the NEW finding, ledger C4). With the module off GHOSTDAG alone converges on the heavier chain and the losing side's records re-determine (F24 works when the chain moves). With the module on the overlay holds during the split (A, with 30% of the frozen table, locks nothing; B locks 7 and 8) and then fails at the heal in the shipped node: B's certificates for blocks off n0's chain are "kept pending until the chain decides (no lock at this index)", n0's chain never decides because GHOSTDAG keeps its heavier tip and nothing turns the certificate into a fork-choice constraint, and once n0's last lock (index 7, DAA 209) is one window old (DAA 329) the frozen table stops applying on A's chain ("no frozen table (no lock on this chain inside the window)"), A's two keys are 100% of A's own window (B's post-cut blocks are red there) and n0 locks 10, 11, 12 alone; B's certificates for 10 and 11 then log CONFLICTING on n0 (n0 log, 17:27:04 to 17:29:54 BST). A finality fork from a 96-s honest partition, no attacker, table intact at the heal; the 150-s run and the v2 control end the same way. The spec's fork choice ("GHOSTDAG among tips through all certified checkpoints", 3.5) is therefore implemented only for certificates over blocks already on the node's chain. Fix named in the ledger entry: verify an off-chain certificate against the table at its own block and let it constrain fork choice (a certificate-driven reorg), then re-determine. Raw: `scratchpad fud-a/c4-results-*.md`, node logs `c4-on90-tmp/`, `c4-v2-control-tmp/`. + +## 5 October 2026, layer 3 scratch soundness (Counter ASIC 2.0 step 4; branch ca2-soundness on readwidth b970dda; cryptographer) + +Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0, other agents' builds and the readwidth measurements running beside (the Mac measure lock was free during the GPU runs; nothing here is a hash-rate figure). Write-up `docs/analysis/scratch-soundness.md`; tests `igneum-pow/tests/scratch.rs`; harness `proto-metal/packbench` built from this branch (`--batch-base` added) into the session scratchpad with `swiftc -O -target arm64-apple-macos11 -framework Metal`. + +CPU, `with-lock.sh build nice -n 19 ~/.cargo/bin/cargo test -j4 --test scratch -- --nocapture` (3.6 s): 7 of 7 pass. Stats, 6 classes x 3 seeds x 2^11 units, every read-modify-write traced (3.1 to 12.6 million per class): written-word bias within 6 sigma (worst 3.63); re-hit rate measured against the uniform birthday rate 12.58 vs 10.91 percent (scr2k32), 21.47 vs 20.83 (scr4k32), 36.99 vs 36.50 (scr8k32), 3.84 vs 2.88 (scr2k128), 6.25 vs 5.82 (scr4k128), 11.89 vs 11.37 (scr8k128); slot histogram non-uniform (hottest slot 1.39x to 5.10x the mean: the slot is a register's low bits); deepest chain 5 to 9. Edge: 7 hand-built programs x 2 geometries x 4 bases against an independent hand model, 56 of 56, and 56 of 56 mismatches with the hand model's rewrite words swapped. Static scratch check: 42 of 42 emitted kernels of the 7 scr packs (regenerated byte for byte from program.json first), 6 deliberate breaks caught. Fuzz: 200 generated scratch programs, contract and acceptance on every instruction, 800 units; `IGNEUM_SCRATCH_PACKS_OUT` wrote 214 packs (57 s, three memory-hard caches). The crate's other tests: 33 of 34 lib tests pass; `verify::tests::fold_and_wide_fetch` fails on the readwidth tip itself (`verify.rs:508`, `k as u32 * 0x9E37_79B1` overflows under the test profile; not touched here). + +Metal, `with-lock.sh run