From 8fcd391270c478f3d056d484fff08be89aa04c2c Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Sat, 3 Oct 2026 15:31:04 +0000 Subject: [PATCH] RTX 5090 first run: 96/96 vectors PASS, 228 Mhash/s at 1 GiB; Windows toolset note Co-Authored-By: Claude Fable 5.1 --- docs/bench-log.md | 14 ++++++++++++++ proto-cuda/build.bat | 5 ++++- 2 files changed, 18 insertions(+), 1 deletion(-) diff --git a/docs/bench-log.md b/docs/bench-log.md index fca790e48..be676ef53 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -70,3 +70,17 @@ Memcheck: every `dataset[` in the MSL is `dataset[rN & MASK]` (13/13 at 3 sizes) Bench re-run after the changes: igneum-genesis 44.56 Mhash/s, epoch1 47.74 Mhash/s at 1 GiB, PASS 3/3 warps each (within 2 percent of the first-run table). `--export-pack igneum-genesis` re-run is byte-identical to the existing pack. SHORTCUT MEASURED: `--inline-dataset` replaces every load with the six-op closed form ds_elem and never reads memory: 4,888 Mhash/s wall (6,274 GPU time) vs 44.6 honest at 1 GiB, about 110x, and about 9x the cache-resident honest rate. With a closed-form dataset the hash is not memory-hard; an expensive dataset derivation is required, not optional. Not demonstrated: cryptographic strength, weak-program frequency and rejection, NVIDIA/AMD bit-exactness (CUDA run still pending), CPU verify gate with an expensive dataset element. Next three tests for the cryptographer are listed in TESTS.md section 8. + +## 3 October 2026, RTX 5090 first run (the project lead's PC, Windows, CUDA 12.8 runtime, driver 13.4, Visual Studio 2026 with the 14.30 toolset selected via vcvarsall -vcvars_ver=14.30) + +Pack igneum-genesis, dataset 1024 MiB, 5 batches x 2^24 hashes, 1 warp per block. + +| Card | Mhash/s at 1 GiB | GB/s useful | random loads/s | dataset fill | vectors | +|---|---|---|---|---|---| +| NVIDIA RTX 5090 (170 SMs, 32 GB) | 228.1 | 94.9 | 23.7 G | 0.66 ms, 1638 GB/s | 96/96 PASS, standalone and in batch | +| Apple M5 Max (40 GPU cores, same program, same day) | 45.2 | 18.8 | 4.6 G | 2.34 ms, 427 GB/s | 96/96 PASS | + +Result: the same hourly program, generated on the Mac, compiled by Apple's Metal and NVIDIA's CUDA, produced identical hashes on both vendors. Vendor independence of the lottery program is demonstrated for one program; igneum-hourly and the dataset sweep are the next runs. The ratio 5090 to M5 Max is about 5x on hashes and on random loads per second, approximate, consistent with a memory-bound program (random-access bound, not bandwidth bound: the 5090 writes the dataset at 1638 GB/s but hashes at 95 GB/s of useful 4-byte loads). Caveat unchanged: the prototype dataset is closed-form and not yet memory-hard (see TESTS.md), so these are prototype numbers, not mining numbers. + +Build note for Windows: CUDA 12.8 crashes (cudafe++ access violation) under Visual Studio 2026's 14.51 toolset even with -allow-unsupported-compiler. Fix: install the MSVC v143 (14.30) component and open the environment with +`"C:\Program Files\Microsoft Visual Studio\18\Community\VC\Auxiliary\Build\vcvarsall.bat" x64 -vcvars_ver=14.30`, then build normally. diff --git a/proto-cuda/build.bat b/proto-cuda/build.bat index 5a004e460..5054ba915 100644 --- a/proto-cuda/build.bat +++ b/proto-cuda/build.bat @@ -1,6 +1,9 @@ @echo off rem Build igneum-bench-cuda for one program pack (Windows). -rem Run from an "x64 Native Tools Command Prompt" (VS 2022 or newer) so cl.exe is on PATH. +rem Run from an "x64 Native Tools Command Prompt" so cl.exe is on PATH. +rem Visual Studio 2026 (toolset 14.5x) crashes CUDA 12.8 (cudafe++ 0xC0000005). Install the "MSVC v143" component and run +rem "C:\Program Files\Microsoft Visual Studio\18\Community\VC\Auxiliary\Build\vcvarsall.bat" x64 -vcvars_ver=14.30 +rem first, then this script. The script passes -allow-unsupported-compiler already. rem with CUDA Toolkit 12.8 or newer installed (nvcc on PATH). rem Usage: build.bat [pack] [arch] rem pack directory name under packs\ (default igneum-genesis)