From ee26e882aaf1042f643e276bb4d79faa058f7a14 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Sun, 4 Oct 2026 20:37:27 +0000 Subject: [PATCH] Variant race: bench-log entry (Metal on the M5 Max, 14 variants, two programs, serve check), the base-only short-circuit, the PC 1 zip's hash in the plan Under emulation (or --race with no other name) the race no longer times base alone: emu/test.sh passes again (9 source checks, the serve protocol with prepare, swap and self-heal, 17 sampled hashes equal to igneum-pow). Mac table: g256 (256 threads per threadgroup) +17.3% and +21.2% over the shipped 32-thread groups on two programs, with the live app's worker sharing the GPU; the serve check found the unfair mutex (fixed in 0845279). Co-Authored-By: Claude Fable 5.1 --- docs/bench-log.md | 78 +++++++++++++++++++++++++++++++++++++ docs/plans/miner-perf.md | 9 +++-- proto-cuda/nvrtc/worker.cpp | 11 ++++++ 3 files changed, 94 insertions(+), 4 deletions(-) diff --git a/docs/bench-log.md b/docs/bench-log.md index 7a70cb436..0df4bed4b 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1051,3 +1051,81 @@ moving (150.8M at the height, then stepping down as PC 2's card paused for a pro the chain; a consensus rule changed under a running network with miners on three platforms. The first measurement of v2 on the devnet's own regime (two large miners, bursty parallel blocks) needs PC 2 back from its job; the cloud numbers stand meanwhile (settle 157 to 272 s, no swing). + +## 4 October 2026, miner performance: variant racing (Metal worker on the M5 Max; the RTX 5090 job is ready, not run) + +Method (`docs/design/miner-tuning.md`): at every hourly prepare the worker compiles the bound kernel in several +variants (unroll, load path, register budget, threads per group, combinations), checks each bit for bit against the +base kernel, times each for 2 s with the job loop paused, and keeps the fastest for the hour. Base is the kernel as +it has always shipped. Code: `proto-metal/main.swift` (`raceProgram`, `--race-test`), `proto-cuda/nvrtc/worker.cpp` +(`racePair`, `--race`), branch `miner-perf`, commit 460a99a. + +Machine: Apple M5 Max, Darwin 25.6.0 (macOS 26.6.2), 64 GiB. CONDITIONS: the live Igneum Miner app's own Metal +worker (`igneum-bench --serve`, pid 14687) was mining on the same GPU throughout, and the load average was 130 at the +build and 14 to 67 during the races (other agents' cargo builds). The absolute MH/s below are therefore about half +of the card's (the app reported 26.7 MH/s on 4 October with the GPU to itself) and each window was contended; the +numbers to read are the ratios, taken as the best of three interleaved rounds per variant so the contention hits +every variant alike. A re-run with the Mac card paused is listed under "next". + +Command (under the measure lock, which holds the build lock too): + + tools/lock/with-lock.sh measure bash scratchpad/metal/measure.sh + = swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal (47 s under load 130) + igneum-bench --race-test --seed igneum-genesis --day 2026-10-04 --race-rounds 3 --race-bench-ms 2000 + igneum-bench --race-test --seed igneum-hourly --day 2026-10-04 --race-rounds 3 --race-bench-ms 2000 + +Dataset 2^28 words (1 GiB, memory-hard, built in 285 and 295 ms), batch 2^22 nonces per launch, programs by the +version-2 generator (128 loads per hash, no wide loads). 14 variants, every one bit-exact with base over 2^16 nonces +(no variant discarded). MH/s = best of 3 rounds, 2 s windows, first launch of each window not counted. + +| variant | threads/group | max threads/group | seed igneum-genesis MH/s | vs base | seed igneum-hourly MH/s | vs base | +|---|---|---|---|---|---|---| +| base | 32 | 1024 | 11.180 | 0 | 10.800 | 0 | +| g64 | 64 | 1024 | 11.582 | +3.6% | 11.801 | +9.3% | +| g128 | 128 | 1024 | 12.736 | +13.9% | 12.306 | +14.0% | +| g256 | 256 | 1024 | **13.114** | **+17.3%** | **13.089** | **+21.2%** | +| u2 | 32 | 1024 | 10.683 | -4.4% | 10.596 | -1.9% | +| u8 | 32 | 1024 | 10.823 | -3.2% | 10.280 | -4.8% | +| mt256 | 32 | 256 | 10.953 | -2.0% | 10.664 | -1.3% | +| mt512 | 32 | 512 | 11.022 | -1.4% | 10.636 | -1.5% | +| mt1024 | 32 | 1024 | 10.666 | -4.6% | 11.150 | +3.2% | +| osize | 32 | 1024 | 10.833 | -3.1% | 10.437 | -3.4% | +| u2-g128 | 128 | 1024 | 12.793 | +14.4% | 12.928 | +19.7% | +| u8-g128 | 128 | 1024 | 12.496 | +11.8% | 11.923 | +10.4% | +| mt256-g128 | 128 | 256 | 12.628 | +13.0% | 12.675 | +17.4% | +| mt512-g256 | 256 | 512 | 13.101 | +17.2% | 12.817 | +18.7% | + +Race cost: compile 1,798 ms (first seed; the Metal compiler cold) and 267 ms, timing 108 s for 14 variants x 3 +rounds (2 s windows plus the 2^16-nonce check); in `--serve` the race runs one round, about 40 s, inside a 600-DAA +lead, with mining paused only inside the windows. + +Reading. On Apple silicon the win is threads per threadgroup: the Metal worker has dispatched one 32-thread group +per 32 nonces since 3 October, and 256-thread groups are 17 to 21% faster on both programs under these conditions, +with 128 close behind; the unroll, register-budget and size-optimisation knobs are within noise or worse on their +own. The winner agrees across the two programs, so a tuning entry `Apple_M5_Max: g256` would be the first +fleet default; the race itself finds it in one round. These two programs are two points, under contention; the +figure for the Mac's own card with the GPU to itself is still to take. Nothing here says anything about NVIDIA: +`w8` (8 warps per block) is the CUDA cousin of `g256`, and whether the 5090 moves at all is what the PC job +(`docs/plans/miner-perf.md`) measures. Range to measure there: from no gain to what the block-size and load-path +variants give on a 1 GiB random-read kernel; no claim. + +Serve-protocol check (the same binary, `--serve --race-rounds 1`, scripted stdin: two inline jobs on pair A, the +deferred race on A, `prepare` of pair B with its race, jobs on A meanwhile, the swap to B, a job across the 32-bit +nonce boundary): 44 jobs done, 0 errors, no found line missed; the inline compile of pair A 220 ms, the deferred +race on A `winner g256 14.207 base 11.404 gain +24.58%` (compile 563 ms, 36 s of windows); `prepared` for B after +35,554 ms = program 58 ms, dataset 524 ms, race 34,972 ms (`winner mt512-g256 12.807 base 10.346 gain +23.79%`), +the swap to B in 0.01 ms, the 64-nonce job across the 32-bit boundary 14.8 ms. Found by this check: the job queued +during a race waited for the whole race (job 2 done after 35,946 ms; the mutex is not fair), so the race now +pauses 150 ms after every window (commit 32d1c01); the re-check of that pause is queued under the `run` lock +behind a 3-hour network run and is reported in the agent's hand-over if it ran. + +NVIDIA side, what the Mac could check: `proto-cuda/nvrtc/emu/test.sh` PASS on the race build (the race off under +emulation, "variants 1 base only, no race (emulation)" logged per pair; 9 source checks PASS, the --serve protocol +with prepare, swap and self-heal unchanged, 17 sampled hashes equal to `igneum-pow hash-bound`); mingw cross-compile of +`igneum-worker-cuda.exe` with the race (`build-windows.sh`, mingw, static): 1,509,376 bytes, the same imports as the +shipped worker (KERNEL32 and the Universal CRT), icon and version block verified; zipped as +`~/Desktop/igneum-worker-cuda-race.zip` (429,387 bytes, sha256 321a086e...c4c049) for the PC 1 job. The race has +not run on a GPU. + +Next: the PC 1 job (ready in `docs/plans/miner-perf.md`); the Mac card paused for a clean absolute table; the +Mac app's own worker on this build (its hourly prepare then races by itself and logs the TUNING record). diff --git a/docs/plans/miner-perf.md b/docs/plans/miner-perf.md index d169d023d..5bb40f113 100644 --- a/docs/plans/miner-perf.md +++ b/docs/plans/miner-perf.md @@ -61,10 +61,11 @@ ForEach-Object { "RESULT $_" }` keeps stderr in the report. ## The publish commands (main session) -The zip: `igneum-worker-cuda-race.zip` (one file, `igneum-worker-cuda.exe`, the `miner-perf` build), produced -by the agent in its scratchpad; copy it somewhere stable first (for example `~/Desktop/`). Its sha256 and size -are in the agent's report and in `packaging/ota/publish-jobs.sh add` output (the script hashes the copy it puts in -the downloads folder; the sha below is the agent's build and must match): +The zip: `~/Desktop/igneum-worker-cuda-race.zip` (429,387 bytes, sha256 +`321a086ef47af05421d530ab251d170443cbafb917093113a316f617a7c4c049`; one file, `igneum-worker-cuda.exe`, 1,509,376 +bytes, built by `proto-cuda/nvrtc/build-windows.sh` from `miner-perf` at 32d1c01 plus the base-only short-circuit, +icon and version block verified). `publish-jobs.sh add` hashes the copy it puts in the downloads folder; it must +print this sha: ```bash cd ~/Projects/igneum # master or the miner-perf worktree: the script is the same diff --git a/proto-cuda/nvrtc/worker.cpp b/proto-cuda/nvrtc/worker.cpp index e8cd73b1c..11597fd2a 100644 --- a/proto-cuda/nvrtc/worker.cpp +++ b/proto-cuda/nvrtc/worker.cpp @@ -629,6 +629,17 @@ static void racePair(Ctx& c, Pair* p, const PfPack& pk, const std::string& bound #ifdef IGNEUM_EMU order.resize(1); // the stand-in checks that the handed-over text is the pack's; no rewrites under emulation #endif + if (order.size() < 2) { + // nothing to race against base (emulation, or --race with no known name): no timing, the base kernel serves + p->blockWarps = c.blockWarps; p->variant = "base"; p->raceMs = wallMs() - t0; + p->raceLine = fmt("race %.16s device %s variants 1 base only, no race (%s)", p->epochHex.c_str(), c.name.c_str(), +#ifdef IGNEUM_EMU + "emulation"); +#else + "no other variant named"); +#endif + return; + } const bool pinnedOnly = !pinned.empty() && order.size() == 2; uint32_t batch = 1u << c.batchLog2; int benchMs = c.raceBenchMs;