Variant race: bench-log entry (Metal on the M5 Max, 14 variants, two programs, serve check), the base-only short-circuit, the PC 1 zip's hash in the plan

Under emulation (or --race with no other name) the race no longer times base alone: emu/test.sh passes again
(9 source checks, the serve protocol with prepare, swap and self-heal, 17 sampled hashes equal to igneum-pow).
Mac table: g256 (256 threads per threadgroup) +17.3% and +21.2% over the shipped 32-thread groups on two
programs, with the live app's worker sharing the GPU; the serve check found the unfair mutex (fixed in 32d1c01).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-04 21:37:27 +01:00
parent 32d1c01478
commit f6a2aac986
3 changed files with 94 additions and 4 deletions

View file

@ -1051,3 +1051,81 @@ moving (150.8M at the height, then stepping down as PC 2's card paused for a pro
the chain; a consensus rule changed under a running network with miners on three platforms. The first measurement of
v2 on the devnet's own regime (two large miners, bursty parallel blocks) needs PC 2 back from its job; the cloud
numbers stand meanwhile (settle 157 to 272 s, no swing).
## 4 October 2026, miner performance: variant racing (Metal worker on the M5 Max; the RTX 5090 job is ready, not run)
Method (`docs/design/miner-tuning.md`): at every hourly prepare the worker compiles the bound kernel in several
variants (unroll, load path, register budget, threads per group, combinations), checks each bit for bit against the
base kernel, times each for 2 s with the job loop paused, and keeps the fastest for the hour. Base is the kernel as
it has always shipped. Code: `proto-metal/main.swift` (`raceProgram`, `--race-test`), `proto-cuda/nvrtc/worker.cpp`
(`racePair`, `--race`), branch `miner-perf`, commit 460a99a.
Machine: Apple M5 Max, Darwin 25.6.0 (macOS 26.6.2), 64 GiB. CONDITIONS: the live Igneum Miner app's own Metal
worker (`igneum-bench --serve`, pid 14687) was mining on the same GPU throughout, and the load average was 130 at the
build and 14 to 67 during the races (other agents' cargo builds). The absolute MH/s below are therefore about half
of the card's (the app reported 26.7 MH/s on 4 October with the GPU to itself) and each window was contended; the
numbers to read are the ratios, taken as the best of three interleaved rounds per variant so the contention hits
every variant alike. A re-run with the Mac card paused is listed under "next".
Command (under the measure lock, which holds the build lock too):
tools/lock/with-lock.sh measure bash scratchpad/metal/measure.sh
= swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal (47 s under load 130)
igneum-bench --race-test --seed igneum-genesis --day 2026-10-04 --race-rounds 3 --race-bench-ms 2000
igneum-bench --race-test --seed igneum-hourly --day 2026-10-04 --race-rounds 3 --race-bench-ms 2000
Dataset 2^28 words (1 GiB, memory-hard, built in 285 and 295 ms), batch 2^22 nonces per launch, programs by the
version-2 generator (128 loads per hash, no wide loads). 14 variants, every one bit-exact with base over 2^16 nonces
(no variant discarded). MH/s = best of 3 rounds, 2 s windows, first launch of each window not counted.
| variant | threads/group | max threads/group | seed igneum-genesis MH/s | vs base | seed igneum-hourly MH/s | vs base |
|---|---|---|---|---|---|---|
| base | 32 | 1024 | 11.180 | 0 | 10.800 | 0 |
| g64 | 64 | 1024 | 11.582 | +3.6% | 11.801 | +9.3% |
| g128 | 128 | 1024 | 12.736 | +13.9% | 12.306 | +14.0% |
| g256 | 256 | 1024 | **13.114** | **+17.3%** | **13.089** | **+21.2%** |
| u2 | 32 | 1024 | 10.683 | -4.4% | 10.596 | -1.9% |
| u8 | 32 | 1024 | 10.823 | -3.2% | 10.280 | -4.8% |
| mt256 | 32 | 256 | 10.953 | -2.0% | 10.664 | -1.3% |
| mt512 | 32 | 512 | 11.022 | -1.4% | 10.636 | -1.5% |
| mt1024 | 32 | 1024 | 10.666 | -4.6% | 11.150 | +3.2% |
| osize | 32 | 1024 | 10.833 | -3.1% | 10.437 | -3.4% |
| u2-g128 | 128 | 1024 | 12.793 | +14.4% | 12.928 | +19.7% |
| u8-g128 | 128 | 1024 | 12.496 | +11.8% | 11.923 | +10.4% |
| mt256-g128 | 128 | 256 | 12.628 | +13.0% | 12.675 | +17.4% |
| mt512-g256 | 256 | 512 | 13.101 | +17.2% | 12.817 | +18.7% |
Race cost: compile 1,798 ms (first seed; the Metal compiler cold) and 267 ms, timing 108 s for 14 variants x 3
rounds (2 s windows plus the 2^16-nonce check); in `--serve` the race runs one round, about 40 s, inside a 600-DAA
lead, with mining paused only inside the windows.
Reading. On Apple silicon the win is threads per threadgroup: the Metal worker has dispatched one 32-thread group
per 32 nonces since 3 October, and 256-thread groups are 17 to 21% faster on both programs under these conditions,
with 128 close behind; the unroll, register-budget and size-optimisation knobs are within noise or worse on their
own. The winner agrees across the two programs, so a tuning entry `Apple_M5_Max: g256` would be the first
fleet default; the race itself finds it in one round. These two programs are two points, under contention; the
figure for the Mac's own card with the GPU to itself is still to take. Nothing here says anything about NVIDIA:
`w8` (8 warps per block) is the CUDA cousin of `g256`, and whether the 5090 moves at all is what the PC job
(`docs/plans/miner-perf.md`) measures. Range to measure there: from no gain to what the block-size and load-path
variants give on a 1 GiB random-read kernel; no claim.
Serve-protocol check (the same binary, `--serve --race-rounds 1`, scripted stdin: two inline jobs on pair A, the
deferred race on A, `prepare` of pair B with its race, jobs on A meanwhile, the swap to B, a job across the 32-bit
nonce boundary): 44 jobs done, 0 errors, no found line missed; the inline compile of pair A 220 ms, the deferred
race on A `winner g256 14.207 base 11.404 gain +24.58%` (compile 563 ms, 36 s of windows); `prepared` for B after
35,554 ms = program 58 ms, dataset 524 ms, race 34,972 ms (`winner mt512-g256 12.807 base 10.346 gain +23.79%`),
the swap to B in 0.01 ms, the 64-nonce job across the 32-bit boundary 14.8 ms. Found by this check: the job queued
during a race waited for the whole race (job 2 done after 35,946 ms; the mutex is not fair), so the race now
pauses 150 ms after every window (commit 32d1c01); the re-check of that pause is queued under the `run` lock
behind a 3-hour network run and is reported in the agent's hand-over if it ran.
NVIDIA side, what the Mac could check: `proto-cuda/nvrtc/emu/test.sh` PASS on the race build (the race off under
emulation, "variants 1 base only, no race (emulation)" logged per pair; 9 source checks PASS, the --serve protocol
with prepare, swap and self-heal unchanged, 17 sampled hashes equal to `igneum-pow hash-bound`); mingw cross-compile of
`igneum-worker-cuda.exe` with the race (`build-windows.sh`, mingw, static): 1,509,376 bytes, the same imports as the
shipped worker (KERNEL32 and the Universal CRT), icon and version block verified; zipped as
`~/Desktop/igneum-worker-cuda-race.zip` (429,387 bytes, sha256 321a086e...c4c049) for the PC 1 job. The race has
not run on a GPU.
Next: the PC 1 job (ready in `docs/plans/miner-perf.md`); the Mac card paused for a clean absolute table; the
Mac app's own worker on this build (its hourly prepare then races by itself and logs the TUNING record).

View file

@ -61,10 +61,11 @@ ForEach-Object { "RESULT $_" }` keeps stderr in the report.
## The publish commands (main session)
The zip: `igneum-worker-cuda-race.zip` (one file, `igneum-worker-cuda.exe`, the `miner-perf` build), produced
by the agent in its scratchpad; copy it somewhere stable first (for example `~/Desktop/`). Its sha256 and size
are in the agent's report and in `packaging/ota/publish-jobs.sh add` output (the script hashes the copy it puts in
the downloads folder; the sha below is the agent's build and must match):
The zip: `~/Desktop/igneum-worker-cuda-race.zip` (429,387 bytes, sha256
`321a086ef47af05421d530ab251d170443cbafb917093113a316f617a7c4c049`; one file, `igneum-worker-cuda.exe`, 1,509,376
bytes, built by `proto-cuda/nvrtc/build-windows.sh` from `miner-perf` at 32d1c01 plus the base-only short-circuit,
icon and version block verified). `publish-jobs.sh add` hashes the copy it puts in the downloads folder; it must
print this sha:
```bash
cd ~/Projects/igneum # master or the miner-perf worktree: the script is the same

View file

@ -629,6 +629,17 @@ static void racePair(Ctx& c, Pair* p, const PfPack& pk, const std::string& bound
#ifdef IGNEUM_EMU
order.resize(1); // the stand-in checks that the handed-over text is the pack's; no rewrites under emulation
#endif
if (order.size() < 2) {
// nothing to race against base (emulation, or --race with no known name): no timing, the base kernel serves
p->blockWarps = c.blockWarps; p->variant = "base"; p->raceMs = wallMs() - t0;
p->raceLine = fmt("race %.16s device %s variants 1 base only, no race (%s)", p->epochHex.c_str(), c.name.c_str(),
#ifdef IGNEUM_EMU
"emulation");
#else
"no other variant named");
#endif
return;
}
const bool pinnedOnly = !pinned.empty() && order.size() == 2;
uint32_t batch = 1u << c.batchLog2;
int benchMs = c.raceBenchMs;