Variant race: bench-log entry (Metal on the M5 Max, 14 variants, two programs, serve check), the base-only short-circuit, the PC 1 zip's hash in the plan
Under emulation (or --race with no other name) the race no longer times base alone: emu/test.sh passes again
(9 source checks, the serve protocol with prepare, swap and self-heal, 17 sampled hashes equal to igneum-pow).
Mac table: g256 (256 threads per threadgroup) +17.3% and +21.2% over the shipped 32-thread groups on two
programs, with the live app's worker sharing the GPU; the serve check found the unfair mutex (fixed in 32d1c01).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
32d1c01478
commit
f6a2aac986
3 changed files with 94 additions and 4 deletions
|
|
@ -1051,3 +1051,81 @@ moving (150.8M at the height, then stepping down as PC 2's card paused for a pro
|
|||
the chain; a consensus rule changed under a running network with miners on three platforms. The first measurement of
|
||||
v2 on the devnet's own regime (two large miners, bursty parallel blocks) needs PC 2 back from its job; the cloud
|
||||
numbers stand meanwhile (settle 157 to 272 s, no swing).
|
||||
|
||||
## 4 October 2026, miner performance: variant racing (Metal worker on the M5 Max; the RTX 5090 job is ready, not run)
|
||||
|
||||
Method (`docs/design/miner-tuning.md`): at every hourly prepare the worker compiles the bound kernel in several
|
||||
variants (unroll, load path, register budget, threads per group, combinations), checks each bit for bit against the
|
||||
base kernel, times each for 2 s with the job loop paused, and keeps the fastest for the hour. Base is the kernel as
|
||||
it has always shipped. Code: `proto-metal/main.swift` (`raceProgram`, `--race-test`), `proto-cuda/nvrtc/worker.cpp`
|
||||
(`racePair`, `--race`), branch `miner-perf`, commit 460a99a.
|
||||
|
||||
Machine: Apple M5 Max, Darwin 25.6.0 (macOS 26.6.2), 64 GiB. CONDITIONS: the live Igneum Miner app's own Metal
|
||||
worker (`igneum-bench --serve`, pid 14687) was mining on the same GPU throughout, and the load average was 130 at the
|
||||
build and 14 to 67 during the races (other agents' cargo builds). The absolute MH/s below are therefore about half
|
||||
of the card's (the app reported 26.7 MH/s on 4 October with the GPU to itself) and each window was contended; the
|
||||
numbers to read are the ratios, taken as the best of three interleaved rounds per variant so the contention hits
|
||||
every variant alike. A re-run with the Mac card paused is listed under "next".
|
||||
|
||||
Command (under the measure lock, which holds the build lock too):
|
||||
|
||||
tools/lock/with-lock.sh measure bash scratchpad/metal/measure.sh
|
||||
= swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal (47 s under load 130)
|
||||
igneum-bench --race-test --seed igneum-genesis --day 2026-10-04 --race-rounds 3 --race-bench-ms 2000
|
||||
igneum-bench --race-test --seed igneum-hourly --day 2026-10-04 --race-rounds 3 --race-bench-ms 2000
|
||||
|
||||
Dataset 2^28 words (1 GiB, memory-hard, built in 285 and 295 ms), batch 2^22 nonces per launch, programs by the
|
||||
version-2 generator (128 loads per hash, no wide loads). 14 variants, every one bit-exact with base over 2^16 nonces
|
||||
(no variant discarded). MH/s = best of 3 rounds, 2 s windows, first launch of each window not counted.
|
||||
|
||||
| variant | threads/group | max threads/group | seed igneum-genesis MH/s | vs base | seed igneum-hourly MH/s | vs base |
|
||||
|---|---|---|---|---|---|---|
|
||||
| base | 32 | 1024 | 11.180 | 0 | 10.800 | 0 |
|
||||
| g64 | 64 | 1024 | 11.582 | +3.6% | 11.801 | +9.3% |
|
||||
| g128 | 128 | 1024 | 12.736 | +13.9% | 12.306 | +14.0% |
|
||||
| g256 | 256 | 1024 | **13.114** | **+17.3%** | **13.089** | **+21.2%** |
|
||||
| u2 | 32 | 1024 | 10.683 | -4.4% | 10.596 | -1.9% |
|
||||
| u8 | 32 | 1024 | 10.823 | -3.2% | 10.280 | -4.8% |
|
||||
| mt256 | 32 | 256 | 10.953 | -2.0% | 10.664 | -1.3% |
|
||||
| mt512 | 32 | 512 | 11.022 | -1.4% | 10.636 | -1.5% |
|
||||
| mt1024 | 32 | 1024 | 10.666 | -4.6% | 11.150 | +3.2% |
|
||||
| osize | 32 | 1024 | 10.833 | -3.1% | 10.437 | -3.4% |
|
||||
| u2-g128 | 128 | 1024 | 12.793 | +14.4% | 12.928 | +19.7% |
|
||||
| u8-g128 | 128 | 1024 | 12.496 | +11.8% | 11.923 | +10.4% |
|
||||
| mt256-g128 | 128 | 256 | 12.628 | +13.0% | 12.675 | +17.4% |
|
||||
| mt512-g256 | 256 | 512 | 13.101 | +17.2% | 12.817 | +18.7% |
|
||||
|
||||
Race cost: compile 1,798 ms (first seed; the Metal compiler cold) and 267 ms, timing 108 s for 14 variants x 3
|
||||
rounds (2 s windows plus the 2^16-nonce check); in `--serve` the race runs one round, about 40 s, inside a 600-DAA
|
||||
lead, with mining paused only inside the windows.
|
||||
|
||||
Reading. On Apple silicon the win is threads per threadgroup: the Metal worker has dispatched one 32-thread group
|
||||
per 32 nonces since 3 October, and 256-thread groups are 17 to 21% faster on both programs under these conditions,
|
||||
with 128 close behind; the unroll, register-budget and size-optimisation knobs are within noise or worse on their
|
||||
own. The winner agrees across the two programs, so a tuning entry `Apple_M5_Max: g256` would be the first
|
||||
fleet default; the race itself finds it in one round. These two programs are two points, under contention; the
|
||||
figure for the Mac's own card with the GPU to itself is still to take. Nothing here says anything about NVIDIA:
|
||||
`w8` (8 warps per block) is the CUDA cousin of `g256`, and whether the 5090 moves at all is what the PC job
|
||||
(`docs/plans/miner-perf.md`) measures. Range to measure there: from no gain to what the block-size and load-path
|
||||
variants give on a 1 GiB random-read kernel; no claim.
|
||||
|
||||
Serve-protocol check (the same binary, `--serve --race-rounds 1`, scripted stdin: two inline jobs on pair A, the
|
||||
deferred race on A, `prepare` of pair B with its race, jobs on A meanwhile, the swap to B, a job across the 32-bit
|
||||
nonce boundary): 44 jobs done, 0 errors, no found line missed; the inline compile of pair A 220 ms, the deferred
|
||||
race on A `winner g256 14.207 base 11.404 gain +24.58%` (compile 563 ms, 36 s of windows); `prepared` for B after
|
||||
35,554 ms = program 58 ms, dataset 524 ms, race 34,972 ms (`winner mt512-g256 12.807 base 10.346 gain +23.79%`),
|
||||
the swap to B in 0.01 ms, the 64-nonce job across the 32-bit boundary 14.8 ms. Found by this check: the job queued
|
||||
during a race waited for the whole race (job 2 done after 35,946 ms; the mutex is not fair), so the race now
|
||||
pauses 150 ms after every window (commit 32d1c01); the re-check of that pause is queued under the `run` lock
|
||||
behind a 3-hour network run and is reported in the agent's hand-over if it ran.
|
||||
|
||||
NVIDIA side, what the Mac could check: `proto-cuda/nvrtc/emu/test.sh` PASS on the race build (the race off under
|
||||
emulation, "variants 1 base only, no race (emulation)" logged per pair; 9 source checks PASS, the --serve protocol
|
||||
with prepare, swap and self-heal unchanged, 17 sampled hashes equal to `igneum-pow hash-bound`); mingw cross-compile of
|
||||
`igneum-worker-cuda.exe` with the race (`build-windows.sh`, mingw, static): 1,509,376 bytes, the same imports as the
|
||||
shipped worker (KERNEL32 and the Universal CRT), icon and version block verified; zipped as
|
||||
`~/Desktop/igneum-worker-cuda-race.zip` (429,387 bytes, sha256 321a086e...c4c049) for the PC 1 job. The race has
|
||||
not run on a GPU.
|
||||
|
||||
Next: the PC 1 job (ready in `docs/plans/miner-perf.md`); the Mac card paused for a clean absolute table; the
|
||||
Mac app's own worker on this build (its hourly prepare then races by itself and logs the TUNING record).
|
||||
|
|
|
|||
|
|
@ -61,10 +61,11 @@ ForEach-Object { "RESULT $_" }` keeps stderr in the report.
|
|||
|
||||
## The publish commands (main session)
|
||||
|
||||
The zip: `igneum-worker-cuda-race.zip` (one file, `igneum-worker-cuda.exe`, the `miner-perf` build), produced
|
||||
by the agent in its scratchpad; copy it somewhere stable first (for example `~/Desktop/`). Its sha256 and size
|
||||
are in the agent's report and in `packaging/ota/publish-jobs.sh add` output (the script hashes the copy it puts in
|
||||
the downloads folder; the sha below is the agent's build and must match):
|
||||
The zip: `~/Desktop/igneum-worker-cuda-race.zip` (429,387 bytes, sha256
|
||||
`321a086ef47af05421d530ab251d170443cbafb917093113a316f617a7c4c049`; one file, `igneum-worker-cuda.exe`, 1,509,376
|
||||
bytes, built by `proto-cuda/nvrtc/build-windows.sh` from `miner-perf` at 32d1c01 plus the base-only short-circuit,
|
||||
icon and version block verified). `publish-jobs.sh add` hashes the copy it puts in the downloads folder; it must
|
||||
print this sha:
|
||||
|
||||
```bash
|
||||
cd ~/Projects/igneum # master or the miner-perf worktree: the script is the same
|
||||
|
|
|
|||
|
|
@ -629,6 +629,17 @@ static void racePair(Ctx& c, Pair* p, const PfPack& pk, const std::string& bound
|
|||
#ifdef IGNEUM_EMU
|
||||
order.resize(1); // the stand-in checks that the handed-over text is the pack's; no rewrites under emulation
|
||||
#endif
|
||||
if (order.size() < 2) {
|
||||
// nothing to race against base (emulation, or --race with no known name): no timing, the base kernel serves
|
||||
p->blockWarps = c.blockWarps; p->variant = "base"; p->raceMs = wallMs() - t0;
|
||||
p->raceLine = fmt("race %.16s device %s variants 1 base only, no race (%s)", p->epochHex.c_str(), c.name.c_str(),
|
||||
#ifdef IGNEUM_EMU
|
||||
"emulation");
|
||||
#else
|
||||
"no other variant named");
|
||||
#endif
|
||||
return;
|
||||
}
|
||||
const bool pinnedOnly = !pinned.empty() && order.size() == 2;
|
||||
uint32_t batch = 1u << c.batchLog2;
|
||||
int benchMs = c.raceBenchMs;
|
||||
|
|
|
|||
Loading…
Reference in a new issue