igneum/docs/plans/miner-perf.md
igneum-labs 382039ea5d Variant race: bench-log entry (Metal on the M5 Max, 14 variants, two programs, serve check), the base-only short-circuit, the PC 1 zip's hash in the plan
Under emulation (or --race with no other name) the race no longer times base alone: emu/test.sh passes again
(9 source checks, the serve protocol with prepare, swap and self-heal, 17 sampled hashes equal to igneum-pow).
Mac table: g256 (256 threads per threadgroup) +17.3% and +21.2% over the shipped 32-thread groups on two
programs, with the live app's worker sharing the GPU; the serve check found the unfair mutex (fixed in 123ee1a).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:37:27 +00:00

123 lines
8.5 KiB
Markdown

# Miner performance: the variant race on PC 1's RTX 5090
4 October 2026, evening. The race is built and measured on the Mac's Metal worker (`docs/bench-log.md`, "miner
performance: variant racing"); the design is `docs/design/miner-tuning.md`. This plan is the same race on the RTX
5090 with the NVRTC worker, as a signed job for PC 1 (machine `ae432dc7`), ready for the main session to publish.
Nothing here was published by the agent: PC 1 mines, PC 2 runs a shard job.
## What the job does
Two jobs in the signed file, in this order (the app runs them one at a time, in file order):
1. `fetch-race-worker-20261004` (kind `fetch`, `--dir jobs --extract`): the race build of
`igneum-worker-cuda.exe` (cross-compiled on the Mac with mingw from the `miner-perf` branch, the same
`build-windows.sh` the package uses) lands in `%LOCALAPPDATA%\igneum\app\jobs\fetch-race-worker-20261004\`.
2. `run-race-5090-20261004` (kind `run`, `--stop-miners`, cap 20 min): `relay/playbooks/race-5090.ps1`. The
miners are stopped first, so the 5090 is the race's alone. The script copies the NVRTC DLLs from the installed
app next to the fetched exe, takes this hour's prepared pack (`packs\prepare\<epoch16>-<day>`, else
`packs\devnet`), prints the card's state from `nvidia-smi`, and runs the race twice
(`--race --pack <dir> --race-rounds 3 --race-bench-ms 2000 --race-budget-s 400`): 17 variants, each
self-tested against the pack's vectors and timed for 2 s, three interleaved rounds, best per variant. Every
line starts with `RESULT`, so the dashboard's job strip and `tools/jobs.mjs` show them. The miners restart
when the job ends (the engine's job path).
Expected wall time: a 1 GiB dataset build (about 10 s on the 5090 from the 4 October numbers), 17 NVRTC
compiles in four threads, then 17 x 3 x about 2.3 s of timing, twice: about 5 min with the miners stopped.
## The script
`relay/playbooks/race-5090.ps1` (in the repo, parse-checked by `windows.yml` with the other playbooks):
```powershell
$ErrorActionPreference = 'Continue'
$app = $env:IGNEUM_APP_DIR
$jobs = Split-Path $env:IGNEUM_JOB_DIR
$fetched = Join-Path $jobs 'fetch-race-worker-20261004'
$exe = Join-Path $fetched 'igneum-worker-cuda.exe'
if (-not (Test-Path $exe)) { Write-Output "RESULT race worker missing at $exe (the fetch job runs first)"; exit 2 }
$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
if (-not $inst) { Write-Output "RESULT no installed igneum-worker-cuda.exe found (the NVRTC DLLs come from there)"; exit 2 }
Get-ChildItem $inst -Filter 'nvrtc*.dll' | Copy-Item -Destination $fetched -Force
Write-Output "RESULT worker $exe with $((Get-ChildItem $fetched -Filter 'nvrtc*.dll').Count) NVRTC DLL(s) from $inst"
$pack = Get-ChildItem "$app\packs\prepare" -Directory -ErrorAction SilentlyContinue | Sort-Object LastWriteTime -Descending | Select-Object -First 1
if ($pack) { $pack = $pack.FullName } else { $pack = "$app\packs\devnet" }
if (-not (Test-Path "$pack\seeds.txt")) { Write-Output "RESULT no pack with seeds.txt under $app\packs"; exit 2 }
Write-Output "RESULT pack $pack"
Get-Content "$pack\seeds.txt" | ForEach-Object { "RESULT seeds $_" }
& nvidia-smi --query-gpu=name,driver_version,power.limit,power.default_limit,clocks.sm,clocks.mem,temperature.gpu --format=csv,noheader | ForEach-Object { "RESULT gpu $_" }
foreach ($run in 1..2) {
Write-Output "RESULT run $run start $(Get-Date -Format HH:mm:ss)"
& $exe --race --pack $pack --race-rounds 3 --race-bench-ms 2000 --race-budget-s 400 2>&1 | ForEach-Object { "RESULT $_" }
Write-Output "RESULT run $run exit $LASTEXITCODE"
}
& nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu --format=csv,noheader | ForEach-Object { "RESULT gpu-after $_" }
exit 0
```
Quoting, as in the jobs that worked on PC 2 (`run-20261004-173115`): a plain PowerShell file, no nested
`bash -c` here because the NVRTC worker is a native Windows exe (no WSL); `$env:IGNEUM_*` come from the app
(`jobrun.rs` `job_env`); the script runs with its job folder as the working directory. `& $exe ... 2>&1 |
ForEach-Object { "RESULT $_" }` keeps stderr in the report.
## The publish commands (main session)
The zip: `~/Desktop/igneum-worker-cuda-race.zip` (429,387 bytes, sha256
`321a086ef47af05421d530ab251d170443cbafb917093113a316f617a7c4c049`; one file, `igneum-worker-cuda.exe`, 1,509,376
bytes, built by `proto-cuda/nvrtc/build-windows.sh` from `miner-perf` at 32d1c01 plus the base-only short-circuit,
icon and version block verified). `publish-jobs.sh add` hashes the copy it puts in the downloads folder; it must
print this sha:
```bash
cd ~/Projects/igneum # master or the miner-perf worktree: the script is the same
# 1. the race worker (the fetch copies the zip into dl/<token>/ and hashes it)
packaging/ota/publish-jobs.sh add --kind fetch --target ae432dc7 --platform windows \
--file ~/Desktop/igneum-worker-cuda-race.zip --dir jobs --extract \
--id fetch-race-worker-20261004 --title "Race build of the NVRTC worker for PC 1 (variant racing)" --expires-hours 48
# 2. the race, miners stopped for its duration (about 5 min)
packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --platform windows \
--script relay/playbooks/race-5090.ps1 --stop-miners --timeout-minutes 20 \
--id run-race-5090-20261004 --title "Variant race on the RTX 5090 (17 NVRTC variants, 3 rounds, twice)" --expires-hours 48 \
--deploy
# read back (the app polls every 10 minutes; Settings > remote jobs > Check now at once)
node tools/jobs.mjs watch run-race-5090-20261004
node tools/jobs.mjs run-race-5090-20261004 | grep '^RESULT'
```
`--deploy` on the second `add` ships both (the first one leaves the file written, not deployed). The
`--target` is PC 1 only. If `relay/playbooks/race-5090.ps1` is not on master yet, give its path in the
`miner-perf` worktree (`~/Projects/igneum-wt-perf/relay/playbooks/race-5090.ps1`).
## What to read in the result
- `RESULT race <epoch16> device NVIDIA_GeForce_RTX_5090 driver ... variants 17 base=<MH/s>/<regs>r/1w w2=... winner <name> <MH/s> base <MH/s> gain <pct>% ...`,
twice (run 1, run 2). The two winners should agree; the gain is the number for the bench log.
- `| <variant>: <why>` after the line names any variant that was discarded (a compile error on sm_120, a vector
mismatch) or not timed (budget).
- `RESULT winner <name>: <regs> registers, <n> blocks/SM at <w> warp(s)/block`.
- `RESULT gpu ...`: power limit (80% cap by the app's default) and clocks before; `gpu-after` after.
The gain on the 5090 is unknown until this runs. The Mac's Metal race (bench log) is a different compiler and a
different memory system; its table says what the method finds there and nothing about NVIDIA. The range to
measure on the 5090: from no gain (base stays, the compiler was already right for a 128-load random kernel) to
whatever the load path and block variants give on a 1 GiB random-read kernel; the memory-hard kernel is
bandwidth-bound by design, so a double-digit gain would be a surprise to check, not a claim.
## After the result
1. Add the two race lines to `docs/bench-log.md` under "miner performance: variant racing" (the PC 1 rows).
2. If a variant wins twice: the race build of the worker goes into the next app version (it is the same
`worker.cpp`; `build-windows.sh` then `packaging/windows/push-inputs.sh` as for any worker change), so every
machine races at every prepare and logs the `TUNING` record; `node tools/tuning.mjs` shows the fleet table
after a day; `node tools/tuning.mjs --write tuning.json` and `packaging/ota/publish-manifest.sh --version
<current> --tuning tuning.json --deploy` send the per-card defaults back.
3. If base wins both runs: the race costs the 5090 about 40 s of paused mining per hour for nothing on this
program class; keep `--race on` for a day of records before deciding, since programs differ hour to hour.
## Untested
- The NVRTC worker's race has never run on a GPU (the Mac has none): the emulation suite covers the protocol and
the base path only; the rewrites (`#pragma unroll`, `__ldg`/`__ldcg`/`__ldcs`, `__launch_bounds__`,
`--maxrregcount`) are first compiled by the real NVRTC in this job. A variant NVRTC refuses is discarded with
its reason in the line; the race still ends with a winner (base at worst).
- The job's folder assumptions (`jobs\fetch-race-worker-20261004`, the per-user install folder for the DLLs) follow
`jobrun.rs` and the 0.3.3 installer; PC 1 has not run a `fetch --dir jobs --extract` job before.