Under emulation (or --race with no other name) the race no longer times base alone: emu/test.sh passes again
(9 source checks, the serve protocol with prepare, swap and self-heal, 17 sampled hashes equal to igneum-pow).
Mac table: g256 (256 threads per threadgroup) +17.3% and +21.2% over the shipped 32-thread groups on two
programs, with the live app's worker sharing the GPU; the serve check found the unfair mutex (fixed in 123ee1a).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
123 lines
8.5 KiB
Markdown
123 lines
8.5 KiB
Markdown
# Miner performance: the variant race on PC 1's RTX 5090
|
|
|
|
4 October 2026, evening. The race is built and measured on the Mac's Metal worker (`docs/bench-log.md`, "miner
|
|
performance: variant racing"); the design is `docs/design/miner-tuning.md`. This plan is the same race on the RTX
|
|
5090 with the NVRTC worker, as a signed job for PC 1 (machine `ae432dc7`), ready for the main session to publish.
|
|
Nothing here was published by the agent: PC 1 mines, PC 2 runs a shard job.
|
|
|
|
## What the job does
|
|
|
|
Two jobs in the signed file, in this order (the app runs them one at a time, in file order):
|
|
|
|
1. `fetch-race-worker-20261004` (kind `fetch`, `--dir jobs --extract`): the race build of
|
|
`igneum-worker-cuda.exe` (cross-compiled on the Mac with mingw from the `miner-perf` branch, the same
|
|
`build-windows.sh` the package uses) lands in `%LOCALAPPDATA%\igneum\app\jobs\fetch-race-worker-20261004\`.
|
|
2. `run-race-5090-20261004` (kind `run`, `--stop-miners`, cap 20 min): `relay/playbooks/race-5090.ps1`. The
|
|
miners are stopped first, so the 5090 is the race's alone. The script copies the NVRTC DLLs from the installed
|
|
app next to the fetched exe, takes this hour's prepared pack (`packs\prepare\<epoch16>-<day>`, else
|
|
`packs\devnet`), prints the card's state from `nvidia-smi`, and runs the race twice
|
|
(`--race --pack <dir> --race-rounds 3 --race-bench-ms 2000 --race-budget-s 400`): 17 variants, each
|
|
self-tested against the pack's vectors and timed for 2 s, three interleaved rounds, best per variant. Every
|
|
line starts with `RESULT`, so the dashboard's job strip and `tools/jobs.mjs` show them. The miners restart
|
|
when the job ends (the engine's job path).
|
|
|
|
Expected wall time: a 1 GiB dataset build (about 10 s on the 5090 from the 4 October numbers), 17 NVRTC
|
|
compiles in four threads, then 17 x 3 x about 2.3 s of timing, twice: about 5 min with the miners stopped.
|
|
|
|
## The script
|
|
|
|
`relay/playbooks/race-5090.ps1` (in the repo, parse-checked by `windows.yml` with the other playbooks):
|
|
|
|
```powershell
|
|
$ErrorActionPreference = 'Continue'
|
|
$app = $env:IGNEUM_APP_DIR
|
|
$jobs = Split-Path $env:IGNEUM_JOB_DIR
|
|
$fetched = Join-Path $jobs 'fetch-race-worker-20261004'
|
|
$exe = Join-Path $fetched 'igneum-worker-cuda.exe'
|
|
if (-not (Test-Path $exe)) { Write-Output "RESULT race worker missing at $exe (the fetch job runs first)"; exit 2 }
|
|
$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
|
|
if (-not $inst) { Write-Output "RESULT no installed igneum-worker-cuda.exe found (the NVRTC DLLs come from there)"; exit 2 }
|
|
Get-ChildItem $inst -Filter 'nvrtc*.dll' | Copy-Item -Destination $fetched -Force
|
|
Write-Output "RESULT worker $exe with $((Get-ChildItem $fetched -Filter 'nvrtc*.dll').Count) NVRTC DLL(s) from $inst"
|
|
$pack = Get-ChildItem "$app\packs\prepare" -Directory -ErrorAction SilentlyContinue | Sort-Object LastWriteTime -Descending | Select-Object -First 1
|
|
if ($pack) { $pack = $pack.FullName } else { $pack = "$app\packs\devnet" }
|
|
if (-not (Test-Path "$pack\seeds.txt")) { Write-Output "RESULT no pack with seeds.txt under $app\packs"; exit 2 }
|
|
Write-Output "RESULT pack $pack"
|
|
Get-Content "$pack\seeds.txt" | ForEach-Object { "RESULT seeds $_" }
|
|
& nvidia-smi --query-gpu=name,driver_version,power.limit,power.default_limit,clocks.sm,clocks.mem,temperature.gpu --format=csv,noheader | ForEach-Object { "RESULT gpu $_" }
|
|
foreach ($run in 1..2) {
|
|
Write-Output "RESULT run $run start $(Get-Date -Format HH:mm:ss)"
|
|
& $exe --race --pack $pack --race-rounds 3 --race-bench-ms 2000 --race-budget-s 400 2>&1 | ForEach-Object { "RESULT $_" }
|
|
Write-Output "RESULT run $run exit $LASTEXITCODE"
|
|
}
|
|
& nvidia-smi --query-gpu=power.draw,clocks.sm,temperature.gpu --format=csv,noheader | ForEach-Object { "RESULT gpu-after $_" }
|
|
exit 0
|
|
```
|
|
|
|
Quoting, as in the jobs that worked on PC 2 (`run-20261004-173115`): a plain PowerShell file, no nested
|
|
`bash -c` here because the NVRTC worker is a native Windows exe (no WSL); `$env:IGNEUM_*` come from the app
|
|
(`jobrun.rs` `job_env`); the script runs with its job folder as the working directory. `& $exe ... 2>&1 |
|
|
ForEach-Object { "RESULT $_" }` keeps stderr in the report.
|
|
|
|
## The publish commands (main session)
|
|
|
|
The zip: `~/Desktop/igneum-worker-cuda-race.zip` (429,387 bytes, sha256
|
|
`321a086ef47af05421d530ab251d170443cbafb917093113a316f617a7c4c049`; one file, `igneum-worker-cuda.exe`, 1,509,376
|
|
bytes, built by `proto-cuda/nvrtc/build-windows.sh` from `miner-perf` at 32d1c01 plus the base-only short-circuit,
|
|
icon and version block verified). `publish-jobs.sh add` hashes the copy it puts in the downloads folder; it must
|
|
print this sha:
|
|
|
|
```bash
|
|
cd ~/Projects/igneum # master or the miner-perf worktree: the script is the same
|
|
# 1. the race worker (the fetch copies the zip into dl/<token>/ and hashes it)
|
|
packaging/ota/publish-jobs.sh add --kind fetch --target ae432dc7 --platform windows \
|
|
--file ~/Desktop/igneum-worker-cuda-race.zip --dir jobs --extract \
|
|
--id fetch-race-worker-20261004 --title "Race build of the NVRTC worker for PC 1 (variant racing)" --expires-hours 48
|
|
# 2. the race, miners stopped for its duration (about 5 min)
|
|
packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --platform windows \
|
|
--script relay/playbooks/race-5090.ps1 --stop-miners --timeout-minutes 20 \
|
|
--id run-race-5090-20261004 --title "Variant race on the RTX 5090 (17 NVRTC variants, 3 rounds, twice)" --expires-hours 48 \
|
|
--deploy
|
|
# read back (the app polls every 10 minutes; Settings > remote jobs > Check now at once)
|
|
node tools/jobs.mjs watch run-race-5090-20261004
|
|
node tools/jobs.mjs run-race-5090-20261004 | grep '^RESULT'
|
|
```
|
|
|
|
`--deploy` on the second `add` ships both (the first one leaves the file written, not deployed). The
|
|
`--target` is PC 1 only. If `relay/playbooks/race-5090.ps1` is not on master yet, give its path in the
|
|
`miner-perf` worktree (`~/Projects/igneum-wt-perf/relay/playbooks/race-5090.ps1`).
|
|
|
|
## What to read in the result
|
|
|
|
- `RESULT race <epoch16> device NVIDIA_GeForce_RTX_5090 driver ... variants 17 base=<MH/s>/<regs>r/1w w2=... winner <name> <MH/s> base <MH/s> gain <pct>% ...`,
|
|
twice (run 1, run 2). The two winners should agree; the gain is the number for the bench log.
|
|
- `| <variant>: <why>` after the line names any variant that was discarded (a compile error on sm_120, a vector
|
|
mismatch) or not timed (budget).
|
|
- `RESULT winner <name>: <regs> registers, <n> blocks/SM at <w> warp(s)/block`.
|
|
- `RESULT gpu ...`: power limit (80% cap by the app's default) and clocks before; `gpu-after` after.
|
|
|
|
The gain on the 5090 is unknown until this runs. The Mac's Metal race (bench log) is a different compiler and a
|
|
different memory system; its table says what the method finds there and nothing about NVIDIA. The range to
|
|
measure on the 5090: from no gain (base stays, the compiler was already right for a 128-load random kernel) to
|
|
whatever the load path and block variants give on a 1 GiB random-read kernel; the memory-hard kernel is
|
|
bandwidth-bound by design, so a double-digit gain would be a surprise to check, not a claim.
|
|
|
|
## After the result
|
|
|
|
1. Add the two race lines to `docs/bench-log.md` under "miner performance: variant racing" (the PC 1 rows).
|
|
2. If a variant wins twice: the race build of the worker goes into the next app version (it is the same
|
|
`worker.cpp`; `build-windows.sh` then `packaging/windows/push-inputs.sh` as for any worker change), so every
|
|
machine races at every prepare and logs the `TUNING` record; `node tools/tuning.mjs` shows the fleet table
|
|
after a day; `node tools/tuning.mjs --write tuning.json` and `packaging/ota/publish-manifest.sh --version
|
|
<current> --tuning tuning.json --deploy` send the per-card defaults back.
|
|
3. If base wins both runs: the race costs the 5090 about 40 s of paused mining per hour for nothing on this
|
|
program class; keep `--race on` for a day of records before deciding, since programs differ hour to hour.
|
|
|
|
## Untested
|
|
|
|
- The NVRTC worker's race has never run on a GPU (the Mac has none): the emulation suite covers the protocol and
|
|
the base path only; the rewrites (`#pragma unroll`, `__ldg`/`__ldcg`/`__ldcs`, `__launch_bounds__`,
|
|
`--maxrregcount`) are first compiled by the real NVRTC in this job. A variant NVRTC refuses is discarded with
|
|
its reason in the line; the race still ends with a winner (base at worst).
|
|
- The job's folder assumptions (`jobs\fetch-race-worker-20261004`, the per-user install folder for the DLLs) follow
|
|
`jobrun.rs` and the 0.3.3 installer; PC 1 has not run a `fetch --dir jobs --extract` job before.
|