the project lead, 4 Oct 2026 evening: "we need to make our miner better than anything else can be". The compile-ahead pipeline built one kernel per program; it now builds a catalogue (unroll 2 or 8; the dataset load path __ldg, __ldcg, __ldcs on NVIDIA; a register budget by --maxrregcount or __launch_bounds__, max_total_threads_per_threadgroup on Apple; 2, 4, 8 warps per block or 64 to 256 threads per threadgroup; combinations), self-tests each against the pack's vector warps (NVIDIA) or the base kernel over 2^16 nonces (Metal), bit for bit or out, and times each for about two seconds with the job loop paused (one mutex, mining resumes between variants). Base is the pack's text as shipped, always first, never discarded; a race has a budget (default 120 s against the 600-DAA lead) and keeps the best so far when it runs out, so the swap is never delayed. One line per race: variants, MH/s each, winner, gain, time. A tuning file (--tuning, IGNEUM_TUNING_FILE) pins a variant or orders the candidates per card model. NVIDIA: textual rewrites on the pack's own kernel_bound.cu with exact anchors from igneum-pow's emitter, so the pack format, the miner and igneum-pow are untouched and old packs race. --race --pack <dir> runs the race alone. Under IGNEUM_EMU the race is off (the stand-in checks the handed-over text is the pack's). Metal: the MSL hooks in generateMSL, a lock-protected kernel slot per program, a pair compiled inline races after its first job, --race-test --seed --day runs the race alone with a table. OpenCL is not raced yet. docs/design/miner-tuning.md. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
10 KiB
Miner tuning: variant racing and fleet learning
4 October 2026, evening. the project lead: "we need to make our miner better than anything else can be". Two levers, both
measured: a race between kernel variants at every hourly swap, and a fleet that remembers which variant each card
model likes. Numbers live in docs/bench-log.md ("miner performance: variant racing"); the PC job is in
docs/plans/miner-perf.md. Nothing here changes the hash: every variant is the same instruction text in a
different shape for the compiler, and a variant that is not bit-exact is discarded before it is timed.
1. Why a race
The lottery program changes every hour. The generator draws 64 instructions with 16 loads; the compiler sees a different straight-line body each time, and what suits one body (full unrolling, a read-only load path, more threads per block) does not suit the next. A fixed compile is a guess. The compile-ahead pipeline already builds the next hour's kernel one lead (600 DAA, about 600 s) before the boundary, so there is time to build several and let the card pick.
2. The variants
| Knob | NVIDIA (NVRTC worker, proto-cuda/nvrtc/worker.cpp) |
Apple (Metal worker, proto-metal/main.swift) |
|---|---|---|
| Unroll | #pragma unroll 2 or 8 before the iteration loop (u2, u8) |
the same pragma (u2, u8) |
| Load path | ds[i] as shipped; __ldg read-only path (ldg); __ldcg L2 only (ldcg); __ldcs streaming (ldcs) |
none (one address space on Apple silicon) |
| Register budget | --maxrregcount=32 or 64 (r32, r64); __launch_bounds__(128, 4) (lb4-w4), (64, 8) (lb8-w2) |
[[max_total_threads_per_threadgroup(N)]] 256, 512, 1024 (mt256 ...) |
| Threads per block | 2, 4, 8 warps (w2, w4, w8) |
64, 128, 256 threads per threadgroup (g64, g128, g256) |
| Compiler | optimizationLevel = .size (osize) |
|
| Combinations | u2-ldg, u2-w4, ldg-w4, ldcg-w4 |
u2-g128, u8-g128, mt256-g128, mt512-g256 |
17 names on NVIDIA, 14 on Metal. base is always the pack's text as shipped with the worker's default block:
the kernel every machine ran before this change. Names are stable; the tuning file and the fleet records use them.
NVIDIA rewrites are textual, on the pack's own kernel_bound.cu, with exact anchors from igneum-pow's emitter
("\n for (uint32_t it = 0u; it < ", " ^ ds[", "__global__ void igneum_hash_bound("); a text without the
anchor refuses the variant instead of guessing. The pack format, the miner and igneum-pow are untouched, so a
new worker races old packs. OpenCL (AMD, Intel) is not raced yet: proto-opencl/host.c builds through
clBuildProgram with -D IGNEUM_GROUP already, so the same catalogue (unroll pragma, group size, -cl- options)
is the next step; see "open".
3. The race inside the prepare
prepare <epoch> <day> <pack> (the miner, one lead before the boundary)
compile kernel.cu, build cache + dataset, self-test base (as before)
compile the variants (NVRTC: 4 threads; Metal: in turn) budget: --race-budget-s, default 120
for each round (1 in --serve):
for each variant: lock the card, self-test, time ~2 s, unlock
winner = fastest; base keeps its place unless beaten by 0.5%
one "race ..." line, then "prepared ..." as before
job on the new pair at the boundary (swaps as before; the winner serves the hour)
Rules that keep the swap safe:
| Rule | Where |
|---|---|
| Base is the first entry and is never discarded; a race that runs out of budget keeps the best so far | racePair, raceProgram |
| Every variant must reproduce the pack's vector warps (NVIDIA) or the base kernel's output over 2^16 nonces (Metal), bit for bit, or it is out | raceTime, raceProgram |
| The card is exclusive while a variant is timed: one mutex, held per chunk by the job loop and per window by the race. Mining pauses about 2 s per variant and resumes between variants | gpuMutex, gpuLock |
--race-budget-s is capped at 540 (the lead is 600 DAA); default 120 |
option parsing |
| A pair compiled inline (nobody prepared it) races after its first job, in the background | Metal raceDue; NVIDIA self-heal path |
--race off restores the old behaviour; --race a,b,c limits the catalogue |
both workers |
Under IGNEUM_EMU (the Mac's emulation test) the race is off: the stand-in checks that the handed-over text is the pack's |
racePair |
Cost per hour: on the 5090 about 17 variants x (2 s window + a self-test) of paused mining, under 1% of the hour, plus the compiles on the CPU. The expected gain is what the race measures; nothing is claimed for it.
The line, one per race (the miner logs it as worker: race ...):
race <epoch16> device <name> driver <d> arch <a> loads <n> wide <n> variants <k> base=<MH/s>/<regs>r/<warps>w u2=... ldg=-
winner <name> <MH/s> base <MH/s> gain <+pct>% compile <ms> bench <ms> total <ms> ms [pinned by tuning|tuned order]
[| <variant>: <why it is out>]
--race --pack <dir> (NVIDIA) and --race-test --seed <s> --day <d> (Metal) run the race alone, three rounds,
and print a table; that is what the bench log and the PC job use.
4. Fleet learning
4.1 The record
The app's engine reads the race line (engine.rs race_line) and writes one TUNING {json} line to the app
log, which the existing intake receives with every upload (Neon miner_logs, the same table tools/logs.mjs
reads). Fields:
| Field | From |
|---|---|
ts, machine (id8), app (version) |
the engine |
card (the worker's device name, spaces as underscores: the key everything else uses), vendor, worker (CUDA, Metal, OpenCL) |
the race line, the card state |
driver, arch (sm_120, metal) |
the race line |
epoch (16 hex), loads, wide (the program class features the generator fixes: loads per hash and wide loads per hash) |
the race line |
variants {name: MH/s or null} |
the race line |
winner, mhs, base_mhs, gain_pct, total_ms, pinned, tuned, notes |
the race line |
power_limit_w, power_w, power_pct, mh_per_w (= mhs / power_w, 0 when the card reports no draw) |
the card state (nvidia-smi telemetry; Apple reports none) |
The card state also carries variant, race_mhs, race_gain_pct, race_variants for the dashboard, and one
event per race ("RTX 5090 kernel race: u2-ldg at 118.3 MH/s (+2.1% over base, 17 variants, 41 s)").
4.2 The aggregation
tools/tuning.mjs on the Mac (or a small job): every app-log upload of the window (default 7 days) with a
TUNING line, de-duplicated on (machine, card, epoch) because the log is re-sent every minute, then per card
model and variant the sample count, the median MH/s and the median MH per watt. The winner is the best median
with at least --min-samples (3) samples; --by mhw ranks by MH per watt instead. --write tuning.json writes:
{"updated": "2026-10-04T21:00:00Z", "window_days": 7,
"cards": {"NVIDIA_GeForce_RTX_5090": {"variant": "u2-ldg", "race": true, "candidates": ["u2-ldg", "ldg", "base"],
"samples": 41, "races": 41, "machines": 2, "mhs": 118.3, "base_mhs": 115.9,
"gain_pct": 2.07, "mh_per_w": 0.254, "worker": "CUDA", "by": "mhs"}}}
A program class split (by loads, wide) is in the record and not yet in the aggregation: the generator fixes
16 loads per instruction block, so every program has 128 loads per hash today; the split starts to matter when the
era draw changes the mix.
4.3 The way back: the manifest
packaging/ota/publish-manifest.sh --tuning tuning.json puts the object under tuning in the signed
igneum-app-latest.json (carried over from the current manifest when not given; --no-tuning drops it; the
same script now also takes --override '{json}' for consensus.override and carries that over too).
manifest.rs parses tuning (an object with a cards object, else the manifest is refused). The updater writes
it as is to <app data>/app/tuning.json (ota.rs write_tuning, next to override.json), logs one event, and
the engine starts every miner with IGNEUM_TUNING_FILE=<that path> (procs::spawn gained an environment
parameter); the miner's child, the worker, reads it at every prepare. No restart for a change: a worker that has
the path reads the file again at its next prepare; a miner started before the file existed is restarted by the
usual hourly path.
What a worker does with its entry (readTuning, readMetalTuning):
| Entry | Behaviour |
|---|---|
| none for this card model | the full race |
race: true with candidates |
the candidates are raced first, then the rest as the budget allows (the fleet keeps learning; the card starts from the known best) |
race: false with variant |
the variant is compiled and self-tested, no timing (about 1 s); a failed self-test falls back to the full race |
--variant <name> on the worker |
the same as a pinned entry, for tests |
4.4 What is implemented and what is not
| Piece | State |
|---|---|
NVRTC worker race, --race mode, tuning file |
code, syntax-checked for mingw and the emulation build; emulation suite; not yet run on a GPU (the PC job does that) |
Metal worker race, --race-test, deferred race, tuning file |
code and the Mac measurement (bench log) |
| OpenCL worker race | not implemented (open) |
Engine: race line to TUNING record, card state, event, IGNEUM_TUNING_FILE |
code, unit tests of the manifest parse |
publish-manifest.sh --tuning, --override, carry-over |
code (dry run against a scratch folder in the plan) |
tools/tuning.mjs |
code; needs records, so no table yet |
| Aggregation as a job on a PC or the observer | not needed yet: the Mac script reads the intake |
| Dashboard field for the variant | state only; the UI does not show it yet |
5. Open
- OpenCL: the same catalogue through
clBuildProgramoptions and the pragma;--group-warpsexists already. - The program class split in the aggregation once eras change the instruction mix.
- A per-card cap on how long a race may pause mining (today the 2 s windows plus the budget); and whether to race only every N hours once a card's winner is stable (the pinned entry does that by hand).
- Power: the record has MH per watt, the race does not touch the power cap; a race across caps is a later lever.