igneum/docs/design/miner-tuning.md
igneum-labs 27bcb81494 Variant racing: the NVRTC and Metal workers compile several kernel variants at every prepare and keep the fastest for the hour
the project lead, 4 Oct 2026 evening: "we need to make our miner better than anything else can be". The compile-ahead pipeline
built one kernel per program; it now builds a catalogue (unroll 2 or 8; the dataset load path __ldg, __ldcg, __ldcs
on NVIDIA; a register budget by --maxrregcount or __launch_bounds__, max_total_threads_per_threadgroup on Apple;
2, 4, 8 warps per block or 64 to 256 threads per threadgroup; combinations), self-tests each against the pack's
vector warps (NVIDIA) or the base kernel over 2^16 nonces (Metal), bit for bit or out, and times each for about two
seconds with the job loop paused (one mutex, mining resumes between variants). Base is the pack's text as shipped,
always first, never discarded; a race has a budget (default 120 s against the 600-DAA lead) and keeps the best so
far when it runs out, so the swap is never delayed. One line per race: variants, MH/s each, winner, gain, time.
A tuning file (--tuning, IGNEUM_TUNING_FILE) pins a variant or orders the candidates per card model.

NVIDIA: textual rewrites on the pack's own kernel_bound.cu with exact anchors from igneum-pow's emitter, so the
pack format, the miner and igneum-pow are untouched and old packs race. --race --pack <dir> runs the race alone.
Under IGNEUM_EMU the race is off (the stand-in checks the handed-over text is the pack's). Metal: the MSL hooks in
generateMSL, a lock-protected kernel slot per program, a pair compiled inline races after its first job,
--race-test --seed --day runs the race alone with a table. OpenCL is not raced yet. docs/design/miner-tuning.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:08:45 +00:00

10 KiB

Miner tuning: variant racing and fleet learning

4 October 2026, evening. the project lead: "we need to make our miner better than anything else can be". Two levers, both measured: a race between kernel variants at every hourly swap, and a fleet that remembers which variant each card model likes. Numbers live in docs/bench-log.md ("miner performance: variant racing"); the PC job is in docs/plans/miner-perf.md. Nothing here changes the hash: every variant is the same instruction text in a different shape for the compiler, and a variant that is not bit-exact is discarded before it is timed.

1. Why a race

The lottery program changes every hour. The generator draws 64 instructions with 16 loads; the compiler sees a different straight-line body each time, and what suits one body (full unrolling, a read-only load path, more threads per block) does not suit the next. A fixed compile is a guess. The compile-ahead pipeline already builds the next hour's kernel one lead (600 DAA, about 600 s) before the boundary, so there is time to build several and let the card pick.

2. The variants

Knob NVIDIA (NVRTC worker, proto-cuda/nvrtc/worker.cpp) Apple (Metal worker, proto-metal/main.swift)
Unroll #pragma unroll 2 or 8 before the iteration loop (u2, u8) the same pragma (u2, u8)
Load path ds[i] as shipped; __ldg read-only path (ldg); __ldcg L2 only (ldcg); __ldcs streaming (ldcs) none (one address space on Apple silicon)
Register budget --maxrregcount=32 or 64 (r32, r64); __launch_bounds__(128, 4) (lb4-w4), (64, 8) (lb8-w2) [[max_total_threads_per_threadgroup(N)]] 256, 512, 1024 (mt256 ...)
Threads per block 2, 4, 8 warps (w2, w4, w8) 64, 128, 256 threads per threadgroup (g64, g128, g256)
Compiler optimizationLevel = .size (osize)
Combinations u2-ldg, u2-w4, ldg-w4, ldcg-w4 u2-g128, u8-g128, mt256-g128, mt512-g256

17 names on NVIDIA, 14 on Metal. base is always the pack's text as shipped with the worker's default block: the kernel every machine ran before this change. Names are stable; the tuning file and the fleet records use them.

NVIDIA rewrites are textual, on the pack's own kernel_bound.cu, with exact anchors from igneum-pow's emitter ("\n for (uint32_t it = 0u; it < ", " ^ ds[", "__global__ void igneum_hash_bound("); a text without the anchor refuses the variant instead of guessing. The pack format, the miner and igneum-pow are untouched, so a new worker races old packs. OpenCL (AMD, Intel) is not raced yet: proto-opencl/host.c builds through clBuildProgram with -D IGNEUM_GROUP already, so the same catalogue (unroll pragma, group size, -cl- options) is the next step; see "open".

3. The race inside the prepare

prepare <epoch> <day> <pack>        (the miner, one lead before the boundary)
  compile kernel.cu, build cache + dataset, self-test base      (as before)
  compile the variants (NVRTC: 4 threads; Metal: in turn)       budget: --race-budget-s, default 120
  for each round (1 in --serve):
    for each variant: lock the card, self-test, time ~2 s, unlock
  winner = fastest; base keeps its place unless beaten by 0.5%
  one "race ..." line, then "prepared ..." as before
job on the new pair at the boundary                              (swaps as before; the winner serves the hour)

Rules that keep the swap safe:

Rule Where
Base is the first entry and is never discarded; a race that runs out of budget keeps the best so far racePair, raceProgram
Every variant must reproduce the pack's vector warps (NVIDIA) or the base kernel's output over 2^16 nonces (Metal), bit for bit, or it is out raceTime, raceProgram
The card is exclusive while a variant is timed: one mutex, held per chunk by the job loop and per window by the race. Mining pauses about 2 s per variant and resumes between variants gpuMutex, gpuLock
--race-budget-s is capped at 540 (the lead is 600 DAA); default 120 option parsing
A pair compiled inline (nobody prepared it) races after its first job, in the background Metal raceDue; NVIDIA self-heal path
--race off restores the old behaviour; --race a,b,c limits the catalogue both workers
Under IGNEUM_EMU (the Mac's emulation test) the race is off: the stand-in checks that the handed-over text is the pack's racePair

Cost per hour: on the 5090 about 17 variants x (2 s window + a self-test) of paused mining, under 1% of the hour, plus the compiles on the CPU. The expected gain is what the race measures; nothing is claimed for it.

The line, one per race (the miner logs it as worker: race ...):

race <epoch16> device <name> driver <d> arch <a> loads <n> wide <n> variants <k> base=<MH/s>/<regs>r/<warps>w u2=... ldg=-
  winner <name> <MH/s> base <MH/s> gain <+pct>% compile <ms> bench <ms> total <ms> ms [pinned by tuning|tuned order]
  [| <variant>: <why it is out>]

--race --pack <dir> (NVIDIA) and --race-test --seed <s> --day <d> (Metal) run the race alone, three rounds, and print a table; that is what the bench log and the PC job use.

4. Fleet learning

4.1 The record

The app's engine reads the race line (engine.rs race_line) and writes one TUNING {json} line to the app log, which the existing intake receives with every upload (Neon miner_logs, the same table tools/logs.mjs reads). Fields:

Field From
ts, machine (id8), app (version) the engine
card (the worker's device name, spaces as underscores: the key everything else uses), vendor, worker (CUDA, Metal, OpenCL) the race line, the card state
driver, arch (sm_120, metal) the race line
epoch (16 hex), loads, wide (the program class features the generator fixes: loads per hash and wide loads per hash) the race line
variants {name: MH/s or null} the race line
winner, mhs, base_mhs, gain_pct, total_ms, pinned, tuned, notes the race line
power_limit_w, power_w, power_pct, mh_per_w (= mhs / power_w, 0 when the card reports no draw) the card state (nvidia-smi telemetry; Apple reports none)

The card state also carries variant, race_mhs, race_gain_pct, race_variants for the dashboard, and one event per race ("RTX 5090 kernel race: u2-ldg at 118.3 MH/s (+2.1% over base, 17 variants, 41 s)").

4.2 The aggregation

tools/tuning.mjs on the Mac (or a small job): every app-log upload of the window (default 7 days) with a TUNING line, de-duplicated on (machine, card, epoch) because the log is re-sent every minute, then per card model and variant the sample count, the median MH/s and the median MH per watt. The winner is the best median with at least --min-samples (3) samples; --by mhw ranks by MH per watt instead. --write tuning.json writes:

{"updated": "2026-10-04T21:00:00Z", "window_days": 7,
 "cards": {"NVIDIA_GeForce_RTX_5090": {"variant": "u2-ldg", "race": true, "candidates": ["u2-ldg", "ldg", "base"],
                                       "samples": 41, "races": 41, "machines": 2, "mhs": 118.3, "base_mhs": 115.9,
                                       "gain_pct": 2.07, "mh_per_w": 0.254, "worker": "CUDA", "by": "mhs"}}}

A program class split (by loads, wide) is in the record and not yet in the aggregation: the generator fixes 16 loads per instruction block, so every program has 128 loads per hash today; the split starts to matter when the era draw changes the mix.

4.3 The way back: the manifest

packaging/ota/publish-manifest.sh --tuning tuning.json puts the object under tuning in the signed igneum-app-latest.json (carried over from the current manifest when not given; --no-tuning drops it; the same script now also takes --override '{json}' for consensus.override and carries that over too). manifest.rs parses tuning (an object with a cards object, else the manifest is refused). The updater writes it as is to <app data>/app/tuning.json (ota.rs write_tuning, next to override.json), logs one event, and the engine starts every miner with IGNEUM_TUNING_FILE=<that path> (procs::spawn gained an environment parameter); the miner's child, the worker, reads it at every prepare. No restart for a change: a worker that has the path reads the file again at its next prepare; a miner started before the file existed is restarted by the usual hourly path.

What a worker does with its entry (readTuning, readMetalTuning):

Entry Behaviour
none for this card model the full race
race: true with candidates the candidates are raced first, then the rest as the budget allows (the fleet keeps learning; the card starts from the known best)
race: false with variant the variant is compiled and self-tested, no timing (about 1 s); a failed self-test falls back to the full race
--variant <name> on the worker the same as a pinned entry, for tests

4.4 What is implemented and what is not

Piece State
NVRTC worker race, --race mode, tuning file code, syntax-checked for mingw and the emulation build; emulation suite; not yet run on a GPU (the PC job does that)
Metal worker race, --race-test, deferred race, tuning file code and the Mac measurement (bench log)
OpenCL worker race not implemented (open)
Engine: race line to TUNING record, card state, event, IGNEUM_TUNING_FILE code, unit tests of the manifest parse
publish-manifest.sh --tuning, --override, carry-over code (dry run against a scratch folder in the plan)
tools/tuning.mjs code; needs records, so no table yet
Aggregation as a job on a PC or the observer not needed yet: the Mac script reads the intake
Dashboard field for the variant state only; the UI does not show it yet

5. Open

  • OpenCL: the same catalogue through clBuildProgram options and the pragma; --group-warps exists already.
  • The program class split in the aggregation once eras change the instruction mix.
  • A per-card cap on how long a race may pause mining (today the 2 s windows plus the budget); and whether to race only every N hours once a card's winner is stable (the pinned entry does that by hand).
  • Power: the record has MH per watt, the race does not touch the power cap; a race across caps is a later lever.