igneum/docs/design/miner-tuning.md
igneum-labs 7eed16a29a Pre-public scrub, the text pass (7 October 2026, 19:5x UK): no founder name, personal login, earlier business or personal address in any tracked text file, and a gate check that keeps it so
The sweep (main's item 1): 199 tracked text files, 783 lines. The founder's full name, first name and possessive become "the founder" (sentence starts capitalised); the lowercase operating-system user name in WSL paths and commands becomes <user>; the second owner login becomes "the second owner login"; the three earlier businesses and the two other brands become "the other business", "the earlier entity", "the earlier business" and "another brand"; the Chrome profile rule names the igneum.network profile, not the profile's label. The standing commit login igneum-labs is not a founder term here: the fresh-repository step renames it in the history (docs/plans/history-rewrite.md, tools/repo/fresh-repo.sh).

The patterns never appear in plain text in the tree (a plaintext list would be the hit): tools/ci/founder-strings.b64 (perl regex, tab, a sample per row) is read by tools/ci/founder-strings-check.sh (every tracked text file, perl, known-failed first: the self-test plants each row's sample in a fixture and the hit must name the file), by tools/community/discord-hooks.mjs (the guard's founder and business rows; the test takes its fixtures from the samples) and by tools/repo/fresh-repo.sh (the business names of the rewrite rules). site/forbidden-strings.txt carries the same patterns as b64: lines, decoded case-insensitive by site/scrub.mjs and tools/ci/launch-gates-check.mjs (whose fixture now plants an encoded made-up name). The check runs in the gate's tree checks on every merge.

Not in this commit, by main's word: the 105 commit messages and 40 personal-identity commits that need the history rewrite (listed, not run), and the secrets found by gitleaks over the history (reported with owners).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 18:39:50 +00:00

10 KiB

Miner tuning: variant racing and fleet learning

4 October 2026, evening. The founder: "we need to make our miner better than anything else can be". Two levers, both measured: a race between kernel variants at every hourly swap, and a fleet that remembers which variant each card model likes. Numbers live in docs/bench-log.md ("miner performance: variant racing"); the PC job is in docs/plans/miner-perf.md. Nothing here changes the hash: every variant is the same instruction text in a different shape for the compiler, and a variant that is not bit-exact is discarded before it is timed.

1. Why a race

The lottery program changes every hour. The generator draws 64 instructions with 16 loads; the compiler sees a different straight-line body each time, and what suits one body (full unrolling, a read-only load path, more threads per block) does not suit the next. A fixed compile is a guess. The compile-ahead pipeline already builds the next hour's kernel one lead (600 DAA, about 600 s) before the boundary, so there is time to build several and let the card pick.

2. The variants

Knob NVIDIA (NVRTC worker, proto-cuda/nvrtc/worker.cpp) Apple (Metal worker, proto-metal/main.swift)
Unroll #pragma unroll 2 or 8 before the iteration loop (u2, u8) the same pragma (u2, u8)
Load path ds[i] as shipped; __ldg read-only path (ldg); __ldcg L2 only (ldcg); __ldcs streaming (ldcs) none (one address space on Apple silicon)
Register budget --maxrregcount=32 or 64 (r32, r64); __launch_bounds__(128, 4) (lb4-w4), (64, 8) (lb8-w2) [[max_total_threads_per_threadgroup(N)]] 256, 512, 1024 (mt256 ...)
Threads per block 2, 4, 8 warps (w2, w4, w8) 64, 128, 256 threads per threadgroup (g64, g128, g256)
Compiler optimizationLevel = .size (osize)
Combinations u2-ldg, u2-w4, ldg-w4, ldcg-w4 u2-g128, u8-g128, mt256-g128, mt512-g256

17 names on NVIDIA, 14 on Metal. base is always the pack's text as shipped with the worker's default block: the kernel every machine ran before this change. Names are stable; the tuning file and the fleet records use them.

NVIDIA rewrites are textual, on the pack's own kernel_bound.cu, with exact anchors from igneum-pow's emitter ("\n for (uint32_t it = 0u; it < ", " ^ ds[", "__global__ void igneum_hash_bound("); a text without the anchor refuses the variant instead of guessing. The pack format, the miner and igneum-pow are untouched, so a new worker races old packs. OpenCL (AMD, Intel) is not raced yet: proto-opencl/host.c builds through clBuildProgram with -D IGNEUM_GROUP already, so the same catalogue (unroll pragma, group size, -cl- options) is the next step; see "open".

3. The race inside the prepare

prepare <epoch> <day> <pack>        (the miner, one lead before the boundary)
  compile kernel.cu, build cache + dataset, self-test base      (as before)
  compile the variants (NVRTC: 4 threads; Metal: in turn)       budget: --race-budget-s, default 120
  for each round (1 in --serve):
    for each variant: lock the card, self-test, time ~2 s, unlock
  winner = fastest; base keeps its place unless beaten by 0.5%
  one "race ..." line, then "prepared ..." as before
job on the new pair at the boundary                              (swaps as before; the winner serves the hour)

Rules that keep the swap safe:

Rule Where
Base is the first entry and is never discarded; a race that runs out of budget keeps the best so far racePair, raceProgram
Every variant must reproduce the pack's vector warps (NVIDIA) or the base kernel's output over 2^16 nonces (Metal), bit for bit, or it is out raceTime, raceProgram
The card is exclusive while a variant is timed: one mutex, held per chunk by the job loop and per window by the race. Mining pauses about 2 s per variant and resumes between variants gpuMutex, gpuLock
--race-budget-s is capped at 540 (the lead is 600 DAA); default 120 option parsing
A pair compiled inline (nobody prepared it) races after its first job, in the background Metal raceDue; NVIDIA self-heal path
--race off restores the old behaviour; --race a,b,c limits the catalogue both workers
Under IGNEUM_EMU (the Mac's emulation test) the race is off: the stand-in checks that the handed-over text is the pack's racePair

Cost per hour: on the 5090 about 17 variants x (2 s window + a self-test) of paused mining, under 1% of the hour, plus the compiles on the CPU. The expected gain is what the race measures; nothing is claimed for it.

The line, one per race (the miner logs it as worker: race ...):

race <epoch16> device <name> driver <d> arch <a> loads <n> wide <n> variants <k> base=<MH/s>/<regs>r/<warps>w u2=... ldg=-
  winner <name> <MH/s> base <MH/s> gain <+pct>% compile <ms> bench <ms> total <ms> ms [pinned by tuning|tuned order]
  [| <variant>: <why it is out>]

--race --pack <dir> (NVIDIA) and --race-test --seed <s> --day <d> (Metal) run the race alone, three rounds, and print a table; that is what the bench log and the PC job use.

4. Fleet learning

4.1 The record

The app's engine reads the race line (engine.rs race_line) and writes one TUNING {json} line to the app log, which the existing intake receives with every upload (Neon miner_logs, the same table tools/logs.mjs reads). Fields:

Field From
ts, machine (id8), app (version) the engine
card (the worker's device name, spaces as underscores: the key everything else uses), vendor, worker (CUDA, Metal, OpenCL) the race line, the card state
driver, arch (sm_120, metal) the race line
epoch (16 hex), loads, wide (the program class features the generator fixes: loads per hash and wide loads per hash) the race line
variants {name: MH/s or null} the race line
winner, mhs, base_mhs, gain_pct, total_ms, pinned, tuned, notes the race line
power_limit_w, power_w, power_pct, mh_per_w (= mhs / power_w, 0 when the card reports no draw) the card state (nvidia-smi telemetry; Apple reports none)

The card state also carries variant, race_mhs, race_gain_pct, race_variants for the dashboard, and one event per race ("RTX 5090 kernel race: u2-ldg at 118.3 MH/s (+2.1% over base, 17 variants, 41 s)").

4.2 The aggregation

tools/tuning.mjs on the Mac (or a small job): every app-log upload of the window (default 7 days) with a TUNING line, de-duplicated on (machine, card, epoch) because the log is re-sent every minute, then per card model and variant the sample count, the median MH/s and the median MH per watt. The winner is the best median with at least --min-samples (3) samples; --by mhw ranks by MH per watt instead. --write tuning.json writes:

{"updated": "2026-10-04T21:00:00Z", "window_days": 7,
 "cards": {"NVIDIA_GeForce_RTX_5090": {"variant": "u2-ldg", "race": true, "candidates": ["u2-ldg", "ldg", "base"],
                                       "samples": 41, "races": 41, "machines": 2, "mhs": 118.3, "base_mhs": 115.9,
                                       "gain_pct": 2.07, "mh_per_w": 0.254, "worker": "CUDA", "by": "mhs"}}}

A program class split (by loads, wide) is in the record and not yet in the aggregation: the generator fixes 16 loads per instruction block, so every program has 128 loads per hash today; the split starts to matter when the era draw changes the mix.

4.3 The way back: the manifest

packaging/ota/publish-manifest.sh --tuning tuning.json puts the object under tuning in the signed igneum-app-latest.json (carried over from the current manifest when not given; --no-tuning drops it; the same script now also takes --override '{json}' for consensus.override and carries that over too). manifest.rs parses tuning (an object with a cards object, else the manifest is refused). The updater writes it as is to <app data>/app/tuning.json (ota.rs write_tuning, next to override.json), logs one event, and the engine starts every miner with IGNEUM_TUNING_FILE=<that path> (procs::spawn gained an environment parameter); the miner's child, the worker, reads it at every prepare. No restart for a change: a worker that has the path reads the file again at its next prepare; a miner started before the file existed is restarted by the usual hourly path.

What a worker does with its entry (readTuning, readMetalTuning):

Entry Behaviour
none for this card model the full race
race: true with candidates the candidates are raced first, then the rest as the budget allows (the fleet keeps learning; the card starts from the known best)
race: false with variant the variant is compiled and self-tested, no timing (about 1 s); a failed self-test falls back to the full race
--variant <name> on the worker the same as a pinned entry, for tests

4.4 What is implemented and what is not

Piece State
NVRTC worker race, --race mode, tuning file code, syntax-checked for mingw and the emulation build; emulation suite; not yet run on a GPU (the PC job does that)
Metal worker race, --race-test, deferred race, tuning file code and the Mac measurement (bench log)
OpenCL worker race not implemented (open)
Engine: race line to TUNING record, card state, event, IGNEUM_TUNING_FILE code, unit tests of the manifest parse
publish-manifest.sh --tuning, --override, carry-over code (dry run against a scratch folder in the plan)
tools/tuning.mjs code; needs records, so no table yet
Aggregation as a job on a PC or the observer not needed yet: the Mac script reads the intake
Dashboard field for the variant state only; the UI does not show it yet

5. Open

  • OpenCL: the same catalogue through clBuildProgram options and the pragma; --group-warps exists already.
  • The program class split in the aggregation once eras change the instruction mix.
  • A per-card cap on how long a race may pause mining (today the 2 s windows plus the budget); and whether to race only every N hours once a card's winner is stable (the pinned entry does that by hand).
  • Power: the record has MH per watt, the race does not touch the power cap; a race across caps is a later lever.