igneum/docs/build/tuning.md

8.7 KiB

Tuning: how Ember reaches a good operating point, and what the kit builds with

8 October 2026. The register row (plan p. 14): "Published optimisation work, compiler settings and safe tuning logic: Ember reaches good operating points without third-party software." Served under igneum.network/miner. Every figure here is a measured row of the day it names, or marked approximate.

1. The knee rule

A card's hash rate on the Igneum hash is bound by dependent memory reads, not by the core clock. The core clock can fall a long way before the rate moves, and the card's power falls with it. The knee is the lowest core clock that holds the rate. Ember locks the card at the knee. Nothing else is touched: no voltage, no fan curve, no memory overclock.

Card Knee (SM MHz) Rate at the knee Watts at the knee Rate at stock Watts at stock Energy per hash, knee against stock Measured
RTX 5090 1,300 132.7 MH/s 314.8 W 140.0 MH/s 456.1 W 2.37 against 3.26 µJ, minus 27 percent 8 Oct 2026, class v5 genesis pack, PC 1, nvidia-smi 1 Hz
RTX 5080 1,100 60.3 MH/s 123 W approximate 61 MH/s approximate 200 W approximate minus 38 percent 8 Oct 2026, Ember climb, PC 1 (the stock row approximate)
RX 9070 XT no core lock lever in 0.3.20; the AMD knob is a clock offset and a power limit 19.0 MH/s at minus 500 MHz and minus 30 percent 149.3 W 18.2 MH/s approximate 300 W 0.127 MH per W at the knob against approximate 0.06 8 Oct 2026, the AMD grid, PC 1
RX 7600 the same knob; grid owed (the card must be mining for the grid) 13.9 MH/s 113 W the same the same 0.123 MH per W at stock 8 Oct 2026, card-in, PC 1

The rule, as the app applies it: on NVIDIA the core clock is locked at the knee with the driver's own lock (nvidia-smi -lgc 0,<knee>), through the Igneum Power Helper task so the app never runs elevated; on AMD the knob is the clock offset and the power limit through the vendor's own interface; on Apple there is no lever and the card runs at stock. A lock is released (-rgc) when the app stops mining, when a job ends, and on every start.

2. The ladders and the priors

Ember does not start from zero. The search has a prior per architecture and walks from it; the measured knee decides.

Architecture Prior knee Where it came from
Blackwell, RTX 5090 1,300 MHz the efficiency passes of 7 and 8 October 2026
Blackwell, RTX 5080 1,000 MHz the same
Ada (RTX 40) 2,400 MHz nine capped-against-uncapped pairs on rented cards, class v4 (a power cap holds the full rate until the SM clock falls under about 2,400)
Ampere (RTX 30, A-series) 1,800 MHz the same pairs (a 3080 Ti capped to 63 percent with the SM at 749 MHz loses 24 to 28 percent)
anything else none the full ladder

The ladders: the power ladder's rungs are 100, 90, 80, 70, 60 and 50 percent of the card's limit; the clock ladder's rungs are 100, 90, 80, 70, 60, 50 and 45 percent of the card's maximum. The power ladder starts at the rung the prior clock implies (the draw follows the clock cubed on the voltage and frequency curve, rounded up to the next rung). Ember 2 adds a hill-climb from the start point: each probe moves the memory clock up by one step or the core clock down by one step; a probe that produces a refused row (an invalid hash, a fault, a rate under the floor) backs that knob off and is never retried in the run.

Every row of a search carries the clock, the power percent, the memory clock, the measured watts, the rate and the rate per watt. Three tiers come out of the rows:

Tier Rule
efficiency the usable row with the most MH per W
balanced the row within 1 percent of the best rate with the fewest watts
max the usable row with the highest rate

A card with no lever has one tier, its stock row, and the note says why. A tier is remeasured when the class changes (the rows are keyed to the class the search ran on) and on the app's own schedule.

3. Safe defaults the app applies per class

Class What the app does on a fresh card Why
v3 (the devnet's class until the class v6 cut) the prior knee as the clock cap, the power ladder from its implied rung the knee rule above
v4 and v5 (the shadow block, the state leaves) the same priors; the v5 knee on the 5090 reads the v4 rate at +2 percent watts (8 Oct 2026) the shadow block runs in the memory shadow; the state leaves cost nothing on the card
v6 (the index fold, the re-weight table, the 64-register window) the same priors; the window halves the hash rate at about 7 percent less card power (a window hash is twice the work) and the tiers are remeasured on the class flip the register budget changes (section 5)

The app never writes a setting it cannot restore. The lock is released on stop and on start. A tuning failure leaves the card at stock. A tuning file (--tuning <file>, or IGNEUM_TUNING_FILE) carries the chosen worker variant per card so the race (section 4) does not rerun on every start.

4. The worker's flags and what each costs

The CUDA worker (igneum-worker-cuda) compiles each pack's kernel at run time with NVRTC. The OpenCL worker (igneum-worker-opencl) builds it with the vendor's OpenCL compiler. The Metal worker builds it with Apple's. The text of the kernel is the pack's; the worker adds nothing to the hash.

Flag Meaning Cost or effect
--device D the card none
--batch-log2 B nonces per dispatch, 2^B (default 22) more nonces per dispatch amortise the launch; 24 is the bench setting
--block-warps W warps per thread block (default 1) the 5090 and 4090 read within 3 percent across 1, 2, 4 and 8 on every class measured today (8 Oct 2026)
--arch sm_XY, compute_XY, auto the NVRTC target (default the device's own) a PTX target is JIT-compiled by the driver once per pack
--race on, off, a,b,c race the variants of section 5 on first use, or none, or a list one race per pack, about 2 seconds per variant, kept in the tuning file
--variant <name> one variant, no race none
--check --pack <dir> compile, build the cache and dataset, run the self-test, print timings, exit 0 or 1 the gate every pack passes before it serves a job
--bench --pack <dir> --batches N time N dispatches, print the fingerprint of the 2^B outputs at base nonce 0 the measurement every row on this site comes from
--memprobe the card's dependent and independent read latencies at 4, 64 and 1024 MiB a diagnostic, no hash

The self-test, run before any job: the cache head, last line and FNV-1a 64; the dataset head, last word and 64 samples; the three vector warps of the pack through the bound kernel. A pack that fails is refused.

5. Compiler settings the kit builds with

Backend Compiler and options Register budget Notes
CUDA (NVRTC in the worker) the device's sm_XY, --std=c++17, -default-device; a variant may add --maxrregcount=N or __launch_bounds__ class v5: 48 registers, 24 blocks per SM on the 5090; the window class: 88 to 96 registers on the 5090 (20 blocks per SM), 87 to 104 on the 4090 (16 to 20), no spill (8 Oct 2026, ptxas) no fast-math, no unsafe flag: the hash is integer only
CUDA (the offline bench, proto-cuda/build.sh) nvcc -O3 -std=c++17 -arch=sm_120 (or native) the same needs CUDA 12.8 or newer
OpenCL clBuildProgram with -cl-std=CL1.2, CL2.0 or CL3.0 by the device's version and exchange mode; --build-opts appends the vendor's AMD reports the card by its gfx name (the RX 7600 is gfx1102, the RX 9070 XT gfx1201)
Metal Apple's compiler, the pack's .metal texts, no options the compiler's the M5 Max and the Mac mini M6 run at stock

The variants the worker races, by name: base (the pack's text, one warp per block), w2, w4, w8 (warps per block), u2, u8 (loop unroll), ldg, ldcg, ldcs (the load path), r32, r64 (a register cap), lb4-w4, lb8-w2 (launch bounds), and the pairs u2-ldg, u2-w4, ldg-w4, ldcg-w4. On the 5090 and 4090 the base variant wins or ties on every class measured on 8 October 2026; the race exists for cards the team does not own.

6. What a miner can check

Every row above is reproducible with the kit's worker and the public packs: the fingerprint printed by --bench is the same on every backend and every card (the class v6 all-together pack reads 59e6708e46f1e87c on CPU, CUDA, Metal, Apple OpenCL, an RTX 5090 and an RX 7600; 8 October 2026). A different fingerprint is a bug, and the card's rate is not a figure of merit until it matches.

Nothing here needs third-party software: the lock is the driver's own, the knob is the vendor's own, the worker is the kit's.