docs/build/tuning.md: the register row on tuning (the knee rule with the 5090, 5080, 9070 XT and 7600 rows, the priors and the ladders, the tiers, the per-class defaults, the worker's flags and their costs, the compiler settings and register budgets, the variant catalogue, the fingerprint check)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
76b4116d0c
commit
7125a10e17
1 changed files with 85 additions and 0 deletions
85
docs/build/tuning.md
vendored
Normal file
85
docs/build/tuning.md
vendored
Normal file
|
|
@ -0,0 +1,85 @@
|
|||
# Tuning: how Ember reaches a good operating point, and what the kit builds with
|
||||
|
||||
8 October 2026. The register row (plan p. 14): "Published optimisation work, compiler settings and safe tuning logic: Ember reaches good operating points without third-party software." Served under igneum.network/miner. Every figure here is a measured row of the day it names, or marked approximate.
|
||||
|
||||
## 1. The knee rule
|
||||
|
||||
A card's hash rate on the Igneum hash is bound by dependent memory reads, not by the core clock. The core clock can fall a long way before the rate moves, and the card's power falls with it. The knee is the lowest core clock that holds the rate. Ember locks the card at the knee. Nothing else is touched: no voltage, no fan curve, no memory overclock.
|
||||
|
||||
| Card | Knee (SM MHz) | Rate at the knee | Watts at the knee | Rate at stock | Watts at stock | Energy per hash, knee against stock | Measured |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| RTX 5090 | 1,300 | 132.7 MH/s | 314.8 W | 140.0 MH/s | 456.1 W | 2.37 against 3.26 µJ, minus 27 percent | 8 Oct 2026, class v5 genesis pack, PC 1, nvidia-smi 1 Hz |
|
||||
| RTX 5080 | 1,100 | 60.3 MH/s | 123 W | approximate 61 MH/s | approximate 200 W | approximate minus 38 percent | 8 Oct 2026, Ember climb, PC 1 (the stock row approximate) |
|
||||
| RX 9070 XT | no core lock lever in 0.3.20; the AMD knob is a clock offset and a power limit | 19.0 MH/s at minus 500 MHz and minus 30 percent | 149.3 W | 18.2 MH/s | approximate 300 W | 0.127 MH per W at the knob against approximate 0.06 | 8 Oct 2026, the AMD grid, PC 1 |
|
||||
| RX 7600 | the same knob; grid owed (the card must be mining for the grid) | 13.9 MH/s | 113 W | the same | the same | 0.123 MH per W at stock | 8 Oct 2026, card-in, PC 1 |
|
||||
|
||||
The rule, as the app applies it: on NVIDIA the core clock is locked at the knee with the driver's own lock (`nvidia-smi -lgc 0,<knee>`), through the Igneum Power Helper task so the app never runs elevated; on AMD the knob is the clock offset and the power limit through the vendor's own interface; on Apple there is no lever and the card runs at stock. A lock is released (`-rgc`) when the app stops mining, when a job ends, and on every start.
|
||||
|
||||
## 2. The ladders and the priors
|
||||
|
||||
Ember does not start from zero. The search has a prior per architecture and walks from it; the measured knee decides.
|
||||
|
||||
| Architecture | Prior knee | Where it came from |
|
||||
|---|---|---|
|
||||
| Blackwell, RTX 5090 | 1,300 MHz | the efficiency passes of 7 and 8 October 2026 |
|
||||
| Blackwell, RTX 5080 | 1,000 MHz | the same |
|
||||
| Ada (RTX 40) | 2,400 MHz | nine capped-against-uncapped pairs on rented cards, class v4 (a power cap holds the full rate until the SM clock falls under about 2,400) |
|
||||
| Ampere (RTX 30, A-series) | 1,800 MHz | the same pairs (a 3080 Ti capped to 63 percent with the SM at 749 MHz loses 24 to 28 percent) |
|
||||
| anything else | none | the full ladder |
|
||||
|
||||
The ladders: the power ladder's rungs are 100, 90, 80, 70, 60 and 50 percent of the card's limit; the clock ladder's rungs are 100, 90, 80, 70, 60, 50 and 45 percent of the card's maximum. The power ladder starts at the rung the prior clock implies (the draw follows the clock cubed on the voltage and frequency curve, rounded up to the next rung). Ember 2 adds a hill-climb from the start point: each probe moves the memory clock up by one step or the core clock down by one step; a probe that produces a refused row (an invalid hash, a fault, a rate under the floor) backs that knob off and is never retried in the run.
|
||||
|
||||
Every row of a search carries the clock, the power percent, the memory clock, the measured watts, the rate and the rate per watt. Three tiers come out of the rows:
|
||||
|
||||
| Tier | Rule |
|
||||
|---|---|
|
||||
| efficiency | the usable row with the most MH per W |
|
||||
| balanced | the row within 1 percent of the best rate with the fewest watts |
|
||||
| max | the usable row with the highest rate |
|
||||
|
||||
A card with no lever has one tier, its stock row, and the note says why. A tier is remeasured when the class changes (the rows are keyed to the class the search ran on) and on the app's own schedule.
|
||||
|
||||
## 3. Safe defaults the app applies per class
|
||||
|
||||
| Class | What the app does on a fresh card | Why |
|
||||
|---|---|---|
|
||||
| v3 (the devnet's class until the class v6 cut) | the prior knee as the clock cap, the power ladder from its implied rung | the knee rule above |
|
||||
| v4 and v5 (the shadow block, the state leaves) | the same priors; the v5 knee on the 5090 reads the v4 rate at +2 percent watts (8 Oct 2026) | the shadow block runs in the memory shadow; the state leaves cost nothing on the card |
|
||||
| v6 (the index fold, the re-weight table, the 64-register window) | the same priors; the window halves the hash rate at about 7 percent less card power (a window hash is twice the work) and the tiers are remeasured on the class flip | the register budget changes (section 5) |
|
||||
|
||||
The app never writes a setting it cannot restore. The lock is released on stop and on start. A tuning failure leaves the card at stock. A tuning file (`--tuning <file>`, or `IGNEUM_TUNING_FILE`) carries the chosen worker variant per card so the race (section 4) does not rerun on every start.
|
||||
|
||||
## 4. The worker's flags and what each costs
|
||||
|
||||
The CUDA worker (`igneum-worker-cuda`) compiles each pack's kernel at run time with NVRTC. The OpenCL worker (`igneum-worker-opencl`) builds it with the vendor's OpenCL compiler. The Metal worker builds it with Apple's. The text of the kernel is the pack's; the worker adds nothing to the hash.
|
||||
|
||||
| Flag | Meaning | Cost or effect |
|
||||
|---|---|---|
|
||||
| `--device D` | the card | none |
|
||||
| `--batch-log2 B` | nonces per dispatch, 2^B (default 22) | more nonces per dispatch amortise the launch; 24 is the bench setting |
|
||||
| `--block-warps W` | warps per thread block (default 1) | the 5090 and 4090 read within 3 percent across 1, 2, 4 and 8 on every class measured today (8 Oct 2026) |
|
||||
| `--arch sm_XY, compute_XY, auto` | the NVRTC target (default the device's own) | a PTX target is JIT-compiled by the driver once per pack |
|
||||
| `--race on, off, a,b,c` | race the variants of section 5 on first use, or none, or a list | one race per pack, about 2 seconds per variant, kept in the tuning file |
|
||||
| `--variant <name>` | one variant, no race | none |
|
||||
| `--check --pack <dir>` | compile, build the cache and dataset, run the self-test, print timings, exit 0 or 1 | the gate every pack passes before it serves a job |
|
||||
| `--bench --pack <dir> --batches N` | time N dispatches, print the fingerprint of the 2^B outputs at base nonce 0 | the measurement every row on this site comes from |
|
||||
| `--memprobe` | the card's dependent and independent read latencies at 4, 64 and 1024 MiB | a diagnostic, no hash |
|
||||
|
||||
The self-test, run before any job: the cache head, last line and FNV-1a 64; the dataset head, last word and 64 samples; the three vector warps of the pack through the bound kernel. A pack that fails is refused.
|
||||
|
||||
## 5. Compiler settings the kit builds with
|
||||
|
||||
| Backend | Compiler and options | Register budget | Notes |
|
||||
|---|---|---|---|
|
||||
| CUDA (NVRTC in the worker) | the device's `sm_XY`, `--std=c++17`, `-default-device`; a variant may add `--maxrregcount=N` or `__launch_bounds__` | class v5: 48 registers, 24 blocks per SM on the 5090; the window class: 88 to 96 registers on the 5090 (20 blocks per SM), 87 to 104 on the 4090 (16 to 20), no spill (8 Oct 2026, ptxas) | no fast-math, no unsafe flag: the hash is integer only |
|
||||
| CUDA (the offline bench, `proto-cuda/build.sh`) | `nvcc -O3 -std=c++17 -arch=sm_120` (or `native`) | the same | needs CUDA 12.8 or newer |
|
||||
| OpenCL | `clBuildProgram` with `-cl-std=CL1.2`, `CL2.0` or `CL3.0` by the device's version and exchange mode; `--build-opts` appends | the vendor's | AMD reports the card by its gfx name (the RX 7600 is gfx1102, the RX 9070 XT gfx1201) |
|
||||
| Metal | Apple's compiler, the pack's `.metal` texts, no options | the compiler's | the M5 Max and the Mac mini M6 run at stock |
|
||||
|
||||
The variants the worker races, by name: `base` (the pack's text, one warp per block), `w2`, `w4`, `w8` (warps per block), `u2`, `u8` (loop unroll), `ldg`, `ldcg`, `ldcs` (the load path), `r32`, `r64` (a register cap), `lb4-w4`, `lb8-w2` (launch bounds), and the pairs `u2-ldg`, `u2-w4`, `ldg-w4`, `ldcg-w4`. On the 5090 and 4090 the base variant wins or ties on every class measured on 8 October 2026; the race exists for cards the team does not own.
|
||||
|
||||
## 6. What a miner can check
|
||||
|
||||
Every row above is reproducible with the kit's worker and the public packs: the fingerprint printed by `--bench` is the same on every backend and every card (the class v6 all-together pack reads `59e6708e46f1e87c` on CPU, CUDA, Metal, Apple OpenCL, an RTX 5090 and an RX 7600; 8 October 2026). A different fingerprint is a bug, and the card's rate is not a figure of merit until it matches.
|
||||
|
||||
Nothing here needs third-party software: the lock is the driver's own, the knob is the vendor's own, the worker is the kit's.
|
||||
Loading…
Reference in a new issue