igneum/docs/plans/ember-tune.md
igneum-labs ebedff9b66 ember-tune.md: the next-cut note names the right commit
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 08:03:19 +00:00

18 KiB

Ember Tune: every card tuned for MH per watt, out of the box

5 October 2026, night. the project lead: "make sure we have ember tuning every single card for efficiency out of the box, the more data = the better the tune, make an awesome system." Branch ember-tune, worktree ../igneum-wt-ember-tune, on top of the AMD telemetry commit (7adcd4c, branch opencl-rdna4-telemetry) and the Power control commit (3562f26, branch job-console), both cherry-picked. Lever 3 of docs/plans/miner-eff.md grows two knobs and a fleet memory; lever 2 (docs/design/miner-tuning.md) carries the priors in the same signed tuning section.

1. What a user sees

Moment The card row says What happened
First 2 minutes of mining tuning: waits for 120 s of steady mining The worker warms up; nothing is touched.
Tuning tuning: holding 2472 MHz · 100% · 41 s (step 7 of 9) with a Stop button One card at a time, on the live kernel, never restarting the worker.
Tuned Tuned: 122.3 MH/s at 290 W (0.422 MH/W), then 2470 MHz at 100%, full tune, 1 h ago The point is pinned on the card; the result went to the fleet.
Known model the same line, from the fleet prior, confirmed, 2 min ago The card started at its model's prior and confirmed it in two steps instead of nine.
Apple silicon Tuned: 26.7 MH/s at 38 W (0.703 MH/W) (measured as it runs), and measure only on Apple silicon: the system sets the clocks and the power; no control exposed Nothing can be set; the number is still reported so the row and the fleet know what the card does.
NVIDIA, Power control off the measured line and measure only until Power control is on in Settings (Windows asks for administrator rights once) The app never raises the prompt by itself (5 October 2026). One switch, one prompt, and the full tune runs.
Slider moved your setting stays pinned A manual point is never overridden; the tune still measures and reports.
Stopped tuning stopped: a remote job took the GPU and the card back where it was Any fault reverts the step and the run.
Fleet pause Settings: tuning paused fleet-wide by the signed manifest The kill switch.

Settings: one switch, "Ember Tune: tune every card for MH per watt out of the box (once after install, then weekly, and after a driver or program change)". AMD needs no rights. NVIDIA needs the Power control switch (one administrator prompt) for both knobs; off, it measures only.

2. The knobs, per vendor

Vendor Power limit Core clock cap Memory clock How Rights
NVIDIA nvidia-smi -pl <W>, percent of the default, inside power.min_limit and power.max_limit nvidia-smi -lgc 0,<MHz>, percent of clocks.max.gr; -rgc = unlocked never touched (-lmc is not used); read back as clocks.mem directly when the engine is elevated, else the one-prompt helper (<seq> pl <W>, <seq> lgc <MHz>, <seq> rgc in sweep/cmd.txt) administrator, so only with Power control on
AMD igneum-gpu-telemetry --card N --set-plimit <offset> (0 = default, -20 = 80%), inside the tune line's plimit_range (PC 1's 9070 XT: -30 to 10, so 70% is the floor) --set-gmax only when the tune line's gmax_range is absolute MHz (floor 0 or above); on RDNA 4 the range is an offset from stock (-500 to 1000 on PC 1) and the clock knob stays closed until the stock clock is known; --reset for the default point not settable through ADLX on RDNA 4; read back as mclk_mhz, and a step whose mean memory clock falls under 95% of the baseline's is marked and cannot win the helper, one process per request, exit 0 and a tune ... ok line none on Windows (ADLX manual tuning); root on Linux, so measure only there
Apple none none none measure only none

Vendor limits are never exceeded and the floor is never undercut: the plan clamps every point (Limits::clamp_clock, Limits::watts_for), and a clock floor the vendor does not report is 60% of the maximum.

3. The plan and the choice

Full plan (a new model, or a prior that lost its confirm check): the power ladder 100, 90, 80, 70, 60, 50% at the unlocked clock (duplicate watts dropped where the card's floor clamps them), then the clock ladder 90, 80, 70, 60% of the maximum at the power point the power ladder chose. 60 s hold after 15 s settle per step; 9 steps on an RTX 5090 (five power, four clock), about 12 minutes.

Confirm plan (the model's prior has 5 or more reports): the prior's point, then one neighbour (the next clock step up when the prior caps the clock, else one power step down). If the neighbour beats the prior by over 1% MH/W, the full plan is queued; else the prior stands. Two steps, about 3 minutes.

Baseline plan (measure only): one step at the card's current point. The "before" number for the row and the fleet.

The choice (ember::choose): among the usable steps whose rate is within the tolerance (1%, settable from the manifest) of the fastest step, the best MH per watt; within 1% on efficiency the higher rate; within 1% on both the lower draw. A card never gives up more than the tolerance in blocks for the saving. A step is unusable when it is marked: faulted (a rejected or mismatched hash during the hold: the step is reverted and marked), hot (the GPU reached 85 C; the run aborts at 90), memory_clock_dropped, unapplied (the readback disagreed with the request), no_readings (under three draw samples or no STATUS line).

4. The data flow

card mines 120 s ──> probe (limits, driver, how to set) ──> plan ──> steps ──> choice ──> point pinned
                                                                              │
          app log: TUNE start / TUNE card=.. step=.. / TUNE chosen / TUNE {json}   (and stdout under --sweep)
                                                                              │
          log upload (every minute, the existing intake, site/api/log.mjs) ──> Neon miner_logs
                                                                              │
          relay/lib/ember.mjs aggregate: per (card model | driver major | program class)
            median clock cap (10 MHz), median power %, median MH/W, MH/s, W, spread (MAD %), samples, machines
                      │                                     │
          console: /r/<token>/c/tuning, `node tools/console.mjs tuning`     site: tools/tuning.mjs --priors --site
                      │                                                     -> site/miner-priors.json -> /miners#priors
          tools/tuning.mjs --priors --write tuning.json  (priors + ember settings beside the kernel-variant cards)
                      │
          packaging/ota/publish-manifest.sh --tuning tuning.json --deploy   (signed; carried over when not given)
                      │
          every app: <app data>/tuning.json  ──> ember::settings_of (kill switch, min samples, tolerance, period)
                                              ──> ember::prior_of(key)   ──> a new card's confirm plan

The record (ember::record_json): ts, machine (the first 8 hex of SHA-256 over the install id; the id itself is random per install and never sent), app, os, card, vendor, driver, driver_major, class, key, plan, steps (the full table: clock, power %, limit, watts, MH/s, MH/W, core and memory clock, hottest reading, faults, mark), chosen, before (the full plan's 100% step), eff, mhs, watts. The key: <card model with underscores>|<driver major>|<program class>, the class from the worker's race line (l128w16 today; v2 before a race has run).

5. Scheduling and safety

Rule Where
One card at a time; the card must have mined 120 s and have a STATUS line tick_sweep
Never under a remote job hold, a pause, inside 600 s of the hour boundary, or while the app quits tick_sweep, sweep_drive
Due once after install, every 7 days (manifest ember.period_s), and when the driver major or the program class changed since the last tune tick_sweep (CardPref.sweep_driver, sweep_class)
A pinned card (the slider) is measured, never changed sweep_finish
Kill switch: tuning.ember.enabled = false in the signed manifest stops every tune fleet-wide; the Settings line says so ember::settings_of, tick_sweep
Faults: a rejected or mismatched hash marks the step; the card leaving mining, a worker error, a job, a pause or 90 C aborts the run and restores the point from before Run::sample_fault, sweep_drive, sweep_abort
Memory clock held: never set; a step that drags it under 95% of the baseline's cannot win Row::from_samples
Vendor limits: every point clamped to the reported range; the clock floor 60% when none is reported Limits
A signed prior is only ever a starting point inside the card's OWN reported limits (power.min_limit to power.max_limit, the clock floor to clocks.max.gr or the ADLX gmax_range), never a memory clock, never a value the card did not report; the confirm step measures it and the full plan replaces it when a neighbour beats it, so a bad prior costs the fleet one confirm step per card, not a setting. The signing key (K1, docs/security/keys.md) therefore cannot push a card past its vendor ceiling or under its floor Plan::confirm clamps through Limits::clamp_clock and power_pct.clamp(50, 100); proven by ember::tests::the_confirm_plan_checks_the_prior_and_its_neighbour (a prior of 9,000 MHz at 30% becomes 3,090 MHz at 50%) and limits_never_exceed_the_vendor_or_undercut_the_floor
No prompt the user did not ask for: the NVIDIA helper starts only with Power control on; the --sweep job never counts as permission sweep_probe_known, sweep_helper_start
The elevated helper restores the limit and resets the clocks by itself after 20 idle minutes sweep::helper_script_*
A playbook that starts a second engine beside the installed app (the PC measurement jobs) gives it NO pipe (its output goes to a file the script tails: a pipe's write end is inherited by the engine's miners, and the installed app's jobs runner then waits forever for EOF after an abort; C35, PC 1 22:31 UTC, a 24-minute hang and orphaned miners), ends the engine's whole process tree at the end and on the budget (taskkill /T /F), and lets the installed app's miners come back only after that relay/playbooks/ember-tune-pc1.ps1, sweep-5090.ps1; CI tools/ci/second-engine-check.sh fails any playbook without both
Every quit: line in the app log names its source (the window host's stdin, the host gone, POST /api/quit, the --sweep run's end) Cmd::Quit(&'static str) (b671c8b)
A second engine never runs the updater: IGNEUM_APP_NO_OTA=1 (implied by --sweep) skips the OTA tick and refuses Check now, whatever the manifest's min_supported_version says (the installer it would launch quits the installed app: PC 1, 22:31 UTC) Engine.no_ota; the playbooks set the variable; tools/ci/second-engine-check.sh demands it

6. Tests

Test What it fixes
ember::tests::the_full_plan_is_the_power_ladder_then_the_clock_ladder_at_the_chosen_power 5 + 4 steps on the 5090's limits, the clamps, the dynamic second half, the 1% and 5% choices
limits_never_exceed_the_vendor_or_undercut_the_floor clamps
the_choice_keeps_the_best_mh_per_watt_within_the_rate_tolerance the rule, the ties, marked rows never win
the_guards_mark_a_step_so_it_cannot_win faulted, hot, memory clock, unapplied, no readings, the line
a_fault_during_a_step_reverts_it_and_the_run_goes_on the state machine with a fake clock: the faulted 70% step is marked and never chosen
the_confirm_plan_checks_the_prior_and_its_neighbour the two steps, Keep against FullDue, a prior outside the range clamped
a_baseline_plan_measures_the_card_as_it_runs no control, still a number and the Tuned line
the_record_and_the_prior_round_trip_through_the_manifest_shape record fields (no address, no host), priors and ember beside cards, the sample floor, the kill switch
control_reasons_per_vendor who measures only and why
sweep::tests::helper_scripts_carry_the_protocol the helper's pl, lgc, rgc
relay/test/ember.test.mjs five samples converge (2,470 MHz at 100%), an outlier (0.908 MH/W at 1,854 MHz) moves nothing, baseline records make no prior, de-duplication, the manifest merge keeps lever 2's cards, the canonical round trip, AMD keys
app/igneum-app/ui/tune-line.test.mjs the row line per state

Run: cargo test -p igneum-app ember sweep (on a PC through the build job, or on the Mac under the build lock), node --test relay/test/ember.test.mjs app/igneum-app/ui/tune-line.test.mjs.

7. The tier consequences

Tier What Ember Tune does for it What it costs
A laptop GPU (NVIDIA, 60 to 115 W) the power ladder usually finds the vendor floor binding; the clock ladder is where a memory-bound program saves watts; the thermal mark keeps a hot chassis from winning a step it cannot hold about 12 minutes once, then 3 minutes a week; under 1% of the hour during the tune (the worker never stops)
One 8 GB card the same two knobs; the 8 GB card is identities-limited (2 by default), the tune does not change that the same
One 12 or 16 GB card the same the same
One 24 or 32 GB card (the 5090) the draw sits far under the cap (290 W under 460 W on PC 1), so the power ladder is flat and the clock ladder is the lever; expected saving from the 4 October stability line: tens of watts at under 1% rate, to be measured the same
A rig (several cards) one card at a time, so a six-card rig takes about 70 minutes to tune once; every card of one model after the first starts at the prior (3 minutes); the tune never touches a card a remote job holds linear in cards once, then the confirm plan
A pool user the same per card; a pool submits the same hashes, so the 1% rate tolerance is the same 1% of shares the same
AMD on Linux measure only (sysfs needs root); the row says so 60 s a week
Apple silicon measure only; the row says so 60 s a week

Privacy line: what is uploaded is the record in section 4 and nothing else: a hash of the random install id, the card model, the driver version, the OS, the program class, the step table and the chosen point. No address, no hostname, no raw machine id, no user name. The public priors table carries only the aggregate per model.

8. Measurements

PC 1, 5 October 2026 (night)

Tonight's constraints, read from PC 1's own uploads: the installed app runs as DESKTOP-KMCV30N\Admin with elevated=False (the account line at 19:02:33 UTC), the two in-app sweep attempts at 20:09 UTC aborted on the cancelled administrator prompt (SWEEP aborted ... the_elevated_helper_did_not_run_(the_administrator_prompt_was_cancelled)), so no stored sweep result exists from today, and the RX 9070 XT left the PCI bus at about 20:40 UTC (eGPU link, not restarted tonight). NVIDIA's -pl and -lgc need administrator rights, the project lead is asleep, and the app never raises the prompt by itself, so tonight's run on PC 1 is the baseline plan on the 5090 through the whole pipeline (probe, measure, TUNE record, upload, aggregation, prior shape in a test manifest). The two-knob tune on the 5090 and the 9070 XT run are owed: the 5090 the moment Power control is switched on (one prompt, then the tune runs by itself within 2 minutes of steady mining), the 9070 XT when the card is back on the bus.

Run 1 (ember-tune-pc1-1, 22:30 UTC): aborted 46 s in by the installed app quitting, named the next morning: the second engine's own updater (0.3.9 under min_supported_version = urgent) ran the per-user installer, whose PrepareToInstall quit the installed app through its api/quit (C35 in the bench log); before any step; nothing set; the "before" snapshots are in the bench log (5090: 450 W of 575, 2,850 MHz core, 3,090 MHz maximum, 14,001 MHz memory; 9070 XT present on bus 98 with OFFSET ranges gmax_range -500 1000, plimit_range -30 10). The offset finding changed the AMD mapping (054e041): an offset clock range closes the clock knob and the power ladder runs on a percent scale bounded by plimit_range. The re-run follows the 0.3.11 rollout.

8a. Next-cut notes (for the 0.3.12 shipper)

Commit What Where
b671c8b every quit: names its source; Power control alone decides; no cap at start under --sweep main.rs, server.rs, engine.rs (separable)
e600e63 a second engine never runs the updater (IGNEUM_APP_NO_OTA, implied by --sweep) engine.rs (6 lines, separable)
1e9550e (this commit, amended) the elevated job path's output file is followed while the script runs, so the 5-minute progress reports carry its lines (a 35-minute run that never mined showed only "script running" on 6 October 2026); the tune playbook's watchdog fails a run that mines nothing within 120 s of its first status line, with the engine's last log line in the RESULT jobrun.rs follow_file, relay/playbooks/ember-tune-pc1.ps1

9. Open

  • The NVIDIA clock readback: nvidia-smi -lgc is confirmed only through the core clock during the hold (a mean over the cap by 5% marks the step unapplied); the first run with Power control on tells whether the driver honours the lock on the 5090 under this kernel.
  • ADLX on RDNA 4 exposes no memory-clock setter; the memory-clock mark is the guard. The telemetry agent's 9070 XT sweep tells whether a core cap drags the memory clock on that card.
  • The confirm plan's neighbour is one step; a second neighbour (the other knob) would cost 75 s more and catch a prior that is wrong on both knobs.
  • Intel: no knob yet; the row says measure only.