From 9c6b62ab5eb7954cfbeb71234189cfc9d7ffc70b Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 13:17:48 +0000 Subject: [PATCH] Ember Tune plan: the fleet priors from rented cards (baseline rows; the containers refuse the knobs) --- docs/plans/ember-tune.md | 221 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 221 insertions(+) create mode 100644 docs/plans/ember-tune.md diff --git a/docs/plans/ember-tune.md b/docs/plans/ember-tune.md new file mode 100644 index 000000000..05ffea371 --- /dev/null +++ b/docs/plans/ember-tune.md @@ -0,0 +1,221 @@ +# Ember Tune: every card tuned for MH per watt, out of the box + +5 October 2026, night. the project lead: "make sure we have ember tuning every single card for efficiency out of the box, the +more data = the better the tune, make an awesome system." Branch `ember-tune`, worktree `../igneum-wt-ember-tune`, +on top of the AMD telemetry commit (7adcd4c, branch `opencl-rdna4-telemetry`) and the Power control commit (3562f26, +branch `job-console`), both cherry-picked. Lever 3 of docs/plans/miner-eff.md grows two knobs and a fleet memory; +lever 2 (docs/design/miner-tuning.md) carries the priors in the same signed `tuning` section. + +## 1. What a user sees + +| Moment | The card row says | What happened | +|---|---|---| +| First 2 minutes of mining | `tuning: waits for 120 s of steady mining` | The worker warms up; nothing is touched. | +| Tuning | `tuning: holding 2472 MHz · 100% · 41 s (step 7 of 9)` with a Stop button | One card at a time, on the live kernel, never restarting the worker. | +| Tuned | **Tuned: 122.3 MH/s at 290 W (0.422 MH/W)**, then `2470 MHz at 100%, full tune, 1 h ago` | The point is pinned on the card; the result went to the fleet. | +| Known model | the same line, `from the fleet prior, confirmed, 2 min ago` | The card started at its model's prior and confirmed it in two steps instead of nine. | +| Apple silicon | **Tuned: 26.7 MH/s at 38 W (0.703 MH/W)** `(measured as it runs)`, and `measure only on Apple silicon: the system sets the clocks and the power; no control exposed` | Nothing can be set; the number is still reported so the row and the fleet know what the card does. | +| NVIDIA, Power control off | the measured line and `measure only until Power control is on in Settings (Windows asks for administrator rights once)` | The app never raises the prompt by itself (5 October 2026). One switch, one prompt, and the full tune runs. | +| Slider moved | `your setting stays pinned` | A manual point is never overridden; the tune still measures and reports. | +| Stopped | `tuning stopped: a remote job took the GPU` and the card back where it was | Any fault reverts the step and the run. | +| Fleet pause | Settings: `tuning paused fleet-wide by the signed manifest` | The kill switch. | + +Settings: one switch, "Ember Tune: tune every card for MH per watt out of the box (once after install, then weekly, +and after a driver or program change)". AMD needs no rights. NVIDIA needs the Power control switch (one administrator +prompt) for both knobs; off, it measures only. + +## 2. The knobs, per vendor + +| Vendor | Power limit | Core clock cap | Memory clock | How | Rights | +|---|---|---|---|---|---| +| NVIDIA | `nvidia-smi -pl `, percent of the default, inside `power.min_limit` and `power.max_limit` | `nvidia-smi -lgc 0,`, percent of `clocks.max.gr`; `-rgc` = unlocked | never touched (`-lmc` is not used); read back as `clocks.mem` | directly when the engine is elevated, else the one-prompt helper (` pl `, ` lgc `, ` rgc` in `sweep/cmd.txt`) | administrator, so only with Power control on | +| AMD | `igneum-gpu-telemetry --card N --set-plimit ` (0 = default, -20 = 80%), inside the `tune` line's `plimit_range` (PC 1's 9070 XT: -30 to 10, so 70% is the floor) | `--set-gmax` only when the `tune` line's `gmax_range` is absolute MHz (floor 0 or above); on RDNA 4 the range is an offset from stock (-500 to 1000 on PC 1) and the clock knob stays closed until the stock clock is known; `--reset` for the default point | not settable through ADLX on RDNA 4; read back as `mclk_mhz`, and a step whose mean memory clock falls under 95% of the baseline's is marked and cannot win | the helper, one process per request, exit 0 and a `tune ... ok` line | none on Windows (ADLX manual tuning); root on Linux, so measure only there | +| Apple | none | none | none | measure only | none | + +Vendor limits are never exceeded and the floor is never undercut: the plan clamps every point (`Limits::clamp_clock`, +`Limits::watts_for`), and a clock floor the vendor does not report is 60% of the maximum. + +## 3. The plan and the choice + +Full plan (a new model, or a prior that lost its confirm check): the power ladder 100, 90, 80, 70, 60, 50% at the +unlocked clock (duplicate watts dropped where the card's floor clamps them), then the clock ladder 90, 80, 70, 60% +of the maximum at the power point the power ladder chose. 60 s hold after 15 s settle per step; 9 steps on an +RTX 5090 (five power, four clock), about 12 minutes. + +Confirm plan (the model's prior has 5 or more reports): the prior's point, then one neighbour (the next clock step up +when the prior caps the clock, else one power step down). If the neighbour beats the prior by over 1% MH/W, the full +plan is queued; else the prior stands. Two steps, about 3 minutes. + +Baseline plan (measure only): one step at the card's current point. The "before" number for the row and the fleet. + +The choice (`ember::choose`): among the usable steps whose rate is within the tolerance (1%, settable from the +manifest) of the fastest step, the best MH per watt; within 1% on efficiency the higher rate; within 1% on both the +lower draw. A card never gives up more than the tolerance in blocks for the saving. A step is unusable when it is +marked: `faulted` (a rejected or mismatched hash during the hold: the step is reverted and marked), `hot` (the GPU +reached 85 C; the run aborts at 90), `memory_clock_dropped`, `unapplied` (the readback disagreed with the request), +`no_readings` (under three draw samples or no STATUS line). + +## 4. The data flow + +``` +card mines 120 s ──> probe (limits, driver, how to set) ──> plan ──> steps ──> choice ──> point pinned + │ + app log: TUNE start / TUNE card=.. step=.. / TUNE chosen / TUNE {json} (and stdout under --sweep) + │ + log upload (every minute, the existing intake, site/api/log.mjs) ──> Neon miner_logs + │ + relay/lib/ember.mjs aggregate: per (card model | driver major | program class) + median clock cap (10 MHz), median power %, median MH/W, MH/s, W, spread (MAD %), samples, machines + │ │ + console: /r//c/tuning, `node tools/console.mjs tuning` site: tools/tuning.mjs --priors --site + │ -> site/miner-priors.json -> /miners#priors + tools/tuning.mjs --priors --write tuning.json (priors + ember settings beside the kernel-variant cards) + │ + packaging/ota/publish-manifest.sh --tuning tuning.json --deploy (signed; carried over when not given) + │ + every app: /tuning.json ──> ember::settings_of (kill switch, min samples, tolerance, period) + ──> ember::prior_of(key) ──> a new card's confirm plan +``` + +The record (`ember::record_json`): `ts`, `machine` (the first 8 hex of SHA-256 over the install id; the id itself +is random per install and never sent), `app`, `os`, `card`, `vendor`, `driver`, `driver_major`, `class`, `key`, +`plan`, `steps` (the full table: clock, power %, limit, watts, MH/s, MH/W, core and memory clock, hottest reading, +faults, mark), `chosen`, `before` (the full plan's 100% step), `eff`, `mhs`, `watts`. The key: `||`, the class from the worker's race line (`l128w16` today; `v2` before a +race has run). + +## 5. Scheduling and safety + +| Rule | Where | +|---|---| +| One card at a time; the card must have mined 120 s and have a STATUS line | `tick_sweep` | +| Never under a remote job hold, a pause, inside 600 s of the hour boundary, or while the app quits | `tick_sweep`, `sweep_drive` | +| Due once after install, every 7 days (manifest `ember.period_s`), and when the driver major or the program class changed since the last tune | `tick_sweep` (`CardPref.sweep_driver`, `sweep_class`) | +| A pinned card (the slider) is measured, never changed | `sweep_finish` | +| Kill switch: `tuning.ember.enabled = false` in the signed manifest stops every tune fleet-wide; the Settings line says so | `ember::settings_of`, `tick_sweep` | +| Faults: a rejected or mismatched hash marks the step; the card leaving `mining`, a worker error, a job, a pause or 90 C aborts the run and restores the point from before | `Run::sample_fault`, `sweep_drive`, `sweep_abort` | +| Memory clock held: never set; a step that drags it under 95% of the baseline's cannot win | `Row::from_samples` | +| Vendor limits: every point clamped to the reported range; the clock floor 60% when none is reported | `Limits` | +| A signed prior is only ever a starting point inside the card's OWN reported limits (`power.min_limit` to `power.max_limit`, the clock floor to `clocks.max.gr` or the ADLX `gmax_range`), never a memory clock, never a value the card did not report; the confirm step measures it and the full plan replaces it when a neighbour beats it, so a bad prior costs the fleet one confirm step per card, not a setting. The signing key (K1, docs/security/keys.md) therefore cannot push a card past its vendor ceiling or under its floor | `Plan::confirm` clamps through `Limits::clamp_clock` and `power_pct.clamp(50, 100)`; proven by `ember::tests::the_confirm_plan_checks_the_prior_and_its_neighbour` (a prior of 9,000 MHz at 30% becomes 3,090 MHz at 50%) and `limits_never_exceed_the_vendor_or_undercut_the_floor` | +| No prompt the user did not ask for: the NVIDIA helper starts only with Power control on; the `--sweep` job never counts as permission | `sweep_probe_known`, `sweep_helper_start` | +| The elevated helper restores the limit and resets the clocks by itself after 20 idle minutes | `sweep::helper_script_*` | +| A playbook that starts a second engine beside the installed app (the PC measurement jobs) gives it NO pipe (its output goes to a file the script tails: a pipe's write end is inherited by the engine's miners, and the installed app's jobs runner then waits forever for EOF after an abort; C35, PC 1 22:31 UTC, a 24-minute hang and orphaned miners), ends the engine's whole process tree at the end and on the budget (`taskkill /T /F`), and lets the installed app's miners come back only after that | `relay/playbooks/ember-tune-pc1.ps1`, `sweep-5090.ps1`; CI `tools/ci/second-engine-check.sh` fails any playbook without both | +| Every `quit:` line in the app log names its source (the window host's stdin, the host gone, `POST /api/quit`, the `--sweep` run's end) | `Cmd::Quit(&'static str)` (b671c8b) | +| A second engine never runs the updater: `IGNEUM_APP_NO_OTA=1` (implied by `--sweep`) skips the OTA tick and refuses Check now, whatever the manifest's `min_supported_version` says (the installer it would launch quits the installed app: PC 1, 22:31 UTC) | `Engine.no_ota`; the playbooks set the variable; `tools/ci/second-engine-check.sh` demands it | + +## 6. Tests + +| Test | What it fixes | +|---|---| +| `ember::tests::the_full_plan_is_the_power_ladder_then_the_clock_ladder_at_the_chosen_power` | 5 + 4 steps on the 5090's limits, the clamps, the dynamic second half, the 1% and 5% choices | +| `limits_never_exceed_the_vendor_or_undercut_the_floor` | clamps | +| `the_choice_keeps_the_best_mh_per_watt_within_the_rate_tolerance` | the rule, the ties, marked rows never win | +| `the_guards_mark_a_step_so_it_cannot_win` | faulted, hot, memory clock, unapplied, no readings, the line | +| `a_fault_during_a_step_reverts_it_and_the_run_goes_on` | the state machine with a fake clock: the faulted 70% step is marked and never chosen | +| `the_confirm_plan_checks_the_prior_and_its_neighbour` | the two steps, Keep against FullDue, a prior outside the range clamped | +| `a_baseline_plan_measures_the_card_as_it_runs` | no control, still a number and the Tuned line | +| `the_record_and_the_prior_round_trip_through_the_manifest_shape` | record fields (no address, no host), `priors` and `ember` beside `cards`, the sample floor, the kill switch | +| `control_reasons_per_vendor` | who measures only and why | +| `sweep::tests::helper_scripts_carry_the_protocol` | the helper's `pl`, `lgc`, `rgc` | +| `relay/test/ember.test.mjs` | five samples converge (2,470 MHz at 100%), an outlier (0.908 MH/W at 1,854 MHz) moves nothing, baseline records make no prior, de-duplication, the manifest merge keeps lever 2's cards, the canonical round trip, AMD keys | +| `app/igneum-app/ui/tune-line.test.mjs` | the row line per state | + +Run: `cargo test -p igneum-app ember sweep` (on a PC through the build job, or on the Mac under the build lock), +`node --test relay/test/ember.test.mjs app/igneum-app/ui/tune-line.test.mjs`. + +## 7. The tier consequences + +| Tier | What Ember Tune does for it | What it costs | +|---|---|---| +| A laptop GPU (NVIDIA, 60 to 115 W) | the power ladder usually finds the vendor floor binding; the clock ladder is where a memory-bound program saves watts; the thermal mark keeps a hot chassis from winning a step it cannot hold | about 12 minutes once, then 3 minutes a week; under 1% of the hour during the tune (the worker never stops) | +| One 8 GB card | the same two knobs; the 8 GB card is identities-limited (2 by default), the tune does not change that | the same | +| One 12 or 16 GB card | the same | the same | +| One 24 or 32 GB card (the 5090) | the draw sits far under the cap (290 W under 460 W on PC 1), so the power ladder is flat and the clock ladder is the lever; expected saving from the 4 October stability line: tens of watts at under 1% rate, to be measured | the same | +| A rig (several cards) | one card at a time, so a six-card rig takes about 70 minutes to tune once; every card of one model after the first starts at the prior (3 minutes); the tune never touches a card a remote job holds | linear in cards once, then the confirm plan | +| A pool user | the same per card; a pool submits the same hashes, so the 1% rate tolerance is the same 1% of shares | the same | +| AMD on Linux | measure only (sysfs needs root); the row says so | 60 s a week | +| Apple silicon | measure only; the row says so | 60 s a week | + +Privacy line: what is uploaded is the record in section 4 and nothing else: a hash of the random install id, the +card model, the driver version, the OS, the program class, the step table and the chosen point. No address, no +hostname, no raw machine id, no user name. The public priors table carries only the aggregate per model. + +## 8. Measurements + +### PC 1, 5 October 2026 (night) + +Tonight's constraints, read from PC 1's own uploads: the installed app runs as `DESKTOP-KMCV30N\Admin` with +`elevated=False` (the account line at 19:02:33 UTC), the two in-app sweep attempts at 20:09 UTC aborted on the +cancelled administrator prompt (`SWEEP aborted ... the_elevated_helper_did_not_run_(the_administrator_prompt_was_cancelled)`), +so no stored sweep result exists from today, and the RX 9070 XT left the PCI bus at about 20:40 UTC (eGPU link, +not restarted tonight). NVIDIA's `-pl` and `-lgc` need administrator rights, the project lead is asleep, and the app never raises +the prompt by itself, so tonight's run on PC 1 is the baseline plan on the 5090 through the whole pipeline (probe, +measure, TUNE record, upload, aggregation, prior shape in a test manifest). The two-knob tune on the 5090 and the +9070 XT run are owed: the 5090 the moment Power control is switched on (one prompt, then the tune runs by itself +within 2 minutes of steady mining), the 9070 XT when the card is back on the bus. + +Run 1 (ember-tune-pc1-1, 22:30 UTC): aborted 46 s in by the installed app quitting, named the next morning: the +second engine's own updater (0.3.9 under min_supported_version = urgent) ran the per-user installer, whose +PrepareToInstall quit the installed app through its api/quit (C35 in the bench log); before any step; nothing set; the "before" snapshots are in the bench log (5090: 450 W of 575, 2,850 MHz core, 3,090 MHz +maximum, 14,001 MHz memory; 9070 XT present on bus 98 with OFFSET ranges `gmax_range -500 1000`, `plimit_range -30 +10`). The offset finding changed the AMD mapping (054e041): an offset clock range closes the clock knob and the power +ladder runs on a percent scale bounded by `plimit_range`. The re-run follows the 0.3.11 rollout. + +## 7a. One administrator approval, ever (0.3.13; the project lead, 6 October 2026, 11:50 UTC) + +What 0.3.12 does: Power control on raises one prompt and sets every cap in that step; every later cap (an app start, a +reboot, a slider move) and every tune's helper is another elevated launch, so another prompt. Not "once, ever". + +What `src/powertask.rs` does: the first approval's elevated step also registers a per-user Windows scheduled task, +`Igneum Power Helper` (principal = the signed-in user, interactive logon, RunLevel Highest, no trigger, hidden, one +hour limit, new starts ignored while one runs), whose action is the app's own exe in the install folder with +`--power-helper`. A task the user owns is started by the user's unelevated engine with `Start-ScheduledTask`, no +prompt, and runs elevated. Every later cap and every tune's helper starts the task and writes the command file +`/app/sweep/cmd.txt` (` dev `, ` pl `, ` lgc `, ` rgc`, `quit`). The task +survives app restarts, updates (the per-user installer replaces the exe in place; the task's action path is the +install folder) and reboots. Power control off starts the task once and sends `remove`: the helper unregisters the +task (elevated) and exits; nothing is left behind. Linux keeps pkexec per step; macOS has no cap. + +Threat note: the helper runs only fixed verbs with digit-only arguments through `Command::new(nvidia-smi).args` +(the driver's own path, never PATH, never a shell); a line that is anything else is ignored; the sequence must rise +(a stale file runs nothing); an attacker running as the user gains the power limit and clock cap of the user's own +NVIDIA cards inside the driver's ranges, which the same user could set with one approved prompt anyway; no file, +process, registry key or other binary is reachable through it. Tests: `powertask::tests` (the parser refuses every +non-digit or extra argument, the arguments reach nvidia-smi as a list, the registration is per-user, highest, +trigger-less and quote-safe, a stale command file runs nothing). + +## 8a. Next-cut notes (for the 0.3.12 shipper) + +| Commit | What | Where | +|---|---|---| +| b671c8b | every `quit:` names its source; Power control alone decides; no cap at start under `--sweep` | main.rs, server.rs, engine.rs (separable) | +| e600e63 | a second engine never runs the updater (`IGNEUM_APP_NO_OTA`, implied by `--sweep`) | engine.rs (6 lines, separable) | +| 1e9550e | the elevated job path's output file is followed while the script runs, so the 5-minute progress reports carry its lines (a 35-minute run that never mined showed only "script running" on 6 October 2026); the tune playbook's watchdog fails a run that mines nothing within 120 s of its first status line, with the engine's last log line in the RESULT | jobrun.rs `follow_file`, relay/playbooks/ember-tune-pc1.ps1 | + +## 9. Open + +- The NVIDIA clock readback: `nvidia-smi -lgc` is confirmed only through the core clock during the hold (a mean over + the cap by 5% marks the step `unapplied`); the first run with Power control on tells whether the driver honours + the lock on the 5090 under this kernel. +- ADLX on RDNA 4 exposes no memory-clock setter; the memory-clock mark is the guard. The telemetry agent's 9070 XT + sweep tells whether a core cap drags the memory clock on that card. +- The confirm plan's neighbour is one step; a second neighbour (the other knob) would cost 75 s more and catch a + prior that is wrong on both knobs. +- Intel: no knob yet; the row says measure only. + + +## Fleet priors from rented cards (6 October 2026, branch gpu-fleet) + +Measured by `tools/fleet/box-ember.sh` on Vast.ai containers, the 0.3.12 CUDA worker against the box's devnet node, 15 s settle and 60 s hold per step. `nvidia-smi -pl` and `-lgc` are refused inside the containers (the host's driver holds the knobs), so every ladder is its baseline step only: the untuned point per model, as `plan: baseline` records in `relay/lib/ember.mjs` terms (they summarise beside a prior, never set one). A tuned prior per model needs bare metal or a VM with the driver inside. + +| Card | Driver | MH/s | W | MH/W | Power default W | Clock max MHz | Plan | +|---|---|---|---|---|---|---|---| +| RTX 3060 | 580.126.20 | 24.59 | 106.1 | 0.2318 | 170.0 | 2100.0 | baseline | +| RTX 4060 Ti | | 17.64 | 73.5 | 0.24 | | | baseline | +| RTX 4070 | 580.159.04 | 24.77 | 91.3 | 0.2713 | 200.0 | 3120.0 | baseline | +| RTX 4090 | 595.71.05 | 52.24 | 179.9 | 0.2904 | 450.0 | 3105.0 | baseline | +| RTX A5000 | 580.82.09 | 47.6 | 222.3 | 0.2141 | 230.0 | 2100.0 | baseline | +| RTX 5090 | 580.173.02 | 98.12 | 258.2 | 0.38 | 575.0 | 3090.0 | baseline | +| RTX 5070 | 595.84 | 56.83 | 136.4 | 0.4166 | 250.0 | 3135.0 | baseline | + +The TUNE records themselves (`ember.json` per instance under `~/Desktop/fleet//`) carry the step, the mean core and memory clock and the hottest reading; `relay/lib/ember.mjs parseRecords` reads the `TUNE {...}` line each ladder prints.