Ember Tune plan: the fleet priors from rented cards (baseline rows; the containers refuse the knobs)
This commit is contained in:
parent
9a8657bb00
commit
9c6b62ab5e
1 changed files with 221 additions and 0 deletions
221
docs/plans/ember-tune.md
Normal file
221
docs/plans/ember-tune.md
Normal file
|
|
@ -0,0 +1,221 @@
|
|||
# Ember Tune: every card tuned for MH per watt, out of the box
|
||||
|
||||
5 October 2026, night. the project lead: "make sure we have ember tuning every single card for efficiency out of the box, the
|
||||
more data = the better the tune, make an awesome system." Branch `ember-tune`, worktree `../igneum-wt-ember-tune`,
|
||||
on top of the AMD telemetry commit (7adcd4c, branch `opencl-rdna4-telemetry`) and the Power control commit (3562f26,
|
||||
branch `job-console`), both cherry-picked. Lever 3 of docs/plans/miner-eff.md grows two knobs and a fleet memory;
|
||||
lever 2 (docs/design/miner-tuning.md) carries the priors in the same signed `tuning` section.
|
||||
|
||||
## 1. What a user sees
|
||||
|
||||
| Moment | The card row says | What happened |
|
||||
|---|---|---|
|
||||
| First 2 minutes of mining | `tuning: waits for 120 s of steady mining` | The worker warms up; nothing is touched. |
|
||||
| Tuning | `tuning: holding 2472 MHz · 100% · 41 s (step 7 of 9)` with a Stop button | One card at a time, on the live kernel, never restarting the worker. |
|
||||
| Tuned | **Tuned: 122.3 MH/s at 290 W (0.422 MH/W)**, then `2470 MHz at 100%, full tune, 1 h ago` | The point is pinned on the card; the result went to the fleet. |
|
||||
| Known model | the same line, `from the fleet prior, confirmed, 2 min ago` | The card started at its model's prior and confirmed it in two steps instead of nine. |
|
||||
| Apple silicon | **Tuned: 26.7 MH/s at 38 W (0.703 MH/W)** `(measured as it runs)`, and `measure only on Apple silicon: the system sets the clocks and the power; no control exposed` | Nothing can be set; the number is still reported so the row and the fleet know what the card does. |
|
||||
| NVIDIA, Power control off | the measured line and `measure only until Power control is on in Settings (Windows asks for administrator rights once)` | The app never raises the prompt by itself (5 October 2026). One switch, one prompt, and the full tune runs. |
|
||||
| Slider moved | `your setting stays pinned` | A manual point is never overridden; the tune still measures and reports. |
|
||||
| Stopped | `tuning stopped: a remote job took the GPU` and the card back where it was | Any fault reverts the step and the run. |
|
||||
| Fleet pause | Settings: `tuning paused fleet-wide by the signed manifest` | The kill switch. |
|
||||
|
||||
Settings: one switch, "Ember Tune: tune every card for MH per watt out of the box (once after install, then weekly,
|
||||
and after a driver or program change)". AMD needs no rights. NVIDIA needs the Power control switch (one administrator
|
||||
prompt) for both knobs; off, it measures only.
|
||||
|
||||
## 2. The knobs, per vendor
|
||||
|
||||
| Vendor | Power limit | Core clock cap | Memory clock | How | Rights |
|
||||
|---|---|---|---|---|---|
|
||||
| NVIDIA | `nvidia-smi -pl <W>`, percent of the default, inside `power.min_limit` and `power.max_limit` | `nvidia-smi -lgc 0,<MHz>`, percent of `clocks.max.gr`; `-rgc` = unlocked | never touched (`-lmc` is not used); read back as `clocks.mem` | directly when the engine is elevated, else the one-prompt helper (`<seq> pl <W>`, `<seq> lgc <MHz>`, `<seq> rgc` in `sweep/cmd.txt`) | administrator, so only with Power control on |
|
||||
| AMD | `igneum-gpu-telemetry --card N --set-plimit <offset>` (0 = default, -20 = 80%), inside the `tune` line's `plimit_range` (PC 1's 9070 XT: -30 to 10, so 70% is the floor) | `--set-gmax` only when the `tune` line's `gmax_range` is absolute MHz (floor 0 or above); on RDNA 4 the range is an offset from stock (-500 to 1000 on PC 1) and the clock knob stays closed until the stock clock is known; `--reset` for the default point | not settable through ADLX on RDNA 4; read back as `mclk_mhz`, and a step whose mean memory clock falls under 95% of the baseline's is marked and cannot win | the helper, one process per request, exit 0 and a `tune ... ok` line | none on Windows (ADLX manual tuning); root on Linux, so measure only there |
|
||||
| Apple | none | none | none | measure only | none |
|
||||
|
||||
Vendor limits are never exceeded and the floor is never undercut: the plan clamps every point (`Limits::clamp_clock`,
|
||||
`Limits::watts_for`), and a clock floor the vendor does not report is 60% of the maximum.
|
||||
|
||||
## 3. The plan and the choice
|
||||
|
||||
Full plan (a new model, or a prior that lost its confirm check): the power ladder 100, 90, 80, 70, 60, 50% at the
|
||||
unlocked clock (duplicate watts dropped where the card's floor clamps them), then the clock ladder 90, 80, 70, 60%
|
||||
of the maximum at the power point the power ladder chose. 60 s hold after 15 s settle per step; 9 steps on an
|
||||
RTX 5090 (five power, four clock), about 12 minutes.
|
||||
|
||||
Confirm plan (the model's prior has 5 or more reports): the prior's point, then one neighbour (the next clock step up
|
||||
when the prior caps the clock, else one power step down). If the neighbour beats the prior by over 1% MH/W, the full
|
||||
plan is queued; else the prior stands. Two steps, about 3 minutes.
|
||||
|
||||
Baseline plan (measure only): one step at the card's current point. The "before" number for the row and the fleet.
|
||||
|
||||
The choice (`ember::choose`): among the usable steps whose rate is within the tolerance (1%, settable from the
|
||||
manifest) of the fastest step, the best MH per watt; within 1% on efficiency the higher rate; within 1% on both the
|
||||
lower draw. A card never gives up more than the tolerance in blocks for the saving. A step is unusable when it is
|
||||
marked: `faulted` (a rejected or mismatched hash during the hold: the step is reverted and marked), `hot` (the GPU
|
||||
reached 85 C; the run aborts at 90), `memory_clock_dropped`, `unapplied` (the readback disagreed with the request),
|
||||
`no_readings` (under three draw samples or no STATUS line).
|
||||
|
||||
## 4. The data flow
|
||||
|
||||
```
|
||||
card mines 120 s ──> probe (limits, driver, how to set) ──> plan ──> steps ──> choice ──> point pinned
|
||||
│
|
||||
app log: TUNE start / TUNE card=.. step=.. / TUNE chosen / TUNE {json} (and stdout under --sweep)
|
||||
│
|
||||
log upload (every minute, the existing intake, site/api/log.mjs) ──> Neon miner_logs
|
||||
│
|
||||
relay/lib/ember.mjs aggregate: per (card model | driver major | program class)
|
||||
median clock cap (10 MHz), median power %, median MH/W, MH/s, W, spread (MAD %), samples, machines
|
||||
│ │
|
||||
console: /r/<token>/c/tuning, `node tools/console.mjs tuning` site: tools/tuning.mjs --priors --site
|
||||
│ -> site/miner-priors.json -> /miners#priors
|
||||
tools/tuning.mjs --priors --write tuning.json (priors + ember settings beside the kernel-variant cards)
|
||||
│
|
||||
packaging/ota/publish-manifest.sh --tuning tuning.json --deploy (signed; carried over when not given)
|
||||
│
|
||||
every app: <app data>/tuning.json ──> ember::settings_of (kill switch, min samples, tolerance, period)
|
||||
──> ember::prior_of(key) ──> a new card's confirm plan
|
||||
```
|
||||
|
||||
The record (`ember::record_json`): `ts`, `machine` (the first 8 hex of SHA-256 over the install id; the id itself
|
||||
is random per install and never sent), `app`, `os`, `card`, `vendor`, `driver`, `driver_major`, `class`, `key`,
|
||||
`plan`, `steps` (the full table: clock, power %, limit, watts, MH/s, MH/W, core and memory clock, hottest reading,
|
||||
faults, mark), `chosen`, `before` (the full plan's 100% step), `eff`, `mhs`, `watts`. The key: `<card model with
|
||||
underscores>|<driver major>|<program class>`, the class from the worker's race line (`l128w16` today; `v2` before a
|
||||
race has run).
|
||||
|
||||
## 5. Scheduling and safety
|
||||
|
||||
| Rule | Where |
|
||||
|---|---|
|
||||
| One card at a time; the card must have mined 120 s and have a STATUS line | `tick_sweep` |
|
||||
| Never under a remote job hold, a pause, inside 600 s of the hour boundary, or while the app quits | `tick_sweep`, `sweep_drive` |
|
||||
| Due once after install, every 7 days (manifest `ember.period_s`), and when the driver major or the program class changed since the last tune | `tick_sweep` (`CardPref.sweep_driver`, `sweep_class`) |
|
||||
| A pinned card (the slider) is measured, never changed | `sweep_finish` |
|
||||
| Kill switch: `tuning.ember.enabled = false` in the signed manifest stops every tune fleet-wide; the Settings line says so | `ember::settings_of`, `tick_sweep` |
|
||||
| Faults: a rejected or mismatched hash marks the step; the card leaving `mining`, a worker error, a job, a pause or 90 C aborts the run and restores the point from before | `Run::sample_fault`, `sweep_drive`, `sweep_abort` |
|
||||
| Memory clock held: never set; a step that drags it under 95% of the baseline's cannot win | `Row::from_samples` |
|
||||
| Vendor limits: every point clamped to the reported range; the clock floor 60% when none is reported | `Limits` |
|
||||
| A signed prior is only ever a starting point inside the card's OWN reported limits (`power.min_limit` to `power.max_limit`, the clock floor to `clocks.max.gr` or the ADLX `gmax_range`), never a memory clock, never a value the card did not report; the confirm step measures it and the full plan replaces it when a neighbour beats it, so a bad prior costs the fleet one confirm step per card, not a setting. The signing key (K1, docs/security/keys.md) therefore cannot push a card past its vendor ceiling or under its floor | `Plan::confirm` clamps through `Limits::clamp_clock` and `power_pct.clamp(50, 100)`; proven by `ember::tests::the_confirm_plan_checks_the_prior_and_its_neighbour` (a prior of 9,000 MHz at 30% becomes 3,090 MHz at 50%) and `limits_never_exceed_the_vendor_or_undercut_the_floor` |
|
||||
| No prompt the user did not ask for: the NVIDIA helper starts only with Power control on; the `--sweep` job never counts as permission | `sweep_probe_known`, `sweep_helper_start` |
|
||||
| The elevated helper restores the limit and resets the clocks by itself after 20 idle minutes | `sweep::helper_script_*` |
|
||||
| A playbook that starts a second engine beside the installed app (the PC measurement jobs) gives it NO pipe (its output goes to a file the script tails: a pipe's write end is inherited by the engine's miners, and the installed app's jobs runner then waits forever for EOF after an abort; C35, PC 1 22:31 UTC, a 24-minute hang and orphaned miners), ends the engine's whole process tree at the end and on the budget (`taskkill /T /F`), and lets the installed app's miners come back only after that | `relay/playbooks/ember-tune-pc1.ps1`, `sweep-5090.ps1`; CI `tools/ci/second-engine-check.sh` fails any playbook without both |
|
||||
| Every `quit:` line in the app log names its source (the window host's stdin, the host gone, `POST /api/quit`, the `--sweep` run's end) | `Cmd::Quit(&'static str)` (b671c8b) |
|
||||
| A second engine never runs the updater: `IGNEUM_APP_NO_OTA=1` (implied by `--sweep`) skips the OTA tick and refuses Check now, whatever the manifest's `min_supported_version` says (the installer it would launch quits the installed app: PC 1, 22:31 UTC) | `Engine.no_ota`; the playbooks set the variable; `tools/ci/second-engine-check.sh` demands it |
|
||||
|
||||
## 6. Tests
|
||||
|
||||
| Test | What it fixes |
|
||||
|---|---|
|
||||
| `ember::tests::the_full_plan_is_the_power_ladder_then_the_clock_ladder_at_the_chosen_power` | 5 + 4 steps on the 5090's limits, the clamps, the dynamic second half, the 1% and 5% choices |
|
||||
| `limits_never_exceed_the_vendor_or_undercut_the_floor` | clamps |
|
||||
| `the_choice_keeps_the_best_mh_per_watt_within_the_rate_tolerance` | the rule, the ties, marked rows never win |
|
||||
| `the_guards_mark_a_step_so_it_cannot_win` | faulted, hot, memory clock, unapplied, no readings, the line |
|
||||
| `a_fault_during_a_step_reverts_it_and_the_run_goes_on` | the state machine with a fake clock: the faulted 70% step is marked and never chosen |
|
||||
| `the_confirm_plan_checks_the_prior_and_its_neighbour` | the two steps, Keep against FullDue, a prior outside the range clamped |
|
||||
| `a_baseline_plan_measures_the_card_as_it_runs` | no control, still a number and the Tuned line |
|
||||
| `the_record_and_the_prior_round_trip_through_the_manifest_shape` | record fields (no address, no host), `priors` and `ember` beside `cards`, the sample floor, the kill switch |
|
||||
| `control_reasons_per_vendor` | who measures only and why |
|
||||
| `sweep::tests::helper_scripts_carry_the_protocol` | the helper's `pl`, `lgc`, `rgc` |
|
||||
| `relay/test/ember.test.mjs` | five samples converge (2,470 MHz at 100%), an outlier (0.908 MH/W at 1,854 MHz) moves nothing, baseline records make no prior, de-duplication, the manifest merge keeps lever 2's cards, the canonical round trip, AMD keys |
|
||||
| `app/igneum-app/ui/tune-line.test.mjs` | the row line per state |
|
||||
|
||||
Run: `cargo test -p igneum-app ember sweep` (on a PC through the build job, or on the Mac under the build lock),
|
||||
`node --test relay/test/ember.test.mjs app/igneum-app/ui/tune-line.test.mjs`.
|
||||
|
||||
## 7. The tier consequences
|
||||
|
||||
| Tier | What Ember Tune does for it | What it costs |
|
||||
|---|---|---|
|
||||
| A laptop GPU (NVIDIA, 60 to 115 W) | the power ladder usually finds the vendor floor binding; the clock ladder is where a memory-bound program saves watts; the thermal mark keeps a hot chassis from winning a step it cannot hold | about 12 minutes once, then 3 minutes a week; under 1% of the hour during the tune (the worker never stops) |
|
||||
| One 8 GB card | the same two knobs; the 8 GB card is identities-limited (2 by default), the tune does not change that | the same |
|
||||
| One 12 or 16 GB card | the same | the same |
|
||||
| One 24 or 32 GB card (the 5090) | the draw sits far under the cap (290 W under 460 W on PC 1), so the power ladder is flat and the clock ladder is the lever; expected saving from the 4 October stability line: tens of watts at under 1% rate, to be measured | the same |
|
||||
| A rig (several cards) | one card at a time, so a six-card rig takes about 70 minutes to tune once; every card of one model after the first starts at the prior (3 minutes); the tune never touches a card a remote job holds | linear in cards once, then the confirm plan |
|
||||
| A pool user | the same per card; a pool submits the same hashes, so the 1% rate tolerance is the same 1% of shares | the same |
|
||||
| AMD on Linux | measure only (sysfs needs root); the row says so | 60 s a week |
|
||||
| Apple silicon | measure only; the row says so | 60 s a week |
|
||||
|
||||
Privacy line: what is uploaded is the record in section 4 and nothing else: a hash of the random install id, the
|
||||
card model, the driver version, the OS, the program class, the step table and the chosen point. No address, no
|
||||
hostname, no raw machine id, no user name. The public priors table carries only the aggregate per model.
|
||||
|
||||
## 8. Measurements
|
||||
|
||||
### PC 1, 5 October 2026 (night)
|
||||
|
||||
Tonight's constraints, read from PC 1's own uploads: the installed app runs as `DESKTOP-KMCV30N\Admin` with
|
||||
`elevated=False` (the account line at 19:02:33 UTC), the two in-app sweep attempts at 20:09 UTC aborted on the
|
||||
cancelled administrator prompt (`SWEEP aborted ... the_elevated_helper_did_not_run_(the_administrator_prompt_was_cancelled)`),
|
||||
so no stored sweep result exists from today, and the RX 9070 XT left the PCI bus at about 20:40 UTC (eGPU link,
|
||||
not restarted tonight). NVIDIA's `-pl` and `-lgc` need administrator rights, the project lead is asleep, and the app never raises
|
||||
the prompt by itself, so tonight's run on PC 1 is the baseline plan on the 5090 through the whole pipeline (probe,
|
||||
measure, TUNE record, upload, aggregation, prior shape in a test manifest). The two-knob tune on the 5090 and the
|
||||
9070 XT run are owed: the 5090 the moment Power control is switched on (one prompt, then the tune runs by itself
|
||||
within 2 minutes of steady mining), the 9070 XT when the card is back on the bus.
|
||||
|
||||
Run 1 (ember-tune-pc1-1, 22:30 UTC): aborted 46 s in by the installed app quitting, named the next morning: the
|
||||
second engine's own updater (0.3.9 under min_supported_version = urgent) ran the per-user installer, whose
|
||||
PrepareToInstall quit the installed app through its api/quit (C35 in the bench log); before any step; nothing set; the "before" snapshots are in the bench log (5090: 450 W of 575, 2,850 MHz core, 3,090 MHz
|
||||
maximum, 14,001 MHz memory; 9070 XT present on bus 98 with OFFSET ranges `gmax_range -500 1000`, `plimit_range -30
|
||||
10`). The offset finding changed the AMD mapping (054e041): an offset clock range closes the clock knob and the power
|
||||
ladder runs on a percent scale bounded by `plimit_range`. The re-run follows the 0.3.11 rollout.
|
||||
|
||||
## 7a. One administrator approval, ever (0.3.13; the project lead, 6 October 2026, 11:50 UTC)
|
||||
|
||||
What 0.3.12 does: Power control on raises one prompt and sets every cap in that step; every later cap (an app start, a
|
||||
reboot, a slider move) and every tune's helper is another elevated launch, so another prompt. Not "once, ever".
|
||||
|
||||
What `src/powertask.rs` does: the first approval's elevated step also registers a per-user Windows scheduled task,
|
||||
`Igneum Power Helper` (principal = the signed-in user, interactive logon, RunLevel Highest, no trigger, hidden, one
|
||||
hour limit, new starts ignored while one runs), whose action is the app's own exe in the install folder with
|
||||
`--power-helper`. A task the user owns is started by the user's unelevated engine with `Start-ScheduledTask`, no
|
||||
prompt, and runs elevated. Every later cap and every tune's helper starts the task and writes the command file
|
||||
`<app data>/app/sweep/cmd.txt` (`<seq> dev <n>`, `<seq> pl <W>`, `<seq> lgc <MHz>`, `<seq> rgc`, `quit`). The task
|
||||
survives app restarts, updates (the per-user installer replaces the exe in place; the task's action path is the
|
||||
install folder) and reboots. Power control off starts the task once and sends `remove`: the helper unregisters the
|
||||
task (elevated) and exits; nothing is left behind. Linux keeps pkexec per step; macOS has no cap.
|
||||
|
||||
Threat note: the helper runs only fixed verbs with digit-only arguments through `Command::new(nvidia-smi).args`
|
||||
(the driver's own path, never PATH, never a shell); a line that is anything else is ignored; the sequence must rise
|
||||
(a stale file runs nothing); an attacker running as the user gains the power limit and clock cap of the user's own
|
||||
NVIDIA cards inside the driver's ranges, which the same user could set with one approved prompt anyway; no file,
|
||||
process, registry key or other binary is reachable through it. Tests: `powertask::tests` (the parser refuses every
|
||||
non-digit or extra argument, the arguments reach nvidia-smi as a list, the registration is per-user, highest,
|
||||
trigger-less and quote-safe, a stale command file runs nothing).
|
||||
|
||||
## 8a. Next-cut notes (for the 0.3.12 shipper)
|
||||
|
||||
| Commit | What | Where |
|
||||
|---|---|---|
|
||||
| b671c8b | every `quit:` names its source; Power control alone decides; no cap at start under `--sweep` | main.rs, server.rs, engine.rs (separable) |
|
||||
| e600e63 | a second engine never runs the updater (`IGNEUM_APP_NO_OTA`, implied by `--sweep`) | engine.rs (6 lines, separable) |
|
||||
| 1e9550e | the elevated job path's output file is followed while the script runs, so the 5-minute progress reports carry its lines (a 35-minute run that never mined showed only "script running" on 6 October 2026); the tune playbook's watchdog fails a run that mines nothing within 120 s of its first status line, with the engine's last log line in the RESULT | jobrun.rs `follow_file`, relay/playbooks/ember-tune-pc1.ps1 |
|
||||
|
||||
## 9. Open
|
||||
|
||||
- The NVIDIA clock readback: `nvidia-smi -lgc` is confirmed only through the core clock during the hold (a mean over
|
||||
the cap by 5% marks the step `unapplied`); the first run with Power control on tells whether the driver honours
|
||||
the lock on the 5090 under this kernel.
|
||||
- ADLX on RDNA 4 exposes no memory-clock setter; the memory-clock mark is the guard. The telemetry agent's 9070 XT
|
||||
sweep tells whether a core cap drags the memory clock on that card.
|
||||
- The confirm plan's neighbour is one step; a second neighbour (the other knob) would cost 75 s more and catch a
|
||||
prior that is wrong on both knobs.
|
||||
- Intel: no knob yet; the row says measure only.
|
||||
|
||||
|
||||
## Fleet priors from rented cards (6 October 2026, branch gpu-fleet)
|
||||
|
||||
Measured by `tools/fleet/box-ember.sh` on Vast.ai containers, the 0.3.12 CUDA worker against the box's devnet node, 15 s settle and 60 s hold per step. `nvidia-smi -pl` and `-lgc` are refused inside the containers (the host's driver holds the knobs), so every ladder is its baseline step only: the untuned point per model, as `plan: baseline` records in `relay/lib/ember.mjs` terms (they summarise beside a prior, never set one). A tuned prior per model needs bare metal or a VM with the driver inside.
|
||||
|
||||
| Card | Driver | MH/s | W | MH/W | Power default W | Clock max MHz | Plan |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| RTX 3060 | 580.126.20 | 24.59 | 106.1 | 0.2318 | 170.0 | 2100.0 | baseline |
|
||||
| RTX 4060 Ti | | 17.64 | 73.5 | 0.24 | | | baseline |
|
||||
| RTX 4070 | 580.159.04 | 24.77 | 91.3 | 0.2713 | 200.0 | 3120.0 | baseline |
|
||||
| RTX 4090 | 595.71.05 | 52.24 | 179.9 | 0.2904 | 450.0 | 3105.0 | baseline |
|
||||
| RTX A5000 | 580.82.09 | 47.6 | 222.3 | 0.2141 | 230.0 | 2100.0 | baseline |
|
||||
| RTX 5090 | 580.173.02 | 98.12 | 258.2 | 0.38 | 575.0 | 3090.0 | baseline |
|
||||
| RTX 5070 | 595.84 | 56.83 | 136.4 | 0.4166 | 250.0 | 3135.0 | baseline |
|
||||
|
||||
The TUNE records themselves (`ember.json` per instance under `~/Desktop/fleet/<instance>/`) carry the step, the mean core and memory clock and the hottest reading; `relay/lib/ember.mjs parseRecords` reads the `TUNE {...}` line each ladder prints.
|
||||
Loading…
Reference in a new issue