igneum/docs/plans/ember-tune.md
igneum-labs fa99216dd3 Ember: the 0.3.16 window in the plan and the bench log (A to E held, the 9070 XT measured, cause F and its fix)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit d5f236afc0)
2026-10-06 23:19:24 +00:00

396 lines
41 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Ember Tune: every card tuned for MH per watt, out of the box
5 October 2026, night. the project lead: "make sure we have ember tuning every single card for efficiency out of the box, the
more data = the better the tune, make an awesome system." Branch `ember-tune`, worktree `../igneum-wt-ember-tune`,
on top of the AMD telemetry commit (7adcd4c, branch `opencl-rdna4-telemetry`) and the Power control commit (3562f26,
branch `job-console`), both cherry-picked. Lever 3 of docs/plans/miner-eff.md grows two knobs and a fleet memory;
lever 2 (docs/design/miner-tuning.md) carries the priors in the same signed `tuning` section.
## 1. What a user sees
| Moment | The card row says | What happened |
|---|---|---|
| First 2 minutes of mining | `tuning: waits for 120 s of steady mining` | The worker warms up; nothing is touched. |
| Tuning | `tuning: holding 2472 MHz · 100% · 41 s (step 7 of 9)` with a Stop button | One card at a time, on the live kernel, never restarting the worker. |
| Tuned | **Tuned: 122.3 MH/s at 290 W (0.422 MH/W)**, then `2470 MHz at 100%, full tune, 1 h ago` | The point is pinned on the card; the result went to the fleet. |
| Known model | the same line, `from the fleet prior, confirmed, 2 min ago` | The card started at its model's prior and confirmed it in two steps instead of nine. |
| Apple silicon | **Tuned: 26.7 MH/s at 38 W (0.703 MH/W)** `(measured as it runs)`, and `measure only on Apple silicon: the system sets the clocks and the power; no control exposed` | Nothing can be set; the number is still reported so the row and the fleet know what the card does. |
| NVIDIA, Power control off | the measured line and `measure only until Power control is on in Settings (Windows asks for administrator rights once)` | The app never raises the prompt by itself (5 October 2026). One switch, one prompt, and the full tune runs. |
| Slider moved | `your setting stays pinned` | A manual point is never overridden; the tune still measures and reports. |
| Stopped | `tuning stopped: a remote job took the GPU` and the card back where it was | Any fault reverts the step and the run. |
| Fleet pause | Settings: `tuning paused fleet-wide by the signed manifest` | The kill switch. |
Settings: one switch, "Ember Tune: tune every card for MH per watt out of the box (once after install, then weekly,
and after a driver or program change)". AMD needs no rights. NVIDIA needs the Power control switch (one administrator
prompt) for both knobs; off, it measures only.
## 2. The knobs, per vendor
| Vendor | Power limit | Core clock cap | Memory clock | How | Rights |
|---|---|---|---|---|---|
| NVIDIA | `nvidia-smi -pl <W>`, percent of the default, inside `power.min_limit` and `power.max_limit` | `nvidia-smi -lgc 0,<MHz>`, percent of `clocks.max.gr`; `-rgc` = unlocked | never touched (`-lmc` is not used); read back as `clocks.mem` | directly when the engine is elevated, else the one-prompt helper (`<seq> pl <W>`, `<seq> lgc <MHz>`, `<seq> rgc` in `sweep/cmd.txt`) | administrator, so only with Power control on |
| AMD | `igneum-gpu-telemetry --card N --set-plimit <offset>` (0 = default, -20 = 80%), inside the `tune` line's `plimit_range` (PC 1's 9070 XT: -30 to 10, so 70% is the floor). RULE (run 6, 6 October 2026): ADLX reports the limit as an OFFSET, never watts, so an AMD step is applied on the helper's acknowledgement (`Run::applied`: a card with no limit readback confirms on `acked`), and the draw comes from the telemetry stream (`sample_telemetry`), never from a limit readback; `limit_w` is 0 for every AMD card by design | `--set-gmax` only when the `tune` line's `gmax_range` is absolute MHz (floor 0 or above); on RDNA 4 the range is an offset from stock (-500 to 1000 on PC 1) and the clock knob stays closed until the stock clock is known; `--reset` for the default point | not settable through ADLX on RDNA 4; read back as `mclk_mhz`, and a step whose mean memory clock falls under 95% of the baseline's is marked and cannot win | the helper, one process per request, exit 0 and a `tune ... ok` line | none on Windows (ADLX manual tuning); root on Linux, so measure only there |
| Apple | none | none | none | measure only | none |
Vendor limits are never exceeded and the floor is never undercut: the plan clamps every point (`Limits::clamp_clock`,
`Limits::watts_for`), and a clock floor the vendor does not report is 60% of the maximum.
## 3. The plan and the choice
Full plan (a new model, or a prior that lost its confirm check): the power ladder 100, 90, 80, 70, 60, 50% at the
unlocked clock (duplicate watts dropped where the card's floor clamps them), then the clock ladder 90, 80, 70, 60%
of the maximum at the power point the power ladder chose. 60 s hold after 15 s settle per step; 9 steps on an
RTX 5090 (five power, four clock), about 12 minutes.
Confirm plan (the model's prior has 5 or more reports): the prior's point, then one neighbour (the next clock step up
when the prior caps the clock, else one power step down). If the neighbour beats the prior by over 1% MH/W, the full
plan is queued; else the prior stands. Two steps, about 3 minutes.
Baseline plan (measure only): one step at the card's current point. The "before" number for the row and the fleet.
The choice (`ember::choose`): among the usable steps whose rate is within the tolerance (1%, settable from the
manifest) of the fastest step, the best MH per watt; within 1% on efficiency the higher rate; within 1% on both the
lower draw. A card never gives up more than the tolerance in blocks for the saving. A step is unusable when it is
marked: `faulted` (a rejected or mismatched hash during the hold: the step is reverted and marked), `hot` (the GPU
reached 85 C; the run aborts at 90), `memory_clock_dropped`, `unapplied` (the readback disagreed with the request),
`no_readings` (under three draw samples or no STATUS line).
## 4. The data flow
```
card mines 120 s ──> probe (limits, driver, how to set) ──> plan ──> steps ──> choice ──> point pinned
│
app log: TUNE start / TUNE card=.. step=.. / TUNE chosen / TUNE {json} (and stdout under --sweep)
│
log upload (every minute, the existing intake, site/api/log.mjs) ──> Neon miner_logs
│
relay/lib/ember.mjs aggregate: per (card model | driver major | program class)
median clock cap (10 MHz), median power %, median MH/W, MH/s, W, spread (MAD %), samples, machines
│ │
console: /r/<token>/c/tuning, `node tools/console.mjs tuning` site: tools/tuning.mjs --priors --site
│ -> site/miner-priors.json -> /miners#priors
tools/tuning.mjs --priors --write tuning.json (priors + ember settings beside the kernel-variant cards)
│
packaging/ota/publish-manifest.sh --tuning tuning.json --deploy (signed; carried over when not given)
│
every app: <app data>/tuning.json ──> ember::settings_of (kill switch, min samples, tolerance, period)
──> ember::prior_of(key) ──> a new card's confirm plan
```
The record (`ember::record_json`): `ts`, `machine` (the first 8 hex of SHA-256 over the install id; the id itself
is random per install and never sent), `app`, `os`, `card`, `vendor`, `driver`, `driver_major`, `class`, `key`,
`plan`, `steps` (the full table: clock, power %, limit, watts, MH/s, MH/W, core and memory clock, hottest reading,
faults, mark), `chosen`, `before` (the full plan's 100% step), `eff`, `mhs`, `watts`. The key: `<card model with
underscores>|<driver major>|<program class>`, the class from the worker's race line (`l128w16` today; `v2` before a
race has run).
## 5. Scheduling and safety
| Rule | Where |
|---|---|
| One card at a time; the card must have mined 120 s and have a STATUS line | `tick_sweep` |
| Never under a remote job hold, a pause, inside 600 s of the hour boundary, or while the app quits | `tick_sweep`, `sweep_drive` |
| Due once after install, every 7 days (manifest `ember.period_s`), and when the driver major or the program class changed since the last tune | `tick_sweep` (`CardPref.sweep_driver`, `sweep_class`) |
| A pinned card (the slider) is measured, never changed | `sweep_finish` |
| Kill switch: `tuning.ember.enabled = false` in the signed manifest stops every tune fleet-wide; the Settings line says so | `ember::settings_of`, `tick_sweep` |
| Faults: a rejected or mismatched hash marks the step; the card leaving `mining`, a worker error, a job, a pause or 90 C aborts the run and restores the point from before | `Run::sample_fault`, `sweep_drive`, `sweep_abort` |
| Memory clock held: never set; a step that drags it under 95% of the baseline's cannot win | `Row::from_samples` |
| Vendor limits: every point clamped to the reported range; the clock floor 60% when none is reported | `Limits` |
| A signed prior is only ever a starting point inside the card's OWN reported limits (`power.min_limit` to `power.max_limit`, the clock floor to `clocks.max.gr` or the ADLX `gmax_range`), never a memory clock, never a value the card did not report; the confirm step measures it and the full plan replaces it when a neighbour beats it, so a bad prior costs the fleet one confirm step per card, not a setting. The signing key (K1, docs/security/keys.md) therefore cannot push a card past its vendor ceiling or under its floor | `Plan::confirm` clamps through `Limits::clamp_clock` and `power_pct.clamp(50, 100)`; proven by `ember::tests::the_confirm_plan_checks_the_prior_and_its_neighbour` (a prior of 9,000 MHz at 30% becomes 3,090 MHz at 50%) and `limits_never_exceed_the_vendor_or_undercut_the_floor` |
| No prompt the user did not ask for: the NVIDIA helper starts only with Power control on; the `--sweep` job never counts as permission | `sweep_probe_known`, `sweep_helper_start` |
| The elevated helper restores the limit and resets the clocks by itself after 20 idle minutes | `sweep::helper_script_*` |
| A playbook that starts a second engine beside the installed app (the PC measurement jobs) gives it NO pipe (its output goes to a file the script tails: a pipe's write end is inherited by the engine's miners, and the installed app's jobs runner then waits forever for EOF after an abort; C35, PC 1 22:31 UTC, a 24-minute hang and orphaned miners), ends the engine's whole process tree at the end and on the budget (`taskkill /T /F`), and lets the installed app's miners come back only after that | `relay/playbooks/ember-tune-pc1.ps1`, `sweep-5090.ps1`; CI `tools/ci/second-engine-check.sh` fails any playbook without both |
| Every `quit:` line in the app log names its source (the window host's stdin, the host gone, `POST /api/quit`, the `--sweep` run's end) | `Cmd::Quit(&'static str)` (b671c8b) |
| A second engine never runs the updater: `IGNEUM_APP_NO_OTA=1` (implied by `--sweep`) skips the OTA tick and refuses Check now, whatever the manifest's `min_supported_version` says (the installer it would launch quits the installed app: PC 1, 22:31 UTC) | `Engine.no_ota`; the playbooks set the variable; `tools/ci/second-engine-check.sh` demands it |
## 6. Tests
| Test | What it fixes |
|---|---|
| `ember::tests::the_full_plan_is_the_power_ladder_then_the_clock_ladder_at_the_chosen_power` | 5 + 4 steps on the 5090's limits, the clamps, the dynamic second half, the 1% and 5% choices |
| `limits_never_exceed_the_vendor_or_undercut_the_floor` | clamps |
| `the_choice_keeps_the_best_mh_per_watt_within_the_rate_tolerance` | the rule, the ties, marked rows never win |
| `the_guards_mark_a_step_so_it_cannot_win` | faulted, hot, memory clock, unapplied, no readings, the line |
| `a_fault_during_a_step_reverts_it_and_the_run_goes_on` | the state machine with a fake clock: the faulted 70% step is marked and never chosen |
| `the_confirm_plan_checks_the_prior_and_its_neighbour` | the two steps, Keep against FullDue, a prior outside the range clamped |
| `a_baseline_plan_measures_the_card_as_it_runs` | no control, still a number and the Tuned line |
| `the_record_and_the_prior_round_trip_through_the_manifest_shape` | record fields (no address, no host), `priors` and `ember` beside `cards`, the sample floor, the kill switch |
| `control_reasons_per_vendor` | who measures only and why |
| `sweep::tests::helper_scripts_carry_the_protocol` | the helper's `pl`, `lgc`, `rgc` |
| `relay/test/ember.test.mjs` | five samples converge (2,470 MHz at 100%), an outlier (0.908 MH/W at 1,854 MHz) moves nothing, baseline records make no prior, de-duplication, the manifest merge keeps lever 2's cards, the canonical round trip, AMD keys |
| `app/igneum-app/ui/tune-line.test.mjs` | the row line per state |
Run: `cargo test -p igneum-app ember sweep` (on a PC through the build job, or on the Mac under the build lock),
`node --test relay/test/ember.test.mjs app/igneum-app/ui/tune-line.test.mjs`.
## 7. The tier consequences
| Tier | What Ember Tune does for it | What it costs |
|---|---|---|
| A laptop GPU (NVIDIA, 60 to 115 W) | the power ladder usually finds the vendor floor binding; the clock ladder is where a memory-bound program saves watts; the thermal mark keeps a hot chassis from winning a step it cannot hold | about 12 minutes once, then 3 minutes a week; under 1% of the hour during the tune (the worker never stops) |
| One 8 GB card | the same two knobs; the 8 GB card is identities-limited (2 by default), the tune does not change that | the same |
| One 12 or 16 GB card | the same | the same |
| One 24 or 32 GB card (the 5090) | the draw sits far under the cap (290 W under 460 W on PC 1), so the power ladder is flat and the clock ladder is the lever; expected saving from the 4 October stability line: tens of watts at under 1% rate, to be measured | the same |
| A rig (several cards) | one card at a time, so a six-card rig takes about 70 minutes to tune once; every card of one model after the first starts at the prior (3 minutes); the tune never touches a card a remote job holds | linear in cards once, then the confirm plan |
| A pool user | the same per card; a pool submits the same hashes, so the 1% rate tolerance is the same 1% of shares | the same |
| AMD on Linux | measure only (sysfs needs root); the row says so | 60 s a week |
| Apple silicon | measure only; the row says so | 60 s a week |
Privacy line: what is uploaded is the record in section 4 and nothing else: a hash of the random install id, the
card model, the driver version, the OS, the program class, the step table and the chosen point. No address, no
hostname, no raw machine id, no user name. The public priors table carries only the aggregate per model.
## 8. Measurements
### PC 1, 5 October 2026 (night)
Tonight's constraints, read from PC 1's own uploads: the installed app runs as `DESKTOP-KMCV30N\Admin` with
`elevated=False` (the account line at 19:02:33 UTC), the two in-app sweep attempts at 20:09 UTC aborted on the
cancelled administrator prompt (`SWEEP aborted ... the_elevated_helper_did_not_run_(the_administrator_prompt_was_cancelled)`),
so no stored sweep result exists from today, and the RX 9070 XT left the PCI bus at about 20:40 UTC (eGPU link,
not restarted tonight). NVIDIA's `-pl` and `-lgc` need administrator rights, the project lead is asleep, and the app never raises
the prompt by itself, so tonight's run on PC 1 is the baseline plan on the 5090 through the whole pipeline (probe,
measure, TUNE record, upload, aggregation, prior shape in a test manifest). The two-knob tune on the 5090 and the
9070 XT run are owed: the 5090 the moment Power control is switched on (one prompt, then the tune runs by itself
within 2 minutes of steady mining), the 9070 XT when the card is back on the bus.
Run 1 (ember-tune-pc1-1, 22:30 UTC): aborted 46 s in by the installed app quitting, named the next morning: the
second engine's own updater (0.3.9 under min_supported_version = urgent) ran the per-user installer, whose
PrepareToInstall quit the installed app through its api/quit (C35 in the bench log); before any step; nothing set; the "before" snapshots are in the bench log (5090: 450 W of 575, 2,850 MHz core, 3,090 MHz
maximum, 14,001 MHz memory; 9070 XT present on bus 98 with OFFSET ranges `gmax_range -500 1000`, `plimit_range -30
10`). The offset finding changed the AMD mapping (054e041): an offset clock range closes the clock knob and the power
ladder runs on a percent scale bounded by `plimit_range`. The re-run follows the 0.3.11 rollout.
## 6a. A card never shows 0 MH/s without a reason word (the project lead, 6 October 2026, watching PC 1 during run 5)
Rule: the Mine page's rate column is blank, never "0", when a card is not mining, and the row's word says why:
`mining`, `tuning: step k of n · measuring W W` (the live rate stays in the rate column), `held for a remote job:
<title>`, `paused`, `waiting for the node`, `restart in N s`, `starting`, `ready`, `stopped after repeated faults`,
`failed`, `off`, `removed`, `not usable (<problem>)`. Checked against `View.cardRow`: every state the engine can set
has a word; `held` and `tuning` were the two that read "off" before (a job's `--stop-miners` left `state = off`).
During a tune the big button reads "Tuning, mining again in about N min" (disabled) and the strip carries one line
("Tuning <card>: the full tune, step k of n. Mining again in about N min."). A tune run by a job beside the installed
app reaches the installed app's rows through `POST /api/tune-progress` (the playbook forwards the measurement
engine's `TUNE progress` lines every 10 s and its `TUNE chosen` line as `done`): the posted state is the step, the
live MH/s and W, the seconds left and the plan; a POST of state, never a quit, pause or resume. At the end the row
reads `Tuned: <MH/s> at <W> W (<MH/W>)` with the point and the plan. Tests: `ui/tune-line.test.mjs` (tuning, held,
tuned rows; the button; the strip line).
## 6b. Ember 2: the clocks, the hill-climb, the goal (6 October 2026)
Why: run 5 measured the 5090's power ladder flat (310 to 314 W under every cap from 575 to 400 W, 0.41 MH/W), so
on this memory-latency-bound hash the power cap is a safety, not a lever; the pair that matters is the memory
clock UP and the core clock DOWN. Ember 2 adds both and replaces the fixed ladder with a hill-climb.
| Piece | What | Where |
|---|---|---|
| The memory knob | NVIDIA: `nvidia-smi -lmc <m>,<m>` locks the memory clock to one value, `-rmc` resets; the range is the driver's own (`clocks.mem` under load = the floor, `clocks.max.mem` = the ceiling; PC 1's 5090: 13,801 to 14,001 MHz); directly from an elevated engine or through the Power Helper's `lmc`/`rmc` verbs. AMD: ADLX on RDNA 4 exposes no memory setter through the helper's path, so an AMD card climbs on core and power only (the row says so). Linux NVIDIA: `-lmc` through pkexec as the cap is | `Point.mem_mhz`, `Limits.mem_default_mhz`, `Limits.mem_max_mhz`, `tune_apply`, `powertask::HelperCmd::MemClock`, `sweep::helper_script_*` |
| The hill-climb | from the start point (the fleet prior for the model, else the card's point): each probe moves memory up by 5% of the range above the default (at least 25 MHz) or core down by 5% of the maximum (at least 25 MHz); a probe that improves the goal's score keeps that direction, else the climb turns to the other knob, then to both; 5 probes of 60 s (10 s settle) converge in under 10 minutes | `Plan::climb`, `Climb`, `Plan::climb_next` |
| The stability guard | a rejected or mismatched hash during a probe marks it `faulted`; a faulted probe backs that knob off for good in this climb (memory never goes above, core never below, the last good point); the hash fingerprint is the worker's own self-test at every pack (a variant that is not bit-exact never serves), so a step that keeps mining with 0 mismatches is stable by the only test that matters | `Row::from_samples`, `climb_next`'s `refused_mem` / `refused_core`, docs/design/miner-tuning.md |
| The goal | Settings > goal: `efficiency` (the most MH/W within 10% of the top rate), `balanced` (within 1%, lever 3's rule), `rate` (the fastest point; MH/W breaks ties); one goal for every card | `Goal`, `Settings.tune_goal`, `POST /api/tune/goal {goal, price_pence, climb}` |
| The price | Settings > electricity, pence per kWh: the card's row reads the chosen point as £ a day (`watts × 24 / 1000 × price / 100`: 310 W at 28.5 p = £2.12) | `pounds_per_day`, `Settings.power_price_pence`, the UI |
| The curve | every row of the last plan on the card (`tune_curve`: point, MH/s, W, MH/W, clocks, hottest reading, mark) so a manual tuner sees the points the climb measured | `CardState.tune_curve`, the card detail |
| The fleet | the record's `steps` already carry the curve; the prior gains `mem_mhz` (the median memory clock, 10 MHz steps); a stranger's card of a known model starts its climb at the prior's point | `relay/lib/ember.mjs`, `ember::prior_of` |
| Re-check | weekly (the manifest's period), after a driver major or program-class change, and when the card's temperature band changes (owed: the band trigger) | `tick_sweep` |
Tests: `ember::tests::the_climb_walks_memory_up_and_core_down_and_converges_in_five_probes` (a synthetic
memory-bound card: every probe moves memory up or core down inside the range, the chosen point beats the start in
five probes), `a_refused_probe_backs_that_knob_off_for_good`, `each_goal_picks_its_point` (and the £/day formula),
`powertask::tests` (the `lmc`/`rmc` verbs reach nvidia-smi as `-lmc m,m` and `-rmc`), `sweep::tests` (the helper
scripts carry them), `relay/test/ember.test.mjs` (climb records aggregate into a prior with the memory clock).
Owed: the measured comparison on PC 1 (one climb per card against run 5's ladder rows, no prompt through the
Power Helper), the temperature-band trigger, NVAPI/nvidia-settings for cards whose driver refuses `-lmc`.
## 7a. One administrator approval, ever
Threat note, `reregister` (6 October 2026): the verb takes no path. The helper, already elevated, picks the installed
exe from the two install folders itself (`powertask::install_candidates`), so the most a writer of `cmd.txt` can do
is point the task back at the installed app. The exposure that remains is the one every per-user install with a
highest-run task has: the install folder is the user's own, so a process running as the user can replace the exe the
task runs. The installer's code signature and the OTA's hash check are the answer to that, not the task. (0.3.13; the project lead, 6 October 2026, 11:50 UTC)
What 0.3.12 does: Power control on raises one prompt and sets every cap in that step; every later cap (an app start, a
reboot, a slider move) and every tune's helper is another elevated launch, so another prompt. Not "once, ever".
What `src/powertask.rs` does: the first approval's elevated step also registers a per-user Windows scheduled task,
`Igneum Power Helper` (principal = the signed-in user, interactive logon, RunLevel Highest, no trigger, hidden, one
hour limit, new starts ignored while one runs), whose action is the app's own exe in the install folder with
`--power-helper`. A task the user owns is started by the user's unelevated engine with `Start-ScheduledTask`, no
prompt, and runs elevated. Every later cap and every tune's helper starts the task and writes the command file
`<app data>/app/sweep/cmd.txt` (`<seq> dev <n>`, `<seq> pl <W>`, `<seq> lgc <MHz>`, `<seq> rgc`, `quit`). The task
survives app restarts, updates (the per-user installer replaces the exe in place; the task's action path is the
install folder) and reboots. Power control off starts the task once and sends `remove`: the helper unregisters the
task (elevated) and exits; nothing is left behind. Linux keeps pkexec per step; macOS has no cap.
Threat note: the helper runs only fixed verbs with digit-only arguments through `Command::new(nvidia-smi).args`
(the driver's own path, never PATH, never a shell); a line that is anything else is ignored; the sequence must rise
(a stale file runs nothing); an attacker running as the user gains the power limit and clock cap of the user's own
NVIDIA cards inside the driver's ranges, which the same user could set with one approved prompt anyway; no file,
process, registry key or other binary is reachable through it. Tests: `powertask::tests` (the parser refuses every
non-digit or extra argument, the arguments reach nvidia-smi as a list, the registration is per-user, highest,
trigger-less and quote-safe, a stale command file runs nothing).
### Run 6 (6 October 2026, 16:01Z, PC 1 on 0.3.13 with kit-6 = 564bdea, elevated, the project lead's one click)
The scheduled task registered inside the run ("helper registered=true"), but its action was the run's scratch copy
(`igneum-tune-20261006-170105\bin\igneum-app.exe`): `task_exe` preferred the installed exe only for a path under
`jobs\`. Fixed the same hour: the running exe stands only when it is itself an install candidate; and the helper
gained a `reregister` verb (no path argument, so a writer of cmd.txt can never choose what runs elevated: the helper
finds the install folder itself, re-registers, and logs the task's action read back). The scratch root stays until
that read-back shows the installed path.
RTX 5090 (driver 617.14, class l128w0, no prior), full plan, every row ok, 65 C at most:
| Step | Clock cap | Power | Draw | Rate | MH/W |
|---|---|---|---|---|---|
| 0 (before) | unlocked | 100% 575 W | 311.0 W | 127.90 | 0.411 |
| 1 to 4 | unlocked | 90, 80, 70, 60% | 313 W | 127.9 | 0.408 to 0.409 (the cap never binds) |
| 5 | 2781 MHz | 100% | 298.8 W | 127.89 | 0.428 |
| 6 | 2472 MHz | 100% | 262.0 W | 127.86 | 0.488 |
| 7 | 2163 MHz | 100% | 239.9 W | 127.82 | 0.533 |
| 8 (chosen) | 1854 MHz | 100% | 226.8 W | 127.71 | 0.563 |
The per-user number: 84 W saved for 0.15% of rate, 37% more hashes per watt. The chosen clock is the ladder's floor
(60% of 3,090), not the optimum: `CLOCK_STEPS_PCT` now ends 60, 50, 45 and `CLOCK_FLOOR_PCT` is 45, with the 1% rate
tolerance as the guard; a chosen point on the floor carries `floor=1 note=floor, not optimum` in the chosen line and
the record. Ember 2's climb starts from the chosen point.
RTX 4070 (same driver and class), 10 rows, every one ok, 55 C at most:
| Step | Clock cap | Power | Draw | Rate | MH/W |
|---|---|---|---|---|---|
| 0 (before) | unlocked | 100% 200 W | 106.0 W | 28.72 | 0.271 |
| 1 to 4 | unlocked | 90, 80, 70, 60% | 106.0 W | 28.71 | 0.271 (the cap never binds) |
| 5 | unlocked | 50% 100 W | 99.1 W | 28.70 | 0.290 |
| 6 | 2794 MHz | 50% | 99.1 W | 28.70 | 0.290 |
| 7 | 2484 MHz | 50% | 81.1 W | 28.73 | 0.354 |
| 8 | 2173 MHz | 50% | 77.7 W | 28.74 | 0.370 |
| 9 (chosen) | 1863 MHz | 50% | 75.6 W | 28.78 | 0.381 |
Per user: 30 W saved for no rate lost (+0.2%), 41% more hashes per watt; the floor again.
RX 9070 XT: no row. "TUNE aborted: the setting (0 MHz, 100% = 100 W) did not take within 30 s (card reports 0 W,
acknowledged true)", twice. The cause was `Run::applied`: a power step demanded the limit read back in watts, and
ADLX reports an offset, never watts, so no AMD power step could ever confirm. Fixed (bd7fcf4): a card with no limit
readback is applied on the acknowledgement. The watts themselves were read on every tick (`amd_watts_source=
engine_telemetry`, 363 nonzero engine samples, 456 nonzero app-state samples). The ladder reruns with kit-7. The
"100 W" in that line is the percent scale printed as watts (owed: a unit word for AMD).
After the run (nvidia-smi, the engine gone, the installed miners not yet back): 5090 1845 MHz locked at 575 W;
4070 at the 100 W limit. The installed app mined again within a minute: 5090 122.9 MH/s, 4070 28.8, 9070 XT 18.9,
170.6 MH/s, 0 faults (console, 16:44Z). The 9070 XT's ADLX state read "gmax 0 plimit 0 factory 0" (before: factory
1): zero offsets, the factory flag cleared by the engine's set of 0; the 3.0 coordinator's reset job puts it back.
### The re-point and the no-prompt proof (6 October 2026, 16:48 to 16:53Z)
`ember-repoint-pc1-1` (unelevated): task action before = the scratch copy; kit-7 (bd7fcf4) put in its place;
`Start-ScheduledTask`; `1 reregister` written; helper.log at 16:48:42Z: "1 reregister ok: the task now runs
C:\Users\Admin\AppData\Local\Programs\Igneum Miner\igneum-app.exe --power-helper"; action read back = the
installed exe. Three seconds later the installed app's own caps went through the task with no prompt (helper.log:
"nvidia-smi -i 0 -pl 460: Power limit for GPU 01:00.0 was set to 460.00 W from 575.00 W"; "-i 1 -pl 160: set to
160.00 W from 100.00 W"). `ember-proof-pc1-2` (reads only): task state Running, run level Highest, user Admin;
5090 221 W at 1845 MHz, 4070 75.8 W at 1860 MHz, 9070 XT 202 W; all three mining through the installed app. The
Security log's process-start audit is off on PC 1, so a consent.exe count is unreadable; the helper's log is the
record. Known in 0.3.13 on PC 1 until 0.3.14: the rows show the app's own earlier baseline (no /api/tune-progress
to receive the run's result), and the 80% caps are the app's setting; the draw is the tuned one regardless because
the driver holds the clock locks and the caps do not bind.
### The 0.3.14 window (6 October 2026, 18:18 to 18:27Z): no tune ran; five causes
The shipped 0.3.14 app, Power control on, the task registered, asked over its own API to tune each card
(`relay/playbooks/ember-installed-tune-pc1.ps1`). From its log and helper.log:
| | When (UTC) | What | Cause | Fix (0.3.15) |
|---|---|---|---|---|
| C | 17:55:25 | the app's own first 5090 tune wrote its first request in the second the helper started; the helper skipped it as "present at my start"; 30 s later "did not take (card reports 460 W, acknowledged true)" | the engine wrote before the helper was up | the helper writes a heartbeat (`helper.alive`) every poll; the engine starts the task and waits for a fresh heartbeat BEFORE writing (`powertask::ensure_running`) |
| A | 17:55:55 | the 4070's helper start read the stale `quit` from the 5090's abort and exited in the same second | quit/remove matched before the sequence check | the helper acts only on lines added after its start (`commands_after`) |
| D | 17:57:27 | the 9070 XT, measure only ("igneum-gpu-telemetry gave no tune line": the 3.0 coordinator's job was on the ADLX helper), was still sent a set; refused; "tuning stopped" | a measure-only card was set | `ember::may_set`: measure only sets nothing |
| E | 18:18 to 18:27 | three "Tune now" requests sat "queued" behind the hour's back-off from the 17:55 failures; none started | `sweep_retry` blocked a forced request silently | a forced request clears the back-off; a held row says "tuning: retry in N min (why)" (`ember::retry_note`) |
| B | (latent) | the engine's helper flag stayed true after the helper's idle exit; a later step would have been acknowledged after a blind 4 s sleep | a flag instead of a fact | the heartbeat decides; the acknowledgement is the helper's own log line for the sequence (up to 15 s), a missing line fails the step with that reason |
What a user on 0.3.14 sees: Power control on, no prompt, and no card tunes (the rows keep the baseline or
"tuning stopped"); the lever-2 TUNING records still flow. 0.3.15 closes it; the window repeats there.
### The 0.3.16 window (6 October 2026, 22:16 to 22:25Z): A to E held; a sixth cause (F) stopped the climbs
`ember-installed-pc1-2`, unelevated, the installed 0.3.16 app (the branch at 5429e82) tuning itself through the task.
Gate ok; prompts 0; the helper started once per NVIDIA card and stayed up (task state Running at the end).
| Card | Plan | Result |
|---|---|---|
| RX 9070 XT | baseline, measure (the probe gave no tune line again) | "Measured: 19.2 MH/s at 202 W (0.095 MH/W)", gclk 3298, mclk 2505, 62 C, ok; no set sent (D held). Open: why the installed app's probe gets no tune line when run 6's kit engine did |
| RTX 5090 | climb, 5 probes, floor 1390 MHz, mode helper | request 1 "the helper did not run sequence 1 within 15 s (no line in helper.log)"; 211.5 W before, 208 W after, unchanged |
| RTX 4070 | climb | the same; 74.8 W before, 75.5 W after |
Cause F (cmd.txt read back: `00 dev 1 / 01 pl 160 / 02 rgc / 03 rmc`): the tune path wrote its request index as the
wire sequence while the cap path had written unix-based numbers earlier in the same helper session; the helper runs
only numbers above every one it has seen, so every tune command read as stale. The honest acknowledgement (B) said
so within 15 s instead of a blind yes. Fix (cbd4d51, 0.3.17): one monotonic wire space for every writer
(`powertask::wire_seq`, the unix time modulo 999,990), the acknowledgement keyed to the wire number; test. The
elevated runs never crossed F (they set limits directly), which is why run 6 tuned and the self-tune did not.
## 8a. Next-cut notes (for the 0.3.12 shipper)
Correction (6 October 2026, 19:0xZ): release-0.3.15 took ember-tune at 5429e82 (merge 886c075), because Miner UI 4's
Cards tab ships in 0.3.15 and reads the per-card goal and the before fields; so the row below is IN 0.3.15 except
`card.tune_floor` (995530b), which rides 0.3.17.
Numbering (the shipper, 6 October 2026, 21:xx UK): 0.3.16 is a corrective cut of 0.3.15 under a new number (the Mac and PC 1 had taken earlier 0.3.15 builds), so everything below written as "0.3.16" ships as 0.3.17; the 0.3.15/0.3.16 content is the branch at 5429e82.
For 0.3.17 (main, 6 October 2026 evening), the engine fields Miner UI 4 reads, on ember-tune past 5429e82:
| Field | What | Where |
|---|---|---|
| `state.mining.watts_total` | the sum of the mining cards' draw (a reading under 60 s old) | `ember::fleet_watts`, engine `derive` + test |
| `state.mining.pounds_per_day` | that draw as £ a day at `settings.power_price_pence` (0 when no price) | `ember::pounds_per_day` |
| `state.address.balance_wei` | the payout address's balance in wei as a decimal string, `eth_getBalance` through the node's own RPC every 30 s while the node runs (60 s after a failure); null until read; `balance_age_s` (-1 until then), `balance_note` (the last error in words) | engine `tick_balance`, `Cmd::BalanceRead`, `ember::wei_from_hex` + test |
| `state.address.price_gbp_per_ign` | null. Its one source will be a SIGNED field of the OTA manifest (`price`: gbp_per_ign, as_of, source), checked like the tuning object; the app never computes or fetches a price itself | state.rs (documented), no code until a market exists |
| (F) the helper's wire sequence | one monotonic space for every writer of cmd.txt (`powertask::wire_seq`); the tune path's two-digit request index read as stale beside the cap path's unix-based numbers, so the 0.3.16 self-tune's climbs stopped at their first request | powertask.rs, engine.rs + test |
| Horizon polish Q4 (updater) | `fork_is_close`: an activation height at or below the DAA has passed, nothing is pending (the old rule read it as close: every update since the 0.3.14 manifest said "0 blocks away, installing now", stripped Later and skipped every guard; PC 1's 17:52:54Z install under a job came through it); `publish-manifest.sh` refuses an activation height at or below the live DAA (/api/live state.daa) unless `--allow-passed-activation` | manifest.rs + test, packaging/ota/publish-manifest.sh |
| Horizon polish Q83/Q84/Q2 (the pause shown) | `state.finality.paused`, `paused_since` (the last lock's time, else the engine's start), `reason`, `held_by`, `line` = "Finality paused since 18:39 UTC: under two thirds of the weight is signing" (the node's cause when it carries one: the engine parses a node log line carrying `finality_reason=<words or "quoted words"> held_by=<id>`, the node lane's to emit); the node line, the Overview's state and the Finality card show the one sentence while paused, the Finality card's age reads "paused", and no surface calls a lock final; `finality.message` carries the sentence too so older UIs show it | ember.rs `finality_paused_line`, `parse_finality_line`, `hhmm_utc` + tests; engine.rs derive and the node-line parse; ui/app.js `View.finalityWords` + test |
| the finality rule (updater) | Horizon frontier lane: the updater installs nothing while the network's finality is paused (a synced node with no checkpoint lock for `manifest::FINALITY_PAUSE_S` = 15 min; the last LOCK line's age, else the engine's uptime); the update card reads "waiting for finality: ..."; slot, catch-up and patience rules unchanged otherwise; only the signed manifest's own `urgent` flag installs through a pause (a fork-close or unsupported urgency does not); the known-failed case is the test | manifest.rs `Moment.finality_paused`, `manifest_urgent`, `Manifest.urgent`, `safe_to_apply` + test; ota.rs `Ctx`; engine.rs |
| `GET /api/live` | the observer's reply shape from a local source: `"source": "node"` when igneumd carries `igneum_getRecentBlocks(seconds)` (the node lane, a283f5f0d364ceef0; the engine computes miners_10m, blocks_10m, blocks_per_minute, the 90 s blocks and the miners list from it and takes the DAG numbers from its node state), else `"source": "site"` (the public reply, fetched by curl, cached 60 s) with `age_s`; `pending: true` before the first fetch | src/live.rs (`shape_from_blocks`, `parse_recent` + tests), server.rs |
For Miner UI 4's Cards tab (Ember as its second layer), 0.3.17, on ember-tune:
| Commit | What | Where |
|---|---|---|
| 5429e82 (in 0.3.15), 995530b (`tune_floor`, next cut) | a per-card goal: `POST /api/tune/goal {key, goal}` (efficiency, balanced, rate; "" or "global" clears), persisted as `CardPref.tune_goal`, echoed as `card.tune_goal` ("" = the global goal), resolved at the tune's start (`ember::goal_for`); `card.tune_floor` (the chosen clock is the ladder's floor; `CardPref.sweep_floor`); the untuned point of the last full plan on the row: `card.tune_before_watts`, `card.tune_before_mhs` (its step 0, kept across confirm plans; `CardPref.sweep_before_*`), so the row can read "saves 84 W, 0.15% of rate" against `sweep_watts` / `sweep_mhs`; Retune = `POST /api/sweep/start {key}` (forced: clears a back-off since 0.3.15) | ember.rs, engine.rs, server.rs, config.rs, state.rs, hotplug.rs + test |
For the 0.3.15 cut (6 October 2026 evening), on ember-tune (taken at 5429e82, see the correction above):
| Commit | What | Where |
|---|---|---|
| a8cab23 | (A) the helper acts only on lines added after its start; the acknowledgement is the helper's log line for the sequence | powertask.rs `commands_after` + test, engine.rs |
| (tip) | (B, C) the heartbeat, `ensure_running` before every task-path write; (D) `may_set`; (E) the back-off yields to a forced request, `retry_note`; the window playbook's gate at 0.3.15; `ember-tune-pc1.ps1` deleted (superseded by the installed-tune playbook) | powertask.rs, engine.rs, ember.rs + tests, relay/playbooks |
| e70ac5e | every task command goes to the helper's fixed folder | platform.rs, powertask.rs, engine.rs, main.rs + test |
For the 0.3.14 cut (6 October 2026, after run 6), on ember-tune:
| Commit | What | Where |
|---|---|---|
| (this cut) | every task command goes to the helper's fixed folder (`powertask::helper_dir` = `<platform data root>/app/sweep`, IGNEUM_APP_DATA ignored): a scratch-root measurement engine talking to the task wrote under its own root, where the task never reads (found 6 October 2026 planning the unelevated climb; the elevated runs set limits directly and never crossed it) | platform.rs `fixed_data_root`, powertask.rs, engine.rs `helper_cmd_dir`, main.rs + test |
| (this cut) | a baseline result's row reads "Measured: ..." not "Tuned: ..." (`ember::result_line`); PC 1's 0.3.13 rows read "Tuned: 114.2 MH/s at 305 W" for a baseline | ember.rs, engine.rs + test |
| bd7fcf4 | AMD: a power step with no limit readback is applied on the acknowledgement (run 6 aborted the 9070 XT's ladder at step 1 on "card reports 0 W") | ember.rs `Run::applied` + test |
| 200362a | `task_exe` prefers the installed exe for every path that is not itself an install candidate; the helper's `reregister` verb; the clock ladder and floor to 45%; a chosen point on the floor says so; the fleet page's team rows | powertask.rs, ember.rs, engine.rs, site |
| 5165de7, 9fba504, 5b0968f | Ember 2's Settings (goal, price, climb), £/day and the curve on the card; the playbook's AMD watts source and kit-7 preference; run 6 notes | ui, relay/playbooks, docs |
Earlier (shipped in 0.3.12 and 0.3.13 unless marked):
| Commit | What | Where |
|---|---|---|
| b671c8b | every `quit:` names its source; Power control alone decides; no cap at start under `--sweep` | main.rs, server.rs, engine.rs (separable) |
| e600e63 | a second engine never runs the updater (`IGNEUM_APP_NO_OTA`, implied by `--sweep`) | engine.rs (6 lines, separable) |
| 1e9550e | the elevated job path's output file is followed while the script runs, so the 5-minute progress reports carry its lines (a 35-minute run that never mined showed only "script running" on 6 October 2026); the tune playbook's watchdog fails a run that mines nothing within 120 s of its first status line, with the engine's last log line in the RESULT | jobrun.rs `follow_file`, relay/playbooks/ember-tune-pc1.ps1 |
## 9. Open
- The NVIDIA clock readback: `nvidia-smi -lgc` is confirmed only through the core clock during the hold (a mean over
the cap by 5% marks the step `unapplied`); the first run with Power control on tells whether the driver honours
the lock on the 5090 under this kernel.
- ADLX on RDNA 4 exposes no memory-clock setter; the memory-clock mark is the guard. The telemetry agent's 9070 XT
sweep tells whether a core cap drags the memory clock on that card.
- The confirm plan's neighbour is one step; a second neighbour (the other knob) would cost 75 s more and catch a
prior that is wrong on both knobs.
- Intel: no knob yet; the row says measure only.