igneum/docs/plans/miner-eff.md
igneum-josh b80d4c9690 Miner efficiency sweep: hash per watt per NVIDIA card (lever 3)
src/sweep.rs (new): the cap steps 100% to 50% in 10% steps clamped to the card's limits, per-step rows (mean draw
from nvidia-smi power.draw, mean worker interval rate), the choice (best MH/W, ties to the higher rate then the lower
cap), the nvidia-smi power parser, a state machine on an explicit clock (15 s settle, 60 s hold, 30 s cap readback
limit), the elevated helper scripts (one administrator prompt per sweep: a command file polled by one elevated
process, self-restoring after 20 idle minutes), the unsupported reasons (Apple silicon, AMD). 9 unit tests with
PC 1's recorded RTX 5090 numbers (575 W default, 460 W cap, 290 W draw, memory temperature [N/A]).

Engine: scheduler (once after install, then weekly; one card at a time; only while the card mines, after 120 s
steady, never under a remote job hold, a pause, or inside 600 s of the hour boundary), the cap-mode probe (direct
when the engine runs elevated, else the helper), abort on any fault (card leaves mining, worker error, GPU 90 C,
job, pause, quit) with the cap restored, the chosen cap held and recorded, SWEEP table lines in the app log,
--sweep mode (sweep every supported card, print the table on stdout, leave the caps, quit). Cap floor 50% (was 60).
A readback that matches the asked cap now counts as applied (PC 1 showed "cap NOT applied" for hours at 460 W).

Dashboard: live eff MH/W on each tile, the sweep line (phase, last result, or why unsupported), Sweep now / Stop /
Unpin, "pinned" and "chosen by the sweep" on the cap line, the Settings toggle, the cards-page note, slider min 50.
A cap moved by hand pins the card: the sweep records but does not change it.

PC 1 measurement: relay/playbooks/sweep-5090.ps1 (a run job, elevated, miners stopped; a second engine with --sweep
in a scratch data folder, RESULT SWEEP lines) and docs/plans/miner-eff.md with the publish command. Not published.
Untested on a card.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 21:38:07 +01:00

9.7 KiB

Hash per watt: the efficiency sweep (lever 3)

4 October 2026, evening. Josh: "make our miner better than anything else can be". Miners pay for electricity; the number they compare is MH per watt. The app already caps NVIDIA cards (default 80% of the default limit, nvidia-smi -pl, readback, Retry). This adds the sweep that finds the best cap per card, the live MH/W on every tile, a --sweep mode for the engine, and the PC 1 measurement job. Branch miner-eff, worktree ../igneum-wt-eff.

What is built

Piece Where State
Sweep logic: step plan (100% to 50% in 10% steps, clamped to the card's min and max, duplicate watts dropped), per-step rows, the choice (best MH/W; within 1% the higher rate; within 1% on both the lower cap), the nvidia-smi power parser, the state machine on an explicit clock, the elevated helper scripts, the unsupported reasons app/igneum-app/src/sweep.rs 9 unit tests pass on the Mac (36 in the crate)
Engine: scheduler (one card at a time, once after install then weekly, never under a remote job hold, a pause, or inside 600 s of the hour boundary; the card must have mined 120 s), the cap-mode probe (direct when the engine is elevated, else one elevated helper polling a command file: one administrator prompt per sweep, not six), fault abort (card leaves "mining", worker error line, GPU 90 C, job, pause, quit), restore on abort, record and hold on finish, --sweep mode, the table in the app log and on stdout app/igneum-app/src/engine.rs (tick_sweep, sweep_*, Cmd::Sweep*) compiles; not run on a card
Settings: sweep (default on), installed_at, per card pinned, sweep_at, sweep_pct, sweep_eff, sweep_watts, sweep_mhs; cap floor 50% (was 60) config.rs, detect.rs, state.rs, server.rs (/api/sweep/start, stop, pin, enable)
Dashboard: eff <MH/W> on the tile's telemetry line (live hash over draw), the sweep line under it (phase while running, else "best N% · x MH/W (rate at watts, when)", or why unsupported), Sweep now / Stop / Unpin, the cap line says "pinned" or "chosen by the sweep"; Settings toggle "Find each NVIDIA card's best MH per watt (once after install, then weekly)"; the cards page note; slider min 50 with a "pinned" / "sweep N%" tag app/igneum-app/ui/
Readback self-heal: a cap that matches what was asked counts as applied when the telemetry reads it back (PC 1's log said "cap NOT applied" for hours while the limit read 460 W) engine.rs telemetry_line
The PC 1 job script relay/playbooks/sweep-5090.ps1 parse-checked by windows.yml; UNTESTED on a PC

The window hosts (Swift, WebView2) are untouched: no host string changes.

The log lines

SWEEP start card=nvidia-ae432dc7-1 name=NVIDIA_GeForce_RTX_5090 steps=100,90,80,70,60 default=575 min=400 max=600 before=460 mode=direct
SWEEP card=nvidia-ae432dc7-1 cap=100 limit=575 watts=290.4 mhs=124.10 eff=0.4273
SWEEP card=nvidia-ae432dc7-1 cap=90 limit=518 watts=290.1 mhs=124.00 eff=0.4274
...
SWEEP chosen card=nvidia-ae432dc7-1 cap=70 limit=403 watts=289.9 mhs=123.90 eff=0.4274
SWEEP aborted card=nvidia-ae432dc7-1 reason=the_card_left_mining_(worker_restart,_likely_the_hour_boundary)

watts is the mean draw during the 60 s hold (what the miner pays for), limit the cap set, mhs the mean of the worker's interval rates (now= in the STATUS lines) during the hold, eff = mhs / watts. A step with under 3 draw readings or no rate prints eff=0 reason=no_readings and cannot win. The numbers above are the shape, not a measurement.

How a sweep runs

  1. Scheduler (every 10 s): a card is due when sweep is on, it is not pinned and its last sweep is older than 7 days (or never). "Sweep now" and --sweep queue a card regardless. It starts only while the card is mining, its worker has run 120 s, no remote job holds the GPU, nothing is paused, and over 600 s remain to the hour boundary (the tile says what it waits for).
  2. Probe: one nvidia-smi -i <dev> -pl <current limit> from the engine. "All done" means this process may set caps (it runs elevated, as the PC job does); otherwise the helper starts: powershell -File <app>\sweep\helper.ps1 through the existing run_elevated (one UAC prompt), polling <app>\sweep\cmd.txt twice a second for <seq> <watts> lines, quit to end, and restoring the entry limit by itself after 20 idle minutes.
  3. Each step: set the cap, wait for the telemetry readback (power.limit, every 5 s) to match (30 s limit), settle 15 s, hold 60 s collecting draws and rates, write the row. Then the chosen cap is set and read back, the card and the settings record it (power_pct becomes the chosen percent unless the card is pinned), the helper is told to quit. The worker is never restarted; the hour's program is never lost.
  4. Abort: the cap goes back to the limit in force before the sweep, SWEEP aborted is logged, the card waits an hour before the scheduler tries again (under --sweep: 30 s, up to 3 attempts).

IGNEUM_APP_SWEEP_FAST=1 makes a step 2 s settle + 6 s hold (plumbing check only; its numbers mean nothing).

The PC 1 measurement (do not publish from this branch; the main session schedules PC jobs)

The job runs relay/playbooks/sweep-5090.ps1 elevated with the installed app's miners stopped and held. The script starts a second engine from the same install with --sweep in %LOCALAPPDATA%\igneum-sweep (the real settings.json, machine-id and wallet.json copied in; remote jobs, auto-update and proving switched off in the copy so the second engine cannot run this job again or update itself). That engine finds the installed app's node on 127.0.0.1:26610 (external), exports the pack, mines on the 5090 with --status-secs 10 (6 rate samples per hold), sweeps, prints the table on stdout, leaves the chosen cap in force and quits. The script re-emits every SWEEP line as RESULT SWEEP ..., adds RESULT SWEEP before ... / after ... (the nvidia-smi limits around the run) and the sweep engine's log tail. Budget 40 minutes (worst case: 10 minutes waiting for the hour boundary, 2 minutes steady, 5 steps of about 80 s). The installed app's miners restart when the job ends.

Needs: the build that carries src/sweep.rs installed on PC 1 (ship it with tools/ship-app.mjs; the OTA applies it, or publish-jobs.sh add --kind update-now --target ae432dc7). An older engine ignores --sweep and the script reports no_rows.

Publish command (from the repo root, the main session, after the version is on PC 1):

packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --platform windows --requires nvidia \
    --script relay/playbooks/sweep-5090.ps1 --elevated --stop-miners --timeout-minutes 45 \
    --title "Efficiency sweep on the RTX 5090 (PC 1)" --deploy

Reading it back:

node tools/jobs.mjs watch <job id>        every 30 s until final; the RESULT SWEEP lines are the table
node tools/jobs.mjs <job id>              the latest report

What PC 1's own log already says (4 October 2026, run win-ae432dc7-20261004-164723)

stability: NVIDIA GeForce RTX 5090: draw p95 290 W, max 295 W, limit 460 W (cap NOT applied), max GPU 65 C, max memory 0 C

The 5090 draws 290 W under a 460 W limit (80% of 575 W). The cap does not bind at any step down to the card's floor (the 5090's power.min_limit is expected near 400 W; the sweep reads the real value and clamps). Expect a flat table: the same draw and rate at every step, and the choice falling to the lowest cap by the tie rule. That is a result in itself: for this kernel the power cap is a safety, not an efficiency lever. The lever would be the clocks (nvidia-smi -lgc / -lmc, core and memory offsets), which this sweep does not touch; a clock sweep is the next step once the table confirms the flat line. "memory 0 C" is temperature.memory reading [N/A] on this driver.

Unit tests (cargo test -p igneum-app sweep)

  • plan_clamps_to_the_card_floor_and_drops_duplicates: 575 / 400 / 600 W plans 100, 90, 80, 70, 60 (60% and 50% both clamp to 400 W; one kept).
  • parses_nvidia_smi_power_draw: the recorded telemetry shape from PC 1 (0, 290.12), CRLF, [N/A], with units and a header, the driver-failure text, empty.
  • chooses_the_best_mh_per_watt_then_rate_then_the_lower_cap: a card that improves then collapses; PC 1's flat case; the 1% ties; unusable rows never win.
  • row_lines_carry_the_table_fields, state_machine_runs_a_sweep (five steps with a 4 s readback lag, 395 to 420 s), state_machine_fails_when_a_cap_never_takes, state_machine_fails_without_readings, unsupported_reasons, helper_scripts_carry_the_protocol.

Untested

  • The real sweep on a card: the probe, the helper and its UAC prompt, the readback timing, the table. Only the PC job above measures it.
  • The second-engine arrangement of the job (two engines from one install, one holding its miners).
  • The dashboard rendering (the UI was edited by hand; node --check passes; the Rust side serialises the new fields).
  • The Windows build: checked on macOS only (cargo test in the crate); the Windows-only paths are runtime cfg! branches, so they compile here too, but the GitHub runner has not built this branch.
  • AMD on Windows: reported unsupported with the reason (no power reading or cap in the app). Apple silicon: unsupported (no cap to set; powermetrics needs root).

Merge notes for the other miner agents

Touched outside the new module: engine.rs (new Cmd variants, engine fields, the sweep block before the tick, one line each in miner_line STATUS and telemetry_line, derive, shutdown, the apply_power_limits guard, the 50% clamps), config.rs, state.rs, detect.rs, server.rs, main.rs, ui/. Nothing in the workers, the manifest, the payout, the template path or the fault guards.