Ember Tune: every card tuned for MH per watt out of the box, the fleet prior per card model in the signed manifest, the console and /miners priors table

the project lead, 5 October 2026, 22:45 BST: "make sure we have ember tuning every single card for efficiency out of the box, the
more data = the better the tune, make an awesome system." Built on lever 3 (docs/plans/miner-eff.md), lever 2's signed
tuning section (docs/design/miner-tuning.md), the AMD telemetry helper (35e3d26, its --tune/--set-gmax/--set-plimit/
--reset contract) and the Power control switch (652e848). Design, data flow, tiers and the privacy line:
docs/plans/ember-tune.md.

- src/ember.rs (new): two knobs per card (power limit %, core clock cap MHz; memory clock never touched), the full plan
  (power ladder 100..50%, then the clock ladder 90..60% at the chosen power), the confirm plan (the fleet prior and one
  neighbour), the baseline plan (measure only), the marks (faulted, hot, memory_clock_dropped, unapplied, no_readings),
  the choice (best MH/W within 1% of the top rate, then rate, then draw), the fleet record (a hash of the install id,
  no address), the prior lookup and the kill switch (tuning.ember), the state machine on a fake clock. 9 unit tests.
- engine.rs: tick_sweep schedules every NVIDIA, AMD and Apple card (120 s steady, 600 s to the boundary, no job hold,
  no pause, weekly, again after a driver major or program-class change, never under the manifest kill switch); the
  probe (nvidia-smi clocks.max.gr + driver_version and the direct/helper mode; igneum-gpu-telemetry --tune for AMD);
  tune_apply (nvidia-smi -pl / -lgc 0,<MHz> / -rgc directly or through the helper; the AMD helper per request);
  Cmd::TuneProbe, Cmd::TuneSet; faults from rejected and mismatched hashes mark the step; the TUNE lines and the TUNE
  {json} record, uploaded with the log; the Tuned line on the card state. The NVIDIA helper starts only with Power
  control on: the --sweep job never counts as permission (no prompt on a PC with nobody there).
- sweep.rs: the helper protocol gains lgc/rgc (clock cap and reset) and resets the clocks after 20 idle minutes.
- state.rs, config.rs: the tune fields (clock cap, driver, class, source, the Tuned line); the nvidia-smi telemetry
  query carries clocks.gr and clocks.mem; the AMD sample line's plimit_pct and gmax_mhz are parsed.
- ui: "Tuned: X MH/s at Y W (Z MH/W)" with the point, the source and when; measure-only cards say why; the Ember Tune
  switch; tune-line.test.mjs.
- relay/lib/ember.mjs + relay/test/ember.test.mjs: the aggregation per (card model | driver major | program class):
  median point, MH/W, spread, samples, machines; five samples converge, an outlier does not move the median, baselines
  make no prior, de-duplication, the manifest merge keeps lever 2's cards. api/console.mjs fn=tuning and
  tools/console.mjs tuning; tools/tuning.mjs --priors [--write tuning.json] [--site] [--tuning-off].
- site: the fleet priors table on /miners (site/miner-priors.json), the lever text.
- relay/playbooks/ember-tune-pc1.ps1: the PC 1 run (second engine with --sweep from a scratch copy of the install).

Measured tonight: see the bench log entry that follows the PC 1 run. The 9070 XT left PC 1's bus at 20:40 UTC and the
5090 needs the administrator prompt the project lead cannot answer asleep, so tonight's PC 1 run is the baseline plan on the 5090
through the whole pipeline; the two-knob tune on both cards is owed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-05 21:21:51 +00:00
parent b9f3becc36
commit 735887a309
22 changed files with 2354 additions and 184 deletions

View file

@ -81,6 +81,6 @@ jobs:
- name: ship tool self-test (version bump, the dl-both and public manifest helpers)
run: node tools/ship-app.mjs --self-test
- name: relay unit tests (parsers, secret compare, the wake endpoint)
run: node --test relay/test/parse.test.mjs relay/test/auth.test.mjs relay/test/wake.test.mjs
run: node --test relay/test/parse.test.mjs relay/test/auth.test.mjs relay/test/wake.test.mjs relay/test/ember.test.mjs
- name: miner app notice strip and update card (ordering, keys, wording, timers, when the card shows)
run: node --test app/igneum-app/ui/notices.test.mjs app/igneum-app/ui/update-card.test.mjs
run: node --test app/igneum-app/ui/notices.test.mjs app/igneum-app/ui/update-card.test.mjs app/igneum-app/ui/tune-line.test.mjs

View file

@ -29,6 +29,16 @@ pub struct CardPref {
pub sweep_watts: f64,
#[serde(default)]
pub sweep_mhs: f64,
/// Ember Tune (src/ember.rs): the clock cap the last tune chose (0 = unlocked), the driver and program class it
/// ran under (a change makes the card due again), and the plan that produced it (full | confirm | baseline)
#[serde(default)]
pub sweep_clock_mhz: u32,
#[serde(default)]
pub sweep_driver: String,
#[serde(default)]
pub sweep_class: String,
#[serde(default)]
pub sweep_source: String,
}
#[derive(Clone, Serialize, Deserialize)]

View file

@ -66,6 +66,7 @@ fn card(index: usize, name: &str, vendor: &str, worker: &str, detail: &str, devi
device: device.into(),
enabled: true,
state: "off".into(),
amd_ordinal: -1,
..Default::default()
}
}

1033
app/igneum-app/src/ember.rs Normal file

File diff suppressed because it is too large Load diff

View file

@ -65,13 +65,13 @@ pub enum Cmd {
SweepEnable(bool),
/// Settings > Power control (config.rs power_control): on asks for administrator rights once, at that moment
PowerControl(bool),
/// the probe that decides how caps are set during a sweep: Ok(true) = nvidia-smi -pl works from this process
/// (the engine runs elevated), Ok(false) = the elevated helper is needed; Err = nvidia-smi did not answer
SweepCapMode(Result<bool, String>),
/// Ember Tune's probe of a card (src/ember.rs): the vendor's clock range, the driver, and how settings reach the
/// card (direct, the elevated helper, the AMD helper, or measure only)
TuneProbe(usize, Result<TuneProbe, String>),
/// the elevated helper process ended (Ok at `quit`, Err when the prompt was cancelled or it failed)
SweepHelperDone(Result<(), String>),
/// a direct `nvidia-smi -pl` for the sweep finished: what it printed
SweepCapSet(String),
/// a tune request (by number) was carried out: what the vendor tool printed, or why it refused
TuneSet(u64, Result<String, String>),
/// restart the node with the verifier decided again (src/verifier.rs): the trust setting changed, or the
/// prover found a host that was not there when the node started
RestartNode(String),
@ -115,7 +115,7 @@ impl Shared {
st.mining.accepted_total = settings.accepted_total;
st.mining.fee_total = settings.fee_total;
st.address = address_state(&settings, &wallet_path);
st.settings = crate::state::SettingsState { identities: settings.identities, vote: settings.vote, start_at_login: crate::platform::start_at_login_is_on(), auto_update: settings.auto_update, remote_jobs: settings.remote_jobs, prove: settings.prove, sweep: settings.sweep, power_control: settings.power_control, power_note: String::new(), dev_fee: settings.dev_fee, proof_verify_trust: settings.proof_verify_trust };
st.settings = crate::state::SettingsState { identities: settings.identities, vote: settings.vote, start_at_login: crate::platform::start_at_login_is_on(), auto_update: settings.auto_update, remote_jobs: settings.remote_jobs, prove: settings.prove, sweep: settings.sweep, power_control: settings.power_control, power_note: String::new(), tuning_off: false, tuning_note: String::new(), dev_fee: settings.dev_fee, proof_verify_trust: settings.proof_verify_trust };
st.dev_fee = crate::state::DevFeeState { on: settings.dev_fee, percent: if settings.dev_fee { 1 } else { 0 }, address: String::new(), line: String::new() };
st.live_page = packaged.live_page.clone();
st.finality.message = "waiting for the miner".into();
@ -462,7 +462,11 @@ pub struct Engine {
last_settings_save: Instant,
last_error_event: Instant,
// the efficiency sweep (src/sweep.rs): one card at a time
sweep: Option<crate::sweep::Run>,
sweep: Option<crate::ember::Run>,
/// the request number the vendor tool last carried out (the run's acknowledgement)
tune_acked: Option<u64>,
/// cards whose confirm check found a better neighbour: the full plan runs next
tune_full_due: std::collections::HashSet<usize>,
/// a sweep waiting for the cap-mode probe: (card index, forced)
sweep_pending: Option<(usize, bool)>,
/// how caps are set: Some(true) directly, Some(false) through the elevated helper, None = not probed yet
@ -555,6 +559,8 @@ impl Engine {
last_settings_save: now,
last_error_event: now - Duration::from_secs(600),
sweep: None,
tune_acked: None,
tune_full_due: std::collections::HashSet::new(),
sweep_pending: None,
sweep_direct: None,
sweep_helper: false,
@ -679,6 +685,12 @@ impl Engine {
c.sweep_watts = p.sweep_watts;
c.sweep_mhs = p.sweep_mhs;
c.sweep_at = p.sweep_at as f64;
c.tune_clock_mhz = p.sweep_clock_mhz;
c.clock_cap_mhz = if p.pinned || p.sweep_source == "baseline" { 0 } else { p.sweep_clock_mhz };
c.tune_source = p.sweep_source.clone();
if p.sweep_mhs > 0.0 && p.sweep_watts > 0.0 {
c.tune_line = crate::ember::tuned_line(p.sweep_mhs, p.sweep_watts, p.sweep_eff);
}
}
}
let supported: Vec<String> = st.mining.cards.iter().filter(|c| c.sweep_supported && c.enabled).map(|c| c.key.clone()).collect();
@ -912,19 +924,18 @@ impl Engine {
let (name, ok, why) = {
let st = self.st();
match st.mining.cards.iter().find(|c| c.key == key) {
Some(c) if !c.sweep_supported => (c.name.clone(), false, c.sweep_note.clone()),
Some(c) if !matches!(c.vendor.as_str(), "nvidia" | "amd" | "apple") => (c.name.clone(), false, "no power or clock control for this card".to_string()),
Some(c) if !c.enabled => (c.name.clone(), false, "the card is switched off".to_string()),
Some(c) if !self.elevation_allowed() => (c.name.clone(), false, "switch Power control on in Settings first (Windows asks for administrator rights once)".to_string()),
Some(c) => (c.name.clone(), true, String::new()),
None => (key.clone(), false, "no such card".to_string()),
}
};
if !ok {
self.shared.event("error", &format!("{name}: no sweep: {why}"));
self.shared.event("error", &format!("{name}: no tune: {why}"));
} else if !self.sweep_queue.contains(&key) {
self.sweep_queue.push(key);
self.sweep_next_check = Instant::now();
self.shared.event("info", &format!("{name}: efficiency sweep queued; it starts once the card has mined for {} s with over {} s to the hour boundary", crate::sweep::STABLE_S, crate::sweep::NEEDS_S));
self.shared.event("info", &format!("{name}: tune queued; it starts once the card has mined for {} s with over {} s to the hour boundary", crate::ember::STABLE_S, crate::ember::NEEDS_S));
}
}
Cmd::SweepStop => {
@ -951,11 +962,10 @@ impl Engine {
}
Cmd::SweepEnable(on) => {
let allowed = self.elevation_allowed();
let on = on && allowed;
self.shared.settings.lock().unwrap().sweep = on;
self.shared.save_settings();
self.st().settings.sweep = on;
self.shared.event("info", if on { "efficiency sweep on: once after install, then weekly" } else if allowed { "efficiency sweep off; Sweep now on a card still runs one" } else { "efficiency sweep stays off: switch Power control on in Settings first" });
self.shared.event("info", if on && allowed { "Ember Tune on: every card is tuned for MH per watt once after install, then weekly" } else if on { "Ember Tune on: AMD cards are tuned; NVIDIA cards measure only until Power control is on in Settings" } else { "Ember Tune off; Tune now on a card still runs one" });
}
Cmd::PowerControl(on) => {
{
@ -983,7 +993,7 @@ impl Engine {
self.power_control_off("power control off; the cap and the sweep do not run and nothing asks for administrator rights");
}
}
Cmd::SweepCapMode(r) => self.sweep_mode_known(r),
Cmd::TuneProbe(idx, r) => self.sweep_probe_known(idx, r),
Cmd::SweepHelperDone(r) => {
self.sweep_helper = false;
if let Err(e) = &r {
@ -997,11 +1007,20 @@ impl Engine {
self.shared.log("sweep helper: exited");
}
}
Cmd::SweepCapSet(text) => {
if !text.contains("All done") {
self.shared.log(&format!("sweep: nvidia-smi -pl answered: {}", short(&text.replace('\n', " "), 200)));
Cmd::TuneSet(seq, r) => match r {
Ok(text) => {
self.tune_acked = Some(seq);
if !text.contains("All done") && text != "helper" {
self.shared.log(&format!("tune: request {seq} answered: {}", short(&text.replace('\n', " "), 200)));
}
}
}
Err(text) => {
self.shared.log(&format!("tune: request {seq} refused: {}", short(&text.replace('\n', " "), 300)));
if self.sweep.as_ref().map(|r| r.seq == seq).unwrap_or(false) {
self.sweep_abort("the card refused the setting (see the log)");
}
}
},
Cmd::Quit => {
self.quitting = true;
self.st().quitting = true;
@ -1561,7 +1580,7 @@ impl Engine {
}
}
if self.telemetry.is_none() && now >= self.telemetry_retry_at {
let args: Vec<String> = ["--query-gpu=index,power.draw,temperature.gpu,temperature.memory,power.limit", "--format=csv,noheader,nounits", "-l", "5"].iter().map(|s| s.to_string()).collect();
let args: Vec<String> = ["--query-gpu=index,power.draw,temperature.gpu,temperature.memory,power.limit,clocks.gr,clocks.mem", "--format=csv,noheader,nounits", "-l", "5"].iter().map(|s| s.to_string()).collect();
let log = self.shared.runtime.log_dir.join(format!("gpu-{}.log", self.stamp));
match procs::spawn(Source::Telemetry, &crate::platform::tool("nvidia-smi"), &args, None, &log, &self.lines_tx, &[]) {
Ok(p) => self.telemetry = Some(p),
@ -1635,8 +1654,20 @@ impl Engine {
c.fan_pct = s.fan_pct.max(0.0);
c.fan_rpm = s.fan_rpm.max(0.0);
c.mclk_mhz = s.mclk_mhz.max(0.0);
c.gclk_mhz = s.gclk_mhz.max(0.0);
c.util_pct = s.util_pct.max(0.0);
c.amd_ordinal = s.ordinal as i64;
if s.gmax_mhz > 0.0 {
c.clock_cap_mhz = s.gmax_mhz as u32;
}
c.telemetry_at = crate::platform::unix_now_f();
let (idx, draw, gclk, mclk, temp) = (c.index, s.watts, s.gclk_mhz, s.mclk_mhz, s.temp_c);
drop(st);
if let Some(run) = self.sweep.as_mut() {
if run.card == idx {
run.sample_telemetry(draw.max(0.0), gclk.max(0.0), mclk.max(0.0), temp.max(0.0));
}
}
}
/// "index, draw, gpu temp, mem temp, limit" every 5 s.
@ -1647,12 +1678,19 @@ impl Engine {
}
let f = |s: &str| s.parse::<f64>().unwrap_or(0.0);
let (draw, tgpu, tmem, limit) = (f(p[1]), f(p[2]), f(p[3]), f(p[4]));
let (gclk, mclk) = (p.get(5).map(|v| f(v)).unwrap_or(0.0), p.get(6).map(|v| f(v)).unwrap_or(0.0));
let sweeping = self.sweep.is_some() || self.sweep_pending.is_some();
let mut st = self.st();
let Some(c) = st.mining.cards.iter_mut().find(|c| c.vendor == "nvidia" && c.device == p[0]) else { return };
c.power_w = draw;
c.temp_gpu = tgpu;
c.temp_mem = tmem;
if gclk > 0.0 {
c.gclk_mhz = gclk;
}
if mclk > 0.0 {
c.mclk_mhz = mclk;
}
if limit > 0.0 {
c.power_limit_w = limit;
// the readback is the truth: a cap the elevated step set after the judgement (or one set by hand
@ -1668,7 +1706,7 @@ impl Engine {
drop(st);
if let Some(run) = self.sweep.as_mut() {
if run.card == idx {
run.sample_draw(draw);
run.sample_telemetry(draw, gclk, mclk, tgpu);
}
}
if mining {
@ -1708,9 +1746,9 @@ impl Engine {
}
}
// ---- the efficiency sweep (src/sweep.rs) ---------------------------------------------------------------------
// ---- Ember Tune (src/ember.rs): every card tuned for MH per watt ---------------------------------------------
/// A SWEEP line: the app log, and stdout under --sweep (the PC job reads it there).
/// A TUNE line: the app log, and stdout under --sweep (the PC job reads it there).
fn sweep_say(&self, line: &str) {
self.shared.log(line);
if self.shared.runtime.sweep_only {
@ -1723,8 +1761,17 @@ impl Engine {
self.shared.runtime.app_dir.join("sweep")
}
/// Every 10 s: drive the running sweep, else start one that is due. One card at a time. Never while a remote
/// job holds the GPU, never under a pause, never inside the last 10 minutes before the hour boundary.
/// The manifest's tuning object (<app data>/tuning.json, written by the updater from the signed manifest), or
/// None: the Ember settings (kill switch, thresholds) and the fleet priors live under it.
fn tuning_object(&self) -> Option<Value> {
let p = self.ota.tuning_path()?;
std::fs::read_to_string(p).ok().and_then(|t| serde_json::from_str(&t).ok())
}
/// Every 10 s: drive the running tune, else start one that is due. One card at a time. Never while a remote
/// job holds the GPU, never under a pause, never inside the last 10 minutes before the hour boundary, never
/// while the manifest's kill switch is off. A card is due once after install, then every period, and again
/// after a driver or program-class change; a pinned card is skipped; "Tune now" queues one regardless.
fn tick_sweep(&mut self, now: Instant) {
if self.sweep.is_some() {
self.sweep_drive(now);
@ -1737,12 +1784,19 @@ impl Engine {
if self.quitting || !self.running || self.power_busy || self.job_hold || self.jobs.holds_miners() {
return;
}
let (auto_on, paused, allowed) = {
let tuning = self.tuning_object();
let ember = crate::ember::settings_of(tuning.as_ref());
let (auto_on, paused, power_control) = {
let s = self.shared.settings.lock().unwrap();
let st = self.st();
(s.sweep, st.mining.paused, elevation_allowed(s.power_control, self.shared.runtime.sweep_only))
};
if paused || !allowed {
{
let mut st = self.st();
st.settings.tuning_off = !ember.enabled;
st.settings.tuning_note = if ember.enabled { String::new() } else { "tuning paused fleet-wide by the signed manifest".into() };
}
if paused || !ember.enabled {
return;
}
let unix = crate::platform::unix_now();
@ -1750,35 +1804,44 @@ impl Engine {
let st = self.st();
(st.node.daa > 0 && st.program.boundary_daa > 0, st.program.eta_s)
};
let boundary_ok = !eta_known || eta_s > crate::sweep::NEEDS_S;
let boundary_ok = !eta_known || eta_s > crate::ember::NEEDS_S;
let prefs = self.shared.settings.lock().unwrap().cards.clone();
let cards = self.st().mining.cards.clone();
let mut pick: Option<(usize, bool)> = None;
for c in cards.iter() {
if !(c.enabled && c.sweep_supported && c.vendor == "nvidia") {
if !(c.enabled && matches!(c.vendor.as_str(), "nvidia" | "amd" | "apple")) {
continue;
}
let forced = self.sweep_queue.contains(&c.key);
let due = prefs.get(&c.key).map(|p| p.sweep_at == 0 || unix.saturating_sub(p.sweep_at) >= crate::sweep::PERIOD_S).unwrap_or(true);
if !forced && !(auto_on && !c.pinned && due) {
let due = match prefs.get(&c.key) {
None => true,
Some(p) => p.sweep_at == 0 || unix.saturating_sub(p.sweep_at) >= ember.period_s || (!c.driver.is_empty() && !p.sweep_driver.is_empty() && crate::ember::driver_major(&c.driver) != crate::ember::driver_major(&p.sweep_driver)) || (!c.program_class.is_empty() && !p.sweep_class.is_empty() && c.program_class != p.sweep_class),
};
let auto = auto_on || !c.tune_control;
// a measure-only card (Apple, NVIDIA without Power control) measures its baseline once a period even
// with the switch off: the row should say what the card does
if !forced && !(auto && !c.pinned && due) {
continue;
}
if c.tune_control && c.vendor == "nvidia" && !power_control && !forced {
// measure-only until Power control is on; the baseline plan covers it below
}
if self.sweep_retry.get(&c.index).map(|t| now < *t).unwrap_or(false) {
continue;
}
let stable = self.miners.iter().find(|m| m.card == c.index).map(|m| m.proc.as_ref().map(|p| p.started.elapsed() >= Duration::from_secs(crate::sweep::STABLE_S)).unwrap_or(false) && m.last_status.is_some()).unwrap_or(false);
let stable = self.miners.iter().find(|m| m.card == c.index).map(|m| m.proc.as_ref().map(|p| p.started.elapsed() >= Duration::from_secs(crate::ember::STABLE_S)).unwrap_or(false) && m.last_status.is_some()).unwrap_or(false);
let wait = if c.state != "mining" {
Some(format!("sweep: waits for the card to mine ({})", c.state))
Some(format!("tuning: waits for the card to mine ({})", c.state))
} else if !stable {
Some(format!("sweep: waits for {} s of steady mining", crate::sweep::STABLE_S))
Some(format!("tuning: waits for {} s of steady mining", crate::ember::STABLE_S))
} else if !boundary_ok {
Some(format!("sweep: starts after the hour boundary ({eta_s} s)"))
Some(format!("tuning: starts after the hour boundary ({eta_s} s)"))
} else {
None
};
match wait {
Some(w) => {
if forced {
if forced || c.sweep_pct == 0 {
if let Some(cc) = self.st().mining.cards.get_mut(c.index) {
cc.sweep_note = w;
}
@ -1797,8 +1860,10 @@ impl Engine {
}
}
/// Step 1 of a sweep: find out how caps can be set from this process (one nvidia-smi -pl at the current limit,
/// a no-op). Elevated engines (the PC job runs with --elevated) set them directly; others need the helper.
/// Step 1 of a tune: the probe. NVIDIA: the vendor's maximum core clock and driver (nvidia-smi query), and
/// whether this process may set limits (one `-pl` at the current limit: "All done" = direct; else the elevated
/// helper, only with Power control on; else measure only). AMD: the helper's `--tune` line (the clock and power
/// ranges in force; no elevation). Apple: measure only.
fn sweep_begin(&mut self, idx: usize, forced: bool) {
let Some(c) = self.st().mining.cards.get(idx).cloned() else { return };
self.sweep_pending = Some((idx, forced));
@ -1806,74 +1871,161 @@ impl Engine {
*self.sweep_attempts.entry(idx).or_default() += 1;
if let Some(cc) = self.st().mining.cards.get_mut(idx) {
cc.sweep_state = "running".into();
cc.sweep_note = "sweep: checking how the cap is set".into();
cc.sweep_note = "tuning: reading what the card allows".into();
}
if let Some(direct) = self.sweep_direct {
self.sweep_mode_known(Ok(direct));
return;
}
let smi = crate::platform::tool("nvidia-smi");
let device = c.device.clone();
let current = if c.power_limit_w > 0.0 { c.power_limit_w } else { c.power_default_w }.round() as u64;
let shared = self.shared.clone();
std::thread::spawn(move || {
let r = match crate::detect::run_timeout(std::process::Command::new(&smi).args(["-i", &device, "-pl", &current.to_string()]), None, Duration::from_secs(20)) {
Some(out) => Ok(out.contains("All done")),
None => Err("nvidia-smi did not answer".to_string()),
};
shared.send(Cmd::SweepCapMode(r));
});
let allowed = self.elevation_allowed();
match c.vendor.as_str() {
"nvidia" => {
let smi = crate::platform::tool("nvidia-smi");
let device = c.device.clone();
let current = if c.power_limit_w > 0.0 { c.power_limit_w } else { c.power_default_w }.round() as u64;
let known_direct = self.sweep_direct;
std::thread::spawn(move || {
let q = crate::detect::run_timeout(std::process::Command::new(&smi).args(["-i", &device, "--query-gpu=clocks.max.gr,driver_version", "--format=csv,noheader,nounits"]), None, Duration::from_secs(10)).unwrap_or_default();
let p: Vec<&str> = q.trim().split(',').map(|s| s.trim()).collect();
let clock_max = p.first().and_then(|s| s.parse::<f64>().ok()).unwrap_or(0.0) as u32;
let driver = p.get(1).map(|s| s.to_string()).unwrap_or_default();
let direct = match known_direct {
Some(d) => Some(d),
None if allowed || current > 0 => crate::detect::run_timeout(std::process::Command::new(&smi).args(["-i", &device, "-pl", &current.to_string()]), None, Duration::from_secs(20)).map(|out| out.contains("All done")),
None => Some(false),
};
shared.send(Cmd::TuneProbe(idx, Ok(TuneProbe { clock_max_mhz: clock_max, clock_min_mhz: 0, driver, direct: direct.unwrap_or(false), amd_ordinal: -1 })));
});
}
"amd" => {
let Some(exe) = self.bins.telemetry.clone() else {
self.sweep_pending = None;
self.shared.send(Cmd::TuneProbe(idx, Ok(TuneProbe { clock_max_mhz: 0, clock_min_mhz: 0, driver: c.driver.clone(), direct: false, amd_ordinal: -1 })));
return;
};
let ordinal = c.amd_ordinal;
let driver = c.driver.clone();
std::thread::spawn(move || {
let out = crate::detect::run_timeout(std::process::Command::new(&exe).arg("--tune"), None, Duration::from_secs(20)).unwrap_or_default();
let t = out.lines().filter_map(parse_amd_tune).find(|t| ordinal < 0 || t.ordinal as i64 == ordinal);
shared.send(Cmd::TuneProbe(idx, Ok(match t {
Some(t) if t.ok => TuneProbe { clock_max_mhz: t.gmax_max as u32, clock_min_mhz: t.gmax_min as u32, driver, direct: true, amd_ordinal: t.ordinal as i64 },
Some(t) => TuneProbe { clock_max_mhz: 0, clock_min_mhz: 0, driver, direct: false, amd_ordinal: t.ordinal as i64 },
None => TuneProbe { clock_max_mhz: 0, clock_min_mhz: 0, driver, direct: false, amd_ordinal: -1 },
})));
});
}
_ => {
self.shared.send(Cmd::TuneProbe(idx, Ok(TuneProbe { clock_max_mhz: 0, clock_min_mhz: 0, driver: c.driver.clone(), direct: false, amd_ordinal: -1 })));
}
}
}
/// Step 2: start the helper when needed, then the run.
fn sweep_mode_known(&mut self, r: Result<bool, String>) {
let Some((idx, forced)) = self.sweep_pending else { return };
let direct = match r {
Ok(d) => d,
/// Step 2: the plan from what the probe found (full, confirm from the fleet prior, or baseline), the helper
/// when NVIDIA needs it, then the run.
fn sweep_probe_known(&mut self, idx: usize, r: Result<TuneProbe, String>) {
let Some((pidx, forced)) = self.sweep_pending else { return };
if pidx != idx {
return;
}
let probe = match r {
Ok(p) => p,
Err(e) => {
self.sweep_abort(&format!("cannot set caps: {e}"));
self.sweep_abort(&format!("cannot read the card's limits: {e}"));
return;
}
};
self.sweep_direct = Some(direct);
let Some(c) = self.st().mining.cards.get(idx).cloned() else {
self.sweep_pending = None;
return;
};
if !direct && !self.sweep_helper {
// NVIDIA control: this process is elevated (direct), or Power control is on so the one-prompt helper may run.
// The --sweep job alone never counts: it must not raise a prompt on a PC with nobody there (5 October 2026).
let power_control = probe.direct || self.shared.settings.lock().unwrap().power_control;
let limits = crate::ember::Limits { power_default_w: c.power_default_w, power_min_w: c.power_min_w, power_max_w: c.power_max_w, clock_max_mhz: probe.clock_max_mhz, clock_min_mhz: probe.clock_min_mhz };
let control = match c.vendor.as_str() {
"nvidia" => crate::ember::control_reason("nvidia", &limits, &c.device, power_control, false),
"amd" => crate::ember::control_reason("amd", &limits, &c.device, power_control, probe.direct && probe.amd_ordinal >= 0),
v => crate::ember::control_reason(v, &limits, &c.device, power_control, false),
};
if c.vendor == "nvidia" {
self.sweep_direct = Some(probe.direct);
}
{
let mut st = self.st();
if let Some(cc) = st.mining.cards.get_mut(idx) {
cc.clock_max_mhz = probe.clock_max_mhz;
cc.clock_min_mhz = probe.clock_min_mhz;
if !probe.driver.is_empty() {
cc.driver = probe.driver.clone();
}
if probe.amd_ordinal >= 0 {
cc.amd_ordinal = probe.amd_ordinal;
}
cc.tune_control = control.is_none();
}
}
let tuning = self.tuning_object();
let ember = crate::ember::settings_of(tuning.as_ref());
let before = crate::ember::Point { clock_mhz: c.clock_cap_mhz, power_pct: if c.power_pct == 0 { 80 } else { c.power_pct } };
let before_w = if c.power_limit_w > 0.0 { c.power_limit_w } else { requested_watts(&c) };
let key = crate::ember::prior_key(&c.name, &probe.driver, &c.program_class);
let prior = crate::ember::prior_of(tuning.as_ref(), &key, ember.min_samples);
let full_due = self.tune_full_due.remove(&idx);
let plan = if let Some(why) = control.as_ref() {
if let Some(cc) = self.st().mining.cards.get_mut(idx) {
cc.sweep_note = why.clone();
}
crate::ember::Plan::baseline(&limits, before, ember.tolerance_pct)
} else if let (Some(p), false, 0) = (prior.as_ref(), full_due, c.sweep_pct) {
crate::ember::Plan::confirm(&limits, p.point, before, ember.tolerance_pct)
} else {
crate::ember::Plan::full(&limits, before, ember.tolerance_pct)
};
if plan.is_empty() {
self.sweep_abort("no steps: the card reported neither a default power limit nor a maximum clock");
return;
}
if c.vendor == "nvidia" && control.is_none() && !probe.direct && !self.sweep_helper {
if let Err(e) = self.sweep_helper_start(&c) {
self.sweep_abort(&format!("the elevated helper could not start: {e}"));
return;
}
}
let steps = crate::sweep::plan_steps(c.power_default_w, c.power_min_w, c.power_max_w);
if steps.is_empty() {
self.sweep_abort("no steps: the card reported no default power limit");
return;
}
let label = self.miners.iter().find(|m| m.card == idx).map(|m| m.label.clone()).unwrap_or_else(|| format!("card-{idx}"));
let before_w = if c.power_limit_w > 0.0 { c.power_limit_w } else { requested_watts(&c) };
let now = Instant::now();
let run = crate::sweep::Run::new(idx, &c.key, &c.device, &label, steps.clone(), c.power_pct, before_w, forced, crate::sweep::Timing::from_env(), now);
let kind = plan.kind;
let steps = plan.len();
let run = crate::ember::Run::new(idx, &c.device, &label, plan, before_w, forced, crate::ember::Timing::from_env(), now);
self.sweep_pending = None;
self.sweep_say(&format!(
"SWEEP start card={label} name={} steps={} default={:.0} min={:.0} max={:.0} before={:.0} mode={}",
"TUNE start card={label} name={} plan={} steps={} clock_max={} clock_min={} default={:.0} min={:.0} max={:.0} before_clock={} before_pct={} before={:.0} mode={} driver={} class={} prior={}",
c.name.replace(' ', "_"),
steps.iter().map(|s| s.pct.to_string()).collect::<Vec<_>>().join(","),
c.power_default_w,
c.power_min_w,
c.power_max_w,
kind.name(),
steps,
limits.clock_max_mhz,
limits.clock_floor(),
limits.power_default_w,
limits.power_min_w,
limits.power_max_w,
before.clock_mhz,
before.power_pct,
before_w,
if direct { "direct" } else { "helper" }
if control.is_some() { "measure" } else if c.vendor == "amd" { "adlx" } else if probe.direct { "direct" } else { "helper" },
probe.driver.replace(' ', "_"),
if c.program_class.is_empty() { "v2" } else { &c.program_class },
prior.as_ref().map(|p| format!("{}mhz/{}pct/{}samples", p.point.clock_mhz, p.point.power_pct, p.samples)).unwrap_or_else(|| "none".into())
));
self.shared.event("info", &format!("{}: efficiency sweep started: {} caps from {}% down, {} s each on the live program", c.name, steps.len(), steps[0].pct, (run.timing.settle + run.timing.hold).as_secs()));
let what = match kind {
crate::ember::PlanKind::Full => format!("{steps} steps over the power limit and the core clock, {} s each on the live program", (run.timing.settle + run.timing.hold).as_secs()),
crate::ember::PlanKind::Confirm => format!("the fleet prior ({} samples) and one neighbour, {} s each", prior.as_ref().map(|p| p.samples).unwrap_or(0), (run.timing.settle + run.timing.hold).as_secs()),
crate::ember::PlanKind::Baseline => format!("measuring the card as it runs ({})", control.clone().unwrap_or_default()),
};
self.shared.event("info", &format!("{}: tuning started: {what}", c.name));
self.sweep = Some(run);
self.sweep_drive(now);
}
/// The elevated helper (src/sweep.rs helper_script_*): one administrator prompt; it polls <app>/sweep/cmd.txt.
fn sweep_helper_start(&mut self, c: &CardState) -> Result<(), String> {
if !self.elevation_allowed() {
if !self.shared.settings.lock().unwrap().power_control {
return Err("Power control is off in Settings".into());
}
let dir = self.sweep_dir();
@ -1890,7 +2042,7 @@ impl Engine {
std::fs::write(&script, crate::sweep::helper_script_unix()).map_err(|e| e.to_string())?;
format!("sh \"{}\" \"{}\" \"{}\" {} {}", script.display(), dir.display(), smi, c.device, restore)
};
self.shared.log(&format!("sweep helper (administrator prompt): {line}"));
self.shared.log(&format!("tune helper (administrator prompt): {line}"));
self.sweep_helper = true;
let shared = self.shared.clone();
std::thread::spawn(move || {
@ -1906,28 +2058,94 @@ impl Engine {
}
}
/// Sets a cap for the sweep: directly (this process is elevated) or through the helper's command file. The
/// telemetry readback (power.limit every 5 s) tells the run when it took.
fn sweep_set_cap(&mut self, device: &str, watts: f64) {
self.sweep_seq += 1;
let w = watts.round() as u64;
if self.sweep_direct == Some(true) {
let smi = crate::platform::tool("nvidia-smi");
let device = device.to_string();
let shared = self.shared.clone();
std::thread::spawn(move || {
let out = crate::detect::run_timeout(std::process::Command::new(&smi).args(["-i", &device, "-pl", &w.to_string()]), None, Duration::from_secs(20)).unwrap_or_else(|| "nvidia-smi did not answer".into());
shared.send(Cmd::SweepCapSet(out));
});
} else {
let _ = std::fs::write(self.sweep_dir().join("cmd.txt"), format!("{} {w}\n", self.sweep_seq));
/// Sets a point on the card. NVIDIA: `-pl <W>` and `-lgc 0,<MHz>` (`-rgc` for unlocked), directly when this
/// process is elevated, else through the helper's command file (`<seq> pl <W>`, `<seq> lgc <MHz>`, `<seq> rgc`).
/// AMD: igneum-gpu-telemetry `--card N --set-plimit <offset>` and `--set-gmax <MHz>` (`--reset` for the default
/// point). The result comes back as Cmd::TuneSet(seq, ..): the acknowledgement the run waits for.
fn tune_apply(&mut self, idx: usize, device: &str, step: &crate::ember::Step) {
let seq = self.sweep.as_ref().map(|r| r.seq).unwrap_or(0);
let card = self.st().mining.cards.get(idx).cloned();
let Some(c) = card else { return };
let w = step.watts.round() as u64;
let clock = step.point.clock_mhz;
let shared = self.shared.clone();
self.tune_acked = None;
match c.vendor.as_str() {
"nvidia" if self.sweep_direct == Some(true) => {
let smi = crate::platform::tool("nvidia-smi");
let device = device.to_string();
let power_only = c.power_default_w > 0.0;
std::thread::spawn(move || {
let mut ok = true;
let mut text = String::new();
if power_only && w > 0 {
let out = crate::detect::run_timeout(std::process::Command::new(&smi).args(["-i", &device, "-pl", &w.to_string()]), None, Duration::from_secs(20)).unwrap_or_else(|| "nvidia-smi did not answer".into());
ok &= out.contains("All done") || out.contains("Power limit for GPU");
text.push_str(out.trim());
}
let out = if clock > 0 {
crate::detect::run_timeout(std::process::Command::new(&smi).args(["-i", &device, "-lgc", &format!("0,{clock}")]), None, Duration::from_secs(20)).unwrap_or_else(|| "nvidia-smi did not answer".into())
} else {
crate::detect::run_timeout(std::process::Command::new(&smi).args(["-i", &device, "-rgc"]), None, Duration::from_secs(20)).unwrap_or_else(|| "nvidia-smi did not answer".into())
};
ok &= out.contains("All done") || out.to_ascii_lowercase().contains("clocks set") || out.to_ascii_lowercase().contains("reset");
text.push(' ');
text.push_str(out.trim());
shared.send(Cmd::TuneSet(seq, if ok { Ok(text) } else { Err(text) }));
});
}
"nvidia" => {
if !self.sweep_helper {
// measure only: nothing is set; the run notices the missing acknowledgement and measures
return;
}
let cmd = format!("{seq} pl {w}\n{seq} {}\n", if clock > 0 { format!("lgc {clock}") } else { "rgc".to_string() });
let _ = std::fs::write(self.sweep_dir().join("cmd.txt"), cmd);
// the helper polls twice a second and nvidia-smi answers within a second or two
std::thread::spawn(move || {
std::thread::sleep(Duration::from_secs(4));
shared.send(Cmd::TuneSet(seq, Ok("helper".into())));
});
}
"amd" => {
let Some(exe) = self.bins.telemetry.clone() else { return };
if c.amd_ordinal < 0 {
return;
}
let n = c.amd_ordinal.to_string();
let offset = step.point.power_pct as i64 - 100;
let unlocked = clock == 0 && offset == 0;
let gmax = if clock > 0 { clock } else { c.clock_max_mhz };
std::thread::spawn(move || {
let run = |args: &[&str]| crate::detect::run_timeout(std::process::Command::new(&exe).args(args), None, Duration::from_secs(20)).unwrap_or_else(|| "igneum-gpu-telemetry did not answer".into());
let mut text = String::new();
let mut ok = true;
if unlocked {
let out = run(&["--card", &n, "--reset"]);
ok &= out.lines().any(|l| l.starts_with("tune ") && l.contains(" ok"));
text.push_str(out.trim());
} else {
if gmax > 0 {
let out = run(&["--card", &n, "--set-gmax", &gmax.to_string()]);
ok &= out.lines().any(|l| l.starts_with("tune ") && l.contains(" ok"));
text.push_str(out.trim());
}
let out = run(&["--card", &n, "--set-plimit", &offset.to_string()]);
ok &= out.lines().any(|l| l.starts_with("tune ") && l.contains(" ok"));
text.push(' ');
text.push_str(out.trim());
}
shared.send(Cmd::TuneSet(seq, if ok { Ok(text) } else { Err(text) }));
});
}
_ => {}
}
self.shared.log(&format!("sweep: cap {w} W requested on device {device}"));
self.shared.log(&format!("tune: {} MHz, {}% ({w} W) requested on device {device} (request {seq})", if clock > 0 { clock.to_string() } else { "unlocked".into() }, step.point.power_pct));
}
/// Faults and conditions first, then the state machine.
fn sweep_drive(&mut self, now: Instant) {
let Some((idx, device, started)) = self.sweep.as_ref().map(|r| (r.card, r.device.clone(), r.started)) else { return };
let Some((idx, device, started, seq)) = self.sweep.as_ref().map(|r| (r.card, r.device.clone(), r.started, r.seq)) else { return };
let card = self.st().mining.cards.get(idx).cloned();
let Some(c) = card else {
self.sweep_abort("the card disappeared");
@ -1953,24 +2171,37 @@ impl Engine {
self.sweep_abort(&why);
return;
}
let outs = self.sweep.as_mut().map(|r| r.tick(now, c.power_limit_w)).unwrap_or_default();
let acked = self.tune_acked == Some(seq);
let rb = crate::ember::Readback { limit_w: if c.vendor == "nvidia" { c.power_limit_w } else { 0.0 }, acked };
// the clock readback during a hold: a core clock well over the cap means the cap did not take
if let Some(run) = self.sweep.as_mut() {
if matches!(run.phase, crate::ember::Phase::Holding { .. }) {
if let Some(step) = run.current.as_ref() {
if step.point.clock_mhz > 0 && c.gclk_mhz > step.point.clock_mhz as f64 * 1.05 && c.vendor == "nvidia" && c.tune_control {
run.mark_unapplied();
}
}
}
}
let outs = self.sweep.as_mut().map(|r| r.tick(now, rb)).unwrap_or_default();
let label = self.sweep.as_ref().map(|r| r.label.clone()).unwrap_or_default();
for o in outs {
match o {
crate::sweep::Out::Apply(w) => self.sweep_set_cap(&device, w),
crate::sweep::Out::Row(row) => {
self.sweep_say(&row.line(&label));
crate::ember::Out::Apply(step) => self.tune_apply(idx, &device, &step),
crate::ember::Out::Row(row) => {
let i = self.sweep.as_ref().map(|r| r.rows.len().saturating_sub(1)).unwrap_or(0);
self.sweep_say(&row.line(&label, i));
if let Some(cc) = self.st().mining.cards.get_mut(idx) {
if row.usable() {
cc.sweep_note = format!("sweep: {}% done · {:.3} MH/W", row.pct, row.eff);
cc.sweep_note = format!("tuning: step {} done · {:.3} MH/W", i + 1, row.eff);
}
}
}
crate::sweep::Out::Finished(row) => {
crate::ember::Out::Finished(row) => {
self.sweep_finish(row);
return;
}
crate::sweep::Out::Failed(e) => {
crate::ember::Out::Failed(e) => {
self.sweep_abort(&e);
return;
}
@ -1983,93 +2214,129 @@ impl Engine {
}
}
/// The chosen cap is in force: record it on the card and in the settings, say so, let the helper go.
fn sweep_finish(&mut self, row: crate::sweep::Row) {
/// The chosen point is in force: record it on the card and in the settings, write the fleet record, say so,
/// let the helper go. A confirm check that found a better neighbour queues the full plan.
fn sweep_finish(&mut self, row: crate::ember::Row) {
let Some(run) = self.sweep.take() else { return };
let idx = run.card;
let unix = crate::platform::unix_now();
let (name, pinned, pinned_w) = {
let kind = run.plan.kind;
let full_due = run.full_due();
let (name, pinned, control, vendor, driver, class, key) = {
let mut settings = self.shared.settings.lock().unwrap();
let mut st = self.st();
let Some(c) = st.mining.cards.get_mut(idx) else { return };
c.sweep_pct = row.pct;
c.sweep_pct = row.point.power_pct;
c.sweep_eff = row.eff;
c.sweep_watts = row.watts;
c.sweep_mhs = row.mhs;
c.sweep_at = unix as f64;
c.sweep_state = "idle".into();
let pinned_w = requested_watts(c);
if c.pinned {
c.sweep_note = format!("sweep: best {}% · {:.3} MH/W; cap pinned at {}%", row.pct, row.eff, c.power_pct);
c.tune_clock_mhz = row.point.clock_mhz;
c.tune_source = kind.name().into();
c.tune_line = crate::ember::tuned_line(row.mhs, row.watts, row.eff);
let control = c.tune_control;
if kind == crate::ember::PlanKind::Baseline {
c.sweep_note = if control { String::new() } else { c.sweep_note.clone() };
} else if c.pinned {
c.sweep_note = format!("tuning: best {} · {:.3} MH/W; your setting stays pinned", point_words(&row.point), row.eff);
} else {
c.power_pct = row.pct;
c.power_limit_w = row.limit;
c.power_applied = true;
c.power_note = format!("power cap: {} W held by the sweep (best MH per watt)", row.limit as u64);
c.sweep_note = format!("sweep: best {}% · {:.3} MH/W", row.pct, row.eff);
c.power_pct = row.point.power_pct;
c.clock_cap_mhz = row.point.clock_mhz;
if c.vendor == "nvidia" {
c.power_limit_w = row.limit;
c.power_applied = true;
c.power_note = format!("power cap: {} W held by the tune (best MH per watt)", row.limit as u64);
}
c.sweep_note = if full_due { "tuning: a neighbour beat the fleet prior; the full tune runs next".into() } else { String::new() };
}
let e = settings.cards.entry(c.key.clone()).or_insert_with(|| crate::config::CardPref { enabled: c.enabled, identities: c.identities, ..Default::default() });
e.sweep_at = unix;
e.sweep_pct = row.pct;
e.sweep_pct = row.point.power_pct;
e.sweep_eff = row.eff;
e.sweep_watts = row.watts;
e.sweep_mhs = row.mhs;
if !c.pinned {
e.power_pct = row.pct;
e.sweep_clock_mhz = row.point.clock_mhz;
e.sweep_driver = c.driver.clone();
e.sweep_class = c.program_class.clone();
e.sweep_source = kind.name().into();
if !c.pinned && kind != crate::ember::PlanKind::Baseline {
e.power_pct = row.point.power_pct;
}
(c.name.clone(), c.pinned, pinned_w)
(c.name.clone(), c.pinned, control, c.vendor.clone(), c.driver.clone(), c.program_class.clone(), c.key.clone())
};
let _ = key;
self.shared.save_settings();
self.sweep_say(&format!("SWEEP chosen card={} cap={} limit={:.0} watts={:.1} mhs={:.2} eff={:.4}{}", run.label, row.pct, row.limit, row.watts, row.mhs, row.eff, if pinned { " pinned=1" } else { "" }));
if pinned {
self.sweep_set_cap(&run.device.clone(), pinned_w);
self.shared.event("ok", &format!("{name}: best efficiency at {}% ({:.0} W cap): {:.1} MH/s at {:.0} W = {:.3} MH/W. The cap stays pinned where you set it.", row.pct, row.limit, row.mhs, row.watts, row.eff));
} else {
self.shared.event("ok", &format!("{name}: best efficiency at {}% ({:.0} W cap): {:.1} MH/s at {:.0} W = {:.3} MH/W; cap held there", row.pct, row.limit, row.mhs, row.watts, row.eff));
self.sweep_say(&format!("TUNE chosen card={} clock={} cap={} limit={:.0} watts={:.1} mhs={:.2} eff={:.4} plan={}{}", run.label, row.point.clock_mhz, row.point.power_pct, row.limit, row.watts, row.mhs, row.eff, kind.name(), if pinned { " pinned=1" } else { "" }));
let before = run.rows.first().filter(|_| kind == crate::ember::PlanKind::Full).cloned();
let record = crate::ember::record_json(crate::platform::unix_now_f(), &crate::config::fingerprint8(&self.shared.runtime.machine_id), VERSION, crate::manifest::platform_name(), &name, &vendor, &driver, &class, kind, &run.rows, Some(&row), before.as_ref());
self.sweep_say(&format!("TUNE {record}"));
let line = crate::ember::tuned_line(row.mhs, row.watts, row.eff);
match kind {
crate::ember::PlanKind::Baseline => self.shared.event("info", &format!("{name}: measured as it runs: {line}{}", if control { String::new() } else { " (measure only; no control on this card)".into() })),
_ if pinned => self.shared.event("ok", &format!("{name}: best point {} ({line}); your setting stays pinned", point_words(&row.point))),
_ => self.shared.event("ok", &format!("{name}: {line}, {} held", point_words(&row.point))),
}
if pinned && kind != crate::ember::PlanKind::Baseline {
let p = self.st().mining.cards.get(idx).map(|c| crate::ember::Point { clock_mhz: c.clock_cap_mhz, power_pct: c.power_pct }).unwrap_or(run.before);
let w = self.st().mining.cards.get(idx).map(requested_watts).unwrap_or(run.before_w);
self.tune_apply(idx, &run.device, &crate::ember::Step { point: p, watts: w, kind: crate::ember::Kind::Confirm });
}
self.sweep_helper_quit();
if full_due && !pinned {
self.tune_full_due.insert(idx);
let k = self.st().mining.cards.get(idx).map(|c| c.key.clone());
if let Some(k) = k {
self.sweep_queue.push(k);
}
self.sweep_retry.insert(idx, Instant::now() + Duration::from_secs(60));
}
self.upload_logs(false);
if self.shared.runtime.sweep_only && self.sweep_queue.is_empty() {
self.shared.send(Cmd::Quit);
}
}
/// Stops a sweep: the cap goes back to what it was, the log says why, the card waits an hour before the
/// scheduler tries again (a forced sweep under --sweep retries twice).
/// Stops a tune: the point goes back to what it was, the log says why, the card waits an hour before the
/// scheduler tries again (a forced tune under --sweep retries twice).
fn sweep_abort(&mut self, why: &str) {
let pending = self.sweep_pending.take();
let run = self.sweep.take();
let (idx, label, device, before_w, before_pct, forced) = match (&run, pending) {
(Some(r), _) => (r.card, r.label.clone(), r.device.clone(), r.before_w, r.before_pct, r.forced),
let (idx, label, device, before, before_w, forced, touched) = match (&run, pending) {
(Some(r), _) => (r.card, r.label.clone(), r.device.clone(), r.before, r.before_w, r.forced, r.seq > 0 && r.plan.kind != crate::ember::PlanKind::Baseline),
(None, Some((idx, forced))) => {
let c = self.st().mining.cards.get(idx).cloned();
let (label, device, w, pct) = c.map(|c| (format!("card-{idx}"), c.device.clone(), if c.power_limit_w > 0.0 { c.power_limit_w } else { requested_watts(&c) }, c.power_pct)).unwrap_or_default();
(idx, label, device, w, pct, forced)
let (label, device, w, p) = c.map(|c| (format!("card-{idx}"), c.device.clone(), if c.power_limit_w > 0.0 { c.power_limit_w } else { requested_watts(&c) }, crate::ember::Point { clock_mhz: c.clock_cap_mhz, power_pct: c.power_pct })).unwrap_or_default();
(idx, label, device, p, w, forced, false)
}
_ => return,
};
self.sweep_say(&format!("SWEEP aborted card={label} reason={}", why.replace(' ', "_")));
self.sweep_say(&format!("TUNE aborted card={label} reason={}", why.replace(' ', "_")));
let name = {
let mut st = self.st();
match st.mining.cards.get_mut(idx) {
Some(c) => {
c.sweep_state = "idle".into();
c.sweep_note = format!("sweep stopped: {why}");
c.power_pct = before_pct;
c.power_applied = false;
c.power_note = format!("cap restored to {} W after the sweep stopped", before_w as u64);
c.sweep_note = format!("tuning stopped: {why}");
if touched {
c.power_pct = before.power_pct;
c.clock_cap_mhz = before.clock_mhz;
c.power_applied = false;
c.power_note = format!("cap restored to {} W after the tune stopped", before_w as u64);
}
c.name.clone()
}
None => label.clone(),
}
};
if run.is_some() && before_w > 0.0 && !device.is_empty() {
self.sweep_set_cap(&device, before_w);
if touched && !device.is_empty() {
self.sweep = None;
self.tune_restore(idx, &device, before, before_w);
}
self.sweep_helper_quit();
self.shared.event("error", &format!("{name}: efficiency sweep stopped ({why}); cap back to {} W", before_w as u64));
self.shared.event(if touched { "error" } else { "info" }, &format!("{name}: tuning stopped ({why}){}", if touched { format!("; back to {} and {} W", point_words(&before), before_w as u64) } else { String::new() }));
let attempts = self.sweep_attempts.get(&idx).copied().unwrap_or(0);
if forced && self.shared.runtime.sweep_only && attempts < 3 && !self.quitting {
// the PC job: the hour boundary or a restart interrupted it; try again once the card is steady
let key = self.st().mining.cards.get(idx).map(|c| c.key.clone());
if let Some(key) = key {
self.sweep_queue.push(key);
@ -2083,6 +2350,14 @@ impl Engine {
}
}
/// Puts a card back on a point outside a run (the abort path): the same commands as a step, with no run to
/// acknowledge them.
fn tune_restore(&mut self, idx: usize, device: &str, point: crate::ember::Point, watts: f64) {
let saved = self.sweep.take();
self.tune_apply(idx, device, &crate::ember::Step { point, watts, kind: crate::ember::Kind::Confirm });
self.sweep = saved;
}
// ---- the tick ----------------------------------------------------------------------------------------------
fn tick(&mut self) {
@ -2786,6 +3061,11 @@ impl Engine {
let card_name = self.st().mining.cards.get(card).map(|c| c.name.clone()).unwrap_or_else(|| label.clone());
if stderr {
if text.contains(" rejected nonce=") {
if let Some(run) = self.sweep.as_mut() {
if run.card == card {
run.sample_fault();
}
}
let mut st = self.st();
st.mining.rejected_session += 1;
if let Some(c) = st.mining.cards.get_mut(card) {
@ -2869,6 +3149,14 @@ impl Engine {
run.sample_rate(now_rate);
}
}
let more_mismatched = self.st().mining.cards.get(card).map(|c| s.mismatched > c.mismatched).unwrap_or(false);
if more_mismatched {
if let Some(run) = self.sweep.as_mut() {
if run.card == card {
run.sample_fault();
}
}
}
let mut st = self.st();
if let Some(c) = st.mining.cards.get_mut(card) {
c.hash_avg = avg;
@ -2974,6 +3262,10 @@ impl Engine {
c.race_mhs = r.mhs;
c.race_gain_pct = r.gain_pct;
c.race_variants = r.variants.len() as u32;
if !r.driver.is_empty() && c.driver.is_empty() {
c.driver = r.driver.clone();
}
c.program_class = crate::ember::program_class(r.loads as u32, r.wide as u32);
(c.vendor.clone(), c.worker.clone(), c.power_limit_w, c.power_w, c.power_pct)
}
None => (String::new(), String::new(), 0.0, 0.0, 0),
@ -3164,6 +3456,59 @@ fn power_cap_plan(cards: &mut [CardState], allowed: bool, smi: &str) -> (Vec<Str
(cmds, what, held)
}
/// What a tune's probe found for a card (Cmd::TuneProbe).
#[derive(Clone, Debug, Default, PartialEq)]
pub struct TuneProbe {
pub clock_max_mhz: u32,
pub clock_min_mhz: u32,
pub driver: String,
/// NVIDIA: this process sets limits itself (elevated); AMD: the helper answered its `--tune` line with ok
pub direct: bool,
pub amd_ordinal: i64,
}
/// One `tune` line of igneum-gpu-telemetry --tune:
/// `tune <ordinal> name "<name>" gmax <MHz> gmax_range <min> <max> plimit <offset %> plimit_range <min> <max> factory 0|1 ok|<error>`.
#[derive(Clone, Debug, Default, PartialEq)]
pub struct AmdTune {
pub ordinal: usize,
pub name: String,
pub gmax: f64,
pub gmax_min: f64,
pub gmax_max: f64,
pub plimit: f64,
pub plimit_min: f64,
pub plimit_max: f64,
pub factory: bool,
pub ok: bool,
pub error: String,
}
pub fn parse_amd_tune(line: &str) -> Option<AmdTune> {
let line = line.trim();
if !line.starts_with("tune ") {
return None;
}
let (head, rest) = line.split_once(" name \"")?;
let (name, tail) = rest.split_once('"')?;
let ordinal: usize = head.split_whitespace().nth(1)?.parse().ok()?;
let tp: Vec<&str> = tail.split_whitespace().collect();
let at = |key: &str| tp.iter().position(|p| *p == key);
let num = |i: usize| tp.get(i).and_then(|v| if *v == "-" { Some(-1.0) } else { v.parse::<f64>().ok() }).unwrap_or(-1.0);
let g = at("gmax")?;
let gr = at("gmax_range")?;
let pl = at("plimit")?;
let pr = at("plimit_range")?;
let fi = at("factory")?;
let verdict = tp.get(fi + 2).copied().unwrap_or("");
Some(AmdTune { ordinal, name: name.to_string(), gmax: num(g + 1), gmax_min: num(gr + 1), gmax_max: num(gr + 2), plimit: num(pl + 1), plimit_min: num(pr + 1), plimit_max: num(pr + 2), factory: tp.get(fi + 1).copied() == Some("1"), ok: verdict == "ok", error: if verdict == "ok" { String::new() } else { tp[fi + 2..].join(" ") } })
}
/// "2,472 MHz at 100%" or "100%" for the feed.
fn point_words(p: &crate::ember::Point) -> String {
if p.clock_mhz > 0 { format!("{} MHz at {}%", p.clock_mhz, p.power_pct) } else { format!("{}% (clock unlocked)", p.power_pct) }
}
fn requested_watts(c: &CardState) -> f64 {
let pct = if c.power_pct == 0 { 80 } else { c.power_pct.clamp(crate::sweep::MIN_PCT, 100) };
let mut w = c.power_default_w * pct as f64 / 100.0;
@ -3490,6 +3835,9 @@ pub struct AmdTelemetry {
pub mclk_mhz: f64,
pub gclk_mhz: f64,
pub util_pct: f64,
/// the limits in force (0.3.10 lines; -1.0 on older helpers): 100 + the power offset, the max GPU clock
pub plimit_pct: f64,
pub gmax_mhz: f64,
pub source: String,
}
@ -3524,12 +3872,14 @@ pub fn parse_amd_telemetry(line: &str) -> Option<AmdTelemetry> {
let mclk_mhz = num("mclk_mhz")?;
let gclk_mhz = num("gclk_mhz")?;
let util_pct = num("util_pct")?;
let plimit_pct = num("plimit_pct").unwrap_or(-1.0);
let gmax_mhz = num("gmax_mhz").unwrap_or(-1.0);
let source = tp.iter().position(|p| *p == "source").and_then(|i| tp.get(i + 1)).map(|s| s.to_string()).unwrap_or_default();
let kind = hp[5].to_string();
// the rank within the kind: the helper lists one integrated card at most and it comes first when present
// (ADLX order on every PC seen so far), so a discrete card's rank is its ordinal minus the integrated ones before it
let ordinal_in_kind = if kind == "discrete" && ordinal > 0 { ordinal - 1 } else if kind == "discrete" { 0 } else { 0 };
Some(AmdTelemetry { ordinal, ordinal_in_kind, bus: hp[3].to_string(), kind, name: name.to_string(), watts, temp_c, fan_rpm, fan_pct, mclk_mhz, gclk_mhz, util_pct, source })
Some(AmdTelemetry { ordinal, ordinal_in_kind, bus: hp[3].to_string(), kind, name: name.to_string(), watts, temp_c, fan_rpm, fan_pct, mclk_mhz, gclk_mhz, util_pct, plimit_pct, gmax_mhz, source })
}
#[cfg(test)]
@ -3557,6 +3907,26 @@ mod amd_telemetry_tests {
assert!(parse_amd_telemetry("amd 0 bus - kind - name \"x\" watts").is_none());
}
/// The 0.3.10 helper's lines (the tuning commands, usage header of proto-opencl/gpu-telemetry.c): the sample
/// line carries the limits in force, the `tune` line the ranges a card allows. An older line (no plimit_pct,
/// no gmax_mhz) still parses with -1 there.
#[test]
fn the_tune_line_and_the_limits_in_force_parse() {
let l = "amd 1 bus 98 kind discrete name \"AMD Radeon RX 9070 XT\" watts 198.9 temp_c 64.0 fan_rpm 657 fan_pct 25 mclk_mhz 2505 gclk_mhz 3290 util_pct 100 plimit_pct 100 gmax_mhz 3100 source adlx";
let s = parse_amd_telemetry(l).unwrap();
assert_eq!((s.plimit_pct, s.gmax_mhz, s.gclk_mhz), (100.0, 3100.0, 3290.0));
let old = "amd 0 bus 0000:0c:00.0 kind discrete name \"AMD Radeon RX 9070 XT\" watts 287.0 temp_c 61.0 fan_rpm 1180 fan_pct 30 mclk_mhz 1258 gclk_mhz 2450 util_pct 90 source sysfs";
assert_eq!(parse_amd_telemetry(old).unwrap().plimit_pct, -1.0);
let t = parse_amd_tune("tune 1 name \"AMD Radeon RX 9070 XT\" gmax 3100 gmax_range 500 3400 plimit 0 plimit_range -30 15 factory 1 ok").unwrap();
assert_eq!((t.ordinal, t.name.as_str(), t.gmax, t.gmax_min, t.gmax_max, t.plimit, t.plimit_min, t.plimit_max, t.factory, t.ok), (1, "AMD Radeon RX 9070 XT", 3100.0, 500.0, 3400.0, 0.0, -30.0, 15.0, true, true));
let e = parse_amd_tune("tune 0 name \"x\" gmax - gmax_range - - plimit - plimit_range - - factory 0 ADLX: GetManualGraphicsTuning returned 2").unwrap();
assert!(!e.ok);
assert_eq!(e.error, "ADLX: GetManualGraphicsTuning returned 2");
assert_eq!(e.gmax_max, -1.0);
assert!(parse_amd_tune("end 3.2 ms 2 card(s)").is_none());
assert!(parse_amd_tune("tune 0 name \"x\" gmax 1").is_none());
}
#[test]
fn the_perfcounter_fallback_line_parses() {
let l = "amd 0 bus luid_0x00000000_0x0000D4E3 kind - name \"-\" watts - temp_c - fan_rpm - fan_pct - mclk_mhz - gclk_mhz - util_pct 100 source perfcounter";

View file

@ -31,6 +31,7 @@ mod prover;
mod verifier;
mod wslhost;
mod sweep;
mod ember;
mod watchdog;
use std::io::{BufRead, Write};

View file

@ -90,6 +90,18 @@ pub struct CardState {
pub sweep_mhs: f64,
pub sweep_at: f64, // unix s of the last sweep
pub pinned: bool, // the user set the cap by hand; the sweep records but does not change it
// Ember Tune (src/ember.rs): the two-knob tune, 5 October 2026
pub clock_max_mhz: u32, // the vendor's maximum core clock (0 = unknown)
pub clock_min_mhz: u32, // the vendor's floor for a cap (0 = 60% of the maximum)
pub gclk_mhz: f64, // core clock now
pub clock_cap_mhz: u32, // the cap in force (0 = unlocked)
pub amd_ordinal: i64, // the `amd N` ordinal of igneum-gpu-telemetry (-1 = unknown)
pub driver: String, // the driver version (nvidia-smi, or the worker's race line)
pub program_class: String, // the program class of the race line (loads and wide loads per hash); "" = unknown
pub tune_control: bool, // both knobs reach the card (else measure only; sweep_note says why)
pub tune_clock_mhz: u32, // the clock cap the last tune chose (0 = unlocked)
pub tune_source: String, // full | confirm | baseline
pub tune_line: String, // "Tuned: 122.3 MH/s at 290 W (0.422 MH/W)" once tuned
// the kernel variant race (docs/design/miner-tuning.md): what the worker's last race chose
pub variant: String,
pub race_mhs: f64,
@ -217,6 +229,9 @@ pub struct SettingsState {
pub power_control: bool,
/// the line beside the Power control switch: why it is off, or that the rights were given
pub power_note: String,
/// Ember Tune is paused fleet-wide by the signed manifest's kill switch (tuning.ember.enabled = false)
pub tuning_off: bool,
pub tuning_note: String,
/// the miner software's dev fee switch (settings; `--dev-fee 0` when off)
pub dev_fee: bool,
/// devnet only: the node trusts proof records without a verifier (`IGNEUM_PROOF_VERIFY=trust`)

View file

@ -340,10 +340,12 @@ impl Run {
}
}
/// The elevated helper that sets caps for a sweep (one administrator prompt per sweep, not one per step). It polls
/// `<dir>/cmd.txt` twice a second: a line `<seq> <watts>` runs `nvidia-smi -i <device> -pl <watts>`, `quit` ends it.
/// After 20 minutes without a new command it restores `<restore watts>` and exits by itself, so an engine that died
/// mid-sweep leaves the card on its old limit. It writes what it ran to `<dir>/helper.log`.
/// The elevated helper that sets limits for a tune (one administrator prompt per tune, not one per step). It polls
/// `<dir>/cmd.txt` twice a second; each line is `<seq> pl <watts>` (`nvidia-smi -i <device> -pl <watts>`; the
/// 0.3.9 form `<seq> <watts>` still works), `<seq> lgc <mhz>` (`-lgc 0,<mhz>`, the core clock cap; the memory clock
/// is never touched) or `<seq> rgc` (`-rgc`, unlocked); `quit` ends it. After 20 minutes without a new command it
/// restores `<restore watts>`, resets the clocks and exits by itself, so an engine that died mid-tune leaves the
/// card on its old limits. It writes what it ran to `<dir>/helper.log`.
pub fn helper_script_windows() -> &'static str {
r#"param([string]$Dir, [string]$Smi, [string]$Device, [string]$Restore)
$ErrorActionPreference = 'Continue'
@ -359,15 +361,27 @@ while ($true) {
$last = $c
$idle = Get-Date
if ($c -eq 'quit') { "$(Get-Date -Format o) quit" | Out-File -FilePath $log -Append -Encoding utf8; break }
$w = ($c -split ' ')[-1]
if ($w -match '^\d+$') {
$out = (& $Smi -i $Device -pl $w 2>&1 | Out-String).Trim()
"$(Get-Date -Format o) -pl $w : $out" | Out-File -FilePath $log -Append -Encoding utf8
foreach ($line in ($c -split "`n")) {
$p = ($line.Trim() -split ' ')
if ($p.Count -lt 2) { continue }
$op = $p[1]; $v = $p[-1]
if ($p.Count -eq 2 -and $v -match '^\d+$') { $op = 'pl' }
if ($op -eq 'pl' -and $v -match '^\d+$') {
$out = (& $Smi -i $Device -pl $v 2>&1 | Out-String).Trim()
"$(Get-Date -Format o) $($p[0]) -pl $v : $out" | Out-File -FilePath $log -Append -Encoding utf8
} elseif ($op -eq 'lgc' -and $v -match '^\d+$') {
$out = (& $Smi -i $Device -lgc "0,$v" 2>&1 | Out-String).Trim()
"$(Get-Date -Format o) $($p[0]) -lgc 0,$v : $out" | Out-File -FilePath $log -Append -Encoding utf8
} elseif ($op -eq 'rgc') {
$out = (& $Smi -i $Device -rgc 2>&1 | Out-String).Trim()
"$(Get-Date -Format o) $($p[0]) -rgc : $out" | Out-File -FilePath $log -Append -Encoding utf8
}
}
}
if (((Get-Date) - $idle).TotalMinutes -gt 20) {
$out = (& $Smi -i $Device -pl $Restore 2>&1 | Out-String).Trim()
"$(Get-Date -Format o) idle 20 min: restored $Restore W and quit: $out" | Out-File -FilePath $log -Append -Encoding utf8
$out2 = (& $Smi -i $Device -rgc 2>&1 | Out-String).Trim()
"$(Get-Date -Format o) idle 20 min: restored $Restore W, clocks reset, and quit: $out / $out2" | Out-File -FilePath $log -Append -Encoding utf8
break
}
Start-Sleep -Milliseconds 500
@ -388,11 +402,20 @@ while true; do
if [ -n "$c" ] && [ "$c" != "$last" ]; then
last="$c"; idle=$(date +%s)
if [ "$c" = "quit" ]; then echo "$(date -u +%FT%TZ) quit" >> "$dir/helper.log"; break; fi
w="${c##* }"
case "$w" in ''|*[!0-9]*) ;; *) echo "$(date -u +%FT%TZ) -pl $w : $("$smi" -i "$dev" -pl "$w" 2>&1)" >> "$dir/helper.log";; esac
printf '%s\n' "$c" | while IFS= read -r line; do
set -- $line
[ $# -ge 2 ] || continue
op="$2"; v="${line##* }"
[ $# -eq 2 ] && op=pl
case "$op" in
pl) case "$v" in ''|*[!0-9]*) ;; *) echo "$(date -u +%FT%TZ) $1 -pl $v : $("$smi" -i "$dev" -pl "$v" 2>&1)" >> "$dir/helper.log";; esac ;;
lgc) case "$v" in ''|*[!0-9]*) ;; *) echo "$(date -u +%FT%TZ) $1 -lgc 0,$v : $("$smi" -i "$dev" -lgc "0,$v" 2>&1)" >> "$dir/helper.log";; esac ;;
rgc) echo "$(date -u +%FT%TZ) $1 -rgc : $("$smi" -i "$dev" -rgc 2>&1)" >> "$dir/helper.log" ;;
esac
done
fi
if [ $(( $(date +%s) - idle )) -gt 1200 ]; then
echo "$(date -u +%FT%TZ) idle 20 min: restored $restore W: $("$smi" -i "$dev" -pl "$restore" 2>&1)" >> "$dir/helper.log"; break
echo "$(date -u +%FT%TZ) idle 20 min: restored $restore W, clocks reset: $("$smi" -i "$dev" -pl "$restore" 2>&1) / $("$smi" -i "$dev" -rgc 2>&1)" >> "$dir/helper.log"; break
fi
sleep 0.5
done
@ -405,8 +428,9 @@ pub fn unsupported_reason(vendor: &str, power_default_w: f64, device: &str) -> O
match vendor {
"nvidia" if power_default_w > 0.0 && !device.is_empty() => None,
"nvidia" => Some("not available: nvidia-smi did not report this card's power limits"),
"apple" => Some("not available on Apple silicon: there is no power cap to set, and powermetrics needs administrator rights for the draw"),
"amd" => Some("not available for AMD in this version: the app has no power reading or cap for AMD cards (nothing like nvidia-smi ships with the driver)"),
// Ember Tune (src/ember.rs, 5 October 2026): AMD is tuned through igneum-gpu-telemetry, Apple measures only;
// the tune itself says which at its start (the card row's note)
"apple" | "amd" => None,
_ => Some("not available: no power reading or cap for this card"),
}
}
@ -580,14 +604,16 @@ mod tests {
fn unsupported_reasons() {
assert!(unsupported_reason("nvidia", 575.0, "0").is_none());
assert!(unsupported_reason("nvidia", 0.0, "0").unwrap().contains("power limits"));
assert!(unsupported_reason("apple", 0.0, "").unwrap().contains("powermetrics"));
assert!(unsupported_reason("amd", 0.0, "1").unwrap().contains("AMD"));
assert!(unsupported_reason("apple", 0.0, "").is_none(), "measure only, said by the tune");
assert!(unsupported_reason("amd", 0.0, "1").is_none(), "tuned through igneum-gpu-telemetry");
assert!(unsupported_reason("other", 0.0, "1").is_some());
}
#[test]
fn helper_scripts_carry_the_protocol() {
for s in [helper_script_windows(), helper_script_unix()] {
assert!(s.contains("cmd.txt") && s.contains("quit") && s.contains("-pl") && s.contains("20 min"));
assert!(s.contains("-lgc") && s.contains("-rgc"), "the clock cap and its reset");
}
}
}

View file

@ -215,7 +215,35 @@ var UpdateCard = (function () {
}
return { NAME: NAME, LINE_MAX: LINE_MAX, LINES: LINES, FIRST_S: FIRST_S, sentences: sentences, splitNotes: splitNotes, size: size, model: model, blocked: blocked, decide: decide };
})();
if (typeof module === 'object' && module && module.exports) { module.exports = Notices; module.exports.UpdateCard = UpdateCard; }
/* ---------- Ember Tune: the card row's tuning line (pure; tune-line.test.mjs loads this block) ----------
One line per card from the card state (src/state.rs, src/ember.rs): running (the phase and the step), tuned
("Tuned: 122.3 MH/s at 290 W (0.422 MH/W)" plus the point and when), measure only (Apple, NVIDIA without Power
control: the measured line and why nothing is set), stopped (the reason), or not run yet. */
var TuneLine = (function () {
'use strict';
function point(cd) {
var pct = cd.sweep_pct || cd.power_pct || 0;
if (cd.tune_clock_mhz > 0) return cd.tune_clock_mhz + ' MHz at ' + pct + '%';
return pct ? pct + '%, clock unlocked' : '';
}
// the model: {kind: running|tuned|measured|stopped|idle|off, text, pin, source}
function model(cd, now) {
cd = cd || {};
if (cd.vendor !== 'nvidia' && cd.vendor !== 'amd' && cd.vendor !== 'apple') return { kind: 'off', text: cd.sweep_note || 'tuning: no power or clock control for this card' };
if (cd.sweep_state === 'running') return { kind: 'running', text: cd.sweep_note || 'tuning: running' };
var when = cd.sweep_at && now ? ', ' + rel(now - cd.sweep_at) + ' ago' : '';
if (cd.tune_line) {
if (cd.tune_source === 'baseline' || !cd.tune_control) return { kind: 'measured', text: cd.tune_line + ' (measured as it runs' + when + ')', note: cd.tune_control ? '' : (cd.sweep_note || '') };
var src = cd.tune_source === 'confirm' ? 'from the fleet prior, confirmed' : 'full tune';
return { kind: 'tuned', text: cd.tune_line, note: point(cd) + ', ' + src + when + (cd.pinned ? '; your setting stays pinned' : '') + (cd.sweep_note && cd.sweep_note.indexOf('stopped') === 0 ? '; ' + cd.sweep_note : '') };
}
if (cd.sweep_note && cd.sweep_note.indexOf('tuning stopped') === 0) return { kind: 'stopped', text: cd.sweep_note };
return { kind: 'idle', text: cd.sweep_note || 'tuning: not run yet (starts after 120 s of steady mining)' };
}
function rel(s) { s = Math.max(0, Math.floor(s)); return s < 60 ? s + ' s' : s < 3600 ? Math.floor(s / 60) + ' min' : s < 86400 ? Math.floor(s / 3600) + ' h' : Math.floor(s / 86400) + ' d'; }
return { model: model, point: point, rel: rel };
})();
if (typeof module === 'object' && module && module.exports) { module.exports = Notices; module.exports.UpdateCard = UpdateCard; module.exports.TuneLine = TuneLine; }
if (typeof document !== 'undefined') (function () {
'use strict';
@ -393,7 +421,7 @@ if (typeof document !== 'undefined') (function () {
$('pv-setup').addEventListener('click', function () { api('api/prove/setup', {}).then(function (r) { toast(r.ok ? 'Setup started in its own window' : (r.error || 'could not start')); }); });
$('s-jobs-allow').addEventListener('change', function () { api('api/jobs/allow', { on: this.checked }); setTimeout(fillSettings, 800); });
$('s-power-control').addEventListener('change', function () { var on = this.checked; api('api/power/control', { on: on }).then(function (r) { if (r.ok) toast(on ? 'Power control on: Windows asks for administrator rights once' : 'Power control off: nothing asks for administrator rights'); else { toast(r.error || 'could not change'); $('s-power-control').checked = !on; } setTimeout(fillSettings, 1500); }); });
$('s-sweep').addEventListener('change', function () { api('api/sweep/enable', { on: this.checked }).then(function (r) { if (r.ok) toast($('s-sweep').checked ? 'Sweep on: once after install, then weekly' : 'Sweep off'); }); });
$('s-sweep').addEventListener('change', function () { api('api/sweep/enable', { on: this.checked }).then(function (r) { if (r.ok) toast($('s-sweep').checked ? 'Ember Tune on: once after install, then weekly' : 'Ember Tune off'); }); });
$('s-jobs-check').addEventListener('click', function () { api('api/jobs/check', {}); $('s-jobs-note').textContent = 'Checking.'; setTimeout(fillSettings, 4000); });
// ---------- bottom bar ----------
@ -762,22 +790,20 @@ if (typeof document !== 'undefined') (function () {
}
return h + sweepHtml(cd);
}
// the efficiency sweep line: running (the phase), the last result (best cap and MH/W), or why it cannot run
// Ember Tune's line on the card row (TuneLine.model): running, tuned, measured, stopped or not run yet
function sweepHtml(cd) {
if (!cd.sweep_supported) return cd.sweep_note ? '<div class="msg dim sweep">' + esc(cd.sweep_note) + '</div>' : '';
var running = cd.sweep_state === 'running';
var t = '';
if (running) t = esc(cd.sweep_note || 'sweep: running');
else if (cd.sweep_pct) t = 'sweep: best <b>' + cd.sweep_pct + '%</b> · <b>' + cd.sweep_eff.toFixed(3) + ' MH/W</b> (' + cd.sweep_mhs.toFixed(1) + ' MH/s at ' + Math.round(cd.sweep_watts) + ' W' + (cd.sweep_at ? ', ' + rel(state.now - cd.sweep_at) + ' ago' : '') + ')' + (cd.sweep_note && cd.sweep_note.indexOf('stopped') === 0 ? ' · ' + esc(cd.sweep_note) : '');
else t = esc(cd.sweep_note || 'sweep: not run yet');
var btn = running ? '<button class="btn tiny" data-sweep-stop="1">Stop</button>' : '<button class="btn tiny" data-sweep-start="' + esc(cd.key) + '">Sweep now</button>';
if (!running && cd.pinned) btn += ' <button class="btn tiny" data-sweep-pin="' + esc(cd.key) + '" data-pinned="0" title="let the sweep choose the cap again">Unpin</button>';
var m = TuneLine.model(cd, state.now);
if (m.kind === 'off') return m.text ? '<div class="msg dim sweep">' + esc(m.text) + '</div>' : '';
var running = m.kind === 'running';
var t = m.kind === 'tuned' || m.kind === 'measured' ? '<b>' + esc(m.text) + '</b>' + (m.note ? ' <span class="dim">' + esc(m.note) + '</span>' : '') : esc(m.text);
var btn = running ? '<button class="btn tiny" data-sweep-stop="1">Stop</button>' : '<button class="btn tiny" data-sweep-start="' + esc(cd.key) + '">Tune now</button>';
if (!running && cd.pinned) btn += ' <button class="btn tiny" data-sweep-pin="' + esc(cd.key) + '" data-pinned="0" title="let the tune choose the point again">Unpin</button>';
return '<div class="msg sweep' + (running ? ' on' : '') + '">' + t + ' ' + btn + '</div>';
}
document.addEventListener('click', function (e) {
var b = e.target.closest('[data-sweep-start],[data-sweep-stop],[data-sweep-pin]'); if (!b) return;
if (b.dataset.sweepStart) api('api/sweep/start', { key: b.dataset.sweepStart }).then(function () { toast('Sweep queued: 100% down to 50%, 75 s a step'); });
else if (b.dataset.sweepStop) api('api/sweep/stop', {}).then(function () { toast('Sweep stopped; cap restored'); });
if (b.dataset.sweepStart) api('api/sweep/start', { key: b.dataset.sweepStart }).then(function () { toast('Tune queued: the power limit and the core clock, 75 s a step'); });
else if (b.dataset.sweepStop) api('api/sweep/stop', {}).then(function () { toast('Tune stopped; the card is back where it was'); });
else api('api/sweep/pin', { key: b.dataset.sweepPin, pinned: b.dataset.pinned === '1' }).then(function () { toast(b.dataset.pinned === '1' ? 'Cap pinned' : 'The sweep chooses the cap again'); });
});
document.addEventListener('click', function (e) {
@ -1059,7 +1085,8 @@ if (typeof document !== 'undefined') (function () {
$('s-power-control').checked = pc;
$('s-power-note').textContent = (state.settings && state.settings.power_note) || '';
$('s-sweep').checked = !!(state.settings && state.settings.sweep);
$('s-sweep').disabled = !pc;
$('s-sweep').disabled = false;
var tn = $('s-tune-note'); if (tn) tn.textContent = state.settings && state.settings.tuning_off ? state.settings.tuning_note : (pc ? '' : 'NVIDIA cards measure only until Power control is on.');
$('s-jobs-key').textContent = j.key_fingerprint ? 'signing key sha256:' + j.key_fingerprint : '';
var parts = [];
if (j.account) parts.push(j.account);

View file

@ -85,7 +85,7 @@
</div>
</div>
<p class="note" id="cards-note" hidden></p>
<p class="note" id="cards-power" hidden>Igneum caps each NVIDIA card's power at 80% of its default limit to keep it stable (an RTX 5090 at full power hard-crashed in the field). This needs administrator rights once, when mining starts; the limit goes back to what it was on quit. The slider sets the cap per card. In the first hour, and then weekly, the efficiency sweep steps the cap from 100% down to 50% on the live program and holds the best MH per watt; a cap you set by hand stays pinned.</p>
<p class="note" id="cards-power" hidden>Igneum caps each NVIDIA card's power at 80% of its default limit to keep it stable (an RTX 5090 at full power hard-crashed in the field). This needs administrator rights once, when mining starts; the limit goes back to what it was on quit. The slider sets the cap per card. In the first hour, and then weekly, Ember Tune steps the power limit and the core clock on the live program and holds the best MH per watt (the memory clock is never touched); a cap you set by hand stays pinned.</p>
<p class="note" id="cards-help" hidden>Each card you switch on gets its own worker. Identities take turns on the card and each one votes and pays separately; 8 suits a big card, 2 a small one, 1 an integrated GPU.</p>
<div class="cta">
<button class="btn primary" id="btn-cards-next" disabled>Continue</button>
@ -301,13 +301,13 @@
</div>
<div class="field">
<div class="k">power control</div>
<label class="switch"><input type="checkbox" id="s-power-control"><span class="track"></span><span>Cap the NVIDIA cards' power and run the efficiency sweep</span></label>
<p class="note">Windows asks for administrator rights once; the cap and the sweep need them. Off, the app never asks. <span id="s-power-note"></span></p>
<label class="switch"><input type="checkbox" id="s-power-control"><span class="track"></span><span>Cap the NVIDIA cards' power and let Ember Tune set their limits</span></label>
<p class="note">Windows asks for administrator rights once; the NVIDIA cap and the tune need them. Off, the app never asks and NVIDIA cards measure only. AMD cards need no rights. <span id="s-power-note"></span></p>
</div>
<div class="field">
<div class="k">efficiency sweep</div>
<label class="switch"><input type="checkbox" id="s-sweep"><span class="track"></span><span>Find each NVIDIA card's best MH per watt (once after install, then weekly)</span></label>
<p class="note">On the live program, never restarting the worker: the cap steps from 100% of the card's default limit down to 50%, 15 s to settle and 60 s to measure per step, then holds the step with the most MH per watt. One administrator prompt per sweep. It stops at once if the card faults, a remote job takes the GPU, or the hour boundary is near. A cap you set with the slider is pinned: the sweep records, but leaves it. "Sweep now" on a card's tile runs one at any time; the table is in the log (SWEEP lines).</p>
<div class="k">Ember Tune</div>
<label class="switch"><input type="checkbox" id="s-sweep"><span class="track"></span><span>Tune every card for MH per watt out of the box (once after install, then weekly, and after a driver or program change)</span></label>
<p class="note">On the live program, never restarting the worker: the power limit steps from 100% of the card's default down to 50%, then the core clock from its maximum down to 60%, 15 s to settle and 60 s to measure per step; the memory clock is never touched. The card keeps the point with the best MH per watt within 1% of its top rate. A step with a rejected or mismatched hash, a hot GPU or a dragged memory clock is reverted and marked. A card whose model the fleet already knows (5 or more reports) starts at that prior and confirms it in two steps. Every result goes back to the fleet without anything that identifies you. A cap you set with the slider is pinned: the tune records, but leaves it. "Tune now" on a card's row runs one at any time; the table is in the log (TUNE lines). <span id="s-tune-note"></span></p>
</div>
<div class="field">
<div class="k">version</div>

View file

@ -0,0 +1,44 @@
// node --test app/igneum-app/ui/tune-line.test.mjs (no dependencies; CI runs it in the site job)
// The card row's Ember Tune line (app.js TuneLine): what a user sees per state, from the card state fields.
import { test } from 'node:test';
import assert from 'node:assert/strict';
import { readFileSync } from 'node:fs';
import { fileURLToPath } from 'node:url';
import { dirname, join } from 'node:path';
const src = readFileSync(join(dirname(fileURLToPath(import.meta.url)), 'app.js'), 'utf8');
const mod = { exports: {} };
new Function('module', src)(mod);
const { model, point } = mod.exports.TuneLine;
const NOW = 1_800_000_000;
const card = over => ({ vendor: 'nvidia', sweep_state: 'idle', sweep_note: '', sweep_pct: 0, power_pct: 80, sweep_at: 0, tune_line: '', tune_source: '', tune_clock_mhz: 0, tune_control: true, pinned: false, ...over });
test('tuned: the line the brief asks for, with the point, the source and when', () => {
const m = model(card({ tune_line: 'Tuned: 122.3 MH/s at 290 W (0.422 MH/W)', tune_source: 'full', tune_clock_mhz: 2470, sweep_pct: 100, sweep_at: NOW - 3600 }), NOW);
assert.equal(m.kind, 'tuned');
assert.equal(m.text, 'Tuned: 122.3 MH/s at 290 W (0.422 MH/W)');
assert.equal(m.note, '2470 MHz at 100%, full tune, 1 h ago');
const c = model(card({ tune_line: 'Tuned: 17.7 MH/s at 177 W (0.100 MH/W)', tune_source: 'confirm', sweep_pct: 90, sweep_at: NOW - 120, vendor: 'amd' }), NOW);
assert.equal(c.note, '90%, clock unlocked, from the fleet prior, confirmed, 2 min ago');
const p = model(card({ tune_line: 'Tuned: 1 MH/s at 1 W (1.000 MH/W)', tune_source: 'full', sweep_pct: 70, pinned: true }), NOW);
assert.match(p.note, /your setting stays pinned$/);
});
test('measure only: Apple and NVIDIA without Power control say so beside the measured line', () => {
const a = model(card({ vendor: 'apple', tune_control: false, tune_line: 'Tuned: 26.7 MH/s at 38 W (0.703 MH/W)', tune_source: 'baseline', sweep_note: 'measure only on Apple silicon: the system sets the clocks and the power; no control exposed', sweep_at: NOW - 60 }), NOW);
assert.equal(a.kind, 'measured');
assert.equal(a.text, 'Tuned: 26.7 MH/s at 38 W (0.703 MH/W) (measured as it runs, 1 min ago)');
assert.match(a.note, /^measure only on Apple silicon/);
const n = model(card({ tune_control: false, tune_line: 'Tuned: 122.3 MH/s at 290 W (0.422 MH/W)', tune_source: 'baseline', sweep_note: 'measure only until Power control is on in Settings (Windows asks for administrator rights once)' }), NOW);
assert.equal(n.kind, 'measured');
assert.match(n.note, /Power control/);
});
test('running, stopped, idle and off', () => {
assert.deepEqual(model(card({ sweep_state: 'running', sweep_note: 'tuning: holding 2472 MHz · 100% · 41 s (step 7 of 9)' }), NOW), { kind: 'running', text: 'tuning: holding 2472 MHz · 100% · 41 s (step 7 of 9)' });
assert.equal(model(card({ sweep_note: 'tuning stopped: a remote job took the GPU' }), NOW).kind, 'stopped');
assert.equal(model(card({}), NOW).text, 'tuning: not run yet (starts after 120 s of steady mining)');
assert.equal(model(card({ sweep_note: 'tuning: waits for 120 s of steady mining' }), NOW).text, 'tuning: waits for 120 s of steady mining');
assert.equal(model(card({ vendor: 'other' }), NOW).kind, 'off');
assert.equal(point({ tune_clock_mhz: 0, sweep_pct: 0, power_pct: 0 }), '');
});

164
docs/plans/ember-tune.md Normal file
View file

@ -0,0 +1,164 @@
# Ember Tune: every card tuned for MH per watt, out of the box
5 October 2026, night. the project lead: "make sure we have ember tuning every single card for efficiency out of the box, the
more data = the better the tune, make an awesome system." Branch `ember-tune`, worktree `../igneum-wt-ember-tune`,
on top of the AMD telemetry commit (7adcd4c, branch `opencl-rdna4-telemetry`) and the Power control commit (3562f26,
branch `job-console`), both cherry-picked. Lever 3 of docs/plans/miner-eff.md grows two knobs and a fleet memory;
lever 2 (docs/design/miner-tuning.md) carries the priors in the same signed `tuning` section.
## 1. What a user sees
| Moment | The card row says | What happened |
|---|---|---|
| First 2 minutes of mining | `tuning: waits for 120 s of steady mining` | The worker warms up; nothing is touched. |
| Tuning | `tuning: holding 2472 MHz · 100% · 41 s (step 7 of 9)` with a Stop button | One card at a time, on the live kernel, never restarting the worker. |
| Tuned | **Tuned: 122.3 MH/s at 290 W (0.422 MH/W)**, then `2470 MHz at 100%, full tune, 1 h ago` | The point is pinned on the card; the result went to the fleet. |
| Known model | the same line, `from the fleet prior, confirmed, 2 min ago` | The card started at its model's prior and confirmed it in two steps instead of nine. |
| Apple silicon | **Tuned: 26.7 MH/s at 38 W (0.703 MH/W)** `(measured as it runs)`, and `measure only on Apple silicon: the system sets the clocks and the power; no control exposed` | Nothing can be set; the number is still reported so the row and the fleet know what the card does. |
| NVIDIA, Power control off | the measured line and `measure only until Power control is on in Settings (Windows asks for administrator rights once)` | The app never raises the prompt by itself (5 October 2026). One switch, one prompt, and the full tune runs. |
| Slider moved | `your setting stays pinned` | A manual point is never overridden; the tune still measures and reports. |
| Stopped | `tuning stopped: a remote job took the GPU` and the card back where it was | Any fault reverts the step and the run. |
| Fleet pause | Settings: `tuning paused fleet-wide by the signed manifest` | The kill switch. |
Settings: one switch, "Ember Tune: tune every card for MH per watt out of the box (once after install, then weekly,
and after a driver or program change)". AMD needs no rights. NVIDIA needs the Power control switch (one administrator
prompt) for both knobs; off, it measures only.
## 2. The knobs, per vendor
| Vendor | Power limit | Core clock cap | Memory clock | How | Rights |
|---|---|---|---|---|---|
| NVIDIA | `nvidia-smi -pl <W>`, percent of the default, inside `power.min_limit` and `power.max_limit` | `nvidia-smi -lgc 0,<MHz>`, percent of `clocks.max.gr`; `-rgc` = unlocked | never touched (`-lmc` is not used); read back as `clocks.mem` | directly when the engine is elevated, else the one-prompt helper (`<seq> pl <W>`, `<seq> lgc <MHz>`, `<seq> rgc` in `sweep/cmd.txt`) | administrator, so only with Power control on |
| AMD | `igneum-gpu-telemetry --card N --set-plimit <offset>` (0 = default, -20 = 80%), inside the `tune` line's `plimit_range` | `--set-gmax <MHz>` inside `gmax_range`; `--reset` for the default point | not settable through ADLX on RDNA 4; read back as `mclk_mhz`, and a step whose mean memory clock falls under 95% of the baseline's is marked and cannot win | the helper, one process per request, exit 0 and a `tune ... ok` line | none on Windows (ADLX manual tuning); root on Linux, so measure only there |
| Apple | none | none | none | measure only | none |
Vendor limits are never exceeded and the floor is never undercut: the plan clamps every point (`Limits::clamp_clock`,
`Limits::watts_for`), and a clock floor the vendor does not report is 60% of the maximum.
## 3. The plan and the choice
Full plan (a new model, or a prior that lost its confirm check): the power ladder 100, 90, 80, 70, 60, 50% at the
unlocked clock (duplicate watts dropped where the card's floor clamps them), then the clock ladder 90, 80, 70, 60%
of the maximum at the power point the power ladder chose. 60 s hold after 15 s settle per step; 9 steps on an
RTX 5090 (five power, four clock), about 12 minutes.
Confirm plan (the model's prior has 5 or more reports): the prior's point, then one neighbour (the next clock step up
when the prior caps the clock, else one power step down). If the neighbour beats the prior by over 1% MH/W, the full
plan is queued; else the prior stands. Two steps, about 3 minutes.
Baseline plan (measure only): one step at the card's current point. The "before" number for the row and the fleet.
The choice (`ember::choose`): among the usable steps whose rate is within the tolerance (1%, settable from the
manifest) of the fastest step, the best MH per watt; within 1% on efficiency the higher rate; within 1% on both the
lower draw. A card never gives up more than the tolerance in blocks for the saving. A step is unusable when it is
marked: `faulted` (a rejected or mismatched hash during the hold: the step is reverted and marked), `hot` (the GPU
reached 85 C; the run aborts at 90), `memory_clock_dropped`, `unapplied` (the readback disagreed with the request),
`no_readings` (under three draw samples or no STATUS line).
## 4. The data flow
```
card mines 120 s ──> probe (limits, driver, how to set) ──> plan ──> steps ──> choice ──> point pinned
│
app log: TUNE start / TUNE card=.. step=.. / TUNE chosen / TUNE {json} (and stdout under --sweep)
│
log upload (every minute, the existing intake, site/api/log.mjs) ──> Neon miner_logs
│
relay/lib/ember.mjs aggregate: per (card model | driver major | program class)
median clock cap (10 MHz), median power %, median MH/W, MH/s, W, spread (MAD %), samples, machines
│ │
console: /r/<token>/c/tuning, `node tools/console.mjs tuning` site: tools/tuning.mjs --priors --site
│ -> site/miner-priors.json -> /miners#priors
tools/tuning.mjs --priors --write tuning.json (priors + ember settings beside the kernel-variant cards)
│
packaging/ota/publish-manifest.sh --tuning tuning.json --deploy (signed; carried over when not given)
│
every app: <app data>/tuning.json ──> ember::settings_of (kill switch, min samples, tolerance, period)
──> ember::prior_of(key) ──> a new card's confirm plan
```
The record (`ember::record_json`): `ts`, `machine` (the first 8 hex of SHA-256 over the install id; the id itself
is random per install and never sent), `app`, `os`, `card`, `vendor`, `driver`, `driver_major`, `class`, `key`,
`plan`, `steps` (the full table: clock, power %, limit, watts, MH/s, MH/W, core and memory clock, hottest reading,
faults, mark), `chosen`, `before` (the full plan's 100% step), `eff`, `mhs`, `watts`. The key: `<card model with
underscores>|<driver major>|<program class>`, the class from the worker's race line (`l128w16` today; `v2` before a
race has run).
## 5. Scheduling and safety
| Rule | Where |
|---|---|
| One card at a time; the card must have mined 120 s and have a STATUS line | `tick_sweep` |
| Never under a remote job hold, a pause, inside 600 s of the hour boundary, or while the app quits | `tick_sweep`, `sweep_drive` |
| Due once after install, every 7 days (manifest `ember.period_s`), and when the driver major or the program class changed since the last tune | `tick_sweep` (`CardPref.sweep_driver`, `sweep_class`) |
| A pinned card (the slider) is measured, never changed | `sweep_finish` |
| Kill switch: `tuning.ember.enabled = false` in the signed manifest stops every tune fleet-wide; the Settings line says so | `ember::settings_of`, `tick_sweep` |
| Faults: a rejected or mismatched hash marks the step; the card leaving `mining`, a worker error, a job, a pause or 90 C aborts the run and restores the point from before | `Run::sample_fault`, `sweep_drive`, `sweep_abort` |
| Memory clock held: never set; a step that drags it under 95% of the baseline's cannot win | `Row::from_samples` |
| Vendor limits: every point clamped to the reported range; the clock floor 60% when none is reported | `Limits` |
| No prompt the user did not ask for: the NVIDIA helper starts only with Power control on; the `--sweep` job never counts as permission | `sweep_probe_known`, `sweep_helper_start` |
| The elevated helper restores the limit and resets the clocks by itself after 20 idle minutes | `sweep::helper_script_*` |
## 6. Tests
| Test | What it fixes |
|---|---|
| `ember::tests::the_full_plan_is_the_power_ladder_then_the_clock_ladder_at_the_chosen_power` | 5 + 4 steps on the 5090's limits, the clamps, the dynamic second half, the 1% and 5% choices |
| `limits_never_exceed_the_vendor_or_undercut_the_floor` | clamps |
| `the_choice_keeps_the_best_mh_per_watt_within_the_rate_tolerance` | the rule, the ties, marked rows never win |
| `the_guards_mark_a_step_so_it_cannot_win` | faulted, hot, memory clock, unapplied, no readings, the line |
| `a_fault_during_a_step_reverts_it_and_the_run_goes_on` | the state machine with a fake clock: the faulted 70% step is marked and never chosen |
| `the_confirm_plan_checks_the_prior_and_its_neighbour` | the two steps, Keep against FullDue, a prior outside the range clamped |
| `a_baseline_plan_measures_the_card_as_it_runs` | no control, still a number and the Tuned line |
| `the_record_and_the_prior_round_trip_through_the_manifest_shape` | record fields (no address, no host), `priors` and `ember` beside `cards`, the sample floor, the kill switch |
| `control_reasons_per_vendor` | who measures only and why |
| `sweep::tests::helper_scripts_carry_the_protocol` | the helper's `pl`, `lgc`, `rgc` |
| `relay/test/ember.test.mjs` | five samples converge (2,470 MHz at 100%), an outlier (0.908 MH/W at 1,854 MHz) moves nothing, baseline records make no prior, de-duplication, the manifest merge keeps lever 2's cards, the canonical round trip, AMD keys |
| `app/igneum-app/ui/tune-line.test.mjs` | the row line per state |
Run: `cargo test -p igneum-app ember sweep` (on a PC through the build job, or on the Mac under the build lock),
`node --test relay/test/ember.test.mjs app/igneum-app/ui/tune-line.test.mjs`.
## 7. The tier consequences
| Tier | What Ember Tune does for it | What it costs |
|---|---|---|
| A laptop GPU (NVIDIA, 60 to 115 W) | the power ladder usually finds the vendor floor binding; the clock ladder is where a memory-bound program saves watts; the thermal mark keeps a hot chassis from winning a step it cannot hold | about 12 minutes once, then 3 minutes a week; under 1% of the hour during the tune (the worker never stops) |
| One 8 GB card | the same two knobs; the 8 GB card is identities-limited (2 by default), the tune does not change that | the same |
| One 12 or 16 GB card | the same | the same |
| One 24 or 32 GB card (the 5090) | the draw sits far under the cap (290 W under 460 W on PC 1), so the power ladder is flat and the clock ladder is the lever; expected saving from the 4 October stability line: tens of watts at under 1% rate, to be measured | the same |
| A rig (several cards) | one card at a time, so a six-card rig takes about 70 minutes to tune once; every card of one model after the first starts at the prior (3 minutes); the tune never touches a card a remote job holds | linear in cards once, then the confirm plan |
| A pool user | the same per card; a pool submits the same hashes, so the 1% rate tolerance is the same 1% of shares | the same |
| AMD on Linux | measure only (sysfs needs root); the row says so | 60 s a week |
| Apple silicon | measure only; the row says so | 60 s a week |
Privacy line: what is uploaded is the record in section 4 and nothing else: a hash of the random install id, the
card model, the driver version, the OS, the program class, the step table and the chosen point. No address, no
hostname, no raw machine id, no user name. The public priors table carries only the aggregate per model.
## 8. Measurements
### PC 1, 5 October 2026 (night)
Tonight's constraints, read from PC 1's own uploads: the installed app runs as `DESKTOP-KMCV30N\Admin` with
`elevated=False` (the account line at 19:02:33 UTC), the two in-app sweep attempts at 20:09 UTC aborted on the
cancelled administrator prompt (`SWEEP aborted ... the_elevated_helper_did_not_run_(the_administrator_prompt_was_cancelled)`),
so no stored sweep result exists from today, and the RX 9070 XT left the PCI bus at about 20:40 UTC (eGPU link,
not restarted tonight). NVIDIA's `-pl` and `-lgc` need administrator rights, the project lead is asleep, and the app never raises
the prompt by itself, so tonight's run on PC 1 is the baseline plan on the 5090 through the whole pipeline (probe,
measure, TUNE record, upload, aggregation, prior shape in a test manifest). The two-knob tune on the 5090 and the
9070 XT run are owed: the 5090 the moment Power control is switched on (one prompt, then the tune runs by itself
within 2 minutes of steady mining), the 9070 XT when the card is back on the bus.
(The run's numbers are appended below when the job reports.)
## 9. Open
- The NVIDIA clock readback: `nvidia-smi -lgc` is confirmed only through the core clock during the hold (a mean over
the cap by 5% marks the step `unapplied`); the first run with Power control on tells whether the driver honours
the lock on the 5090 under this kernel.
- ADLX on RDNA 4 exposes no memory-clock setter; the memory-clock mark is the guard. The telemetry agent's 9070 XT
sweep tells whether a core cap drags the memory clock on that card.
- The confirm plan's neighbour is one step; a second neighbour (the other knob) would cost 75 s more and catch a
prior that is wrong on both knobs.
- Intel: no knob yet; the row says measure only.

View file

@ -9,10 +9,13 @@
// GET chain igneum.network/api/live trimmed + the Hetzner results item
// GET log?limit=&since= work-log items (kinds log, build, note), newest first
// GET results bench entries (synced from docs/bench-log.md) + the FUD ledger counts
// GET tuning?days=30&min=5 Ember Tune: the fleet priors per (card model, driver major, program class) from the
// TUNE records in miner_logs (relay/lib/ember.mjs), with the sample counts and MH/W
// POST post {kind,title,body,who,key?,meta?} one item; with key it upserts
// POST sync {items:[...]} bulk upsert by key
import { neon, authed, readJson, str, iso } from '../lib/relay.mjs';
import { kv, kvNum, lastMatch, FAULT, parseLabel, parseMinerTail, parseHeader, parseAppTail, STALE_S, markStale } from '../lib/parse.mjs';
import { parseRecords, aggregate } from '../lib/ember.mjs';
const json = (res, status, obj) => { res.status(status).setHeader('Content-Type', 'application/json; charset=utf-8'); res.end(JSON.stringify(obj)); };
const CACHE_MS = 10_000;
@ -212,6 +215,13 @@ async function chain(sql) {
hetzner: het.length ? itemOut(het[0]) : null,
};
}
/// Ember Tune: every TUNE record of the window, folded into priors (the same aggregation the publisher uses).
async function tuning(sql, days, min) {
const rows = await sql(`SELECT lines FROM miner_logs WHERE received_at > now() - ($1 || ' days')::interval AND lines LIKE '%TUNE {%' ORDER BY received_at DESC LIMIT 2000`, [String(days)]);
const records = rows.flatMap(r => parseRecords(r.lines));
const { priors, table } = aggregate(records, { minSamples: min });
return { days, min_samples: min, records: records.length, priors, table };
}
async function results(sql) {
const [bench, ledger] = await Promise.all([
sql(`SELECT * FROM console_items WHERE kind = 'bench' ORDER BY (meta->>'date') DESC NULLS LAST, (meta->>'pos')::int DESC LIMIT 60`),
@ -252,6 +262,11 @@ export default async function handler(req, res) {
if (fn === 'builds') return json(res, 200, { ok: true, now: new Date().toISOString(), ...(await cached('builds', () => builds(sql))) });
if (fn === 'chain') return json(res, 200, { ok: true, ...(await cached('chain', () => chain(sql))) });
if (fn === 'results') return json(res, 200, { ok: true, ...(await cached('results', () => results(sql))) });
if (fn === 'tuning') {
const days = Math.min(365, Math.max(1, Number(q.days) || 30));
const min = Math.min(100, Math.max(1, Number(q.min) || 5));
return json(res, 200, { ok: true, now: new Date().toISOString(), ...(await cached(`tuning-${days}-${min}`, () => tuning(sql, days, min))) });
}
if (fn === 'log') {
const limit = Math.min(300, Math.max(1, Number(q.limit) || 100));
const params = []; let where = `kind IN ('log','build','note')`;

131
relay/lib/ember.mjs Normal file
View file

@ -0,0 +1,131 @@
// Ember Tune, the fleet side (docs/plans/ember-tune.md): the TUNE records every app uploads with its log are folded
// into one prior per (card model, driver major, program class): the median chosen point, its spread and the sample
// count. The publisher writes the priors into the signed manifest's `tuning` section beside the kernel-variant
// cards (tools/tuning.mjs --write), the console shows them (api/console.mjs fn=tuning, tools/console.mjs tuning),
// and the public bench table lists them per model (site/miner-priors.json). No dependencies; the tests in
// relay/test/ember.test.mjs drive these functions with a fixture of captured records.
//
// A record (app/igneum-app/src/ember.rs record_json): {ts, machine (a hash of the install id), app, os, card, vendor,
// driver, driver_major, class, key, plan: full|confirm|baseline, steps: [{clock_mhz, power_pct, limit_w, watts, mhs,
// eff, gclk, mclk, tmax, faults, mark}], chosen: {...}, before: {...}|null, eff, mhs, watts}. Nothing identifies the
// owner: no address, no hostname, no raw machine id.
/// The TUNE records inside uploaded log text, de-duplicated on (machine, card, ts) because the log is re-sent every
/// minute. Baseline records (measure only) are kept apart: they say what a card does untuned, never what to set.
export function parseRecords(text) {
const out = [];
for (const line of String(text || '').split('\n')) {
const i = line.indexOf('TUNE {');
if (i < 0) continue;
let rec;
try { rec = JSON.parse(line.slice(i + 5)); } catch { continue; }
if (!rec || !rec.card || !rec.key || !rec.plan) continue;
out.push(rec);
}
return out;
}
export function dedupe(records) {
const seen = new Set();
const out = [];
for (const r of records) {
const k = `${r.machine}|${r.card}|${r.ts}`;
if (seen.has(k)) continue;
seen.add(k);
out.push(r);
}
return out;
}
export const median = xs => { const s = xs.filter(x => Number.isFinite(x)).sort((a, b) => a - b); return s.length ? (s.length % 2 ? s[(s.length - 1) / 2] : (s[s.length / 2 - 1] + s[s.length / 2]) / 2) : 0; };
/// The median absolute deviation as a percent of the median (0 for one sample or a zero median).
export function spreadPct(xs) {
const m = median(xs);
if (!m || xs.length < 2) return 0;
return Number((median(xs.map(x => Math.abs(x - m))) / m * 100).toFixed(2));
}
const usable = r => r && r.chosen && r.chosen.mark === 'ok' && r.chosen.eff > 0 && (r.plan === 'full' || r.plan === 'confirm');
/// Folds records into priors: one per key, from the full and confirm records with a usable chosen point. The point
/// is the median clock cap and the median power percent (each rounded to the step the apps use: 10 MHz, 1%), the
/// efficiency, rate and draw are medians, the spread is the MAD of the efficiency in percent, `samples` counts the
/// records and `machines` the distinct install hashes. An outlier (one bad card, one hot room) moves the median by
/// at most one rank, never by its size. Baseline records are summarised beside the prior as `baseline` (median
/// MH/W untuned) so the console can show the gain.
export function aggregate(records, { minSamples = 1 } = {}) {
const byKey = new Map();
for (const r of dedupe(records)) {
const g = byKey.get(r.key) || { key: r.key, card: r.card, vendor: r.vendor || '', driver_major: r.driver_major || '', class: r.class || 'v2', tuned: [], baseline: [], machines: new Set() };
byKey.set(r.key, g);
g.machines.add(r.machine);
if (usable(r)) g.tuned.push(r);
else if (r.plan === 'baseline' && r.chosen && r.chosen.eff > 0) g.baseline.push(r);
}
const priors = {};
const table = [];
for (const g of byKey.values()) {
const t = g.tuned;
const row = {
key: g.key, card: g.card, vendor: g.vendor, driver_major: g.driver_major, class: g.class,
samples: t.length, machines: g.machines.size,
baseline_samples: g.baseline.length,
baseline_eff: g.baseline.length ? Number(median(g.baseline.map(r => r.chosen.eff)).toFixed(4)) : null,
baseline_mhs: g.baseline.length ? Number(median(g.baseline.map(r => r.chosen.mhs)).toFixed(2)) : null,
baseline_watts: g.baseline.length ? Number(median(g.baseline.map(r => r.chosen.watts)).toFixed(1)) : null,
};
if (t.length) {
const effs = t.map(r => r.chosen.eff);
const prior = {
clock_mhz: Math.round(median(t.map(r => r.chosen.clock_mhz)) / 10) * 10,
power_pct: Math.round(median(t.map(r => r.chosen.power_pct))),
eff: Number(median(effs).toFixed(4)),
mhs: Number(median(t.map(r => r.chosen.mhs)).toFixed(2)),
watts: Number(median(t.map(r => r.chosen.watts)).toFixed(1)),
spread_pct: spreadPct(effs),
samples: t.length,
machines: g.machines.size,
card: g.card,
vendor: g.vendor,
driver_major: g.driver_major,
class: g.class,
updated: new Date(Math.max(...t.map(r => Number(r.ts) || 0)) * 1000).toISOString().replace(/\.\d{3}Z$/, 'Z'),
};
// the untuned reference: the full plan's first step (the power ladder's 100%), else the baseline records
const befores = t.map(r => r.before && r.before.eff > 0 ? r.before.eff : null).filter(x => x !== null);
if (befores.length) prior.before_eff = Number(median(befores).toFixed(4));
else if (row.baseline_eff) prior.before_eff = row.baseline_eff;
if (prior.before_eff) prior.gain_pct = Number(((prior.eff / prior.before_eff - 1) * 100).toFixed(1));
Object.assign(row, prior);
if (t.length >= minSamples) priors[g.key] = prior;
}
table.push(row);
}
table.sort((a, b) => (b.samples - a.samples) || (a.key < b.key ? -1 : 1));
return { priors, table };
}
/// The manifest's tuning section with the priors folded in: the kernel-variant `cards` object is kept as is,
/// `priors` replaces the previous priors (a key that lost its samples drops out), `ember` carries the settings.
export function mergeTuning(existing, priors, ember = {}) {
const base = existing && typeof existing === 'object' ? existing : {};
const cards = base.cards && typeof base.cards === 'object' && !Array.isArray(base.cards) ? base.cards : {};
const settings = { enabled: true, min_samples: 5, rate_tolerance_pct: 1, ...(base.ember && typeof base.ember === 'object' ? base.ember : {}), ...ember };
return { ...base, updated: new Date().toISOString().replace(/\.\d{3}Z$/, 'Z'), cards, ember: settings, priors: priors || {} };
}
/// A prior as a card starts from it (app/igneum-app/src/ember.rs prior_of): None under the sample floor.
export function priorFor(tuning, key, minSamples) {
const p = tuning && tuning.priors && tuning.priors[key];
const floor = Number.isFinite(minSamples) ? minSamples : (tuning && tuning.ember && tuning.ember.min_samples) || 5;
if (!p || !(p.samples >= floor)) return null;
return { clock_mhz: p.clock_mhz || 0, power_pct: Math.min(100, Math.max(50, p.power_pct || 100)), eff: p.eff, samples: p.samples };
}
/// One text line per prior for the console and the CLI.
export function priorLine(p) {
const point = p.clock_mhz ? `${p.clock_mhz} MHz at ${p.power_pct}%` : `${p.power_pct}% (clock unlocked)`;
const gain = p.gain_pct != null ? ` (${p.gain_pct >= 0 ? '+' : ''}${p.gain_pct}% over untuned ${p.before_eff} MH/W)` : '';
return `${p.card.replace(/_/g, ' ')} | driver ${p.driver_major} | ${p.class}: ${point}, ${p.eff} MH/W${gain}, ${p.mhs} MH/s at ${p.watts} W, spread ${p.spread_pct}%, ${p.samples} sample(s) from ${p.machines} machine(s)`;
}

View file

@ -0,0 +1,143 @@
# Igneum run job: Ember Tune end to end on PC 1 (machine ae432dc7), unattended. 5 October 2026.
# Published as a `run` job with --stop-miners (docs/plans/ember-tune.md): the installed app stops its miners and
# holds them; this script takes the engine that carries src/ember.rs (the one the fetch job put in
# <app dir>\jobs\ember-kit-1\igneum-app-ember.exe, else the installed one; the AMD helper with the --tune and --set
# commands from the telemetry agent's fetch job, <app dir>\jobs\amd-kit-1\kit\igneum-gpu-telemetry.exe), copies the install folder to a scratch
# folder beside it, swaps the engine in, and starts that SECOND engine with `--sweep` in a scratch data folder (the
# real settings.json, machine-id and wallet.json copied in; remote jobs, auto-update and proving switched off
# there). That engine finds the installed app's node on 127.0.0.1:26610, mines on every card with the live program,
# tunes them one after the other (the full two-knob plan where the card can be controlled, the baseline measurement
# where it cannot: NVIDIA without administrator rights, Apple), prints every TUNE line on stdout, uploads its log
# (the TUNE {json} record reaches the intake) and quits. Every TUNE line is re-emitted as a RESULT line, so
# `node tools/jobs.mjs <job id>` shows the table. Before and after, nvidia-smi's limits and clocks and the AMD
# helper's `--tune` lines are printed, so the restore can be read. The installed app's miners restart when the job
# ends. Not elevated: nothing asks for administrator rights (the project lead asleep, 5 October 2026); the NVIDIA card is
# therefore measure only tonight unless the engine finds itself elevated.
$ErrorActionPreference = 'Continue'
$budgetMinutes = 35
$started = Get-Date
$deadline = $started.AddMinutes($budgetMinutes)
function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
# the installed engine: the per-user install (0.3.3+), else Program Files
$installDir = $null
foreach ($d in @((Join-Path $env:LOCALAPPDATA 'Programs\Igneum Miner'), (Join-Path $env:ProgramFiles 'Igneum Miner'))) {
if (Test-Path (Join-Path $d 'igneum-app.exe')) { $installDir = $d; break }
}
if (-not $installDir) { Write-Output 'RESULT TUNE error=no_engine reason=igneum-app.exe_not_found'; exit 2 }
$appData = $env:IGNEUM_APP_DATA
if (-not $appData) { $appData = Join-Path $env:LOCALAPPDATA 'igneum' }
$appDir = $env:IGNEUM_APP_DIR
if (-not $appDir) { $appDir = Join-Path $appData 'app' }
# the scratch install: the whole folder (workers, node, helper, DLLs) with the Ember engine swapped in
$root = Join-Path $env:LOCALAPPDATA 'igneum-tune'
$bin = Join-Path $root 'bin'
$sApp = Join-Path $root 'app'
$sLogs = Join-Path $root 'logs'
New-Item -ItemType Directory -Force -Path $root, $sApp, $sLogs | Out-Null
if (Test-Path $bin) { Remove-Item -LiteralPath $bin -Recurse -Force -ErrorAction SilentlyContinue }
Copy-Item -LiteralPath $installDir -Destination $bin -Recurse -Force
$ember = Join-Path $appDir 'jobs\ember-kit-1\igneum-app-ember.exe'
if (Test-Path $ember) {
Copy-Item -LiteralPath $ember -Destination (Join-Path $bin 'igneum-app.exe') -Force
Say ("engine: the Ember build from " + $ember)
} else {
Say 'engine: the installed one (no jobs\ember-kit-1\igneum-app-ember.exe); an older engine ignores the tune and reports no_rows'
}
$helper = Join-Path $appDir 'jobs\amd-kit-1\kit\igneum-gpu-telemetry.exe'
if (Test-Path $helper) {
Copy-Item -LiteralPath $helper -Destination (Join-Path $bin 'igneum-gpu-telemetry.exe') -Force
Say ("helper: the Ember build of igneum-gpu-telemetry from " + $helper + " sha256=" + (Get-FileHash -LiteralPath $helper -Algorithm SHA256).Hash.ToLower())
} else { Say 'helper: the installed igneum-gpu-telemetry (no jobs\amd-kit-1\kit\igneum-gpu-telemetry.exe); without --tune the AMD card measures only' }
$exe = Join-Path $bin 'igneum-app.exe'
$ver = (& $exe --version 2>&1 | Out-String).Trim()
Say ("engine: " + $exe + " (" + $ver + ")")
Write-Output ("RESULT TUNE engine " + $ver + " sha256=" + (Get-FileHash -LiteralPath $exe -Algorithm SHA256).Hash.ToLower())
if ($ver -notmatch 'igneum-app (\d+)\.(\d+)\.(\d+)') { Write-Output 'RESULT TUNE error=version_unknown'; exit 2 }
foreach ($f in @('settings.json', 'machine-id', 'wallet.json', 'tuning.json')) {
$src = Join-Path $appDir $f
if (Test-Path $src) { Copy-Item -LiteralPath $src -Destination (Join-Path $sApp $f) -Force }
}
# the second engine must not poll jobs (it would see this one), update itself, or prove; the tune is on
$sj = Join-Path $sApp 'settings.json'
if (Test-Path $sj) {
try {
$j = Get-Content -LiteralPath $sj -Raw | ConvertFrom-Json
$j.remote_jobs = $false; $j.auto_update = $false; $j.prove = $false; $j.paused = $false; $j.setup_done = $true; $j.sweep = $true
# every card is due: the stored results are cleared in the COPY only
if ($j.cards) { foreach ($p in $j.cards.PSObject.Properties) { $p.Value.sweep_at = 0; $p.Value.pinned = $false } }
$j | ConvertTo-Json -Depth 8 | Set-Content -LiteralPath $sj -Encoding utf8
} catch { Say ("settings.json: " + $_.Exception.Message) }
} else { Write-Output 'RESULT TUNE error=no_settings reason=the_installed_app_has_no_settings.json'; exit 2 }
Remove-Item -LiteralPath (Join-Path $sApp 'app.url') -Force -ErrorAction SilentlyContinue
# the state before, for the report
$smi = Join-Path $env:ProgramFiles 'NVIDIA Corporation\NVSMI\nvidia-smi.exe'
if (-not (Test-Path $smi)) { $smi = Join-Path $env:SystemRoot 'System32\nvidia-smi.exe' }
$tele = Join-Path $bin 'igneum-gpu-telemetry.exe'
function Snapshot([string] $tag) {
if (Test-Path $smi) {
$q = (& $smi --query-gpu=index,name,driver_version,power.draw,power.limit,power.default_limit,power.min_limit,power.max_limit,clocks.gr,clocks.max.gr,clocks.mem --format=csv,noheader 2>&1 | Out-String).Trim()
Write-Output ("RESULT TUNE " + $tag + " nvidia " + ($q -replace "`r?`n", ' | '))
}
if (Test-Path $tele) {
$t = (& $tele --tune 2>&1 | Out-String).Trim()
Write-Output ("RESULT TUNE " + $tag + " amd " + ($t -replace "`r?`n", ' | '))
} else { Write-Output ("RESULT TUNE " + $tag + " amd no_helper") }
}
Snapshot 'before'
# the tune engine: status every 10 s (6 rate samples per 60 s hold)
$env:IGNEUM_APP_DATA = $root
$env:IGNEUM_APP_LOGS = $sLogs
$env:IGNEUM_APP_STATUS_SECS = '10'
$psi = New-Object System.Diagnostics.ProcessStartInfo
$psi.FileName = $exe
$psi.Arguments = '--sweep'
$psi.WorkingDirectory = $bin
$psi.UseShellExecute = $false
$psi.RedirectStandardOutput = $true
$psi.RedirectStandardError = $true
$psi.CreateNoWindow = $true
$p = New-Object System.Diagnostics.Process
$p.StartInfo = $psi
$lines = New-Object System.Collections.ArrayList
$h = { if ($EventArgs.Data) { [void]$Event.MessageData.Add($EventArgs.Data) } }
Register-ObjectEvent -InputObject $p -EventName OutputDataReceived -Action $h -MessageData $lines | Out-Null
Register-ObjectEvent -InputObject $p -EventName ErrorDataReceived -Action $h -MessageData $lines | Out-Null
[void]$p.Start()
$p.BeginOutputReadLine(); $p.BeginErrorReadLine()
Say ("tune engine started, pid " + $p.Id + ", data " + $root)
$seen = 0
$rows = 0
while (-not $p.HasExited) {
Start-Sleep -Seconds 5
while ($seen -lt $lines.Count) {
$l = [string]$lines[$seen]; $seen++
if ($l -match '^TUNE ') { Write-Output ('RESULT ' + $l); if ($l -match '^TUNE card=') { $rows++ } }
elseif ($l -match '^SWEEP ') { Write-Output ('RESULT ' + $l) }
elseif ($l -match '^(URL|STATE) ') { }
else { Say $l }
}
if ((Get-Date) -gt $deadline) {
Say ("budget of " + $budgetMinutes + " min spent; asking the tune engine to quit")
$u = Join-Path $sApp 'app.url'
if (Test-Path $u) { try { Invoke-WebRequest -Uri ((Get-Content -LiteralPath $u -Raw).Trim() + 'api/quit') -Method POST -Body '{}' -ContentType 'application/json' -UseBasicParsing -TimeoutSec 5 | Out-Null } catch { } }
Start-Sleep -Seconds 20
if (-not $p.HasExited) { $p.Kill() }
Write-Output 'RESULT TUNE error=budget_exceeded'
}
}
while ($seen -lt $lines.Count) { $l = [string]$lines[$seen]; $seen++; if ($l -match '^TUNE ') { Write-Output ('RESULT ' + $l); if ($l -match '^TUNE card=') { $rows++ } } }
Say ("tune engine exited " + $p.ExitCode + " after " + [int]((Get-Date) - $started).TotalSeconds + " s, " + $rows + " table rows")
Snapshot 'after'
# the tune engine's own log: the TUNE lines and what happened around them
$log = Get-ChildItem -Path $sLogs -Filter 'app-*.log' -ErrorAction SilentlyContinue | Sort-Object LastWriteTime -Descending | Select-Object -First 1
if ($log) {
Say ("engine log " + $log.FullName + ":")
Get-Content -LiteralPath $log.FullName | Where-Object { $_ -match 'TUNE|tune|power cap|GPUs:|worker ready|STATUS|exited|upload' } | Select-Object -Last 100 | ForEach-Object { Say (' ' + $_) }
}
if ($rows -eq 0) { Write-Output 'RESULT TUNE error=no_rows'; exit 1 }
exit 0

110
relay/test/ember.test.mjs Normal file
View file

@ -0,0 +1,110 @@
// node --test relay/test/ember.test.mjs (no dependencies; CI runs it in the site job)
// The fleet aggregation of Ember Tune records (relay/lib/ember.mjs) on a fixture of records in the shape
// app/igneum-app/src/ember.rs record_json writes: known-good (five samples converge on one point), known-bad (an
// outlier does not move the median), the de-duplication of re-sent logs, the manifest merge and the prior lookup.
import { test } from 'node:test';
import assert from 'node:assert/strict';
import { parseRecords, dedupe, aggregate, mergeTuning, priorFor, priorLine, median, spreadPct } from '../lib/ember.mjs';
const step = (clock_mhz, power_pct, watts, mhs, mark = 'ok') => ({ clock_mhz, power_pct, limit_w: 575 * power_pct / 100, watts, mhs, eff: Number((mhs / watts).toFixed(4)), gclk: clock_mhz || 2800, mclk: 10500, tmax: 68, faults: 0, mark });
const rec = (machine, ts, chosen, over = {}) => ({
ts, machine, app: '0.3.10', os: 'windows', card: 'NVIDIA_GeForce_RTX_5090', vendor: 'nvidia', driver: '581.57', driver_major: '581', class: 'l128w16',
key: 'NVIDIA_GeForce_RTX_5090|581|l128w16', plan: 'full', steps: [step(0, 100, 290, 124.0), chosen], chosen, before: step(0, 100, 290, 124.0), eff: chosen.eff, mhs: chosen.mhs, watts: chosen.watts, ...over,
});
// five machines, each landing near 2,470 MHz at 100%: 0.55 to 0.57 MH/W
const good = [
rec('a1', 1000, step(2472, 100, 220, 123.5)),
rec('b2', 1001, step(2472, 100, 222, 123.1)),
rec('c3', 1002, step(2781, 100, 236, 123.8)),
rec('d4', 1003, step(2472, 100, 218, 123.6)),
rec('e5', 1004, step(2163, 100, 212, 122.9)),
];
const outlier = rec('f6', 1005, step(1854, 50, 130, 118.0)); // 0.908 MH/W: a card with a broken draw reading
const baseline = rec('g7', 1006, step(0, 80, 290, 122.3), { plan: 'baseline', steps: [step(0, 80, 290, 122.3)], before: null });
const amd = (machine, ts, chosen) => rec(machine, ts, chosen, { card: 'AMD_Radeon_RX_9070_XT', vendor: 'amd', driver: '32.0.15801.1', driver_major: '32', key: 'AMD_Radeon_RX_9070_XT|32|l128w16', plan: 'confirm', before: null });
test('five samples converge on the median point and an outlier does not move it', () => {
const { priors, table } = aggregate(good, { minSamples: 5 });
const p = priors['NVIDIA_GeForce_RTX_5090|581|l128w16'];
assert.ok(p, 'a prior at the sample floor');
assert.equal(p.clock_mhz, 2470, 'the median clock cap, rounded to 10 MHz');
assert.equal(p.power_pct, 100);
assert.equal(p.samples, 5);
assert.equal(p.machines, 5);
assert.ok(p.eff > 0.55 && p.eff < 0.57, `eff ${p.eff}`);
assert.ok(p.spread_pct >= 0 && p.spread_pct < 3, `spread ${p.spread_pct}`);
assert.equal(p.before_eff, Number((124 / 290).toFixed(4)));
assert.ok(p.gain_pct > 25, `gain ${p.gain_pct}% over the untuned 100% point`);
assert.equal(p.updated, '1970-01-01T00:16:44Z');
// the outlier: 0.908 MH/W at 1,854 MHz joins; the median moves by one rank at most
const with6 = aggregate(good.concat(outlier), { minSamples: 5 }).priors['NVIDIA_GeForce_RTX_5090|581|l128w16'];
assert.equal(with6.samples, 6);
assert.equal(with6.clock_mhz, 2470);
assert.equal(with6.power_pct, 100);
assert.ok(with6.eff < 0.58, `the outlier's 0.908 MH/W did not drag the median: ${with6.eff}`);
assert.ok(with6.spread_pct < 5, `spread ${with6.spread_pct}`);
assert.equal(table[0].key, 'NVIDIA_GeForce_RTX_5090|581|l128w16');
});
test('under the floor there is no prior, and baseline records never make one', () => {
const { priors, table } = aggregate(good.slice(0, 4), { minSamples: 5 });
assert.deepEqual(priors, {});
assert.equal(table[0].samples, 4, 'the table still shows the count');
const b = aggregate([baseline, baseline], { minSamples: 1 });
assert.deepEqual(b.priors, {}, 'measure-only records say what a card does, never what to set');
assert.equal(b.table[0].baseline_samples, 1, 'the duplicate upload counted once');
assert.equal(b.table[0].baseline_eff, Number((122.3 / 290).toFixed(4)));
// a marked chosen step (a faulted or hot winner cannot exist, but a record with one is ignored)
const bad = rec('h8', 1007, step(2000, 100, 200, 120, 'faulted'));
assert.deepEqual(aggregate([bad], { minSamples: 1 }).priors, {});
});
test('records are parsed out of log text and de-duplicated on machine, card and time', () => {
const line = `1791230000 TUNE ${JSON.stringify(good[0])}`;
const text = ['1791229999 status: x', line, line, `1791230001 TUNE ${JSON.stringify(good[1])}`, '1791230002 TUNE {not json'].join('\n');
const rs = parseRecords(text);
assert.equal(rs.length, 3);
assert.equal(dedupe(rs).length, 2);
assert.equal(parseRecords('').length, 0);
});
test('the manifest merge keeps the kernel-variant cards and carries the settings', () => {
const existing = { updated: '2026-10-04T21:00:00Z', window_days: 7, cards: { NVIDIA_GeForce_RTX_5090: { variant: 'u2-ldg', race: true, candidates: ['u2-ldg', 'ldg', 'base'] } }, priors: { 'old|1|v2': { samples: 9 } } };
const { priors } = aggregate(good, { minSamples: 5 });
const t = mergeTuning(existing, priors, { rate_tolerance_pct: 1 });
assert.equal(t.cards.NVIDIA_GeForce_RTX_5090.variant, 'u2-ldg', 'lever 2 untouched');
assert.equal(t.window_days, 7);
assert.deepEqual(t.ember, { enabled: true, min_samples: 5, rate_tolerance_pct: 1 });
assert.ok(!t.priors['old|1|v2'], 'a key without samples in the window drops out');
assert.ok(t.priors['NVIDIA_GeForce_RTX_5090|581|l128w16']);
// the kill switch rides the same section
assert.equal(mergeTuning(existing, {}, { enabled: false }).ember.enabled, false);
assert.deepEqual(mergeTuning(null, {}).cards, {});
// the round trip: canonical JSON (what publish-manifest.sh signs) parses back to the same prior
const back = JSON.parse(JSON.stringify(t));
assert.deepEqual(priorFor(back, 'NVIDIA_GeForce_RTX_5090|581|l128w16'), { clock_mhz: 2470, power_pct: 100, eff: priors['NVIDIA_GeForce_RTX_5090|581|l128w16'].eff, samples: 5 });
assert.equal(priorFor(back, 'NVIDIA_GeForce_RTX_5090|581|l128w16', 6), null, 'six wanted, five there');
assert.equal(priorFor(back, 'nothing|0|v2'), null);
assert.equal(priorFor(null, 'x'), null);
});
test('AMD confirm records aggregate by their own key, and the line reads', () => {
const rs = [amd('p1', 2000, step(2600, 90, 177, 17.7)), amd('p2', 2001, step(2600, 90, 180, 17.6)), amd('p3', 2002, step(2500, 90, 170, 17.4))];
const { priors, table } = aggregate(rs, { minSamples: 3 });
const p = priors['AMD_Radeon_RX_9070_XT|32|l128w16'];
assert.equal(p.clock_mhz, 2600);
assert.equal(p.power_pct, 90);
assert.equal(p.vendor, 'amd');
assert.equal(p.gain_pct, undefined, 'confirm records carry no before step and no baseline was uploaded');
assert.match(priorLine(p), /^AMD Radeon RX 9070 XT \| driver 32 \| l128w16: 2600 MHz at 90%, 0\.\d+ MH\/W, 17\.6 MH\/s at 177 W, spread \d+(\.\d+)?%, 3 sample\(s\) from 3 machine\(s\)$/);
assert.equal(table.length, 1);
});
test('median and spread', () => {
assert.equal(median([3, 1, 2]), 2);
assert.equal(median([4, 1, 2, 3]), 2.5);
assert.equal(median([]), 0);
assert.equal(spreadPct([1, 1, 1]), 0);
assert.equal(spreadPct([10]), 0);
assert.equal(spreadPct([9, 10, 11]), 10);
});

View file

@ -340,6 +340,21 @@ for (const [file, active] of PAGES) {
];
const table = '<div class="tbl"><table><thead><tr>' + ['Card', 'Generator', 'Best MH/s', 'MH per watt', 'Miner', 'Date', 'Source', 'Who measured it'].map(h => `<th>${h}</th>`).join('') + '</tr></thead><tbody>' +
rows.map(r => '<tr>' + cell(r).map(c => `<td>${esc(String(c))}</td>`).join('') + '</tr>').join('') + '</tbody></table></div>';
// Ember Tune's fleet priors (site/miner-priors.json, tools/tuning.mjs --priors --site): one row per card model,
// driver major and program class; a row under the sample floor shows its count and no point
const pj = JSON.parse(readFileSync(join(here, 'miner-priors.json'), 'utf8'));
const prows = (pj.rows || []).slice().sort((a, b) => (b.samples - a.samples) || (a.card < b.card ? -1 : 1));
const pcell = r => [
r.card, r.driver_major, r.class, r.samples + (r.machines ? ' from ' + r.machines + ' machine' + (r.machines === 1 ? '' : 's') : ''),
r.prior ? (r.clock_mhz ? fmt(r.clock_mhz) + ' MHz at ' + r.power_pct + '%' : r.power_pct + '%, clock unlocked') : 'under the floor (' + pj.min_samples + ' needed)',
r.mh_per_w == null ? 'not yet' : Number(r.mh_per_w).toFixed(3) + (r.spread_pct != null ? ' (spread ' + r.spread_pct + '%)' : ''),
r.mh_s == null ? '' : fmt(r.mh_s) + ' MH/s at ' + fmt(r.watts) + ' W',
r.untuned_mh_per_w == null ? 'not measured' : Number(r.untuned_mh_per_w).toFixed(3) + (r.gain_pct != null ? ' (' + (r.gain_pct >= 0 ? '+' : '') + r.gain_pct + '%)' : ''),
];
const ptable = prows.length
? '<div class="tbl"><table><thead><tr>' + ['Card', 'Driver', 'Program class', 'Samples', 'Tuned point', 'MH per watt', 'Rate and draw', 'Untuned MH per watt (gain)'].map(h => `<th>${h}</th>`).join('') + '</tr></thead><tbody>' +
prows.map(r => '<tr>' + pcell(r).map(c => `<td>${esc(String(c))}</td>`).join('') + '</tr>').join('') + '</tbody></table></div>'
: '<p>No tune reports yet. The first rows appear once five machines with the same card model have reported.</p>';
const body = scrubBench([
'<h2 id="table">The table</h2>',
'<p>One row per card, generator version and miner version. The rate is the best one measured. Integrated GPUs are not listed. Prototype rows are bench numbers from before the devnet and say so in the miner column.</p>',
@ -349,8 +364,12 @@ for (const [file, active] of PAGES) {
'<p>MH per watt needs the card\'s power draw during the run. The app reads it on NVIDIA cards through the driver. Rows get the figure when a run records it.</p>',
'<p>There is no other Igneum miner to compare with yet, so this table compares cards, not miners. The app that produces these rows: <a href="/miner">the miner page</a>.</p>',
`<p>Rows: ${rows.length}. Source file: <code>site/miner-bench.json</code> in the repository.</p>`,
'<h2 id="priors">Fleet tuning priors</h2>',
'<p>Ember Tune runs on every card the app mines with: the power limit and the core clock are stepped on the live program and the card keeps the point with the best MH per watt within 1% of its top rate. Every finished tune is reported back without anything that identifies the owner, and the fleet\'s median point per card model, driver major and program class comes back down inside the signed update manifest as the starting point for the next card of that model. A model needs ' + pj.min_samples + ' reports before its prior is used.</p>',
ptable,
`<p>Rows: ${prows.length}${pj.generated ? ', generated ' + pj.generated : ''}. Source file: <code>site/miner-priors.json</code> in the repository, written from the fleet records by <code>tools/tuning.mjs --priors --site</code>.</p>`,
].join('\n'));
const toc = [{ lvl: 2, t: 'The table', id: 'table' }, { lvl: 2, t: 'How a row gets here', id: 'how' }];
const toc = [{ lvl: 2, t: 'The table', id: 'table' }, { lvl: 2, t: 'How a row gets here', id: 'how' }, { lvl: 2, t: 'Fleet tuning priors', id: 'priors' }];
writeFileSync(join(here, 'miners.html'), page('Igneum GPU bench table', 'Measured Igneum hash rates per GPU: card, generator version, best MH/s, MH per watt where measured, miner version, date and the log entry each number came from.', body, toc,
'Measured hash rates per card on the Igneum lottery hash, with the generator version, the miner version, the date and the log entry behind each number.',
{ path: '/miners', heading: 'GPU bench table', active: 'miner' }));

6
site/miner-priors.json Normal file
View file

@ -0,0 +1,6 @@
{
"_about": "Rows of the fleet priors table at /miners (site/build.mjs), written by tools/tuning.mjs --priors --site from the TUNE records every Igneum Miner uploads. One row per card model, driver major and program class: the median tuned point, MH per watt, the spread and the sample count. No machine names, no addresses.",
"generated": null,
"min_samples": 5,
"rows": []
}

View file

@ -353,7 +353,7 @@ pre b{color:var(--molten);font-weight:500}
<div class="feat"><b>Variant racing every hour</b><p>At every hourly prepare the worker compiles the program in several shapes (unroll, load path, register budget, threads per group), times each for 2 seconds and keeps the fastest for the hour. Base keeps its place unless beaten.</p><a class="src" href="/bench#4-october-2026-miner-performance-variant-racing-metal-worker-on-the-m5-max-the-rtx-5090-job-is-ready-not-run">log · 4 Oct 2026</a></div>
<div class="feat"><b>Measured on Apple silicon</b><p>On the M5 Max, 256-thread groups ran 17.3% and 21.2% faster than base on two programs, under load from other work. The NVIDIA race is built and has not yet run on a GPU.</p><a class="src" href="/bench#4-october-2026-miner-performance-variant-racing-metal-worker-on-the-m5-max-the-rtx-5090-job-is-ready-not-run">log · 4 Oct 2026, ratios under contention</a></div>
<div class="feat"><b>Fleet tuning manifest</b><p>Every race is one record in the app log. The fleet's best variant per card model goes back out inside the signed update manifest, so a card starts from the known best and keeps racing.</p></div>
<div class="feat"><b>Efficiency mode</b><p>Hash per watt, the number miners compare. The app steps an NVIDIA card's power cap from 100% to 50%, holds each step for 60 seconds on the live kernel and leaves the cap at the best MH per watt. Built and unit-tested; not yet run on a card.</p></div>
<div class="feat"><b>Ember Tune</b><p>Hash per watt, the number miners compare. Out of the box the app steps every card's power limit and core clock on the live kernel, 60 seconds a step, and keeps the point with the best MH per watt within 1% of the card's top rate. The memory clock is never touched; a step with a rejected hash, a hot GPU or a dragged memory clock is reverted and marked. Every result feeds a fleet prior per card model that the next card of that model starts from. Built and unit-tested; the first measured tune is owed.</p></div>
<div class="feat"><b>Latency work</b><p>A block built on a stale tip earns less. The miner is moving from polling to a template subscription and shorter jobs, so a new tip reaches the card in milliseconds. In progress, no number published yet.</p></div>
<div class="feat"><b>Remote signed jobs, opt in</b><p>Our own fleet only. A switch, "Allow remote jobs from Igneum (signed)", with the key's fingerprint beside it. Off aborts the running job and stops polling. Jobs run at most once each and only on the machines they name.</p></div>
</div>
@ -403,8 +403,8 @@ pre b{color:var(--molten);font-weight:500}
<thead><tr><th>Lever</th><th>The idea</th><th>Measured state</th></tr></thead>
<tbody>
<tr><td><b>1 · Race the compiler every hour</b></td><td>For each new program, five to ten kernel variants are compiled, benched for two seconds each, and the winner is kept for the hour.</td><td>Measured on the Mac's Metal worker: the winning variant +17.3% on the genesis seed and +21.2% on the hourly seed over the base compile, under load from other work. The RTX 5090 race is prepared and not yet run. <a href="/bench#4-october-2026-miner-performance-variant-racing-metal-worker-on-the-m5-max-the-rtx-5090-job-is-ready-not-run">Log, 4 Oct 2026</a></td></tr>
<tr><td><b>2 · Auto-tune that learns from the fleet</b></td><td>Apps report the race per card and program class. The best settings come back down through the signed update manifest as defaults, so the miner gets faster for everyone as the fleet grows.</td><td>Shipped. The fleet is still small, so no fleet table yet.</td></tr>
<tr><td><b>3 · Hash per watt, not hash</b></td><td>A sweep per card finds the power point with the best MH per watt and holds it. Miners pay for electricity; that is the number they compare.</td><td>Shipped for NVIDIA cards through the power cap. Not yet run on a card; the first sweep on a 5090 is the owed measurement. Clocks are the next lever.</td></tr>
<tr><td><b>2 · Auto-tune that learns from the fleet</b></td><td>Apps report the race per card and program class, and every finished Ember Tune. The best settings come back down through the signed update manifest as defaults, so the miner gets faster and more efficient for everyone as the fleet grows.</td><td>Shipped. The fleet is still small, so no fleet table yet; the priors table on <a href="/miners#priors">/miners</a> fills as machines report.</td></tr>
<tr><td><b>3 · Hash per watt, not hash</b></td><td>Ember Tune: two knobs per card (power limit, core clock), the memory clock held, the point with the best MH per watt within 1% of the top rate kept and pinned. Miners pay for electricity; that is the number they compare.</td><td>Built for NVIDIA (through the driver, with Power control on) and AMD (through the app's own helper, no administrator rights), measure only on Apple silicon. Unit-tested on every rule; the first tune measured on a card is the owed number.</td></tr>
<tr><td><b>4 · Template latency</b></td><td>Solo against the local node, a new template within 50 ms of a new tip, because a late block on a BlockDAG goes red and earns nothing.</td><td>Approximate, from the 0.3.6 release plan and not yet in the engineering log: switched p50 46 to 52 ms on a three-node CPU run, 5 Oct 2026. Ships in 0.3.6.</td></tr>
<tr><td><b>5 · Never lose a second</b></td><td>Zero-loss hourly program swaps, fault guards, automatic restart, a CPU re-check of every found hash, per-worker health.</td><td>Shipped. Swap 0.01 ms on Metal and 0.00 ms on CUDA with 0 rejected blocks; recoveries in the <a href="#reliability">table above</a>. <a href="/bench#4-october-2026-first-hourly-program-swap-on-the-live-devnet-compile-ahead-no-pause-two-cards">Log, 4 Oct 2026</a></td></tr>
<tr><td><b>6 · Prove it in public</b></td><td>The bench table per card, fed from the job channel. We claim fastest only when the table says so.</td><td>Live at <a href="/miners">/miners</a>, one row per card, generator version and miner version, each with its log entry.</td></tr>

View file

@ -172,12 +172,12 @@ th{font-family:var(--f-mono);font-size:12px;letter-spacing:.12em;text-transform:
<!-- nav:end -->
<main id="main" class="wrap">
<div class="head">
<div class="eyebrow">2 entries, newest at the bottom</div>
<div class="eyebrow">3 entries, newest at the bottom</div>
<h1>GPU bench table</h1>
<p class="note">Measured hash rates per card on the Igneum lottery hash, with the generator version, the miner version, the date and the log entry behind each number.</p>
</div>
<div class="layout">
<nav class="toc" aria-label="Contents"><div class="lbl">Contents</div><ol><li><a href="#table">The table</a></li><li><a href="#how">How a row gets here</a></li></ol></nav>
<nav class="toc" aria-label="Contents"><div class="lbl">Contents</div><ol><li><a href="#table">The table</a></li><li><a href="#how">How a row gets here</a></li><li><a href="#priors">Fleet tuning priors</a></li></ol></nav>
<article><h2 id="table">The table</h2>
<p>One row per card, generator version and miner version. The rate is the best one measured. Integrated GPUs are not listed. Prototype rows are bench numbers from before the devnet and say so in the miner column.</p>
<div class="tbl"><table><thead><tr><th>Card</th><th>Generator</th><th>Best MH/s</th><th>MH per watt</th><th>Miner</th><th>Date</th><th>Source</th><th>Who measured it</th></tr></thead><tbody><tr><td>Apple M5 Max (40 GPU cores, Metal)</td><td>v1</td><td>45.2</td><td>not measured</td><td>proto-metal bench (prototype, not mining)</td><td>2026-10-03</td><td>bench log: 3 October 2026, RTX 5090 first run (the Apple row of the same table)</td><td>measured by the team. genesis program, 1 GiB dataset</td></tr><tr><td>Apple M5 Max (40 GPU cores, Metal)</td><td>v2</td><td>26.7</td><td>not measured</td><td>igneum-miner devnet v4, Metal worker with prepare</td><td>2026-10-04</td><td>bench log: 4 October 2026, first hourly program swap on the live devnet: compile-ahead, no pause, two cards</td><td>measured by the team. live devnet v4, unbroken through the hour boundary</td></tr><tr><td>Apple silicon laptop (model not reported)</td><td>v2</td><td>24.3</td><td>not measured</td><td>Igneum Miner 0.3.1 (DMG)</td><td>2026-10-04</td><td>bench log: 4 October 2026, first outside machine on the devnet: an Apple silicon laptop through the Igneum Miner app</td><td>reported by the fleet. 21.0 MH/s average over 7 minutes, 24.3 MH/s at the moment of the report, 33 accepted blocks</td></tr><tr><td>NVIDIA RTX 5090 (32 GB)</td><td>v1</td><td>229</td><td>not measured</td><td>proto-cuda bench (prototype, not mining)</td><td>2026-10-03</td><td>bench log: 3 October 2026, RTX 5090, memory-hard dataset (pack igneum-genesis-mh)</td><td>measured by the team. genesis program, 104 loads per hash, 1 GiB dataset</td></tr><tr><td>NVIDIA RTX 5090 (32 GB)</td><td>v1</td><td>185.3</td><td>not measured</td><td>proto-cuda bench (prototype, not mining)</td><td>2026-10-03</td><td>bench log: 3 October 2026, RTX 5090 first run, dataset sweep and second program</td><td>measured by the team. hourly program, 128 loads per hash, 1 GiB dataset</td></tr><tr><td>NVIDIA RTX 5090 (32 GB)</td><td>v2</td><td>124.2</td><td>not measured</td><td>Igneum Miner 0.3.0 package, prebuilt NVRTC worker</td><td>2026-10-04</td><td>bench log: 4 October 2026, the gfx1036 worker fault and what the Apple M5 Max could and could not reproduce</td><td>measured by the team. live devnet v4, 128 loads per hash, CPU re-check clean, 0 rejected</td></tr></tbody></table></div>
@ -185,7 +185,11 @@ th{font-family:var(--f-mono);font-size:12px;letter-spacing:.12em;text-transform:
<p>Every row names the engineering log entry or the job it came from. "Measured by the team" means our own hardware and our own log. "Reported by the fleet" means a machine we do not own, read from the status lines its miner uploads.</p>
<p>MH per watt needs the card's power draw during the run. The app reads it on NVIDIA cards through the driver. Rows get the figure when a run records it.</p>
<p>There is no other Igneum miner to compare with yet, so this table compares cards, not miners. The app that produces these rows: <a href="/miner">the miner page</a>.</p>
<p>Rows: 6. Source file: <code>site/miner-bench.json</code> in the repository.</p></article>
<p>Rows: 6. Source file: <code>site/miner-bench.json</code> in the repository.</p>
<h2 id="priors">Fleet tuning priors</h2>
<p>Ember Tune runs on every card the app mines with: the power limit and the core clock are stepped on the live program and the card keeps the point with the best MH per watt within 1% of its top rate. Every finished tune is reported back without anything that identifies the owner, and the fleet's median point per card model, driver major and program class comes back down inside the signed update manifest as the starting point for the next card of that model. A model needs 5 reports before its prior is used.</p>
<p>No tune reports yet. The first rows appear once five machines with the same card model have reported.</p>
<p>Rows: 0. Source file: <code>site/miner-priors.json</code> in the repository, written from the fleet records by <code>tools/tuning.mjs --priors --site</code>.</p></article>
</div>
<p class="gen">Generated from the repository at build time. Times are UTC. Machine names are model names.</p>
</main>

View file

@ -5,6 +5,7 @@
// node tools/console.mjs machines the machine cards as text
// node tools/console.mjs chain the chain numbers
// node tools/console.mjs jobs | builds | results the other tabs as text
// node tools/console.mjs tuning [--days 30] [--min 5] Ember Tune's fleet priors per card model (samples, MH/W)
// node tools/console.mjs sync-bench push docs/bench-log.md entries and the FUD ledger counts
// node tools/console.mjs sync-dl push the downloads folder listing (names, sizes, times)
// node tools/console.mjs sync-hetzner push the newest infra/cloud-devnet/results/ summary
@ -121,6 +122,12 @@ try {
}
else if (cmd === 'jobs') { const j = await api('jobs'); if (!j.jobs.length) console.log(j.note || 'no jobs'); for (const jb of j.jobs) console.log(`${jb.id} ${jb.title} | ${jb.runs.map(r => `${r.machine} ${r.status}${r.exit_code != null ? ' exit ' + r.exit_code : ''}${r.summary ? ': ' + r.summary.slice(0, 80) : ''}`).join(' | ') || 'queued everywhere'}`); }
else if (cmd === 'builds') { const j = await api('builds'); console.log(`manifest ${j.manifest ? j.manifest.version + ' (' + j.manifest.channel + ') ' + Object.keys(j.manifest.platforms || {}).join('+') + ': ' + j.manifest.notes : 'none'}`); if (j.ci) console.log(`ci ${j.ci.installer} fetched ${j.ci.fetched_at} ${j.ci.run}`); for (const b of j.builds) console.log(`${when(b.ts)} ${b.who.padEnd(8)} ${b.title}${b.body ? ': ' + b.body.split('\n')[0].slice(0, 100) : ''}`); }
else if (cmd === 'tuning') {
const j = await api('tuning', { q: { days: flags.days || 30, min: flags.min || 5 } });
console.log(`${j.records} tune record(s) in ${j.days} day(s); a prior needs ${j.min_samples} sample(s)`);
if (!j.table.length) console.log('no tune records yet (the apps log one per finished tune, 0.3.10 and later)');
for (const t of j.table) console.log(`${t.card.replace(/_/g, ' ')} | driver ${t.driver_major} | ${t.class}: ${t.samples} sample(s) from ${t.machines} machine(s)${t.samples ? `, ${t.clock_mhz ? t.clock_mhz + ' MHz at ' : ''}${t.power_pct}%${t.clock_mhz ? '' : ' (clock unlocked)'}, ${t.eff} MH/W (${t.mhs} MH/s at ${t.watts} W, spread ${t.spread_pct}%)` : ''}${t.baseline_eff ? `; untuned ${t.baseline_eff} MH/W from ${t.baseline_samples} baseline(s)` : ''}${t.gain_pct != null ? `; gain ${t.gain_pct >= 0 ? '+' : ''}${t.gain_pct}%` : ''}; prior ${t.key in j.priors ? 'yes' : 'no'}`);
}
else if (cmd === 'results') { const j = await api('results'); if (j.ledger) console.log(`${j.ledger.title}: ${j.ledger.body}`); for (const b of j.bench) console.log(`${b.meta.date || '-'} ${b.title}`); }
else if (cmd === 'sync-bench' || cmd === 'sync-dl' || cmd === 'sync-hetzner' || cmd === 'sync') {
const items = [];

View file

@ -14,10 +14,17 @@
// variant without a race; the record then carries only the self-test), the default keeps racing (the fleet keeps
// learning while the card starts from the known best).
// node tools/tuning.mjs --records [--days 7] [--card <model>] the raw records, newest first
// node tools/tuning.mjs --priors [--days 30] [--min-samples 5] Ember Tune (docs/plans/ember-tune.md): the fleet
// priors per (card model, driver major, program class) from the TUNE records (relay/lib/ember.mjs aggregate);
// --write adds them to the tuning file under "priors" with the "ember" settings (kill switch --tuning-off,
// --rate-tolerance N), beside the kernel-variant cards; --site writes site/miner-priors.json for /miners.
//
// Reads DATABASE_URL from ~/.config/igneum/env. No dependencies: Neon HTTP SQL over fetch.
import { readFileSync, writeFileSync } from 'node:fs';
import { readFileSync, writeFileSync, existsSync } from 'node:fs';
import { homedir } from 'node:os';
import { join, dirname, resolve } from 'node:path';
import { fileURLToPath } from 'node:url';
import { parseRecords, aggregate, mergeTuning, priorLine } from '../relay/lib/ember.mjs';
process.stdout.on('error', e => { if (e.code === 'EPIPE') process.exit(0); throw e; });
@ -46,6 +53,43 @@ const minSamples = Number(opt('--min-samples', 3));
const by = opt('--by', 'mhs');
const outFile = opt('--write', '');
const onlyCard = opt('--card', '');
const ROOT = resolve(dirname(fileURLToPath(import.meta.url)), '..');
// ---- Ember Tune priors (--priors): the TUNE records of the window, folded per key --------------------------------
if (flag('--priors')) {
const pdays = Number(opt('--days', 30));
const pmin = Number(opt('--min-samples', 5));
const prow = await sql(
`SELECT machine, run_id, lines FROM miner_logs WHERE received_at > now() - ($1 || ' days')::interval AND lines LIKE '%TUNE {%' ORDER BY received_at DESC`,
[String(pdays)]);
const recs = prow.flatMap(r => parseRecords(r.lines)).filter(r => !onlyCard || r.card === onlyCard);
const { priors, table } = aggregate(recs, { minSamples: pmin });
if (!recs.length) console.log(`No TUNE records in the last ${pdays} day(s). The apps log one per finished tune (0.3.10 and later).`);
else {
console.log(`${recs.length} tune record(s) in the last ${pdays} day(s); a prior needs ${pmin} sample(s)`);
console.table(table.map(t => ({ key: t.key, samples: t.samples, machines: t.machines, 'clock MHz': t.clock_mhz ?? '-', 'power %': t.power_pct ?? '-', 'MH/W': t.eff ?? '-', 'MH/s': t.mhs ?? '-', W: t.watts ?? '-', 'spread %': t.spread_pct ?? '-', 'untuned MH/W': t.before_eff ?? t.baseline_eff ?? '-', 'gain %': t.gain_pct ?? '-', prior: t.key in priors ? 'yes' : 'no' })));
for (const p of Object.values(priors)) console.log(priorLine(p));
}
if (outFile) {
const existing = existsSync(outFile) ? JSON.parse(readFileSync(outFile, 'utf8')) : {};
const ember = {};
if (flag('--tuning-off')) ember.enabled = false;
if (flag('--tuning-on')) ember.enabled = true;
if (opt('--rate-tolerance', '')) ember.rate_tolerance_pct = Number(opt('--rate-tolerance', '1'));
ember.min_samples = pmin;
const merged = mergeTuning(existing, priors, ember);
writeFileSync(outFile, JSON.stringify(merged, null, 2) + '\n');
console.log(`written ${outFile}: ${Object.keys(merged.cards).length} kernel-variant card(s) kept, ${Object.keys(priors).length} prior(s), ember ${JSON.stringify(merged.ember)}; publish with: packaging/ota/publish-manifest.sh --version <current> --tuning ${outFile} [--deploy]`);
}
if (flag('--site')) {
const site = join(ROOT, 'site', 'miner-priors.json');
const rows = table.map(t => ({ card: t.card.replace(/_/g, ' '), vendor: t.vendor, driver_major: t.driver_major, class: t.class, clock_mhz: t.clock_mhz ?? null, power_pct: t.power_pct ?? null, mh_per_w: t.eff ?? null, mh_s: t.mhs ?? null, watts: t.watts ?? null, spread_pct: t.spread_pct ?? null, samples: t.samples, machines: t.machines, untuned_mh_per_w: t.before_eff ?? t.baseline_eff ?? null, gain_pct: t.gain_pct ?? null, prior: t.key in priors, updated: t.updated || null }));
const about = existsSync(site) ? JSON.parse(readFileSync(site, 'utf8'))._about : undefined;
writeFileSync(site, JSON.stringify({ _about: about || 'Rows of the fleet priors table at /miners (site/build.mjs), written by tools/tuning.mjs --priors --site from the TUNE records every Igneum Miner uploads. One row per card model, driver major and program class: the median tuned point, MH per watt, the spread and the sample count. No machine names, no addresses.', generated: new Date().toISOString().replace(/\.\d{3}Z$/, 'Z'), min_samples: pmin, rows }, null, 2) + '\n');
console.log(`written ${site} (${rows.length} row(s))`);
}
process.exit(0);
}
// Every app-log upload of the window; the TUNING lines out of them. The same race is uploaded many times (the log
// is re-sent every minute), so records are de-duplicated on (machine, card, epoch).