the project lead, 5 October 2026, 22:45 BST: "make sure we have ember tuning every single card for efficiency out of the box, the more data = the better the tune, make an awesome system." Built on lever 3 (docs/plans/miner-eff.md), lever 2's signed tuning section (docs/design/miner-tuning.md), the AMD telemetry helper (423936b, its --tune/--set-gmax/--set-plimit/ --reset contract) and the Power control switch (057f0ec). Design, data flow, tiers and the privacy line: docs/plans/ember-tune.md. - src/ember.rs (new): two knobs per card (power limit %, core clock cap MHz; memory clock never touched), the full plan (power ladder 100..50%, then the clock ladder 90..60% at the chosen power), the confirm plan (the fleet prior and one neighbour), the baseline plan (measure only), the marks (faulted, hot, memory_clock_dropped, unapplied, no_readings), the choice (best MH/W within 1% of the top rate, then rate, then draw), the fleet record (a hash of the install id, no address), the prior lookup and the kill switch (tuning.ember), the state machine on a fake clock. 9 unit tests. - engine.rs: tick_sweep schedules every NVIDIA, AMD and Apple card (120 s steady, 600 s to the boundary, no job hold, no pause, weekly, again after a driver major or program-class change, never under the manifest kill switch); the probe (nvidia-smi clocks.max.gr + driver_version and the direct/helper mode; igneum-gpu-telemetry --tune for AMD); tune_apply (nvidia-smi -pl / -lgc 0,<MHz> / -rgc directly or through the helper; the AMD helper per request); Cmd::TuneProbe, Cmd::TuneSet; faults from rejected and mismatched hashes mark the step; the TUNE lines and the TUNE {json} record, uploaded with the log; the Tuned line on the card state. The NVIDIA helper starts only with Power control on: the --sweep job never counts as permission (no prompt on a PC with nobody there). - sweep.rs: the helper protocol gains lgc/rgc (clock cap and reset) and resets the clocks after 20 idle minutes. - state.rs, config.rs: the tune fields (clock cap, driver, class, source, the Tuned line); the nvidia-smi telemetry query carries clocks.gr and clocks.mem; the AMD sample line's plimit_pct and gmax_mhz are parsed. - ui: "Tuned: X MH/s at Y W (Z MH/W)" with the point, the source and when; measure-only cards say why; the Ember Tune switch; tune-line.test.mjs. - relay/lib/ember.mjs + relay/test/ember.test.mjs: the aggregation per (card model | driver major | program class): median point, MH/W, spread, samples, machines; five samples converge, an outlier does not move the median, baselines make no prior, de-duplication, the manifest merge keeps lever 2's cards. api/console.mjs fn=tuning and tools/console.mjs tuning; tools/tuning.mjs --priors [--write tuning.json] [--site] [--tuning-off]. - site: the fleet priors table on /miners (site/miner-priors.json), the lever text. - relay/playbooks/ember-tune-pc1.ps1: the PC 1 run (second engine with --sweep from a scratch copy of the install). Measured tonight: see the bench log entry that follows the PC 1 run. The 9070 XT left PC 1's bus at 20:40 UTC and the 5090 needs the administrator prompt the project lead cannot answer asleep, so tonight's PC 1 run is the baseline plan on the 5090 through the whole pipeline; the two-knob tune on both cards is owed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
6 lines
385 B
JSON
6 lines
385 B
JSON
{
|
|
"_about": "Rows of the fleet priors table at /miners (site/build.mjs), written by tools/tuning.mjs --priors --site from the TUNE records every Igneum Miner uploads. One row per card model, driver major and program class: the median tuned point, MH per watt, the spread and the sample count. No machine names, no addresses.",
|
|
"generated": null,
|
|
"min_samples": 5,
|
|
"rows": []
|
|
}
|