The known-failed case, from PC 2's 0.3.9 log (run win-1ccfe586-20261005-200114): 1791234223 pause -> 'stopping the
miners (paused)' (every slot's restart_at cleared, the 5090 'off'); 1791235511 '[ok] mining resumed'; then
'0.00 MH/s, waiting' at every 30-s status line until the 0.3.10 restart at 21:49:41Z. Cause: Cmd::Resume re-armed
only slots whose watchdog said faulted; the 5090's slot was healthy and stopped, so nothing restarted it. The test
the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old encodes that slot (faulted false, live
false): the old rule returns [] (the defect), the new rule [0]. cargo test -p igneum-app resume: 3 passed;
provedefault: 6 passed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Twice on 5 October 2026 a PowerShell job script carried a bash body inside a string, a quote was lost on the way
through PowerShell, and bash refused the body: pc1-cpu-prove.ps1 (first version) reported exit 0 having done
nothing, the 0.3.10 installer job failed in 4 s. tools/amd-prove/check-job-bash.sh covered only its own here-string.
tools/ci/bash-body-check.sh reads every *.ps1 under relay/playbooks/ and tools/, finds each bash body however it is
handed over (bash -c "...", bash -lc '...', bash -c $var, a + concatenation in parentheses, the Start-Process argument
list, a here-string written to a file that is later run with bash), unescapes it the way PowerShell would (backtick
escapes and "" in double-quoted strings, '' in single-quoted strings, here-strings verbatim; $var left as-is, a $(...)
subexpression replaced by ${PS_SUBEXPR}), and runs bash -n on it. One line per body with the file line of the error.
A body it sees but cannot read is "unextractable body" and fails too: a skip would be a hole in the class check.
bash 3.2 compatible; python3 for the extractor.
--self-test runs three fixtures under tools/ci/fixtures/: the correct shapes (8 bodies, must pass), the lost quotes
(the awk apostrophe, a dropped closing quote in a literal and in a variable; must fail with the line), and three
unreadable bodies (must fail). Wired into ci.yml next to the copied-sources check, self-test first. The current tree:
7 inline bodies in 3 playbooks, all parse. packaging/README-ship.md: the job-script rule (body to a file, bash <file>).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
the project lead, 5 October 2026, 22:45 BST: "make sure we have ember tuning every single card for efficiency out of the box, the
more data = the better the tune, make an awesome system." Built on lever 3 (docs/plans/miner-eff.md), lever 2's signed
tuning section (docs/design/miner-tuning.md), the AMD telemetry helper (423936b, its --tune/--set-gmax/--set-plimit/
--reset contract) and the Power control switch (057f0ec). Design, data flow, tiers and the privacy line:
docs/plans/ember-tune.md.
- src/ember.rs (new): two knobs per card (power limit %, core clock cap MHz; memory clock never touched), the full plan
(power ladder 100..50%, then the clock ladder 90..60% at the chosen power), the confirm plan (the fleet prior and one
neighbour), the baseline plan (measure only), the marks (faulted, hot, memory_clock_dropped, unapplied, no_readings),
the choice (best MH/W within 1% of the top rate, then rate, then draw), the fleet record (a hash of the install id,
no address), the prior lookup and the kill switch (tuning.ember), the state machine on a fake clock. 9 unit tests.
- engine.rs: tick_sweep schedules every NVIDIA, AMD and Apple card (120 s steady, 600 s to the boundary, no job hold,
no pause, weekly, again after a driver major or program-class change, never under the manifest kill switch); the
probe (nvidia-smi clocks.max.gr + driver_version and the direct/helper mode; igneum-gpu-telemetry --tune for AMD);
tune_apply (nvidia-smi -pl / -lgc 0,<MHz> / -rgc directly or through the helper; the AMD helper per request);
Cmd::TuneProbe, Cmd::TuneSet; faults from rejected and mismatched hashes mark the step; the TUNE lines and the TUNE
{json} record, uploaded with the log; the Tuned line on the card state. The NVIDIA helper starts only with Power
control on: the --sweep job never counts as permission (no prompt on a PC with nobody there).
- sweep.rs: the helper protocol gains lgc/rgc (clock cap and reset) and resets the clocks after 20 idle minutes.
- state.rs, config.rs: the tune fields (clock cap, driver, class, source, the Tuned line); the nvidia-smi telemetry
query carries clocks.gr and clocks.mem; the AMD sample line's plimit_pct and gmax_mhz are parsed.
- ui: "Tuned: X MH/s at Y W (Z MH/W)" with the point, the source and when; measure-only cards say why; the Ember Tune
switch; tune-line.test.mjs.
- relay/lib/ember.mjs + relay/test/ember.test.mjs: the aggregation per (card model | driver major | program class):
median point, MH/W, spread, samples, machines; five samples converge, an outlier does not move the median, baselines
make no prior, de-duplication, the manifest merge keeps lever 2's cards. api/console.mjs fn=tuning and
tools/console.mjs tuning; tools/tuning.mjs --priors [--write tuning.json] [--site] [--tuning-off].
- site: the fleet priors table on /miners (site/miner-priors.json), the lever text.
- relay/playbooks/ember-tune-pc1.ps1: the PC 1 run (second engine with --sweep from a scratch copy of the install).
Measured tonight: see the bench log entry that follows the PC 1 run. The 9070 XT left PC 1's bus at 20:40 UTC and the
5090 needs the administrator prompt the project lead cannot answer asleep, so tonight's PC 1 run is the baseline plan on the 5090
through the whole pipeline; the two-knob tune on both cards is owed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>