C35 named: PC 1's 22:31 UTC quit was the per-user installer launched by the second engine's own updater (0.3.9 under min_supported_version = urgent, beating auto_update = false); a second engine never runs the updater (IGNEUM_APP_NO_OTA=1, implied by --sweep; the playbooks set it; the CI check demands it); bench log and plan carry the named source

Source: the scratch engine's own log in collect ember-c35-collect-1 (06:59Z): 22:31:02Z '0.3.10 is available: downloading',
22:31:05Z 'update: starting the installer first ... ota-apply.ps1', and the installed app's 'quit:' at 22:31:06Z.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 07:01:57 +00:00
parent f9d9805ae8
commit 3d505766d2
6 changed files with 33 additions and 7 deletions

View file

@ -466,6 +466,8 @@ pub struct Engine {
sweep: Option<crate::ember::Run>,
/// who asked for the quit (the log's `quit:` line names it)
quit_source: &'static str,
/// the over-the-air updater never runs: IGNEUM_APP_NO_OTA=1 or --sweep (a second engine beside the installed app)
no_ota: bool,
/// the request number the vendor tool last carried out (the run's acknowledgement)
tune_acked: Option<u64>,
/// cards whose confirm check found a better neighbour: the full plan runs next
@ -491,6 +493,7 @@ pub struct Engine {
impl Engine {
pub fn new(shared: Arc<Shared>, bins: Bins, rx: Receiver<Cmd>, wrapper: bool) -> Engine {
let no_ota = shared.runtime.sweep_only || std::env::var("IGNEUM_APP_NO_OTA").map(|v| v == "1").unwrap_or(false);
let (lines_tx, lines_rx) = channel();
let now = Instant::now();
let stamp = {
@ -563,6 +566,7 @@ impl Engine {
last_error_event: now - Duration::from_secs(600),
sweep: None,
quit_source: "unknown",
no_ota,
tune_acked: None,
tune_full_due: std::collections::HashSet::new(),
sweep_pending: None,
@ -629,6 +633,9 @@ impl Engine {
if self.shared.runtime.sweep_only {
self.shared.log("--sweep: the efficiency sweep runs on every supported card as soon as it mines; the table goes to stdout and this log; the engine quits after it");
}
if self.no_ota {
self.shared.log("updates: off for this engine (IGNEUM_APP_NO_OTA or --sweep): it never downloads or installs, whatever the manifest says (C35)");
}
if (setup_done || self.shared.runtime.sweep_only) && !self.quitting {
self.shared.send(Cmd::Start);
}
@ -761,7 +768,13 @@ impl Engine {
self.restart_node(&why, Duration::from_secs(2));
}
}
Cmd::CheckUpdate => self.ota.check_now(&self.shared),
Cmd::CheckUpdate => {
if self.no_ota {
self.shared.event("info", "updates are off for this engine (a measurement run: IGNEUM_APP_NO_OTA or --sweep)");
} else {
self.ota.check_now(&self.shared);
}
}
Cmd::InstallUpdate => self.ota.install_now(&self.shared),
Cmd::AutoUpdate(on) => self.ota.set_auto(&self.shared, on),
Cmd::OpenUpdateFile => {
@ -2439,8 +2452,13 @@ impl Engine {
daa: st.node.daa,
}
};
if let Some(crate::ota::Action::Apply) = self.ota.tick(&self.shared, &ctx) {
self.apply_update();
// C35 (PC 1, 5 October 2026, 22:31 UTC): a second engine started by a measurement job found itself under the
// manifest's min_supported_version ("urgent" beats auto_update = false), ran the per-user installer, and the
// installer's PrepareToInstall quit the INSTALLED app through its api/quit. A second engine never updates.
if !self.no_ota {
if let Some(crate::ota::Action::Apply) = self.ota.tick(&self.shared, &ctx) {
self.apply_update();
}
}
if let Some(p) = self.ota.take_override_change() {
self.shared.log(&format!("consensus override changed ({}); the node restarts with it at a safe moment", p.display()));

View file

@ -1624,7 +1624,7 @@ Branch `ember-tune` (54ff1bc), docs/plans/ember-tune.md. Every card tuned for MH
**Tier consequences** (docs/plans/ember-tune.md section 7): a 9-step full tune costs about 12 minutes once and 3 minutes a week per card, under 1% of the hour, the worker never stops; a rig tunes one card at a time and every card of a known model after the first takes the 3-minute confirm; a pool user gives up the same 1% of shares at most; Apple silicon and AMD on Linux measure only and the row says so.
**The PC 1 run, 22:30 UTC (job ember-tune-pc1-1, engine aeea3228..., PC 1 on 0.3.10):** the job published at 22:29:40Z, the installed app stopped its miners and started the second engine at 22:30:21Z, and at 22:31:06Z the installed app quit (its log: `quit: stopping the miners, then the node`, then `job ember-tune-pc1-1: aborted (the app is quitting)`), 46 s in, before any step. Nothing was set. Corrected the same night (C35): the first reading, that the 0.3.11 update caused the quit, was wrong; no update, restart or relay task reached PC 1 then (its own jobs lines and the relay feed), and the quit's source is not in the log because the app did not name it (fixed at b671c8b: every `quit:` line now names its sender). What is established: the window host's tray quit is excluded (a host that sent the quit terminates the engine 45 s later, and the engine lived on until 22:55Z), leaving stdin EOF (the host process gone) or `POST /api/quit`; the engine's quit then HUNG for 24 minutes in the jobs runner's abort, waiting for EOF on the script's stdout pipe whose write end the second engine and its miners had inherited, and those miners (2 igneum-miner, 2 CUDA workers, 1 OpenCL worker) mined on, orphaned, until the relay lane killed them at about 23:00Z; the second engine also raised one administrator prompt at about 22:30:25Z (`apply_power_limits` at start counted `--sweep` as Power control), 41 s before the quit; PC 2's unexplained quit at 20:01:09Z came 20 s after a cancelled prompt of the same class, so the prompt is the common factor and the morning's test (one prompt raised beside the mining app on PC 2, the stamped quit line read). Fixed on the branch: b671c8b (quit sources, Power control alone decides, no cap at start under `--sweep`), 8ab9068 (no pipe into a second engine, its tree ended, the CI check). What the run did record, the "before" snapshots with the miners stopped:
**The PC 1 run, 22:30 UTC (job ember-tune-pc1-1, engine aeea3228..., PC 1 on 0.3.10):** the job published at 22:29:40Z, the installed app stopped its miners and started the second engine at 22:30:21Z, and at 22:31:06Z the installed app quit (its log: `quit: stopping the miners, then the node`, then `job ember-tune-pc1-1: aborted (the app is quitting)`), 46 s in, before any step. Nothing was set. Corrected the same night (C35), then named the next morning from the second engine's own log (collect ember-c35-collect-1, 06:59Z): the second engine, reporting 0.3.9 (the branch's Cargo version) under the manifest's `min_supported_version`, took the 0.3.10 update as urgent (the "urgent" rule beats the copied `auto_update = false`), downloaded it at 22:31:02Z and started `ota-apply.ps1` with the per-user installer at 22:31:05Z; the installer's PrepareToInstall sent `POST /api/quit` to the installed app, which logged `quit:` at 22:31:06Z. So the source was my own second engine's updater, through the installer, one second before. The first reading (the 0.3.11 rollout) was wrong in the cause and right in the class: an installer. What else is established: the engine's quit then HUNG for 24 minutes in the jobs runner's abort, waiting for EOF on the script's stdout pipe whose write end the second engine and its miners had inherited, and those miners (2 igneum-miner, 2 CUDA workers, 1 OpenCL worker) mined on, orphaned, until the relay lane killed them at about 23:00Z; the second engine also raised one administrator prompt at about 22:30:25Z (`apply_power_limits` at start counted `--sweep` as Power control), 41 s before the quit; PC 2's unexplained quit at 20:01:09Z came 20 s after a cancelled prompt of the same class, so the prompt is the common factor and the morning's test (one prompt raised beside the mining app on PC 2, the stamped quit line read). Fixed on the branch: b671c8b (quit sources, Power control alone decides, no cap at start under `--sweep`), 8ab9068 (no pipe into a second engine, its tree ended, the CI check), and the third close: a second engine never runs the updater (`IGNEUM_APP_NO_OTA=1`, implied by `--sweep`; the playbooks set it; the CI check demands it). What the run did record, the "before" snapshots with the miners stopped:
| Card | Read back at 22:30:20Z | Meaning |
|---|---|---|

View file

@ -101,6 +101,7 @@ race has run).
| The elevated helper restores the limit and resets the clocks by itself after 20 idle minutes | `sweep::helper_script_*` |
| A playbook that starts a second engine beside the installed app (the PC measurement jobs) gives it NO pipe (its output goes to a file the script tails: a pipe's write end is inherited by the engine's miners, and the installed app's jobs runner then waits forever for EOF after an abort; C35, PC 1 22:31 UTC, a 24-minute hang and orphaned miners), ends the engine's whole process tree at the end and on the budget (`taskkill /T /F`), and lets the installed app's miners come back only after that | `relay/playbooks/ember-tune-pc1.ps1`, `sweep-5090.ps1`; CI `tools/ci/second-engine-check.sh` fails any playbook without both |
| Every `quit:` line in the app log names its source (the window host's stdin, the host gone, `POST /api/quit`, the `--sweep` run's end) | `Cmd::Quit(&'static str)` (b671c8b) |
| A second engine never runs the updater: `IGNEUM_APP_NO_OTA=1` (implied by `--sweep`) skips the OTA tick and refuses Check now, whatever the manifest's `min_supported_version` says (the installer it would launch quits the installed app: PC 1, 22:31 UTC) | `Engine.no_ota`; the playbooks set the variable; `tools/ci/second-engine-check.sh` demands it |
## 6. Tests
@ -153,8 +154,9 @@ measure, TUNE record, upload, aggregation, prior shape in a test manifest). The
9070 XT run are owed: the 5090 the moment Power control is switched on (one prompt, then the tune runs by itself
within 2 minutes of steady mining), the 9070 XT when the card is back on the bus.
Run 1 (ember-tune-pc1-1, 22:30 UTC): aborted 46 s in by the installed app quitting (source unnamed by the 0.3.10 app;
not an update, not a job, not a relay task: C35 in the bench log), before any step; nothing set; the "before" snapshots are in the bench log (5090: 450 W of 575, 2,850 MHz core, 3,090 MHz
Run 1 (ember-tune-pc1-1, 22:30 UTC): aborted 46 s in by the installed app quitting, named the next morning: the
second engine's own updater (0.3.9 under min_supported_version = urgent) ran the per-user installer, whose
PrepareToInstall quit the installed app through its api/quit (C35 in the bench log); before any step; nothing set; the "before" snapshots are in the bench log (5090: 450 W of 575, 2,850 MHz core, 3,090 MHz
maximum, 14,001 MHz memory; 9070 XT present on bus 98 with OFFSET ranges `gmax_range -500 1000`, `plimit_range -30
10`). The offset finding changed the AMD mapping (054e041): an offset clock range closes the clock knob and the power
ladder runs on a percent scale bounded by `plimit_range`. The re-run follows the 0.3.11 rollout.

View file

@ -95,6 +95,7 @@ Snapshot 'before'
$env:IGNEUM_APP_DATA = $root
$env:IGNEUM_APP_LOGS = $sLogs
$env:IGNEUM_APP_STATUS_SECS = '10'
$env:IGNEUM_APP_NO_OTA = '1' # C35: a second engine never runs the updater (the installer would quit the installed app)
# C35 (5 October 2026): the engine's output goes to a FILE, never a pipe. A pipe's write end is inherited by every
# process the engine starts (its miners and workers), so after an abort the installed app's jobs runner waits for an
# EOF that never comes and hangs in its own quit; and the engine's whole tree is killed at the end (nothing orphaned).

View file

@ -60,6 +60,7 @@ if (Test-Path $smi) {
$env:IGNEUM_APP_DATA = $root
$env:IGNEUM_APP_LOGS = $sLogs
$env:IGNEUM_APP_STATUS_SECS = '10'
$env:IGNEUM_APP_NO_OTA = '1' # C35: a second engine never runs the updater (the installer would quit the installed app)
# C35 (5 October 2026): the engine's output goes to a FILE, never a pipe. A pipe's write end is inherited by every
# process the engine starts (its miners and workers), so after an abort the installed app's jobs runner waits for an
# EOF that never comes and hangs in its own quit; and the engine's whole tree is killed at the end (nothing orphaned).

View file

@ -7,7 +7,8 @@
# output goes to a FILE (Start-Process -RedirectStandardOutput <file>), never a pipe into the script; (2) the engine's
# whole process tree is ended at the end and on the budget (taskkill /T /F), so nothing is orphaned; the installed
# app's miners come back only after that (the job runner restarts them when the script ends). This check fails CI when
# a playbook starts an engine without both.
# a playbook starts an engine without both, or without IGNEUM_APP_NO_OTA = '1' (the third line, same night: the second
# engine's updater found itself under min_supported_version and ran the installer, which quit the installed app).
set -euo pipefail
cd "$(dirname "$0")/../.."
fail=0
@ -19,6 +20,9 @@ while IFS= read -r f; do
if ! grep -qE 'taskkill /T /F' "$f"; then
echo "second-engine: $f starts an engine without ending its process tree (taskkill /T /F) at the end"; fail=1
fi
if ! grep -qE "IGNEUM_APP_NO_OTA *= *'1'" "$f"; then
echo "second-engine: $f starts an engine without IGNEUM_APP_NO_OTA = '1' (its updater would run the installer, which quits the installed app: PC 1, 5 October 2026, 22:31 UTC)"; fail=1
fi
done < <(git ls-files 'relay/playbooks/**' 'tools/windows/**' 'packaging/**' | grep -E '\.ps1$')
[ "$fail" = 0 ] && echo "second-engine: every playbook that starts an engine logs to a file and ends its tree"
exit $fail