From aabe260948b090de2ea3b935175b0d866f417828 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Wed, 7 Oct 2026 15:09:50 +0000 Subject: [PATCH] miner-faults: MF-11 in the update-return lane's words (the helper owns the return, the relay service, the ping, the tuner's ceiling) and MF-13 (the Power Helper's stale-command count); the external card kind noted Co-Authored-By: Claude Fable 5.1 --- docs/plans/miner-faults.md | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/docs/plans/miner-faults.md b/docs/plans/miner-faults.md index e832b4b8c..d61d7aed8 100644 --- a/docs/plans/miner-faults.md +++ b/docs/plans/miner-faults.md @@ -32,10 +32,13 @@ Standing rules behind every row (branch `miner-reliability`, off `release-0.3.19 | MF-8 | 7 Oct 2026, every node build from 10db4b61 on a kept 0.3.17 datadir | "Node refuses its own kept datadir after an update": the app updates, the node dies at once (`DeserializationError(Io(Kind(UnexpectedEof)))` at `consensus/src/model/stores/virtual_state.rs:250`), the app restarts it in a loop, the miner never starts | `silent: bool` was added to `BlockRewardData` under `serde(default)` (10db4b61, the 0.3.16 vote-or-burn commit) and bincode ignores serde defaults, so the old row reads short; no canary saw it because every canary wiped. Fixed as ledger N13 on `release-0.3.20-node` b7cc37e7 (the store reads the v1 row and rewrites it) | Every schema change to a stored row ships with a versioned read path (the old shape read, rewritten in the new one); every release's gate starts the pinned binary on a COPY of a standing box's datadir, Linux-shaped and Windows-shaped (a copy of PC 2's), with the rewrite line or the synced line as the pass; the app never asks the user to wipe: a node that dies inside 10 s of its start is reported home (`node-exit` FAULT line) and the row says "the node cannot read its data after the update; the team has the report" | The node line's N13 test (the v1 row read and rewritten); the gate's kept-datadir start on both shapes; the injector step `kept-datadir` (`app-run.mjs --node-old `: the previous release's node writes the datadir to 150 blocks and stops, the engine's new node must open it, read synced, with one start and no exit line; its pod seconds are appended with the 0.3.21 run) | `canary kept-datadir start PASS (linux, windows)` in every cut's plan; `LG-4` row kept=true beside the nine wiped | | MF-9 | 7 Oct 2026, release-0.3.20-node b7cc37e7, the mixed-version gate | A node that stops slower than its restart: the restart died at once on the datadir LOCK of the stopping node | The listener watchdog slept its whole 10 s poll before it checked the shutdown flag, so a stop took up to 10 s while the app's restart followed inside it | Every poll loop in the node returns the moment shutdown is set (a `select` on the shutdown signal, never a sleep then a check); "a shutdown returns within one second" is a named gate of every release beside the kept-datadir start; the app's `stop_node` waits for the exit before the restart (`restart_node` schedules the start after the stop returns) | The node line's shutdown-latency test (stop at a random moment of the poll, return under 1 s); the mixed-version gate's restart case; `app-run.mjs` step `node-silent` reads the restart-to-synced seconds | `node shutdown under 1 s PASS` in every cut's plan | | MF-10 | 7 Oct 2026, the prover roll (3080, 3090, 4070, 5090) | Every proof fails in 12 s with `CudaRustError: named symbol not found`; the miner never notices and the box proves nothing for hours | The prover's `sp1-gpu-server` was built for another card's compute capability (sm_86 on the 3080 and 3090, sm_89 on the 4070, sm_120 on the 5090), so every kernel load fails the same way | The prover reads the card's compute capability at start (`nvidia-smi --query-gpu=compute_cap`) and refuses to start a server that does not match, with the reason on the prover tile and a FAULT line home ("the proving server is built for sm_89, this card is sm_120"); a proof that fails inside 15 s with the symbol error marks the server mismatched the same way; the 0.3.21 kit ships one server per architecture chosen at install, or a fat binary | `provedefault` test `a_server_for_another_card_is_refused` (sm_86 on 8.9 and 12.0 refused, 8.6 passes, a fat binary passes, unknowns pass); the fleet roll's paired line per box | `prover server matches the card PASS` per box in the roll, beside the kept-datadir and the shutdown gates | -| MF-11 | 7 Oct 2026, PC 2 (1ccfe586), silent 10:46Z to 13:38Z (Kernel-Power 41 and 6008 "unexpected shutdown", no bugcheck, no dump; two such events on that box today; the 0.3.19 update-now at 10:34:38Z had returned cleanly, so the update is not the cause) | The machine goes silent and nothing tells anyone: no job, no relay task, no line reaches the team for three hours; the restart path today is a hand at the PC | A power loss took the whole PC; the relay agent runs under the app's lifetime instead of beside it, and no watcher raises a line when a machine stops reporting | An app or agent that does not report within 10 minutes raises a line somewhere that still runs (the intake's own silence watch per machine id: a FAULT line "no report from for N min" to the team's channel); the relay agent runs as a service that survives the app and restarts the app on boot (the relay lane, 0.3.21); the app's first act after any start is the read-back line "app up, node " to the intake | the intake's silence-watch test (a machine whose last upload is older than 10 min is listed once); the relay lane's service test (the app killed, the agent still answers; the box rebooted, the app back) | `silence watch` line in the team channel per quiet box; PC 2's own line when it is back. Note for MF-8: PC 2's datadir, cut off mid-write twice by power loss, started clean on c4459193 (the PC 2 job's PASS) | +| MF-11 | 7 Oct 2026, PC 2 (1ccfe586), silent from 10:46Z after the 0.3.19 update-now (the update-return lane's row; code on update-return 0b423697, the app half on update-return-21b a64c193f) | The app did not come back and nothing reached the PC for hours; the only recovery was a hand on the power button; the stale relay logon task popped "Windows cannot find 'igneum-agent'" at every boot | A power loss or hard reset of the whole PC (Kernel-Power 41, EventLog 6008, no BugCheck 1001, no minidump; three such events that day, the third with the 5060 Ti enclosure attached and no TDR, WHEA or Thunderbolt trace), while the update itself had returned in 10 s (10:34:28Z quit, 10:34:38Z "[ok] updated to 0.3.19"); the relay agent dead since 6 October behind a UAC prompt; the Windows helper's return path checked nothing after the installer's exit; the tuner at 575 W with proving on the same card nine minutes before the first drop | (1) the helper owns the return: exe set kept beside the app, the app launched by the helper (/IGNOTA=2), api/state polled 120 s, the kept set restored, one intake line either way; the host restarts a dead engine and answers the Restart Manager; the first act after an update is the read-back line; a boot after a power loss posts FAULT pc-restart; (2) the relay agent as the per-user logon task IgneumRelayService (LeastPrivilege, no UAC, restart on failure, full path, stale IgneumRelayAgent* removed), the start-app kind, the per-install hostname; (3) every wake request carries the ping; the console says "job channel silent since