docs/plans/miner-faults.md: MF-1 to MF-7, each with its rule, test and gate line.
- MF-1/MF-2: a worker starts and is judged only when the node is READY (synced and igneum_getExecStatus reports an
executed tip; execrpc::probe every 5 s off the engine thread); the node watchdog never counts the catch-up (settled
once read synced; 30 min cap before that; any RPC answer is a sign of life); the watchdog restarts on a ladder 10 s,
30 s, 2 min, 5 min, then every 5 min for ever (watchdog::RETRY_LADDER_S); the faulted state and the one-restart
budget are gone (tools/ci/permanent-fault-check.sh in the gate); a node-caused restart resets the ladder at sync.
- MF-3: the hot-plug pass starts a recovered or revived card's worker (unchanged rule, now in the register).
- MF-4: the status clock starts at ready (program loaded), loading bounded by 300 s; a self-test failure holds the
card 30 min with the reason on its row, released on a driver change; a crash loop climbs the ladder; the pack is
exported once a minute for every card (a refused pack forces one).
- MF-5: the app reads template_wait=, template_ms=, identities_active= from the 0.3.20 miner's STATUS; waiting on
the node is never the card's fault; the row says node slow; every node-wait label clears on the first rate.
- MF-6: a miners hold belongs to the job that took it and releases when that job is gone or at its own cap.
- MF-7: the engine owns every igneum-miner it started: an untracked one on this engine's node RPC is killed at start,
after every stop and every minute, one line and one fault report per kill; a restart kills the old process first.
- Every fault line posts one FAULT line to the log intake (label fault-<id8>, app and node version, 60/h cap).
- The signed cards job kind (per card enabled, identities, power_pct; refused for a card the machine lacks; applied
through the app's own card path, persisted, read back): packaging/ota/publish-jobs.sh add --kind cards.
- LG-4 as a job: relay/playbooks/first-share.ps1 and tools/fleet/first-share-gate.mjs (no Windows box yet).
- tools/reliability: the fault injector with one step per class (catch-up, card-appears, own-restart, zero-ladder,
no-status, node-silent, one-card-fails, orphan-miner); fake-worker.mjs lists devices and fails self-tests on command.
- master's build tooling (97255a4e) and release-0.3.20's igneum-pow taken into the worktree for the box routes.
Box: app 198 + 27 + 8 tests green on igneum-build-2; the tree gate green (33 checks).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
docs/bench-log.md: the fake-worker measurements (slow start one trip and 2.0 s restart, fake-fast guard in under
0.1 s, exit 43 at 8.8 s, CPU re-check stop at 0.5 s, stall guard at 60.1 s with STATUS lines through the silence,
one prepare per epoch with refused retries held; app: zero-rate restart at 79.6 s and faulted at 75.4 s on the
repeat, no-status restart at 90.4 s, silent node restarted at 150.7 s and synced 7.2 s later; no double restart on
the miner's own worker restart). Two defects the harness found are named with their fork commits.
docs/fud-ledger.md and the round-4 review table: M26, M27, X21 Fixed with the commits.
engine.rs: the miner's restart note no longer hides the fault reason on the card.
tools/reliability: the harness matches the miner's stderr lines where they are printed there.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Review round 4 X21 and the project lead's levers 5 and 6 (4 October 2026 evening):
- src/watchdog.rs: CardWatch (no status line for 90 s, or hash rate 0 for 60 s while the node is synced: one
restart, then the card is marked faulted with the reason in the UI and the log; the miner's own worker restarts
suspend both rules until ready, so there are no double restarts; exit 43 counts as a watchdog restart) and
NodeWatch (our node silent for 120 s is restarted in-process through the restart kind the remote jobs use, with a
growing delay). State machines with no clock; 11 unit tests replay recorded STATUS and WORKER FAULT lines.
- engine.rs reads STATUS mismatched=, faults= and the WORKER FAULT lines onto the card (faults, mismatched, message);
a faulted card is not restarted until the user changes its settings or resumes mining; other cards keep mining.
- site/miners.html (build.mjs, scrubbed like /bench, rows from site/miner-bench.json): card, generator version, best
MH/s, MH per watt where measured, miner version, date, source, measured by the team or reported by the fleet. In the
navigation beside the engineering log on every page and in the sitemap. One line says there is no other miner to
compare with.
- tools/reliability: fake-worker.mjs (the serve protocol, misbehaving on command), run.mjs (miner guards on a private
test network, ports 29950+) and app-run.mjs (the app watchdog end to end, ports 29960+).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>