docs/plans/miner-faults.md: MF-1 to MF-7, each with its rule, test and gate line.
- MF-1/MF-2: a worker starts and is judged only when the node is READY (synced and igneum_getExecStatus reports an
executed tip; execrpc::probe every 5 s off the engine thread); the node watchdog never counts the catch-up (settled
once read synced; 30 min cap before that; any RPC answer is a sign of life); the watchdog restarts on a ladder 10 s,
30 s, 2 min, 5 min, then every 5 min for ever (watchdog::RETRY_LADDER_S); the faulted state and the one-restart
budget are gone (tools/ci/permanent-fault-check.sh in the gate); a node-caused restart resets the ladder at sync.
- MF-3: the hot-plug pass starts a recovered or revived card's worker (unchanged rule, now in the register).
- MF-4: the status clock starts at ready (program loaded), loading bounded by 300 s; a self-test failure holds the
card 30 min with the reason on its row, released on a driver change; a crash loop climbs the ladder; the pack is
exported once a minute for every card (a refused pack forces one).
- MF-5: the app reads template_wait=, template_ms=, identities_active= from the 0.3.20 miner's STATUS; waiting on
the node is never the card's fault; the row says node slow; every node-wait label clears on the first rate.
- MF-6: a miners hold belongs to the job that took it and releases when that job is gone or at its own cap.
- MF-7: the engine owns every igneum-miner it started: an untracked one on this engine's node RPC is killed at start,
after every stop and every minute, one line and one fault report per kill; a restart kills the old process first.
- Every fault line posts one FAULT line to the log intake (label fault-<id8>, app and node version, 60/h cap).
- The signed cards job kind (per card enabled, identities, power_pct; refused for a card the machine lacks; applied
through the app's own card path, persisted, read back): packaging/ota/publish-jobs.sh add --kind cards.
- LG-4 as a job: relay/playbooks/first-share.ps1 and tools/fleet/first-share-gate.mjs (no Windows box yet).
- tools/reliability: the fault injector with one step per class (catch-up, card-appears, own-restart, zero-ladder,
no-status, node-silent, one-card-fails, orphan-miner); fake-worker.mjs lists devices and fails self-tests on command.
- master's build tooling (97255a4e) and release-0.3.20's igneum-pow taken into the worktree for the box routes.
Box: app 198 + 27 + 8 tests green on igneum-build-2; the tree gate green (33 checks).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The fleet's 22:09 UK incident (a Mac-side pkill -f <log file name> matched nothing, the roll-everything script lived on and wiped a held box) and the day's two pgrep self-matches are one class. The check flags pgrep -f / pkill -f with a plain literal (every one on a line), any pgrep/pkill on a file-name shape, and ps | grep with a literal; it allows the bracket form, -x, -F pidfile, kill $(cat pidfile), a variable and a full path; 11 banned and 16 allowed shapes in its self-test; 0.15 s over the tree. The 25 pkill -f sp1-gpu-server inside bash -c bodies (which matched the calling bash) are pkill -x; the other 11 literals take the bracket form; prover-socket-check accepts both. Row R in the record; the CLAUDE.md rule names the check and covers pkill and file names.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The dated entries of the 0.3.14 gate expired at 0.3.15. pc2-agg-cost.ps1 no longer switches the 5090 or falls back to /api/pause: the job is published with --cards-off <nvidia key>, the runner switches the card off before the script and puts it back on any exit, and the script fails loud (exit 3) if a CUDA worker is still running. pc2-agg-cost-restore.ps1 keeps the stray and socket clean-up and the read-back only. ember-tune-pc1.ps1 is superseded by the installed-tune playbook on ember-tune.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The node pin moves to 4c6b129d (release-0.3.14-node), which is what the live payload inputs carry, so windows-ci's payload-inputs step is green again (red since e9eca23 on master's 0.3.13 pin against the 0.3.14 inputs).
Conflicts: infra/fast-time/override-60x.json keeps master's side (the fresh-rule field was already there at u64::MAX; the release line would be a duplicate key); tools/ci/playbook-quit-check.sh keeps master's rule 2 and pre-rule list with the release side's Ember allow entry and the 0.3.15 expiry check.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Real screenshots of the live 0.3.14 app on the Apple M5 Max (paused, Ember Tune off, as the project lead left it), shot headless
at 1440x900 and 390 px, dark and light, the rail foot, the address and the machine id masked: miner-dashboard.webp
(the hero, same name), miner-earnings, miner-prove, miner-settings, miner-dashboard-light, miner-phone.
miner.html: a page-by-page section (the card row cell by cell, three page shots, light and the phone), the lead and
the meta copy name the four pages and pounds a day at the user's price, the Ember Tune feat carries the app's one
sentence and the measured PC 1 rows from the 6 Oct log (5090: 311.0 to 226.8 W for 0.15% of rate; 4070: 106.0 to
75.6 W for none; the floor, not the optimum), the Power control feat (one approval, then the helper task, no prompts),
the dev fee moved to Earnings as in the app. index.html: the same hero shot, two pills. build.mjs: the caption no
longer says 'until 0.3.6' (0.3.14 still says Igneum Miner). The 9070 XT row, the market price for pounds earned and
the bench anchor (the run 6 entry is on the ship0314 branch) are marked as owed or coming.
relay/playbooks/site-capture-pc1.ps1: a read-only run job that shoots the installed app's UI on PC 1 with Edge
headless and drops the PNGs on the relay (no quit, pause, resume or settings change; passes the playbook checks).
docs/plans/site-miner-2026-10-06.png: before and after.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>