docs/plans/miner-faults.md: MF-1 to MF-7, each with its rule, test and gate line.
- MF-1/MF-2: a worker starts and is judged only when the node is READY (synced and igneum_getExecStatus reports an
executed tip; execrpc::probe every 5 s off the engine thread); the node watchdog never counts the catch-up (settled
once read synced; 30 min cap before that; any RPC answer is a sign of life); the watchdog restarts on a ladder 10 s,
30 s, 2 min, 5 min, then every 5 min for ever (watchdog::RETRY_LADDER_S); the faulted state and the one-restart
budget are gone (tools/ci/permanent-fault-check.sh in the gate); a node-caused restart resets the ladder at sync.
- MF-3: the hot-plug pass starts a recovered or revived card's worker (unchanged rule, now in the register).
- MF-4: the status clock starts at ready (program loaded), loading bounded by 300 s; a self-test failure holds the
card 30 min with the reason on its row, released on a driver change; a crash loop climbs the ladder; the pack is
exported once a minute for every card (a refused pack forces one).
- MF-5: the app reads template_wait=, template_ms=, identities_active= from the 0.3.20 miner's STATUS; waiting on
the node is never the card's fault; the row says node slow; every node-wait label clears on the first rate.
- MF-6: a miners hold belongs to the job that took it and releases when that job is gone or at its own cap.
- MF-7: the engine owns every igneum-miner it started: an untracked one on this engine's node RPC is killed at start,
after every stop and every minute, one line and one fault report per kill; a restart kills the old process first.
- Every fault line posts one FAULT line to the log intake (label fault-<id8>, app and node version, 60/h cap).
- The signed cards job kind (per card enabled, identities, power_pct; refused for a card the machine lacks; applied
through the app's own card path, persisted, read back): packaging/ota/publish-jobs.sh add --kind cards.
- LG-4 as a job: relay/playbooks/first-share.ps1 and tools/fleet/first-share-gate.mjs (no Windows box yet).
- tools/reliability: the fault injector with one step per class (catch-up, card-appears, own-restart, zero-ladder,
no-status, node-silent, one-card-fails, orphan-miner); fake-worker.mjs lists devices and fails self-tests on command.
- master's build tooling (97255a4e) and release-0.3.20's igneum-pow taken into the worktree for the box routes.
Box: app 198 + 27 + 8 tests green on igneum-build-2; the tree gate green (33 checks).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
docs/plans/miner-ui-5.md (what is built, gate results, owed, the RPC fields the node does not expose yet, what the
site lane takes); docs/plans/miner-ui-5-shots/ (light and dark at 1440 and 390, the saved block card PNG);
docs/plans/miner-ui-5-first-share-runbook.md; docs/community/discord/ladder.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
PC 1's read-back: the node reported synced while its finality replay blocked the template RPC for three minutes; the miners printed only timeouts and the watchdog faulted every card. The timeout line is the miner's heartbeat: silence clocks start over from it, the card says it waits for templates, one Activity line per episode. Recorded sequence as the test; 173 box tests.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The card watchdog judges a miner only while the node is synced; a silence fault from the node's sync is released at the synced transition and the card starts again with an Activity line; a worker silent from its start is faulted at 60 s; one Activity line per card at its first start. Known-failed test from PC 1's 04:51Z relaunch; 173 box tests.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
src/merge.rs, pure: every network sink more than the merge depth ahead and none of our blocks in the network's view for 120 s. Engine polls the observer's view every 30 s while mining and runs the check after sync_decision_v2; state behind, sync_cause not merging, one error event. Known-failed test first; 172 box tests, 42 UI tests.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The mode (own | external | none) is re-decided every 5 s while the app reads or refuses another node; gone for 60 s, the app starts its own node and says so in Activity. node.mode and node.mode_reason in api/state and on the Node details. Pure extnode::step with the goes-away test first.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
View.finalityWords with Ember's test (41 UI tests). extnode reads the node lane's reply shape: params compared value by value, network must match, a stub engine is refused and a fault on the app's own node. 168 box tests green.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>