igneum/docs/plans/miner-faults.md
igneum-labs f19bffed1a miner-faults: MF-8's injector step and MF-10's capability check are code now, not owed
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 13:45:41 +00:00

24 KiB

Miner fault-class register

Started 7 October 2026 (the project lead, 11:2x UK: "we cannot have issues like this with the miner, it needs to be truly plug, tune, play"). Every fault class found in the field gets a row the same day: the rule that makes the class impossible, the test that proves the rule, and the gate line a cut must show before it publishes. A class is closed when all three exist and the gate has fired once on a known-bad case and once on a known-good case (the watcher rule of 4 October).

Standing rules behind every row (branch miner-reliability, off release-0.3.19 db6f0964, for the 0.3.20 app; the miner rows on the fork branch miner-reliability-20 off release-0.3.20-node):

Rule Where
A worker never starts before the node is synced AND igneum_getExecStatus reports an executed tip; the node's catch-up never counts against its watchdog app/igneum-app/src/engine.rs (tick_exec_probe, node_ready, node_settled), src/watchdog.rs (NodeWatch::tick with settled, NODE_STARTUP_CAP_S)
No fault is permanent: a restart waits 10 s, 30 s, 2 min, 5 min, then every 5 min, for ever; the reason stays on the card row in plain words; the hash resumes on its own; there is no "restarted once already" state src/watchdog.rs (RETRY_LADDER_S, retry_delay_s, Action::Restart { delay_s, attempt }), engine.rs (watchdog_verdict, the pack give-up)
A card swap, a driver install or a restart needs no tap: the hot-plug enumeration every 60 s (src/hotplug.rs) starts the worker of a card that appears, recovers from a problem code or revives engine.rs merge_detection → plan_miners
Every fault line reports to the log intake the moment it happens, with the card, the class, the reason and the app version engine.rs fault_report → update::upload_text; label fault-<platform>-<id8>; read with node tools/logs.mjs
Nothing on a user's machine is changed by a one-off script: card settings travel as the signed cards job kind (per card enabled, identities, power_pct), applied through the app's own card path, persisted, read back in the report, refused for a card the machine does not have src/jobs.rs (KINDS, validate_params), src/jobrun.rs (cards_job_choices, cards_applied), engine.rs (Action::ApplyCards), packaging/ota/publish-jobs.sh add --kind cards --cards "key=on:8"
A cut never dies on a kept datadir: every stored-row schema change carries a versioned read path, and every release's gate starts the pinned binary on a copy of a standing box's datadir (Linux and Windows shapes) beside the wiped canary; LG-4's tenth install keeps the ninth's datadir (MF-8) the node line's N13 (b7cc37e7); the canary form; tools/fleet/first-share-gate.mjs row kept=true
A node stops faster than its restart: every poll loop returns the moment shutdown is set, and "a shutdown returns within one second" is a named release gate beside the kept-datadir start (MF-9) the node line's listener watchdog fix (b7cc37e7's follow-up); the mixed-version gate; engine.rs restart_node (the start is scheduled after stop_node returns)
A prover never runs a server built for another card: the compute capability is read at start and a mismatch is refused with the reason shown and reported home; the kit ships one server per architecture or a fat binary (MF-10, 0.3.21) src/provedefault.rs server_mismatch and is_arch_failure (tested), src/prover.rs server_arch_check (the card's compute_cap and the server's sm_ words read through the tools' own shell at the first probe; a mismatch refuses the GPU path with the reason on the tile and a FAULT class=prover-arch line; a proof failing with the symbol error is logged as the same class), the kit's per-architecture servers (0.3.21 kit, owed)
The fresh-install claim (LG-4) is a job, not a runbook: tools/fleet/first-share-gate.mjs on rented Windows boxes, on every cut, its line read by the shipper's publish relay/playbooks/first-share.ps1, tools/fleet/first-share-gate.mjs, site/evidence/first-share-<version>.json

The register

Id Found Class (what the user saw) Root cause Rule Test Gate line
MF-1 7 Oct 2026, PC 1, 0.3.17 and 0.3.18 The node restarted every 40 s; the card said "waiting for the node" for an hour The app's clock sample, then its prover loop, called records-indexing exec RPCs (eth_getBlockByNumber, igneum_getAssignedShards) while the 0.3.17 node's exec follower held no record; the node panicked (rpc.rs:591, rpc.rs:808) Every exec RPC call the app makes goes through execrpc::call; a method is safe on an empty state or gated on an executed tip, and an unclassified method is refused (shipper, 2dfb2e0c). The node's whole records-indexing class is bounds-checked in the 0.3.20 node execrpc unit test: every method literal in the tree is classified, no other file builds an exec request; tools/reliability/app-run.mjs step catch-up: a node with no executed tip gets no gated call and no worker for 120 s, no node restart app tests: execrpc callers classified in the cut's plan; app-run catch-up PASS
MF-2 7 Oct 2026, PC 1, 0.3.18 After a restart all three cards said "no status line from the miner for 90 s (restarted once already)" and never mined again until a cards-API bounce The workers started the moment synced read true, before the node's catch-up (finality replay, exec follower) let templates flow; the watchdog's one-restart budget then marked them faulted for good Workers start only when the node is READY (synced and an executed tip); silence while the node is not ready never counts; a restart the node's readiness explained resets the ladder at sync; the ladder never ends; the faulted state is gone watchdog tests a_worker_started_while_the_node_catches_up_is_never_faulted, zero_rate_restarts_on_the_ladder_for_ever_and_the_hash_resumes, node_watch_waits_out_a_catch_up...; app-run.mjs steps zero-ladder (rungs 10, 30, 120 s observed, then the hash back on its own) and catch-up app tests: watchdog 16 green; app-run zero-ladder PASS; the words "restarted once already" absent from src/ (tools/ci/forbidden-strings.txt)
MF-4 7 Oct 2026, PC 1, Arc B580 beside a 5090 and a 9070 XT One card's worker failed its self-test and restarted every few seconds; the other two cards sat in "loading the program" past the 90-second watchdog and were faulted Every restart of the failing card exported the pack again under the export lock; the healthy cards' exports queued behind it; the watchdog's status clock ran from the process start, not from "program loaded" A card's worker failure never blocks another card: the pack is exported once per minute for every card (EXPORT_REUSE_S; a refused pack forces one); a worker that fails its self-test is held 30 minutes with "not usable on this driver: " on its row and tried again on a driver change (SelfTestFailed, DriverChanged); a miner that exits inside 120 s of its start restarts on the ladder (10 s, 30 s, 2 min, 5 min), never every few seconds; the status clock starts at ready (program loaded), and loading itself is bounded by 300 s watchdog tests self_test_failure_is_held_and_a_crash_loop_climbs_the_ladder, no_status_from_the_start (the clock from ready); app-run.mjs step one-card-fails (three cards, one failing for ever: the two mine on time, the third is held, at most 6 exports) app-run one-card-fails PASS
MF-5 7 Oct 2026, PC 1, 0.3.17 node, 24 identities (the real cause of the 11:2x faults; MF-4 withdrawn as the cause) Every card faulted "no status line from the miner for 90 s (restarted once already)" after the restart Evidence (PC 1, 7 Oct 2026 11:4x UK): with 8 identities a card (24 template fetches a round) the 0.3.17 node answered no template in 5 s; with 2 a card (4 fetches) both cards mined at full rate (5090 122.4 MH/s, 9070 XT 18.9) within two minutes. The node's getBlockTemplate answered past 5 s with 24 identities fetching; the miners waited for a template inside their job-fill loop and printed no STATUS at all; the watchdog read the silence as the worker's; one restart, then faulted for good Miner: STATUS every interval whatever the template state (template_wait=<s> while it waits, template_ms= the node's last template time, identities_active=); the feed fetches only as many identities as fit one pass inside 8 s at the node's measured template time (identities_for: all of them when the node answers under 1 s; 24 at 8 s per template becomes 1), raised again when it answers faster; a pass whose fetches all fail backs off 2, 4, 8, 10 s and retries for ever with a NODE SLOW line once per 30 s. App: a STATUS with template_wait>0 is the miner's heartbeat and the node's latency, never the card's fault (no zero-rate clock, no restart); the card reads "node slow: waiting for a block template for N s; the worker is kept" or "node slow: a template takes N s; k of n identities active"; the mitigation of the day (a one-off script POSTing /api/cards) is closed by the signed cards job kind watchdog tests a_slow_node_never_faults_the_card, parses_status_and_fault_lines (the 0.3.20 line); jobs test for the cards kind; injector step slow-node (a template stub answering in 8 s while three cards run: open, needs the stub) app-run slow-node PASS (0.3.20)
MF-6 7 Oct 2026, PC 1, 0.3.19 After a --stop-miners job, a following read-only job kept both cards "off, held for a remote job" for its whole three minutes The engine released the hold only when no job held the miners; the next job's active state hid the release (engine.rs 3032 class) A hold belongs to the job that took it (job_hold_owner) and releases the moment that job is no longer the running one, whatever runs next, or when its own cap passes (logged); a read-only job never holds (jobrun::hold_release) jobrun test a_hold_belongs_to_the_job_that_took_it (owner running, another job, no job, cap passed, no owner) app tests: hold rule green
MF-7 7 Oct 2026, PC 1, 0.3.19 Orphan igneum-miner.exe processes the app no longer tracked (two alive under --stop-miners with their rows at pid 0, one after) hammered the node's template RPC beside the tracked miners The engine lost track of miners it had started (a stop that timed out, a restart over a live process) and never looked for them again The engine owns every miner it started: at start, after every stop and every minute it kills any igneum-miner whose command line carries THIS engine's node RPC (the fence) and whose pid it does not track, one log line and one fault report per kill, never by name alone (sweep_orphan_miners, platform::miner_processes, kill_pid); a restart kills the slot's old process before the new one starts injector step orphan-miner (a stray miner on the engine's node is killed inside the minute, the engine's own miner left alone) app-run orphan-miner PASS
MF-8 7 Oct 2026, every node build from 10db4b61 on a kept 0.3.17 datadir "Node refuses its own kept datadir after an update": the app updates, the node dies at once (DeserializationError(Io(Kind(UnexpectedEof))) at consensus/src/model/stores/virtual_state.rs:250), the app restarts it in a loop, the miner never starts silent: bool was added to BlockRewardData under serde(default) (10db4b61, the 0.3.16 vote-or-burn commit) and bincode ignores serde defaults, so the old row reads short; no canary saw it because every canary wiped. Fixed as ledger N13 on release-0.3.20-node b7cc37e7 (the store reads the v1 row and rewrites it) Every schema change to a stored row ships with a versioned read path (the old shape read, rewritten in the new one); every release's gate starts the pinned binary on a COPY of a standing box's datadir, Linux-shaped and Windows-shaped (a copy of PC 2's), with the rewrite line or the synced line as the pass; the app never asks the user to wipe: a node that dies inside 10 s of its start is reported home (node-exit FAULT line) and the row says "the node cannot read its data after the update; the team has the report" The node line's N13 test (the v1 row read and rewritten); the gate's kept-datadir start on both shapes; the injector step kept-datadir (app-run.mjs --node-old <previous igneumd>: the previous release's node writes the datadir to 150 blocks and stops, the engine's new node must open it, read synced, with one start and no exit line; its pod seconds are appended with the 0.3.21 run) canary kept-datadir start PASS (linux, windows) in every cut's plan; LG-4 row kept=true beside the nine wiped
MF-9 7 Oct 2026, release-0.3.20-node b7cc37e7, the mixed-version gate A node that stops slower than its restart: the restart died at once on the datadir LOCK of the stopping node The listener watchdog slept its whole 10 s poll before it checked the shutdown flag, so a stop took up to 10 s while the app's restart followed inside it Every poll loop in the node returns the moment shutdown is set (a select on the shutdown signal, never a sleep then a check); "a shutdown returns within one second" is a named gate of every release beside the kept-datadir start; the app's stop_node waits for the exit before the restart (restart_node schedules the start after the stop returns) The node line's shutdown-latency test (stop at a random moment of the poll, return under 1 s); the mixed-version gate's restart case; app-run.mjs step node-silent reads the restart-to-synced seconds node shutdown under 1 s PASS in every cut's plan
MF-10 7 Oct 2026, the prover roll (3080, 3090, 4070, 5090) Every proof fails in 12 s with CudaRustError: named symbol not found; the miner never notices and the box proves nothing for hours The prover's sp1-gpu-server was built for another card's compute capability (sm_86 on the 3080 and 3090, sm_89 on the 4070, sm_120 on the 5090), so every kernel load fails the same way The prover reads the card's compute capability at start (nvidia-smi --query-gpu=compute_cap) and refuses to start a server that does not match, with the reason on the prover tile and a FAULT line home ("the proving server is built for sm_89, this card is sm_120"); a proof that fails inside 15 s with the symbol error marks the server mismatched the same way; the 0.3.21 kit ships one server per architecture chosen at install, or a fat binary provedefault test a_server_for_another_card_is_refused (sm_86 on 8.9 and 12.0 refused, 8.6 passes, a fat binary passes, unknowns pass); the fleet roll's paired line per box prover server matches the card PASS per box in the roll, beside the kept-datadir and the shutdown gates
MF-11 7 Oct 2026, PC 2 (1ccfe586), silent 10:46Z to 13:38Z (Kernel-Power 41 and 6008 "unexpected shutdown", no bugcheck, no dump; two such events on that box today; the 0.3.19 update-now at 10:34:38Z had returned cleanly, so the update is not the cause) The machine goes silent and nothing tells anyone: no job, no relay task, no line reaches the team for three hours; the restart path today is a hand at the PC A power loss took the whole PC; the relay agent runs under the app's lifetime instead of beside it, and no watcher raises a line when a machine stops reporting An app or agent that does not report within 10 minutes raises a line somewhere that still runs (the intake's own silence watch per machine id: a FAULT line "no report from for N min" to the team's channel); the relay agent runs as a service that survives the app and restarts the app on boot (the relay lane, 0.3.21); the app's first act after any start is the read-back line "app up, node " to the intake the intake's silence-watch test (a machine whose last upload is older than 10 min is listed once); the relay lane's service test (the app killed, the agent still answers; the box rebooted, the app back) silence watch line in the team channel per quiet box; PC 2's own line when it is back. Note for MF-8: PC 2's datadir, cut off mid-write twice by power loss, started clean on c4459193 (the PC 2 job's PASS)
MF-12 7 Oct 2026, the fleet lane's ten-member pool window (ten rented 3070s, igneum-miner mine none ... --pool pool-1:4463, daemon 03457d96, miner 9829bdf7); the pool lane's row (its branch pool-mf-row 343dd83b called it MF-11; renumbered here so the register has one number per class) A pool member without a node of its own stopped hashing at the epoch boundary 78 to 79 and never resumed: every member printed POOL SEEDS epoch c1fc2c7c... class 3 and worker: info prepare started for epoch c1fc2c7c160f8a19 ... (NVRTC sm_86 in the background), that prepare never answered prepared or prepare-failed, the worker sat at 0 percent GPU serving the epoch-78 pack, the daemon's STATUS read workers=9 accepted=0, vardiff eased every member from shift 10 to 27 with no share, and no error line was printed anywhere (12:50Z to 13:08Z; cleared only by a restart on a pack exported from pool-1's node) Open (the pool lane): the member's prepare line is the solo miner's shape (igneum/miner/src/pool.rs prepare_line), so the fault is either the member's pack write under --prepare-packs or the worker's NVRTC prepare on that pack; the fleet lane holds the worker lines and a solo control on the same box is the split A member that sent a prepare and heard nothing for 120 s treats it as prepare-failed: it re-exports the pack, re-sends the prepare once, and on a second silence restarts its worker with the reason on its STATUS line and in the pool's stats; the pool daemon flags a member whose accepted count stays 0 across an epoch roll (EPOCH STALL line, the member's online false on the page); no pool member is ever silent across a boundary owed (the pool lane): the pool's measure harness crosses two epoch boundaries with a member on none and a worker whose prepare is held (the stand-in worker of tools/reliability/fake-worker.mjs), and reads shares on both sides the ten-member window's capture: every member's accepted count rises across every boundary the window crosses
MF-3 7 Oct 2026, PC 1, Intel Arc The Intel driver's first install did not bind: the device sat in Code 12 at install time; the card never mined until a reboot A driver installed while the device reports a problem code (12, 43, 31) does not bind; nothing re-scanned the device afterwards, and the app only re-enumerates The app re-enumerates every 60 s and starts the worker the minute the OS drives the card (hotplug::diff recovered / revived, settle_new); the row says what to do while it does not ("reboot with the card attached; if it persists, reinstall the driver with the card attached"); a Windows host asks for a re-scan (pnputil /scan-devices) after a problem code is seen, every 5 minutes, at most 6 times (follow-up, host side) hotplug test a_driven_card_that_turns_faulty_is_errored_and_recovers_later; app-run.mjs step card-appears (a card listed after 2 minutes starts without a tap) app-run card-appears PASS

Commits (7 October 2026)

Repo Branch Commit What
igneum miner-reliability (off release-0.3.19 db6f0964, with release-0.3.20's igneum-pow and master's build tooling) 38a30397 to 9d2c267a (the engine at 89d117ce) the app side of every row, the register, the injector, the cards job kind, the LG-4 job, the CI check
igneum-node miner-reliability-20 (off release-0.3.20-node dc141409) f067f7c1 MF-5's miner side: STATUS while waiting, identities from the template time, the fetch back-off

Box lines: app tests 198 + 27 + 8 green on igneum-build-2 (the tree gate green, 33 checks on the Mac); igneum-miner 19 green on igneum-build-2; both Linux binaries built on igneum-build-1 (app sha256 28650614 at 89d117ce, miner b3bf3590 at f067f7c1).

Injector, run 4 on a one-shot RunPod pod (RTX 3070, Ubuntu 24.04, 7 October 2026 12:19Z to 12:48Z, engine 89d117ce, the 0.3.17-line miner, the stand-in worker; the pod's real GPU switched off through the app's own card path; the private node on devnet suffix 9960 with a 2-thread CPU block producer): seven of eight steps PASS, the eighth on its rerun (run 5, the stand-in's listing fixed at ecd79874): every class's step is green.

Step Class Result Seconds
catch-up MF-1, MF-2 no worker started while the execution layer held no record (120 s), the card said it waits for the executed tip, the node was not restarted, the worker started on its own once the record existed synced after 4.1; blocks to mining 18.0
own-restart (the miner's own worker restart) the card showed the worker fault; the app did not restart the miner (same pid) fault on the card 1.0; mining again 4.0
zero-ladder MF-2 three watchdog restarts at 10, 30, 120 s on the row, the reason in plain words, no faulted state, no "once already" words, mining back on its own zero to rung 1 77.3; rung gaps 81.3, 101.9; healthy to mining 129.8
no-status MF-2 the miner stopped with SIGSTOP: restarted for "no status line for 90 s", the stopped process killed, mining on a new process quiet to restart 92.2; restart to mining 311.7 (the ladder's 300 s rung, carried from the step before: no five healthy minutes between them)
node-silent node watchdog the node stopped with SIGSTOP: restarted by the app (120 s rule plus the 30 s terminate grace), synced, mining again quiet to restart 150.3; restart to synced 7.0; quiet to mining 170.4
one-card-fails MF-4 two new cards listed, two healthy ones mining, the third held 30 minutes with "not usable on this driver: ", the healthy two still mining 90 s on, the pack exported once and reused 7 times listed 48.1; two mining 48.1
orphan-miner MF-7 a stray igneum-miner on the engine's node killed by the minute sweep, one log line per kill, the engine's own miner left alone inject to kill 21.1
card-appears MF-3 run 5 (12:48Z to 12:53Z, the stand-in's listing with the real worker's header, ecd79874): a second card appeared and was listed, its worker started with no tap; it left and was marked removed while the other card kept mining; it came back and mined again with no tap. Run 4 had failed the removal because the stand-in's --list lacked the header line, which the engine reads as "the tool did not answer" and removes nothing on (the right call for a crashed driver) listed 40.1 after it appeared; mining 50.1; marked removed 50.1 after it left; mining again 72.2 after it came back

FAULT lines the engine posted to the intake in run 4: 13 (watchdog, worker-fault, node-exit, orphan-miner classes).

Which side: MF-1, MF-2, MF-3, MF-4, MF-6, MF-7 and the cards kind are app-side (0.3.20's app). MF-5 is both: the miner (the fork branch) prints the fields and caps the identities; the app reads them and keeps the worker. An app without the new miner still gets MF-5's app half (a template timeout line is the heartbeat; the zero-rate clock holds).

How a row is added

  1. The day the class is seen: the row with Found, Class, Root cause. The rule, the test and the gate line the same day where they exist; "open" where they do not, with the owner.
  2. The test is a unit test on recorded lines (src/watchdog.rs, src/hotplug.rs, src/execrpc.rs) or a step of the fault injector (tools/reliability/app-run.mjs against fake-worker.mjs), run on the box.
  3. The gate line is what the cut's plan (docs/plans/release-<v>.md) must carry before publish; the shipper reads it.

The fault injector (box)

tools/reliability/app-run.mjs drives a scratch engine (its own private node, the fake worker in place of the GPU worker, Linux or macOS) through one step per class and prints PASS/FAIL with the seconds. The box runs it with the Linux engine and node from target-remote/: tools/reliability/box-run.sh (the lock is the box's run slot).