PC 2, 7 October 2026 18:34 BST: two rows sat on "exporting this hour's program" for 18 minutes while a third card mined. The export ran on a thread per card under one lock with no bound on the wait, and a result that never came left the slot in the building state for ever. - watchdog::export_wait: inside one interval (the ladder rung, at least 30 s) the row keeps its word; past it the row names what blocks the export; past two intervals the wait ends and the worker retries on its own interval. - engine: EXPORT_HOLDER names the card whose export holds the pack lock and for how long, EXPORT_LAST_ERROR the last failure (the node not at the epoch, no seeds.txt); export_blocker() reads them for the row; build_seq ignores a late result after the wait ended; a failed export retries on the ladder, never a flat five minutes. - docs/plans/miner-faults.md: MF-14 with rule, test and gate line. Test known-failed first: an export that never returns is named at one interval and given up at two. Gates: app tests 260 + 33 + 8 green on igneum-build-2; the tree gate GREEN, 56 checks. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
30 KiB
Miner fault-class register
Started 7 October 2026 (the project lead, 11:2x UK: "we cannot have issues like this with the miner, it needs to be truly plug, tune, play"). Every fault class found in the field gets a row the same day: the rule that makes the class impossible, the test that proves the rule, and the gate line a cut must show before it publishes. A class is closed when all three exist and the gate has fired once on a known-bad case and once on a known-good case (the watcher rule of 4 October).
Standing rules behind every row (branch miner-reliability, off release-0.3.19 db6f0964, for the 0.3.20 app; the miner rows on the fork branch miner-reliability-20 off release-0.3.20-node):
| Rule | Where |
|---|---|
A worker never starts before the node is synced AND igneum_getExecStatus reports an executed tip; the node's catch-up never counts against its watchdog |
app/igneum-app/src/engine.rs (tick_exec_probe, node_ready, node_settled), src/watchdog.rs (NodeWatch::tick with settled, NODE_STARTUP_CAP_S) |
| No fault is permanent: a restart waits 10 s, 30 s, 2 min, 5 min, then every 5 min, for ever; the reason stays on the card row in plain words; the hash resumes on its own; there is no "restarted once already" state | src/watchdog.rs (RETRY_LADDER_S, retry_delay_s, Action::Restart { delay_s, attempt }), engine.rs (watchdog_verdict, the pack give-up) |
A card swap, a driver install or a restart needs no tap: the hot-plug enumeration every 60 s (src/hotplug.rs) starts the worker of a card that appears, recovers from a problem code or revives |
engine.rs merge_detection → plan_miners |
| Every fault line reports to the log intake the moment it happens, with the card, the class, the reason and the app version | engine.rs fault_report → update::upload_text; label fault-<platform>-<id8>; read with node tools/logs.mjs |
Nothing on a user's machine is changed by a one-off script: card settings travel as the signed cards job kind (per card enabled, identities, power_pct), applied through the app's own card path, persisted, read back in the report, refused for a card the machine does not have |
src/jobs.rs (KINDS, validate_params), src/jobrun.rs (cards_job_choices, cards_applied), engine.rs (Action::ApplyCards), packaging/ota/publish-jobs.sh add --kind cards --cards "key=on:8" |
| A cut never dies on a kept datadir: every stored-row schema change carries a versioned read path, and every release's gate starts the pinned binary on a copy of a standing box's datadir (Linux and Windows shapes) beside the wiped canary; LG-4's tenth install keeps the ninth's datadir (MF-8) | the node line's N13 (b7cc37e7); the canary form; tools/fleet/first-share-gate.mjs row kept=true |
| A node stops faster than its restart: every poll loop returns the moment shutdown is set, and "a shutdown returns within one second" is a named release gate beside the kept-datadir start (MF-9) | the node line's listener watchdog fix (b7cc37e7's follow-up); the mixed-version gate; engine.rs restart_node (the start is scheduled after stop_node returns) |
| A prover never runs a server built for another card: the compute capability is read at start and a mismatch is refused with the reason shown and reported home; the kit ships one server per architecture or a fat binary (MF-10, 0.3.21) | src/provedefault.rs server_mismatch and is_arch_failure (tested), src/prover.rs server_arch_check (the card's compute_cap and the server's sm_ words read through the tools' own shell at the first probe; a mismatch refuses the GPU path with the reason on the tile and a FAULT class=prover-arch line; a proof failing with the symbol error is logged as the same class), the kit's per-architecture servers (0.3.21 kit, owed) |
The fresh-install claim (LG-4) is a job, not a runbook: tools/fleet/first-share-gate.mjs on rented Windows boxes, on every cut, its line read by the shipper's publish |
relay/playbooks/first-share.ps1, tools/fleet/first-share-gate.mjs, site/evidence/first-share-<version>.json |
The register
| Id | Found | Class (what the user saw) | Root cause | Rule | Test | Gate line |
|---|---|---|---|---|---|---|
| MF-1 | 7 Oct 2026, PC 1, 0.3.17 and 0.3.18 | The node restarted every 40 s; the card said "waiting for the node" for an hour | The app's clock sample, then its prover loop, called records-indexing exec RPCs (eth_getBlockByNumber, igneum_getAssignedShards) while the 0.3.17 node's exec follower held no record; the node panicked (rpc.rs:591, rpc.rs:808) |
Every exec RPC call the app makes goes through execrpc::call; a method is safe on an empty state or gated on an executed tip, and an unclassified method is refused (shipper, 2dfb2e0c). The node's whole records-indexing class is bounds-checked in the 0.3.20 node |
execrpc unit test: every method literal in the tree is classified, no other file builds an exec request; tools/reliability/app-run.mjs step catch-up: a node with no executed tip gets no gated call and no worker for 120 s, no node restart |
app tests: execrpc callers classified in the cut's plan; app-run catch-up PASS |
| MF-2 | 7 Oct 2026, PC 1, 0.3.18 | After a restart all three cards said "no status line from the miner for 90 s (restarted once already)" and never mined again until a cards-API bounce | The workers started the moment synced read true, before the node's catch-up (finality replay, exec follower) let templates flow; the watchdog's one-restart budget then marked them faulted for good |
Workers start only when the node is READY (synced and an executed tip); silence while the node is not ready never counts; a restart the node's readiness explained resets the ladder at sync; the ladder never ends; the faulted state is gone | watchdog tests a_worker_started_while_the_node_catches_up_is_never_faulted, zero_rate_restarts_on_the_ladder_for_ever_and_the_hash_resumes, node_watch_waits_out_a_catch_up...; app-run.mjs steps zero-ladder (rungs 10, 30, 120 s observed, then the hash back on its own) and catch-up |
app tests: watchdog 16 green; app-run zero-ladder PASS; the words "restarted once already" absent from src/ (tools/ci/forbidden-strings.txt) |
| MF-4 | 7 Oct 2026, PC 1, Arc B580 beside a 5090 and a 9070 XT | One card's worker failed its self-test and restarted every few seconds; the other two cards sat in "loading the program" past the 90-second watchdog and were faulted | Every restart of the failing card exported the pack again under the export lock; the healthy cards' exports queued behind it; the watchdog's status clock ran from the process start, not from "program loaded" | A card's worker failure never blocks another card: the pack is exported once per minute for every card (EXPORT_REUSE_S; a refused pack forces one); a worker that fails its self-test is held 30 minutes with "not usable on this driver: " on its row and tried again on a driver change (SelfTestFailed, DriverChanged); a miner that exits inside 120 s of its start restarts on the ladder (10 s, 30 s, 2 min, 5 min), never every few seconds; the status clock starts at ready (program loaded), and loading itself is bounded by 300 s |
watchdog tests self_test_failure_is_held_and_a_crash_loop_climbs_the_ladder, no_status_from_the_start (the clock from ready); app-run.mjs step one-card-fails (three cards, one failing for ever: the two mine on time, the third is held, at most 6 exports) |
app-run one-card-fails PASS |
| MF-5 | 7 Oct 2026, PC 1, 0.3.17 node, 24 identities (the real cause of the 11:2x faults; MF-4 withdrawn as the cause) | Every card faulted "no status line from the miner for 90 s (restarted once already)" after the restart | Evidence (PC 1, 7 Oct 2026 11:4x UK): with 8 identities a card (24 template fetches a round) the 0.3.17 node answered no template in 5 s; with 2 a card (4 fetches) both cards mined at full rate (5090 122.4 MH/s, 9070 XT 18.9) within two minutes. The node's getBlockTemplate answered past 5 s with 24 identities fetching; the miners waited for a template inside their job-fill loop and printed no STATUS at all; the watchdog read the silence as the worker's; one restart, then faulted for good | Miner: STATUS every interval whatever the template state (template_wait=<s> while it waits, template_ms= the node's last template time, identities_active=); the feed fetches only as many identities as fit one pass inside 8 s at the node's measured template time (identities_for: all of them when the node answers under 1 s; 24 at 8 s per template becomes 1), raised again when it answers faster; a pass whose fetches all fail backs off 2, 4, 8, 10 s and retries for ever with a NODE SLOW line once per 30 s. App: a STATUS with template_wait>0 is the miner's heartbeat and the node's latency, never the card's fault (no zero-rate clock, no restart); the card reads "node slow: waiting for a block template for N s; the worker is kept" or "node slow: a template takes N s; k of n identities active"; the mitigation of the day (a one-off script POSTing /api/cards) is closed by the signed cards job kind |
watchdog tests a_slow_node_never_faults_the_card, parses_status_and_fault_lines (the 0.3.20 line); jobs test for the cards kind; injector step slow-node (a template stub answering in 8 s while three cards run: open, needs the stub) |
app-run slow-node PASS (0.3.20) |
| MF-6 | 7 Oct 2026, PC 1, 0.3.19 | After a --stop-miners job, a following read-only job kept both cards "off, held for a remote job" for its whole three minutes |
The engine released the hold only when no job held the miners; the next job's active state hid the release (engine.rs 3032 class) | A hold belongs to the job that took it (job_hold_owner) and releases the moment that job is no longer the running one, whatever runs next, or when its own cap passes (logged); a read-only job never holds (jobrun::hold_release) |
jobrun test a_hold_belongs_to_the_job_that_took_it (owner running, another job, no job, cap passed, no owner) |
app tests: hold rule green |
| MF-7 | 7 Oct 2026, PC 1, 0.3.19 | Orphan igneum-miner.exe processes the app no longer tracked (two alive under --stop-miners with their rows at pid 0, one after) hammered the node's template RPC beside the tracked miners |
The engine lost track of miners it had started (a stop that timed out, a restart over a live process) and never looked for them again | The engine owns every miner it started: at start, after every stop and every minute it kills any igneum-miner whose command line carries THIS engine's node RPC (the fence) and whose pid it does not track, one log line and one fault report per kill, never by name alone (sweep_orphan_miners, platform::miner_processes, kill_pid); a restart kills the slot's old process before the new one starts |
injector step orphan-miner (a stray miner on the engine's node is killed inside the minute, the engine's own miner left alone) |
app-run orphan-miner PASS |
| MF-8 | 7 Oct 2026, every node build from 10db4b61 on a kept 0.3.17 datadir | "Node refuses its own kept datadir after an update": the app updates, the node dies at once (DeserializationError(Io(Kind(UnexpectedEof))) at consensus/src/model/stores/virtual_state.rs:250), the app restarts it in a loop, the miner never starts |
silent: bool was added to BlockRewardData under serde(default) (10db4b61, the 0.3.16 vote-or-burn commit) and bincode ignores serde defaults, so the old row reads short; no canary saw it because every canary wiped. Fixed as ledger N13 on release-0.3.20-node b7cc37e7 (the store reads the v1 row and rewrites it) |
Every schema change to a stored row ships with a versioned read path (the old shape read, rewritten in the new one); every release's gate starts the pinned binary on a COPY of a standing box's datadir, Linux-shaped and Windows-shaped (a copy of PC 2's), with the rewrite line or the synced line as the pass; the app never asks the user to wipe: a node that dies inside 10 s of its start is reported home (node-exit FAULT line) and the row says "the node cannot read its data after the update; the team has the report" |
The node line's N13 test (the v1 row read and rewritten); the gate's kept-datadir start on both shapes; the injector step kept-datadir (app-run.mjs --node-old <previous igneumd>: the previous release's node writes the datadir to 150 blocks and stops, the engine's new node must open it, read synced, with one start and no exit line; its pod seconds are appended with the 0.3.21 run) |
canary kept-datadir start PASS (linux, windows) in every cut's plan; LG-4 row kept=true beside the nine wiped |
| MF-9 | 7 Oct 2026, release-0.3.20-node b7cc37e7, the mixed-version gate | A node that stops slower than its restart: the restart died at once on the datadir LOCK of the stopping node | The listener watchdog slept its whole 10 s poll before it checked the shutdown flag, so a stop took up to 10 s while the app's restart followed inside it | Every poll loop in the node returns the moment shutdown is set (a select on the shutdown signal, never a sleep then a check); "a shutdown returns within one second" is a named gate of every release beside the kept-datadir start; the app's stop_node waits for the exit before the restart (restart_node schedules the start after the stop returns) |
The node line's shutdown-latency test (stop at a random moment of the poll, return under 1 s); the mixed-version gate's restart case; app-run.mjs step node-silent reads the restart-to-synced seconds |
node shutdown under 1 s PASS in every cut's plan |
| MF-10 | 7 Oct 2026, the prover roll (3080, 3090, 4070, 5090) | Every proof fails in 12 s with CudaRustError: named symbol not found; the miner never notices and the box proves nothing for hours |
The prover's sp1-gpu-server was built for another card's compute capability (sm_86 on the 3080 and 3090, sm_89 on the 4070, sm_120 on the 5090), so every kernel load fails the same way |
The prover reads the card's compute capability at start (nvidia-smi --query-gpu=compute_cap) and refuses to start a server that does not match, with the reason on the prover tile and a FAULT line home ("the proving server is built for sm_89, this card is sm_120"); a proof that fails inside 15 s with the symbol error marks the server mismatched the same way; the 0.3.21 kit ships one server per architecture chosen at install, or a fat binary |
provedefault test a_server_for_another_card_is_refused (sm_86 on 8.9 and 12.0 refused, 8.6 passes, a fat binary passes, unknowns pass); the fleet roll's paired line per box |
prover server matches the card PASS per box in the roll, beside the kept-datadir and the shutdown gates |
| MF-11 | 7 Oct 2026, PC 2 (1ccfe586), silent from 10:46Z after the 0.3.19 update-now (the update-return lane's row; code on update-return 0b423697, the app half on update-return-21b a64c193f) | The app did not come back and nothing reached the PC for hours; the only recovery was a hand on the power button; the stale relay logon task popped "Windows cannot find 'igneum-agent'" at every boot | A power loss or hard reset of the whole PC (Kernel-Power 41, EventLog 6008, no BugCheck 1001, no minidump; three such events that day, the third with the 5060 Ti enclosure attached and no TDR, WHEA or Thunderbolt trace), while the update itself had returned in 10 s (10:34:28Z quit, 10:34:38Z "[ok] updated to 0.3.19"); the relay agent dead since 6 October behind a UAC prompt; the Windows helper's return path checked nothing after the installer's exit; the tuner at 575 W with proving on the same card nine minutes before the first drop | (1) the helper owns the return: exe set kept beside the app, the app launched by the helper (/IGNOTA=2), api/state polled 120 s, the kept set restored, one intake line either way; the host restarts a dead engine and answers the Restart Manager; the first act after an update is the read-back line; a boot after a power loss posts FAULT pc-restart; (2) the relay agent as the per-user logon task IgneumRelayService (LeastPrivilege, no UAC, restart on failure, full path, stale IgneumRelayAgent* removed), the start-app kind, the per-install hostname; (3) every wake request carries the ping; the console says "job channel silent since | ota::return_tests (the legacy known-failed first, ok, rolled-back, rolled-back-silent, relaunched, installer-failed, every sequence, the helper carries every step); bootcheck (pc2_boot_reads_as_a_power_loss, a_6008_alone, a_clean_boot); ember (without_a_ceiling_the_full_plan_asks_the_5090_for_575_w, the_tuner_never_asks_above_the_measured_efficient_point, power_control_on_with_a_cap_the_user_raised, a_card_not_in_the_table); engine a_refused_cap_climbs_the_retry_ladder_then_faults; jobrun wake_query/ping_query; relay service.test.mjs, handler start-app, wake ping | "update-return: ok" per machine in the rollout table, "job channel polled" on every console card, PC 2's own line when it is back. Note for MF-8: PC 2's datadir, cut off mid-write by power loss, started clean on c4459193 (the PC 2 job's PASS) |
| MF-12 | 7 Oct 2026, the fleet lane's ten-member pool window (ten rented 3070s, igneum-miner mine none ... --pool pool-1:4463, daemon 03457d96, miner 9829bdf7); the pool lane's row (its branch pool-mf-row 343dd83b called it MF-11; renumbered here so the register has one number per class) |
A pool member without a node of its own stopped hashing at the epoch boundary 78 to 79 and never resumed: every member printed POOL SEEDS epoch c1fc2c7c... class 3 and worker: info prepare started for epoch c1fc2c7c160f8a19 ... (NVRTC sm_86 in the background), that prepare never answered prepared or prepare-failed, the worker sat at 0 percent GPU serving the epoch-78 pack, the daemon's STATUS read workers=9 accepted=0, vardiff eased every member from shift 10 to 27 with no share, and no error line was printed anywhere (12:50Z to 13:08Z; cleared only by a restart on a pack exported from pool-1's node) |
Open (the pool lane): the member's prepare line is the solo miner's shape (igneum/miner/src/pool.rs prepare_line), so the fault is either the member's pack write under --prepare-packs or the worker's NVRTC prepare on that pack; the fleet lane holds the worker lines and a solo control on the same box is the split |
A member that sent a prepare and heard nothing for 120 s treats it as prepare-failed: it re-exports the pack, re-sends the prepare once, and on a second silence restarts its worker with the reason on its STATUS line and in the pool's stats; the pool daemon flags a member whose accepted count stays 0 across an epoch roll (EPOCH STALL line, the member's online false on the page); no pool member is ever silent across a boundary |
owed (the pool lane): the pool's measure harness crosses two epoch boundaries with a member on none and a worker whose prepare is held (the stand-in worker of tools/reliability/fake-worker.mjs), and reads shares on both sides |
the ten-member window's capture: every member's accepted count rises across every boundary the window crosses |
| MF-13 | 7 Oct 2026, PC 2, every tune since the 14:35Z boot (and the 11:37 local tune of the 0.3.19 run before it); the update-return lane's row | "tune: request N refused: the helper did not run sequence N within 15 s (no line in helper.log)" on every request; the ladder never runs; the 5090 sits at "cap NOT applied" | The Power Helper's stale-command rule counted lines: the lines present at its start are skipped, the count reset only when the file shrank; the engine writes its four command lines with fs::write over a stale four-line cmd.txt, so the count never dropped and every new command read as "present at start"; helper.log shows "helper started" three times after the boot with no command executed; the task itself was Running with its exe on disk and a fresh heartbeat (not the stale-task class) | A rewrite is a new command: the skipped prefix must still read as it did at the start (powertask::effective_skip), else the skip is 0; a helper that gives no heartbeat or no line inside its window is a "FAULT power-helper:" line with its registration re-read; the registered probe wants the task enabled and its action exe on disk; a refused restore after a tune re-applies the cap on the FAULT power-cap ladder | powertask a_stale_command_file_does_not_run_at_start (extended: the same-length rewrite is the known-failed case first, then effective_skip resets it; a grown file keeps the stale prefix skipped); engine a_helper_that_does_not_answer_is_a_fault_line; powertask the_registration_is_per_user_highest_no_trigger_fixed_action (the probe's Test-Path) | the tune's first request acknowledged ("helper: nvidia-smi ...") on PC 1 and PC 2 after the 0.3.21 install; no "did not run sequence" line in a cut's PC runs |
| MF-14 | 7 Oct 2026, PC 2 after its 18:34 BST logon, 0.3.21 | The 5090 and Arc rows sat on "exporting this hour's program" for 18 minutes while the 5060 Ti mined; nothing said what blocked them | The program export runs on a thread per card under one lock, with a 120 s timeout per export and no bound on the wait behind the lock; a card whose export never answered kept its row on "exporting" for as long as the lock was held or the thread was gone (no result ever reached the engine), and its slot stayed in the building state for ever | A worker never waits on the export longer than one retry interval (its ladder rung, at least 30 s) without the reason shown: past it the row names what blocks the export (the card whose export holds the pack lock and for how long, this card's own export still running, the last export's error such as the node not yet at the epoch); past two intervals the wait ends, the slot leaves the building state, a late result is ignored by its sequence, and the worker retries on its own ladder interval, the other cards untouched; an export that fails retries on the ladder, never a flat five minutes (watchdog::export_wait, engine export_blocker, EXPORT_HOLDER, build_seq) |
watchdog test an_export_that_never_returns_is_named_and_given_up_within_two_intervals (known-failed first: 18 minutes behind another card's export is named at one interval and given up at two; a normal 5 to 25 s export says nothing; the 30 s floor on the 10 s rung) |
app tests green with the MF-14 test; owed: an injector step with an export that never returns (the node stopped during a card's start) |
| MF-3 | 7 Oct 2026, PC 1, Intel Arc | The Intel driver's first install did not bind: the device sat in Code 12 at install time; the card never mined until a reboot | A driver installed while the device reports a problem code (12, 43, 31) does not bind; nothing re-scanned the device afterwards, and the app only re-enumerates | The app re-enumerates every 60 s and starts the worker the minute the OS drives the card (hotplug::diff recovered / revived, settle_new); the row says what to do while it does not ("reboot with the card attached; if it persists, reinstall the driver with the card attached"); a Windows host asks for a re-scan (pnputil /scan-devices) after a problem code is seen, every 5 minutes, at most 6 times (follow-up, host side) |
hotplug test a_driven_card_that_turns_faulty_is_errored_and_recovers_later; app-run.mjs step card-appears (a card listed after 2 minutes starts without a tap) |
app-run card-appears PASS |
A kind, not a fault (the update-return lane): the "external" card kind, a USB4 or Thunderbolt router in the device's parent chain, so an eGPU reads eGPU on the Cards page and in the cards line (detect.rs ADAPTERS_SCRIPT, classify_kind; the view test "a card behind a USB4 or Thunderbolt router is an eGPU").
Commits (7 October 2026)
| Repo | Branch | Commit | What |
|---|---|---|---|
| igneum | miner-reliability (off release-0.3.19 db6f0964, with release-0.3.20's igneum-pow and master's build tooling) |
38a30397 to 9d2c267a (the engine at 89d117ce) |
the app side of every row, the register, the injector, the cards job kind, the LG-4 job, the CI check |
| igneum-node | miner-reliability-20 (off release-0.3.20-node dc141409) |
f067f7c1 |
MF-5's miner side: STATUS while waiting, identities from the template time, the fetch back-off |
Box lines: app tests 198 + 27 + 8 green on igneum-build-2 (the tree gate green, 33 checks on the Mac); igneum-miner 19 green on igneum-build-2; both Linux binaries built on igneum-build-1 (app sha256 28650614 at 89d117ce, miner b3bf3590 at f067f7c1).
Injector, run 4 on a one-shot RunPod pod (RTX 3070, Ubuntu 24.04, 7 October 2026 12:19Z to 12:48Z, engine 89d117ce, the 0.3.17-line miner, the stand-in worker; the pod's real GPU switched off through the app's own card path; the private node on devnet suffix 9960 with a 2-thread CPU block producer): seven of eight steps PASS, the eighth on its rerun (run 5, the stand-in's listing fixed at ecd79874): every class's step is green.
| Step | Class | Result | Seconds |
|---|---|---|---|
| catch-up | MF-1, MF-2 | no worker started while the execution layer held no record (120 s), the card said it waits for the executed tip, the node was not restarted, the worker started on its own once the record existed | synced after 4.1; blocks to mining 18.0 |
| own-restart | (the miner's own worker restart) | the card showed the worker fault; the app did not restart the miner (same pid) | fault on the card 1.0; mining again 4.0 |
| zero-ladder | MF-2 | three watchdog restarts at 10, 30, 120 s on the row, the reason in plain words, no faulted state, no "once already" words, mining back on its own | zero to rung 1 77.3; rung gaps 81.3, 101.9; healthy to mining 129.8 |
| no-status | MF-2 | the miner stopped with SIGSTOP: restarted for "no status line for 90 s", the stopped process killed, mining on a new process | quiet to restart 92.2; restart to mining 311.7 (the ladder's 300 s rung, carried from the step before: no five healthy minutes between them) |
| node-silent | node watchdog | the node stopped with SIGSTOP: restarted by the app (120 s rule plus the 30 s terminate grace), synced, mining again | quiet to restart 150.3; restart to synced 7.0; quiet to mining 170.4 |
| one-card-fails | MF-4 | two new cards listed, two healthy ones mining, the third held 30 minutes with "not usable on this driver: ", the healthy two still mining 90 s on, the pack exported once and reused 7 times | listed 48.1; two mining 48.1 |
| orphan-miner | MF-7 | a stray igneum-miner on the engine's node killed by the minute sweep, one log line per kill, the engine's own miner left alone | inject to kill 21.1 |
| card-appears | MF-3 | run 5 (12:48Z to 12:53Z, the stand-in's listing with the real worker's header, ecd79874): a second card appeared and was listed, its worker started with no tap; it left and was marked removed while the other card kept mining; it came back and mined again with no tap. Run 4 had failed the removal because the stand-in's --list lacked the header line, which the engine reads as "the tool did not answer" and removes nothing on (the right call for a crashed driver) |
listed 40.1 after it appeared; mining 50.1; marked removed 50.1 after it left; mining again 72.2 after it came back |
FAULT lines the engine posted to the intake in run 4: 13 (watchdog, worker-fault, node-exit, orphan-miner classes).
Injector on 0.3.21's own binaries (7 October 2026 15:44Z to 16:11Z, a one-shot RunPod 3070: app 0.3.21 sha256 786c3d37
built on igneum-build-2 from release-0.3.21 8ed08dcf, node c4459193 sha256 48acf4aa with N13 and the shutdown watchdog,
the 0.3.17-line igneumd 5899f603 as the previous release for --node-old): nine steps, eight PASS in the run and the
ninth (catch-up) on its fresh-datadir rerun (CATCHUP021).
| Step | Class | Result on 0.3.21 | Seconds |
|---|---|---|---|
| kept-datadir | MF-8 | the previous release's node wrote 151 blocks and stopped in 1.0 s (MF-9's sub-second shutdown, read on the OLD node); the new node opened the kept datadir, read synced, one start, no exit line | new node synced 4.1 after the engine started |
| catch-up | MF-1, MF-2 | in the run the datadir carried records from the first second (kept), so the worker started at once, which is right; the fresh-datadir rerun is the class's own read: CATCHUP021 | |
| card-appears | MF-3 | the card that appeared was listed and mined with no tap; the one that left was marked removed; back, it mined again | listed 58.2, mining 69.2; removed 49.1 after it left; mining again 72.2 after it came back |
| own-restart | (the miner's own) | the fault on the card, no app restart | 1.0 / 3.0 |
| zero-ladder | MF-2 | rungs 10, 30, 120 s, the reason on the row, mining back on its own | zero to rung 1 76.8; gaps 81.3, 101.9; healthy to mining 129.8 |
| no-status | MF-2 | the stopped miner restarted at 90 s, mining on a new process | quiet to restart 93.2; restart to mining 310.8 (the carried 300 s rung) |
| node-silent | node watchdog | the stopped node restarted by the app, synced, mining again | quiet to restart 151.4; restart to synced 3.0; quiet to mining 170.4 |
| one-card-fails | MF-4 | two healthy cards mined, the third held 30 min with the self-test reason, 2 exports with 2 reuses | listed 47.1; two mining 57.1 |
| orphan-miner | MF-7 | a stray miner killed by the sweep, the engine's own left alone | inject to kill 14.0 |
FAULT lines posted in the 0.3.21 run: 9.
Which side: MF-1, MF-2, MF-3, MF-4, MF-6, MF-7 and the cards kind are app-side (0.3.20's app). MF-5 is both: the miner (the fork branch) prints the fields and caps the identities; the app reads them and keeps the worker. An app without the new miner still gets MF-5's app half (a template timeout line is the heartbeat; the zero-rate clock holds).
How a row is added
- The day the class is seen: the row with Found, Class, Root cause. The rule, the test and the gate line the same day where they exist; "open" where they do not, with the owner.
- The test is a unit test on recorded lines (
src/watchdog.rs,src/hotplug.rs,src/execrpc.rs) or a step of the fault injector (tools/reliability/app-run.mjsagainstfake-worker.mjs), run on the box. - The gate line is what the cut's plan (
docs/plans/release-<v>.md) must carry before publish; the shipper reads it.
The fault injector (box)
tools/reliability/app-run.mjs drives a scratch engine (its own private node, the fake worker in place of the GPU
worker, Linux or macOS) through one step per class and prints PASS/FAIL with the seconds. The box runs it with
the Linux engine and node from target-remote/: tools/reliability/box-run.sh (the lock is the box's run slot).