Commit graph

9 commits

Author SHA1 Message Date
igneum-labs
4da5edf24f MF-10: the prover reads the card's compute capability and the GPU server's sm_ words and refuses a mismatch with the reason (and reads a named-symbol proof failure as the same class); MF-8: the injector's kept-datadir step (a previous release's node writes the datadir, the new node must open it)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 13:25:08 +00:00
igneum-labs
cb8a265431 fake worker: the OpenCL listing carries the real worker's header line, so the engine reads an answered enumeration and removes a card that left
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ecd798747a26d339ff514da04773be858780636c)
2026-10-07 13:12:08 +00:00
igneum-labs
942f537ff6 reliability injector: the stale duplicate step blocks removed (the card-appears rewrite had sliced before an earlier definition); the rung is read over two seconds
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 10badc02a3)
2026-10-07 13:12:08 +00:00
igneum-labs
fd4c3e72a7 reliability injector: card-appears uses a second fake device (added, removed, revived); an empty device list is read by the engine as no answer
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit df2617bb43)
2026-10-07 13:12:08 +00:00
igneum-labs
e464a5fb06 reliability injector: reads the fake cards by name and switches a box's real GPU off through the app's own card path
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit b9c70528fa)
2026-10-07 13:12:08 +00:00
igneum-labs
1506f6c003 reliability injector: the catch-up step's flag no longer shadows the process list
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 2efe419521)
2026-10-07 13:12:08 +00:00
igneum-labs
089b078a3a Miner app: plug, tune, play (the project lead, 7 October 2026): node readiness gate, the retry ladder, no permanent fault, fault lines to the intake, the signed cards job, the fault-class register
docs/plans/miner-faults.md: MF-1 to MF-7, each with its rule, test and gate line.
- MF-1/MF-2: a worker starts and is judged only when the node is READY (synced and igneum_getExecStatus reports an
  executed tip; execrpc::probe every 5 s off the engine thread); the node watchdog never counts the catch-up (settled
  once read synced; 30 min cap before that; any RPC answer is a sign of life); the watchdog restarts on a ladder 10 s,
  30 s, 2 min, 5 min, then every 5 min for ever (watchdog::RETRY_LADDER_S); the faulted state and the one-restart
  budget are gone (tools/ci/permanent-fault-check.sh in the gate); a node-caused restart resets the ladder at sync.
- MF-3: the hot-plug pass starts a recovered or revived card's worker (unchanged rule, now in the register).
- MF-4: the status clock starts at ready (program loaded), loading bounded by 300 s; a self-test failure holds the
  card 30 min with the reason on its row, released on a driver change; a crash loop climbs the ladder; the pack is
  exported once a minute for every card (a refused pack forces one).
- MF-5: the app reads template_wait=, template_ms=, identities_active= from the 0.3.20 miner's STATUS; waiting on
  the node is never the card's fault; the row says node slow; every node-wait label clears on the first rate.
- MF-6: a miners hold belongs to the job that took it and releases when that job is gone or at its own cap.
- MF-7: the engine owns every igneum-miner it started: an untracked one on this engine's node RPC is killed at start,
  after every stop and every minute, one line and one fault report per kill; a restart kills the old process first.
- Every fault line posts one FAULT line to the log intake (label fault-<id8>, app and node version, 60/h cap).
- The signed cards job kind (per card enabled, identities, power_pct; refused for a card the machine lacks; applied
  through the app's own card path, persisted, read back): packaging/ota/publish-jobs.sh add --kind cards.
- LG-4 as a job: relay/playbooks/first-share.ps1 and tools/fleet/first-share-gate.mjs (no Windows box yet).
- tools/reliability: the fault injector with one step per class (catch-up, card-appears, own-restart, zero-ladder,
  no-status, node-silent, one-card-fails, orphan-miner); fake-worker.mjs lists devices and fails self-tests on command.
- master's build tooling (97255a4e) and release-0.3.20's igneum-pow taken into the worktree for the box routes.
Box: app 198 + 27 + 8 tests green on igneum-build-2; the tree gate green (33 checks).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 13:11:48 +00:00
igneum-labs
cad8aa0d96 Reliability measured: miner guards and the app watchdog on private test networks; M26, M27, X21 fixed in the ledger
docs/bench-log.md: the fake-worker measurements (slow start one trip and 2.0 s restart, fake-fast guard in under
0.1 s, exit 43 at 8.8 s, CPU re-check stop at 0.5 s, stall guard at 60.1 s with STATUS lines through the silence,
one prepare per epoch with refused retries held; app: zero-rate restart at 79.6 s and faulted at 75.4 s on the
repeat, no-status restart at 90.4 s, silent node restarted at 150.7 s and synced 7.2 s later; no double restart on
the miner's own worker restart). Two defects the harness found are named with their fork commits.
docs/fud-ledger.md and the round-4 review table: M26, M27, X21 Fixed with the commits.
engine.rs: the miner's restart note no longer hides the fault reason on the card.
tools/reliability: the harness matches the miner's stderr lines where they are printed there.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 22:22:16 +00:00
igneum-labs
a601321481 App engine: watchdog per card and per node, WORKER FAULT and mismatched= read from the miner; public GPU bench table and the reliability test harness
Review round 4 X21 and the project lead's levers 5 and 6 (4 October 2026 evening):
- src/watchdog.rs: CardWatch (no status line for 90 s, or hash rate 0 for 60 s while the node is synced: one
  restart, then the card is marked faulted with the reason in the UI and the log; the miner's own worker restarts
  suspend both rules until ready, so there are no double restarts; exit 43 counts as a watchdog restart) and
  NodeWatch (our node silent for 120 s is restarted in-process through the restart kind the remote jobs use, with a
  growing delay). State machines with no clock; 11 unit tests replay recorded STATUS and WORKER FAULT lines.
- engine.rs reads STATUS mismatched=, faults= and the WORKER FAULT lines onto the card (faults, mismatched, message);
  a faulted card is not restarted until the user changes its settings or resumes mining; other cards keep mining.
- site/miners.html (build.mjs, scrubbed like /bench, rows from site/miner-bench.json): card, generator version, best
  MH/s, MH per watt where measured, miner version, date, source, measured by the team or reported by the fleet. In the
  navigation beside the engineering log on every page and in the sitemap. One line says there is no other miner to
  compare with.
- tools/reliability: fake-worker.mjs (the serve protocol, misbehaving on command), run.mjs (miner guards on a private
  test network, ports 29950+) and app-run.mjs (the app watchdog end to end, ports 29960+).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:20:49 +00:00