Igneum bench log
Append-only. Every number here was measured on the machine named, on the date given.
2026-10-03 proto-metal / igneum-bench, first run
@@ -442,7 +442,15 @@ th{font-family:var(--f-mono);font-size:12px;letter-spacing:.12em;text-transform:Two defects the harness found on the way, both fixed on the branch: after a guard kill with the job queue not full, the fill loop's continue re-ran the failed stdin write forever (140,000 lines in 45 s, no restart, no STATUS; 501363e0); and the 60 s line-read timeout meant no STATUS line at all while a worker was silent (945153ab: the read now waits one status interval). The "first healthy STATUS" figures above are bounded by the 10 s status interval: the worker is back about 2.5 s after a trip, the next STATUS line reports it.
App watchdog (src/watchdog.rs), one engine run, the fake worker as the Metal worker, steps in order:
| Step | What happened | Result | Seconds |
|---|---|---|---|
| own-restart | the worker reports jobs done in 0.3 ms | the miner's guard restarted the worker; the card showed the fault and worker faults 1; the app did NOT restart the miner (same pid, app restarts 0) | fault on the card 1.0 after injection, mining again 9.1 after the fault |
| zero-once | jobs complete with 0 hashes | watchdog: hash rate 0 for 60 s while the node is synced; one app restart; mining on the new miner process | zero to restart 79.6 (60 s rule plus the status interval and the quit grace), restart to mining 11.3, zero to mining 91.0 |
| zero-faulted | the same again inside five minutes | card faulted: hash rate 0 for 60 s while the node is synced (restarted once already); 45 s later still faulted, no miner process, no further restart; the node kept running and the app stayed up; resume cleared it and mining resumed | zero to faulted 75.4 |
| no-status | the miner process stopped with SIGSTOP | watchdog: no status line from the miner for 90 s; the stopped process killed; mining on a new process | quiet to restart 90.4, restart to mining 11.0 |
| node-silent | the node process stopped with SIGSTOP | the app restarted the node in-process (the remote-job restart kind), the new node synced, mining resumed | quiet to restart 150.7 (120 s rule plus the 30 s terminate grace on a process that cannot answer SIGTERM), restart to synced 7.2, quiet to mining 160.0 |
Not measured: any real GPU. The job-time and interval guards, exit 43 and the app's faulted card have not run on an RTX 5090, the gfx1036 or a Metal card; the first real run is owed from the fleet logs. The "no status" rule has not been tried against a hung RPC (only a stopped process). Unit tests: cargo test -p igneum-miner -- guard (7) and cargo test -- watchdog in app/igneum-app (11), both replaying recorded STATUS and WORKER FAULT lines.