From fcdb04bb416ea54a00acbed8eb37ecab06c8808f Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Sun, 4 Oct 2026 11:04:54 +0000 Subject: [PATCH] Bench log: the gfx1036 worker fault, the two Apple OpenCL soaks (no reproduction, no leak) and the injected-fault check; ignore the soak outputs Co-Authored-By: Claude Fable 5.1 --- .gitignore | 5 ++++- docs/bench-log.md | 33 +++++++++++++++++++++++++++++++++ 2 files changed, 37 insertions(+), 1 deletion(-) diff --git a/.gitignore b/.gitignore index f4d63d7d..4fa1a60f 100644 --- a/.gitignore +++ b/.gitignore @@ -21,6 +21,9 @@ proto-cuda/nvrtc/redist/ proto-cuda/nvrtc/emu/build/ proto-cuda/nvrtc/*.exe proto-opencl/*.exe -proto-opencl/igneum-bench-cl-generic-test proto-opencl/generic-test*.log proto-cuda/nvrtc/emu/test-run.log +proto-opencl/igneum-bench-cl-generic-test* +proto-opencl/soak-*.log +proto-opencl/hardened-check*.log +vendor/igneum-node-ship/ diff --git a/docs/bench-log.md b/docs/bench-log.md index 06fcd1b7..d3236c4e 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -719,3 +719,36 @@ each, the Mac's Metal miner, the integrated AMD chip's identities), floor 2/3. T carried in block `1e2439a1`; `observer.mjs` logged `checkpoint_locked` 0.7 s after the miner's own `LOCK` line. The floor was raised to 2/3 this morning (O-3.15); this is its first live lock. Difficulty at the moment of the lock was mid-oscillation (93M to 99M, see the oscillation finding), which did not affect voting. + +## 4 October 2026, the gfx1036 worker fault and what the Mac could and could not reproduce + +PC 2 (RTX 5090 plus the Ryzen's integrated gfx1036), package 0.3.0 prebuilt workers. The CUDA worker compiled the pack +with NVRTC (after `-default-device`) and mined at 124.2 MH/s, equal to the nvcc-built worker, 0 rejected, CPU re-check +clean. The OpenCL worker (`igneum-worker-opencl.exe --pack`, path `prebuilt-generic`) self-tested PASS and mined +correctly at 3.3 MH/s for 577 s (8 accepted blocks), then from about 600 s every job "completed" in 0.5 ms with no +hash: 906 jobs became 56,384 within 30 s, the miner reported 4.3 GH/s inside jobs and the dashboard over 1 GH/s, with +no error line, no exit and no restart. PC 1's cl.exe-built worker on the same host.c serve loop ran over an hour +without this. + +Root cause, as far as it can be stated: the AMD runtime kept answering `clEnqueueNDRangeKernel`, `clWaitForEvents` and +the blocking `clEnqueueReadBuffer` with CL_SUCCESS while running nothing, so the loop walked its chunks at memory speed +and reported the stale output buffer as a finished job. What flipped the runtime into that state at 600 s is not +visible in the logs and the job path itself leaks nothing (one event per chunk, created and released; verified below). +The two plausible triggers are a device reset of the integrated GPU with the runtime swallowing it (the generic path +is the only one that self-tests, which reads 256 MiB back and runs the three vector warps at start; an hourly +`prepare` would do the same work again on a second queue while jobs run) and a runtime limit reached after about 900 +jobs. Neither reproduces on Apple OpenCL: + +| Run on the M5 Max (Apple OpenCL 1.2, pack-a, 2^22 nonces per job) | Jobs | Job time ms (mean, min, max) | Faults | Live objects at the end | +|---|---|---|---|---| +| 20-minute soak of the shipped generic worker | 3,365 | 412 / 305 / 591 | 0 | not counted (that build had no counters) | +| 1,200-job soak of the hardened worker (events and buffers counted) | 1,200 | 411 / 315 / 549 | 0 | 0 events, 4 buffers (cache, dataset, out, init words); 1,200 events created and released | +| the same worker with `IGNEUM_FAULT_TEST=6` (the dispatch skipped from chunk 6 on, the runtime "succeeding") | 6 real + 1 | | 1: `the output buffer is unchanged since the previous dispatch`, exit 3 | | + +So the fix is defensive at three levels (commit 112acf6 and vendor devnet-v4 f9392600): the worker treats every +OpenCL error in the job path as fatal, requires CL_COMPLETE on the dispatch event, refuses a chunk 20x faster per nonce +than the running mean or an output buffer unchanged since the previous dispatch, prints live object counts every 200 +jobs and exits 3 on any of these; the miner kills and restarts a worker whose job time per hash drops under 1/20 of +the mean or whose interval rate exceeds 10x the mean before it, rolls its counters back to the last report and prints +`WORKER FAULT`; the launcher shows `worker fault` and `restarting` for that card instead of a rate. The next gfx1036 +run says which guard fires first; that line is the diagnosis the Mac cannot give.