Bench log: the gfx1036 worker fault, the two Apple OpenCL soaks (no reproduction, no leak) and the injected-fault check; ignore the soak outputs
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
f4668da116
commit
ef951cbbd6
2 changed files with 37 additions and 1 deletions
5
.gitignore
vendored
5
.gitignore
vendored
|
|
@ -21,6 +21,9 @@ proto-cuda/nvrtc/redist/
|
|||
proto-cuda/nvrtc/emu/build/
|
||||
proto-cuda/nvrtc/*.exe
|
||||
proto-opencl/*.exe
|
||||
proto-opencl/igneum-bench-cl-generic-test
|
||||
proto-opencl/generic-test*.log
|
||||
proto-cuda/nvrtc/emu/test-run.log
|
||||
proto-opencl/igneum-bench-cl-generic-test*
|
||||
proto-opencl/soak-*.log
|
||||
proto-opencl/hardened-check*.log
|
||||
vendor/igneum-node-ship/
|
||||
|
|
|
|||
|
|
@ -719,3 +719,36 @@ each, the Mac's Metal miner, the integrated AMD chip's identities), floor 2/3. T
|
|||
carried in block `1e2439a1`; `observer.mjs` logged `checkpoint_locked` 0.7 s after the miner's own `LOCK` line. The floor was
|
||||
raised to 2/3 this morning (O-3.15); this is its first live lock. Difficulty at the moment of the lock was mid-oscillation
|
||||
(93M to 99M, see the oscillation finding), which did not affect voting.
|
||||
|
||||
## 4 October 2026, the gfx1036 worker fault and what the Mac could and could not reproduce
|
||||
|
||||
PC 2 (RTX 5090 plus the Ryzen's integrated gfx1036), package 0.3.0 prebuilt workers. The CUDA worker compiled the pack
|
||||
with NVRTC (after `-default-device`) and mined at 124.2 MH/s, equal to the nvcc-built worker, 0 rejected, CPU re-check
|
||||
clean. The OpenCL worker (`igneum-worker-opencl.exe --pack`, path `prebuilt-generic`) self-tested PASS and mined
|
||||
correctly at 3.3 MH/s for 577 s (8 accepted blocks), then from about 600 s every job "completed" in 0.5 ms with no
|
||||
hash: 906 jobs became 56,384 within 30 s, the miner reported 4.3 GH/s inside jobs and the dashboard over 1 GH/s, with
|
||||
no error line, no exit and no restart. PC 1's cl.exe-built worker on the same host.c serve loop ran over an hour
|
||||
without this.
|
||||
|
||||
Root cause, as far as it can be stated: the AMD runtime kept answering `clEnqueueNDRangeKernel`, `clWaitForEvents` and
|
||||
the blocking `clEnqueueReadBuffer` with CL_SUCCESS while running nothing, so the loop walked its chunks at memory speed
|
||||
and reported the stale output buffer as a finished job. What flipped the runtime into that state at 600 s is not
|
||||
visible in the logs and the job path itself leaks nothing (one event per chunk, created and released; verified below).
|
||||
The two plausible triggers are a device reset of the integrated GPU with the runtime swallowing it (the generic path
|
||||
is the only one that self-tests, which reads 256 MiB back and runs the three vector warps at start; an hourly
|
||||
`prepare` would do the same work again on a second queue while jobs run) and a runtime limit reached after about 900
|
||||
jobs. Neither reproduces on Apple OpenCL:
|
||||
|
||||
| Run on the M5 Max (Apple OpenCL 1.2, pack-a, 2^22 nonces per job) | Jobs | Job time ms (mean, min, max) | Faults | Live objects at the end |
|
||||
|---|---|---|---|---|
|
||||
| 20-minute soak of the shipped generic worker | 3,365 | 412 / 305 / 591 | 0 | not counted (that build had no counters) |
|
||||
| 1,200-job soak of the hardened worker (events and buffers counted) | 1,200 | 411 / 315 / 549 | 0 | 0 events, 4 buffers (cache, dataset, out, init words); 1,200 events created and released |
|
||||
| the same worker with `IGNEUM_FAULT_TEST=6` (the dispatch skipped from chunk 6 on, the runtime "succeeding") | 6 real + 1 | | 1: `the output buffer is unchanged since the previous dispatch`, exit 3 | |
|
||||
|
||||
So the fix is defensive at three levels (commit 112acf6 and vendor devnet-v4 f9392600): the worker treats every
|
||||
OpenCL error in the job path as fatal, requires CL_COMPLETE on the dispatch event, refuses a chunk 20x faster per nonce
|
||||
than the running mean or an output buffer unchanged since the previous dispatch, prints live object counts every 200
|
||||
jobs and exits 3 on any of these; the miner kills and restarts a worker whose job time per hash drops under 1/20 of
|
||||
the mean or whose interval rate exceeds 10x the mean before it, rolls its counters back to the last report and prints
|
||||
`WORKER FAULT`; the launcher shows `worker fault` and `restarting` for that card instead of a rate. The next gfx1036
|
||||
run says which guard fires first; that line is the diagnosis the Mac cannot give.
|
||||
|
|
|
|||
Loading…
Reference in a new issue