bench-log: the root-socket recurrence of 21:25Z and the fix job of 22:01Z (the prover proved again at 22:02Z)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-05 22:03:06 +00:00
parent 23ab08202f
commit 4f0b9f5475

View file

@ -1666,7 +1666,7 @@ Reading. The GPU memory of a compressed shard proof is **13.9 GB for a shard of
Job `memminer-pc2-pv1` (`tools/proving-v1/pc2-memory-miner-on.ps1`), 20:28 to 20:30Z, the miner at full rate on the card, the live prover off for the run, the same 1-s sampler: the full shard at `S_p` (60.4 M cycles) peaked at **30,039 MiB** and took 33.0 s (28,295 MiB and 11.4 s with the card to itself: the miner costs the prover 2.9x in time and 1.7 GB of memory); the empty shard **15,670 MiB** and 7.7 s (13,863 and 2.3 s alone). So a 32 GB card mines and proves the prototype shard with 2.5 GB to spare; a 24 GB card cannot prove it even alone (28.3 GB); a 16 GB card cannot hold even the empty shard beside the miner (15.7 GB, the display and driver on top). `sp1-gpu-server --help` prints only `--version`: it has no options of its own, and its strings carry no memory setting (`CUDA_OUT_OF_MEMORY` is an error name). The shard SIZE is therefore the only lever left on this build, measured next as the S_p curve.
The root-socket fault (the class, fixed the same evening). The chain and memory jobs ran the host as root inside WSL2; the first `sp1-gpu-server` they started left `/tmp/sp1-cuda-0.sock` owned by root, and the live prover (the app's own WSL user) then failed every shard with `CudaClientError: Connect(Os { code: 13, kind: PermissionDenied })` (PC 2 app log 1791230456, 20:00:56Z) until the socket was gone. Every pv1 playbook now kills the server and unlinks `/tmp/sp1-cuda-*.sock` at its start and end, `tools/ci/prover-socket-check.sh` fails CI on any playbook that runs a prove mode as root without both lines (shown failing on `pc2-prover-cost.ps1` before its `--mode id`-only exemption, passing after), and the plan carries the rule: a prover job on a shared card runs as the app's user or cleans its socket. Playbooks that run the host: `tools/proving-v1/pc2-chain.ps1`, `pc2-memory-sweep.ps1`, `pc2-memory-miner-on.ps1`, `pc2-sp-curve.ps1` (all root, all with the cleanup now; the first two chain and sweep runs had none), `pc2-prover-cost.ps1` (`--mode id` only), `relay/playbooks/shard-test.ps1` and `proving/windows-wsl2/prove-shard.sh`, `prove-block.sh` (the app's user, not root), `tools/proving-v0/run.mjs` (the Mac, no server).
The root-socket fault (the class, fixed the same evening). The chain and memory jobs ran the host as root inside WSL2; the first `sp1-gpu-server` they started left `/tmp/sp1-cuda-0.sock` owned by root, and the live prover (the app's own WSL user) then failed every shard with `CudaClientError: Connect(Os { code: 13, kind: PermissionDenied })` (PC 2 app log 1791230456, 20:00:56Z) until the socket was gone. Every pv1 playbook now kills the server and unlinks `/tmp/sp1-cuda-*.sock` at its start and end, `tools/ci/prover-socket-check.sh` fails CI on any playbook that runs a prove mode as root without both lines (shown failing on `pc2-prover-cost.ps1` before its `--mode id`-only exemption, passing after), and the plan carries the rule: a prover job on a shared card runs as the app's user or cleans its socket. It recurred at 21:25Z from another agent's job (agg-cost-pc2-1, the same root-run shape) and survived the 0.3.10 restart at 21:49:41Z; the fix job `socketfix-pc2-pv1` (`tools/proving-v1/pc2-socket-fix.ps1`, 22:01:14 to 22:02:12Z) found `/tmp/sp1-cuda-0.sock` owned by root, removed it, switched the prover off and on, and the app's next shard (block 89011 shard 0) was proven and submitted in 34 s and paid 0.93 IGN at 22:02:24Z. Playbooks that run the host: `tools/proving-v1/pc2-chain.ps1`, `pc2-memory-sweep.ps1`, `pc2-memory-miner-on.ps1`, `pc2-sp-curve.ps1` (all root, all with the cleanup now; the first two chain and sweep runs had none), `pc2-prover-cost.ps1` (`--mode id` only), `relay/playbooks/shard-test.ps1` and `proving/windows-wsl2/prove-shard.sh`, `prove-block.sh` (the app's user, not root), `tools/proving-v0/run.mjs` (the Mac, no server).
### The S_p curve: peak GPU memory against shard size against time, the card to itself (the first of the two curve jobs)