diff --git a/docs/fud-ledger.md b/docs/fud-ledger.md index b09381ab5..03a425d70 100644 --- a/docs/fud-ledger.md +++ b/docs/fud-ledger.md @@ -1472,7 +1472,7 @@ Evidence: `docs/bench-log.md`, 4 October 2026 "execution layer attack fixes" (be ### P20. The SP1 GPU client panics on shutdown and the compressed stage waited ten minutes "Your first GPU proof run aborted with a core dump. What else aborts?" -Status: (1) fixed in the host; (2) found on the second RTX 5090 run (4 October 2026, evening): the silence sits BEFORE the `STAGE compressed` line, so it is not the recursion setup. The only work there is saving the core proof, and SP1's `save` serialises straight into an unbuffered file: on WSL2 with the results folder under /mnt/c every 4-byte field element is one round trip across the Windows file bridge, and the 18.1 MB core proof of a full shard had not finished saving 25 minutes after the proof verified (the block-78 run's 10 minutes were the same save of a smaller proof). Fix: the host writes proofs through a 4 MB buffer and prints `saved in N s`, and `prove-shard.sh` writes results on the Linux side and copies them to the Windows folder once per stage. Pending the timed line from the next PC run. Was: Fixed in the host, to be confirmed on the PC (4 October 2026, afternoon). (1) `igneum-prove-host` holds a Tokio runtime for the whole run (`main` enters it before anything else) and drops the proof system inside it, then shuts the runtime down with a 10-s grace, so `sp1-cuda`'s `tokio::spawn` in `Drop` finds a runtime. (2) Every stage prints a `STAGE ... start` line and a `RESULT` line with a UTC timestamp, and the setup line splits client creation from the two key setups (on the Mac CPU the client creation alone is 30 to 50 s of a 40 to 60 s setup; bench-log 4 October 2026, shards). The next PC run (PROVE-SHARD.bat) shows whether the gap sits before the first compressed stage and how long it is on the card. Was: Open, found by the team (4 October 2026, morning, first RTX 5090 run through WSL2, SP1 6.8.1). +Status: (1) fixed in the host; (2) found on the second RTX 5090 run (4 October 2026, evening): the silence sits BEFORE the `STAGE compressed` line, so it is not the recursion setup. The only work there is saving the core proof, and SP1's `save` serialises straight into an unbuffered file: on WSL2 with the results folder under /mnt/c every 4-byte field element is one round trip across the Windows file bridge, and the 18.1 MB core proof of a full shard had not finished saving 25 minutes after the proof verified (the block-78 run's 10 minutes were the same save of a smaller proof). Fix: the host writes proofs through a 4 MB buffer and prints `saved in N s`, and `prove-shard.sh` writes results on the Linux side and copies them to the Windows folder once per stage. Confirmed by the stage timestamps of the second run (job run-20261004-173115): core proof verified 17:40:29 UTC, `STAGE compressed` 18:04:33 UTC, so the 18.1 MB core proof took 24 min 4 s to save; the 1.27 MB compressed proof took 1 min 42 s (18:04:44 to 18:06:26); the proving itself was 8.3 s and 10.9 s. The buffered save ships in the package rebuilt 18:10 UTC; the timed `saved` line on the next run closes this. Was: Fixed in the host, to be confirmed on the PC (4 October 2026, afternoon). (1) `igneum-prove-host` holds a Tokio runtime for the whole run (`main` enters it before anything else) and drops the proof system inside it, then shuts the runtime down with a 10-s grace, so `sp1-cuda`'s `tokio::spawn` in `Drop` finds a runtime. (2) Every stage prints a `STAGE ... start` line and a `RESULT` line with a UTC timestamp, and the setup line splits client creation from the two key setups (on the Mac CPU the client creation alone is 30 to 50 s of a 40 to 60 s setup; bench-log 4 October 2026, shards). The next PC run (PROVE-SHARD.bat) shows whether the gap sits before the first compressed stage and how long it is on the card. Was: Open, found by the team (4 October 2026, morning, first RTX 5090 run through WSL2, SP1 6.8.1). Answer: Two separate things, neither in the proof. (1) After every proof was written, verified and uploaded, `sp1-cuda`'s client dropped its session key outside a Tokio runtime and panicked in its destructor (`sp1-cuda-6.8.1/src/pk.rs:63`, `client.rs:221`), so the host exited 134 with the results already on disk. Fix in our host: hold a runtime for the client's lifetime or drop the proof system inside one. (2) Between the core proof (08:49:10 UTC) and "Proving with mode: Compressed" (08:59:02 UTC) the host was silent for ten minutes while the card was idle; the compressed mode's recursion setup on first use is the suspect, and the second run must time it. Both go on the proving e2e benchmark standard as fixed overheads to measure, not hide.