From 73358ca92d0c1bb53bb892c3d525180f8d5ccd41 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 20:27:45 +0000 Subject: [PATCH] Proving v1: the 12 GB memory sweep on the 5090 (13.9 GB floor on an empty shard, 28.3 GB on a full one, no knob moves the floor) Co-Authored-By: Claude Fable 5.1 --- docs/bench-log.md | 20 ++++++++++++++++++++ 1 file changed, 20 insertions(+) diff --git a/docs/bench-log.md b/docs/bench-log.md index 4fced0354..d3de8413d 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1642,6 +1642,26 @@ Inputs, all RTX 5090 (PC 2), SP1 6.8.1 cuda: a full shard at the provisional `S_ Reading. The card count is the sum of card-seconds of work per block-second, rounded up, with no slack for the exclusive window, the relay or a card's idle gaps; the devnet's own numbers tonight (one card, 1.4 to 1.6 shards a minute, 2.4 to 4.7% of blocks) are the first row. Two levers, both measured tonight: the loop (a shard's carriage through export, cut and a 12-s key setup is 25 s on top of a 7-s proof; the host's `--mode aggregate` and `--mode chain` hold one key setup per process and the prover loop should do the same, the 0.3.11 item in the plan) and the card's other job (a mining card proves 3 to 4x slower than an idle one, `chain-pc2-pv1c` against 4 October; the prover's cost to mining is 4%). A fleet of 18 mining 5090s, or 6 proving-only ones, covers an empty-block chain at 1 block/s through the chain mode; the mandatory rule waits for the measured share to reach one, not for these rows. +### The 12 GB requirement (the project lead, 20:1xZ: "make sure we can prove on 12gb cards"): the GPU memory peak against SP1's knobs + +Job `memsweep-pc2-pv1` (`tools/proving-v1/pc2-memory-sweep.ps1`), PC 2's RTX 5090 (32,607 MiB), the miners STOPPED by the job and the live prover switched off (its `sp1-gpu-server` would otherwise be the one the client connects to), every row: the server killed first, a 1-s `nvidia-smi memory.used` sampler, one `--mode compressed --shard 0` run of the pv1 host (`/opt/igneum-pv1`, SP1 6.8.1 cuda, `sp1-gpu-server` 6.8.1), 20:19 to 20:25Z. The knobs are the environment the GPU server inherits from the host process (`sp1-core-executor-6.8.1/src/opts.rs`: `SHARD_SIZE`, `ELEMENT_THRESHOLD`, `HEIGHT_THRESHOLD`, `MINIMAL_TRACE_CHUNK_THRESHOLD`, `TRACE_CHUNK_SLOTS`; `sp1-prover-6.8.1/src/worker/config.rs`: the `SP1_WORKER_NUM_*` and `*_BUFFER_SIZE` counts, defaults 4 core workers, 8 recursion prover workers). Idle card before the sweep: 1,732 MiB. + +| Config (environment) | Fixture | Cycles | Peak MiB | Compressed prove s | Verified | +|---|---|---|---|---|---| +| baseline (no knob) | block-338-shard1, a full shard at `S_p` (6.75 M pgas) | 60,415,376 | **28,295** | 11.4 | yes | +| baseline | block-83616, an empty live shard | 280,706 | **13,863** | 2.3 | yes | +| ELEMENT_THRESHOLD 2^27 | full shard | 60.4 M | 28,326 | 10.9 | yes | +| ELEMENT_THRESHOLD 2^26, HEIGHT_THRESHOLD 2^21 | full shard | 60.4 M | 28,326 | 10.7 | yes | +| every worker count and buffer 1 | full shard | 60.4 M | 28,326 | 20.8 | yes | +| every worker count and buffer 2 | full shard | 60.4 M | 28,327 | 12.9 | yes | +| workers 1 + ELEMENT 2^27 | full shard | 60.4 M | 28,263 | 20.3 | yes | +| workers 1 + ELEMENT 2^26 + HEIGHT 2^21 | full shard | 60.4 M | 28,326 | 20.6 | yes | +| workers 1 + ELEMENT 2^26 + HEIGHT 2^21 + trace chunks 4 M x 2 slots | full shard | 60.4 M | 28,358 | 22.6 | yes | +| workers 1 + ELEMENT 2^25 + HEIGHT 2^20 | full shard | 60.4 M | 22,919 | 22.2 | yes | +| workers 1 + ELEMENT 2^26 + HEIGHT 2^21 | empty shard | 280,706 | 13,861 | 2.6 | yes | + +Reading. The GPU memory of a compressed shard proof is **13.9 GB for a shard of 280,000 cycles and 28.3 GB for one of 60 M cycles**, and no knob the environment carries moves the floor: the worker counts only slow the proof (11.4 s to 20.8 s), the trace thresholds at 2^26 and 2^27 change nothing, and the smallest trace threshold tried (2^25 elements, 2^20 rows) takes 5.4 GB off the full shard (22.9 GB) at twice the time. The floor sits in the GPU server's own allocation, not in the shard: an empty shard with every knob at its minimum still takes 13.9 GB. So on SP1 6.8.1's `sp1-gpu-server` as shipped, **a 12 GB card cannot prove even an empty shard** (13.9 GB), and the 11.0 GB target of tonight's requirement is out of reach from the environment. The S_p/2 and S_p/4 cuts of block 344 did not run: the package carries no `tools/prove-fixtures/seq.json` (the cut rows need the export; they would sit between the two measured points, and the floor is the binding number anyway). What is left to try, in order: the server's own options (its `--help` and the option names in its strings: the miner-on job prints them), SP1's core-only proof (the node needs the compressed proof, so this changes the protocol), and an SP1 release built for smaller cards (the 6.8.1 release notes are not read here; approximate: the project's documentation names 24 GB as the GPU requirement, `proving/windows-wsl2/setup-wsl.sh` quotes it). + ### Step 4, the rule | What | Measured |