152 lines
15 KiB
Markdown
152 lines
15 KiB
Markdown
# The prover floor: why a 12 GB card cannot prove on SP1 6.8.1's GPU server, and the patch
|
|
|
|
5 October 2026, 22:00 UTC on (the project lead: "execute if it will solve the issue"). Branch `prover-floor`
|
|
(worktree `igneum-wt-prover-floor`). The measured facts this starts from: `docs/plans/proving-v1.md` and the
|
|
bench-log entry "proving v1" (branch proving-v1): the GPU server holds 13.9 GB for an empty shard, 20.4 GB at the
|
|
adopted v1 shard, 28.3 GB flat from 20 M to 60 M cycles, and no environment knob moved the floor. Every figure
|
|
below is from the source at tag v6.8.1 (cloned to `vendor/sp1-6.8.1`, gitignored; the fork is the patch
|
|
`proving/prover-floor/sp1-gpu-6.8.1-floor.patch`) or from a PC 2 run named in `docs/bench-log.md`
|
|
("prover floor"). Sizes in GiB are computed from the source constants (4-byte field elements); sizes in MiB are
|
|
measured by `nvidia-smi` at 1 s.
|
|
|
|
## Where the server is built and what it reads
|
|
|
|
The SDK downloads `sp1_gpu_server_v6.8.1_x86_64.tar.gz` (133,750,780 bytes) from the SP1 release and runs it from
|
|
`$HOME/.sp1/bin/sp1-gpu-server` (`crates/cuda/src/server.rs` 19 to 30, 80 to 99). The source is in the same
|
|
repository: `sp1-gpu/crates/server` (the binary), built by `.github/workflows/release.yml` 234 to 314 on CUDA
|
|
12.8.1 with Go and protoc (`cargo build --release --bin sp1-gpu-server`). The binary takes no options
|
|
(`sp1-gpu/crates/server/src/main.rs` 15 to 18: `--version` only) and reads `CUDA_VISIBLE_DEVICES` (32 to 35);
|
|
everything else comes from the environment the host process passes it, through `SP1CoreOpts::default()`
|
|
(`crates/core/executor/src/opts.rs` 99 to 140: `SHARD_SIZE`, `ELEMENT_THRESHOLD`, `HEIGHT_THRESHOLD`,
|
|
`MINIMAL_TRACE_CHUNK_THRESHOLD`, `TRACE_CHUNK_SLOTS`, `FULL_SIZE_SHARDS`) and the worker counts
|
|
(`crates/prover/src/worker/config.rs`).
|
|
|
|
## The memory model, term by term
|
|
|
|
Every device buffer is sized at construction from constants, not from the shard. The server builds the prover at
|
|
the first `Setup` request (`sp1-gpu/crates/server/src/server.rs` 126 to 137) through
|
|
`cuda_worker_builder_with_machine` (`sp1-gpu/crates/prover_components/src/builder.rs` 102 to 148):
|
|
|
|
| Term | Where | Size | On the device | Moves with the shard |
|
|
|---|---|---|---|---|
|
|
| The gate | `builder.rs` 35 to 39: `gpu_memory_gb = ceil(total / GiB) + 4`; `panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB")` when under 24 | a 12 GB card reads 16, a 16 GB card 20: both refused before any allocation | | no |
|
|
| The core element threshold | `builder.rs` 41 to 48: `ELEMENT_THRESHOLD` = 2^28 + 2^27 = 402,653,184 elements (`opts.rs` 12) on a card reading over 30 (a 32 GB card reads 36); minus 2^26 + 2^25 + 2^24 = 285,212,672 on a card reading 24 to 30 (a 24 GB card). The environment's `ELEMENT_THRESHOLD` is read at `opts.rs` 129 and then OVERWRITTEN at `builder.rs` 48, which is why the sweep's `ELEMENT_THRESHOLD` rows changed nothing; `HEIGHT_THRESHOLD` survives (it is not overwritten), which is why the 2^25 + 2^20 row did | sets the next two terms | | no |
|
|
| The core trace area, one per shard in flight | `builder.rs` 70 to 71: `num_elts = element_threshold + 2^21` (`CORE_LOG_STACKING_HEIGHT` 21, `crates/prover/src/components.rs` 16) = 404,750,336; allocated on the device at `sp1-gpu/crates/jagged_tracegen/src/lib.rs` 484 to 500 (`allocate_and_initialize_traces`: `max_trace_size` felts + `max_trace_size / 2` u32 column index + 2^14 u32) | 6 bytes an element: **2.26 GiB** for the full threshold, 1.59 GiB for the 24 GB threshold | yes, in full, whatever the shard holds | no |
|
|
| The program's preprocessed traces (the proving key) | `sp1-gpu/crates/shard_prover/src/setup.rs` 47 to 58 and 107: the same `allocate_and_initialize_traces(max_trace_size)` at `Setup`, kept in the key cache for the connection's life (`server.rs` 139 to 143) | another **2.26 GiB**, held from `Setup` on | yes | no |
|
|
| The pinned host trace buffers | `sp1-gpu/crates/prover_components/src/components.rs` 99 to 103: 4 `PinnedBuffer` of `max_trace_size` felts per prover (core 4 x 1.51 GiB, recursion 4 x 0.5 GiB, shrink 4 x 0.125 GiB, wrap 4 x 0.32 GiB) | 9.8 GiB of pinned host RAM, not device memory (the WSL2 working set the bench saw) | no | no |
|
|
| The recursion trace area | `builder.rs` 15 and 95: `RECURSION_TRACE_ALLOCATION` = 2^27 elements, one per recursion tracegen (the recursion program's key at setup and its shard at prove) | 0.75 GiB each | yes | no |
|
|
| The shrink and wrap provers | `builder.rs` 16, 19, 117 to 127: 2^25 and 85,376,340 elements, built at `Setup` for every proof mode, used only by the Groth16 and PLONK path | host pinned at build; device only when a wrap runs (never, for a compressed proof) | no | no |
|
|
| The codewords (LDE) and the Merkle trees | `sp1-gpu/crates/basefold/src/fri.rs` 92 to 97: every stacked column of 2^21 rows encoded to 2^(21 + 1) rows (`log_blowup` 1); kept until the query phase unless `drop_ldes` (`builder.rs` 52: only on a 24 GB card with `FULL_SIZE_SHARDS`); the preprocessed codewords live in the key | 2 x the padded trace, so up to 2 x the term above | yes | yes, with the padded trace |
|
|
| The LogUp GKR layers | `builder.rs` 51: `recompute_gkr_trace = false`, so the first layer stays materialised (`sp1-gpu/crates/logup_gkr/src/tracegen.rs` 169 to 224) | of the order of the interaction count | yes | yes |
|
|
| The allocator | `sp1-gpu/crates/cuda/src/task.rs` 152 and 196: the device's default `cudaMallocAsync` pool with its release threshold at `u64::MAX`, so nothing freed is ever returned to the driver: `nvidia-smi` reads the high-water mark of everything live at once | | | |
|
|
|
|
So at zero cycles the server already holds the proving key's 2.26 GiB, the shard's 2.26 GiB (both allocated at the
|
|
threshold, not at the shard's rows), their codewords and trees, and the recursion program's key and traces (2 x
|
|
0.75 GiB and their codewords): the 13.9 GB floor. The shard's own content only adds to the codewords, the GKR
|
|
layers and the working buffers, which is the 13.9 to 20.4 GB step from 0.3 M to 4.7 M cycles, and the flat 28.3 GB
|
|
from 20 M cycles is the threshold's padded area reached. The witness (5 to 22 KB) never appears.
|
|
|
|
## What the patch does (`proving/prover-floor/sp1-gpu-6.8.1-floor.patch`, three files)
|
|
|
|
1. `builder.rs`: the panic is gone; the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks the element threshold
|
|
from a tier table (`element_threshold_for_budget`: over 30 as read, the full 402.6 M; 24 to 30, upstream's 24 GB
|
|
figure; 18 to 24 (a 16 GB card), 2^27 + 2^26 = 201.3 M; under 18 (a 12 GB card), 2^27 = 134.2 M);
|
|
`SP1_GPU_ELEMENT_THRESHOLD` sets it directly and `SP1_GPU_RECURSION_TRACE_ALLOCATION` the recursion buffer.
|
|
The chosen numbers are printed as a `FLOOR opts` line. Every other option is as upstream.
|
|
2. `jagged_tracegen/src/lib.rs`: with `SP1_GPU_FLOOR_LOG` set, every trace allocation prints its capacity and,
|
|
after the shard's traces are in, the elements actually used and the device memory in use.
|
|
3. `server.rs`: a `FLOOR memory` line (device used, free, total) after `Setup` and after every proof, with the
|
|
proof's time.
|
|
|
|
Nothing in the proof changes: the element threshold only moves where the executor splits shards, exactly what
|
|
upstream's own 24 GB tier does with the same verifier and the same keys; the recursion program, the verifying key
|
|
and the pinned guest ids are untouched. The unpatched verifier (the pv1 host's SDK) is the one that verifies every
|
|
measured proof below.
|
|
|
|
## The build (PC 2, WSL2 Ubuntu-24.04, job `floor-toolchain-1` then the build job)
|
|
|
|
Toolchain found 22:10Z (job `floor-toolchain-1`, 4 s): nvcc 12.8 at `/usr/local/cuda-12.8`, cmake 3.28.3, gcc 13.3,
|
|
clang 18, protoc 3.21.12, cargo 1.99.0, no Go. The release workflow installs Go for the server's `native-gnark`
|
|
feature (the Groth16 and PLONK wrap through gnark), which a compressed proof never runs, so the build drops that
|
|
feature from `sp1-gpu/crates/server/Cargo.toml` and nothing else. `CUDA_ARCHS=86,89,120` (consequences reviewer
|
|
C26): the 12 GB tier is sm_86 (RTX 3060) and sm_89 (RTX 4070), the 16 GB tier sm_89 and sm_120 (RTX 5080), PC 2's
|
|
5090 is sm_120; the stock server lists sm_80, 86, 89, 90, 100 and 120, which a shipped build repeats. The recipe:
|
|
`tools/prover-floor/pc2-build-server.ps1` (generated by `make-build-playbook.sh` from the patch, so the two cannot
|
|
drift): clone the tag, `git apply` the patch, `touch` the three files, `cargo build --release --bin sp1-gpu-server`
|
|
niced with 8 jobs into `/opt/igneum-floor/target`, the binary copied to `/opt/igneum-floor/home/.sp1/bin/` (the SDK
|
|
spawns the server it finds under `$HOME/.sp1/bin`, so `HOME=/opt/igneum-floor/home` selects it and the live
|
|
`/root/.sp1/bin/sp1-gpu-server` stays as it is).
|
|
|
|
What a measurement on PC 2 can and cannot say (C26). The server's allocation pattern is deterministic in the
|
|
budget it is given, so a run with `SP1_GPU_MEMORY_BUDGET_GB=12` on the 5090 shows the peak a 12 GB card's build
|
|
would ask for; it does not show that a 3060 proves it in time, nor what the card's display and driver hold. The
|
|
public line keeps "24 GB" until the on-order 12 GB card runs the same fixture. Every row names the arch list and
|
|
the card.
|
|
|
|
What shipping it costs (C26). A patched server means the project signs and distributes its own build of SP1's
|
|
prover: the WSL2 package, the DMG's prover inputs, the K1-signed inputs and `evidence.md` carry it, and every SP1
|
|
upgrade repeats the clone, patch, build and measurement. The verifying key and the pinned guest ids do not move
|
|
(the patch changes buffer sizes and the shard split, not the circuits), which the `verify-segment` and `--mode
|
|
compressed` VERIFIED lines of the unpatched host show on every row below. The packaging path is a row for the
|
|
proving plan before 0.3.12, not this branch.
|
|
|
|
## Step 4 contingency, read not measured: RISC Zero's CUDA prover and its memory per segment
|
|
|
|
If SP1 could not be brought under 11 GB, the alternative's floor is read from its operators' documentation (not
|
|
measured here; a PC 2 run would be the measurement): Boundless' prover guide
|
|
(https://docs.boundless.network/provers/performance-optimization) sets the segment size cap by VRAM as 8 GB:
|
|
po2 19, 16 GB: po2 20, 20 GB: po2 21, 40 GB: po2 22, with measured peaks po2 20: 13,835 MiB, po2 21: 22,905 MiB,
|
|
po2 22: 41,089 MiB; RISC Zero's PR 3761 adds `low_vram` and `pinned_witgen` to fit po2 22 on a 24 GB 4090. So
|
|
RISC Zero proves a 2^19-cycle segment inside 8 GB and a 2^20 one inside 16 GB, and a shard of 4.7 M cycles is
|
|
9 segments at po2 19 plus lift and join steps (times not on the page). Adopting it would cost a second guest (the
|
|
chain rule in the RISC Zero zkVM), a second pinned program id, a second verifier in the node and no shared
|
|
aggregation between the two formats: `docs/analysis/amd-proving.md` and the proving plan carry that row already.
|
|
|
|
### The build, as it ran (job `floor-build-3`, 22:28:24 to 22:32:29Z)
|
|
|
|
Three runs: `floor-build-1` (22:17Z) and `floor-build-2` (22:24Z) failed in 2 to 4 minutes on
|
|
`crates/recursion/gnark-ffi/build.rs:70`, "Failed to build Go library: NotFound" (no `go` on PC 2; the first run's
|
|
playbook lost its own log, a bug fixed before the second). `floor-build-3` fetched go1.27.1 (tarball sha256
|
|
`63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445`, checked before unpacking under
|
|
`/opt/igneum-floor/go`, on the job's PATH only) and built in **240 s** (46 crates on the warm target of run 2, 8
|
|
niced jobs, 16 cores). The binary: `/opt/igneum-floor/bin/sp1-gpu-server`, **166,768,224 bytes, sha256
|
|
`5568108bf7fb9b0e525d8a08926b7046e51136ffaea53f0ca858631d0e938878`**, `--version` 6.8.1, `cuobjdump --list-elf`
|
|
sm_86, sm_89, sm_120 (the stock 251,306,680-byte server lists sm_80, 86, 89, 90, 100, 120 and compute_120 PTX).
|
|
The live `/root/.sp1/bin/sp1-gpu-server` (c2642ad1...) was never touched; the miners mined throughout.
|
|
|
|
## Sweep 1 (job `floor-sweep-1`, 22:34:56 to 22:37:55Z): the shard term gone, a second floor found
|
|
|
|
PC 2's RTX 5090 (32,607 MiB, idle 1,755 MiB with the miners stopped and the live prover off), the patched server
|
|
`5568108b...` (sm_86, sm_89, sm_120; the build above), the unpatched pv1 host `dae6b006...` as client and
|
|
verifier, one `--mode compressed --shard 0` per point, every server killed and its socket unlinked around every
|
|
point, peak = `nvidia-smi memory.used` at 1 s (the idle 1,755 MiB inside it), time = the compressed proof.
|
|
Every proof VERIFIED (1,272,897 bytes, verify 0.037 to 0.040 s), so the unpatched verifier accepts every proof
|
|
of the patched server.
|
|
|
|
| Config (environment to the patched server) | Fixture | Cycles | Peak MiB | Prove s | Verified |
|
|
|---|---|---|---|---|---|
|
|
| control: `SP1_GPU_MEMORY_BUDGET_GB=32` (upstream's sizes) | empty live shard (block 83616) | 280,706 | 13,892 | 2.2 | yes |
|
|
| control | v1 shard (fees-v1-shards2 shard 0) | 4,717,439 | 20,516 | 4.2 | yes |
|
|
| 12 GB tier: budget 12 (threshold 2^27) | empty | 280,706 | 12,740 | 2.4 | yes |
|
|
| 12 GB tier | v1 shard | 4.7 M | 15,396 | 4.1 | yes |
|
|
| 12 GB tier + `SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1` | v1 shard | 4.7 M | 15,428 | 4.0 | yes |
|
|
| 16 GB tier: budget 16 (2^27 + 2^26) | v1 shard | 4.7 M | 18,628 | 3.8 | yes |
|
|
| `SP1_GPU_ELEMENT_THRESHOLD=67108864` (2^26) | empty | 280,706 | 12,772 | 3.1 | yes |
|
|
| 2^26 | v1 shard (split into 4 core shards) | 4.7 M | **12,708** | 5.3 | yes |
|
|
| `SP1_GPU_ELEMENT_THRESHOLD=33554432` (2^25) | v1 shard | 4.7 M | 12,836 | 8.5 | yes |
|
|
|
|
Reading. The control reproduces the proving agent's curve (13.9 and 20.4 GB), so the patched server behaves as
|
|
the stock one at the stock sizes. The shard's term follows the threshold as the model says (20.5 GB at 402 M
|
|
elements, 15.4 at 134 M, 12.7 at 67 M), and then stops: 2^26 and 2^25 both sit at 12.7 to 12.8 GB for the empty
|
|
shard and the v1 shard alike. The `FLOOR memory after setup` line names the rest: **9,703 MiB in use before the
|
|
first shard** (2^26; 11,623 at the stock sizes), and the `FLOOR tracegen alloc` lines at Setup are five
|
|
allocations of 134,217,728 elements (the recursion keys, 0.75 GB each, each using 90,177,536 elements: 35.6 M
|
|
preprocessed and 54.5 M main at prove time), one of 33,554,432 (the shrink key, 0.19 GB) and one core key at the
|
|
threshold. The `NORMALIZE_PROGRAM_CACHE_SIZE` knob does not reach them (they are keys built at `Setup`, not the
|
|
program LRU). The time cost of the split: the v1 shard at 2^26 is 4 core shards and 5.3 s against 4.2 s (1.26x);
|
|
at 2^27 it is 4.1 s with no split.
|
|
|
|
So after sweep 1 the binding term is the Setup-time keys allocated at full capacity, and patch v2 sizes every
|
|
trace buffer (keys and shards) to its padded need: `padded_trace_elements` in `jagged_tracegen/src/lib.rs`
|
|
(each phase pads to the next multiple of 2^21 rows, `generate_jagged_traces`'s "final padding"), applied in
|
|
`setup_tracegen` and `full_tracegen`, one stacking height of slack, `SP1_GPU_FLOOR_EXACT=0` restoring upstream.
|