15 KiB
The prover floor: why a 12 GB card cannot prove on SP1 6.8.1's GPU server, and the patch
5 October 2026, 22:00 UTC on (the project lead: "execute if it will solve the issue"). Branch prover-floor
(worktree igneum-wt-prover-floor). The measured facts this starts from: docs/plans/proving-v1.md and the
bench-log entry "proving v1" (branch proving-v1): the GPU server holds 13.9 GB for an empty shard, 20.4 GB at the
adopted v1 shard, 28.3 GB flat from 20 M to 60 M cycles, and no environment knob moved the floor. Every figure
below is from the source at tag v6.8.1 (cloned to vendor/sp1-6.8.1, gitignored; the fork is the patch
proving/prover-floor/sp1-gpu-6.8.1-floor.patch) or from a PC 2 run named in docs/bench-log.md
("prover floor"). Sizes in GiB are computed from the source constants (4-byte field elements); sizes in MiB are
measured by nvidia-smi at 1 s.
Where the server is built and what it reads
The SDK downloads sp1_gpu_server_v6.8.1_x86_64.tar.gz (133,750,780 bytes) from the SP1 release and runs it from
$HOME/.sp1/bin/sp1-gpu-server (crates/cuda/src/server.rs 19 to 30, 80 to 99). The source is in the same
repository: sp1-gpu/crates/server (the binary), built by .github/workflows/release.yml 234 to 314 on CUDA
12.8.1 with Go and protoc (cargo build --release --bin sp1-gpu-server). The binary takes no options
(sp1-gpu/crates/server/src/main.rs 15 to 18: --version only) and reads CUDA_VISIBLE_DEVICES (32 to 35);
everything else comes from the environment the host process passes it, through SP1CoreOpts::default()
(crates/core/executor/src/opts.rs 99 to 140: SHARD_SIZE, ELEMENT_THRESHOLD, HEIGHT_THRESHOLD,
MINIMAL_TRACE_CHUNK_THRESHOLD, TRACE_CHUNK_SLOTS, FULL_SIZE_SHARDS) and the worker counts
(crates/prover/src/worker/config.rs).
The memory model, term by term
Every device buffer is sized at construction from constants, not from the shard. The server builds the prover at
the first Setup request (sp1-gpu/crates/server/src/server.rs 126 to 137) through
cuda_worker_builder_with_machine (sp1-gpu/crates/prover_components/src/builder.rs 102 to 148):
| Term | Where | Size | On the device | Moves with the shard |
|---|---|---|---|---|
| The gate | builder.rs 35 to 39: gpu_memory_gb = ceil(total / GiB) + 4; panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB") when under 24 |
a 12 GB card reads 16, a 16 GB card 20: both refused before any allocation | no | |
| The core element threshold | builder.rs 41 to 48: ELEMENT_THRESHOLD = 2^28 + 2^27 = 402,653,184 elements (opts.rs 12) on a card reading over 30 (a 32 GB card reads 36); minus 2^26 + 2^25 + 2^24 = 285,212,672 on a card reading 24 to 30 (a 24 GB card). The environment's ELEMENT_THRESHOLD is read at opts.rs 129 and then OVERWRITTEN at builder.rs 48, which is why the sweep's ELEMENT_THRESHOLD rows changed nothing; HEIGHT_THRESHOLD survives (it is not overwritten), which is why the 2^25 + 2^20 row did |
sets the next two terms | no | |
| The core trace area, one per shard in flight | builder.rs 70 to 71: num_elts = element_threshold + 2^21 (CORE_LOG_STACKING_HEIGHT 21, crates/prover/src/components.rs 16) = 404,750,336; allocated on the device at sp1-gpu/crates/jagged_tracegen/src/lib.rs 484 to 500 (allocate_and_initialize_traces: max_trace_size felts + max_trace_size / 2 u32 column index + 2^14 u32) |
6 bytes an element: 2.26 GiB for the full threshold, 1.59 GiB for the 24 GB threshold | yes, in full, whatever the shard holds | no |
| The program's preprocessed traces (the proving key) | sp1-gpu/crates/shard_prover/src/setup.rs 47 to 58 and 107: the same allocate_and_initialize_traces(max_trace_size) at Setup, kept in the key cache for the connection's life (server.rs 139 to 143) |
another 2.26 GiB, held from Setup on |
yes | no |
| The pinned host trace buffers | sp1-gpu/crates/prover_components/src/components.rs 99 to 103: 4 PinnedBuffer of max_trace_size felts per prover (core 4 x 1.51 GiB, recursion 4 x 0.5 GiB, shrink 4 x 0.125 GiB, wrap 4 x 0.32 GiB) |
9.8 GiB of pinned host RAM, not device memory (the WSL2 working set the bench saw) | no | no |
| The recursion trace area | builder.rs 15 and 95: RECURSION_TRACE_ALLOCATION = 2^27 elements, one per recursion tracegen (the recursion program's key at setup and its shard at prove) |
0.75 GiB each | yes | no |
| The shrink and wrap provers | builder.rs 16, 19, 117 to 127: 2^25 and 85,376,340 elements, built at Setup for every proof mode, used only by the Groth16 and PLONK path |
host pinned at build; device only when a wrap runs (never, for a compressed proof) | no | no |
| The codewords (LDE) and the Merkle trees | sp1-gpu/crates/basefold/src/fri.rs 92 to 97: every stacked column of 2^21 rows encoded to 2^(21 + 1) rows (log_blowup 1); kept until the query phase unless drop_ldes (builder.rs 52: only on a 24 GB card with FULL_SIZE_SHARDS); the preprocessed codewords live in the key |
2 x the padded trace, so up to 2 x the term above | yes | yes, with the padded trace |
| The LogUp GKR layers | builder.rs 51: recompute_gkr_trace = false, so the first layer stays materialised (sp1-gpu/crates/logup_gkr/src/tracegen.rs 169 to 224) |
of the order of the interaction count | yes | yes |
| The allocator | sp1-gpu/crates/cuda/src/task.rs 152 and 196: the device's default cudaMallocAsync pool with its release threshold at u64::MAX, so nothing freed is ever returned to the driver: nvidia-smi reads the high-water mark of everything live at once |
So at zero cycles the server already holds the proving key's 2.26 GiB, the shard's 2.26 GiB (both allocated at the threshold, not at the shard's rows), their codewords and trees, and the recursion program's key and traces (2 x 0.75 GiB and their codewords): the 13.9 GB floor. The shard's own content only adds to the codewords, the GKR layers and the working buffers, which is the 13.9 to 20.4 GB step from 0.3 M to 4.7 M cycles, and the flat 28.3 GB from 20 M cycles is the threshold's padded area reached. The witness (5 to 22 KB) never appears.
What the patch does (proving/prover-floor/sp1-gpu-6.8.1-floor.patch, three files)
builder.rs: the panic is gone; the card's memory (orSP1_GPU_MEMORY_BUDGET_GB) picks the element threshold from a tier table (element_threshold_for_budget: over 30 as read, the full 402.6 M; 24 to 30, upstream's 24 GB figure; 18 to 24 (a 16 GB card), 2^27 + 2^26 = 201.3 M; under 18 (a 12 GB card), 2^27 = 134.2 M);SP1_GPU_ELEMENT_THRESHOLDsets it directly andSP1_GPU_RECURSION_TRACE_ALLOCATIONthe recursion buffer. The chosen numbers are printed as aFLOOR optsline. Every other option is as upstream.jagged_tracegen/src/lib.rs: withSP1_GPU_FLOOR_LOGset, every trace allocation prints its capacity and, after the shard's traces are in, the elements actually used and the device memory in use.server.rs: aFLOOR memoryline (device used, free, total) afterSetupand after every proof, with the proof's time.
Nothing in the proof changes: the element threshold only moves where the executor splits shards, exactly what upstream's own 24 GB tier does with the same verifier and the same keys; the recursion program, the verifying key and the pinned guest ids are untouched. The unpatched verifier (the pv1 host's SDK) is the one that verifies every measured proof below.
The build (PC 2, WSL2 Ubuntu-24.04, job floor-toolchain-1 then the build job)
Toolchain found 22:10Z (job floor-toolchain-1, 4 s): nvcc 12.8 at /usr/local/cuda-12.8, cmake 3.28.3, gcc 13.3,
clang 18, protoc 3.21.12, cargo 1.99.0, no Go. The release workflow installs Go for the server's native-gnark
feature (the Groth16 and PLONK wrap through gnark), which a compressed proof never runs, so the build drops that
feature from sp1-gpu/crates/server/Cargo.toml and nothing else. CUDA_ARCHS=86,89,120 (consequences reviewer
C26): the 12 GB tier is sm_86 (RTX 3060) and sm_89 (RTX 4070), the 16 GB tier sm_89 and sm_120 (RTX 5080), PC 2's
5090 is sm_120; the stock server lists sm_80, 86, 89, 90, 100 and 120, which a shipped build repeats. The recipe:
tools/prover-floor/pc2-build-server.ps1 (generated by make-build-playbook.sh from the patch, so the two cannot
drift): clone the tag, git apply the patch, touch the three files, cargo build --release --bin sp1-gpu-server
niced with 8 jobs into /opt/igneum-floor/target, the binary copied to /opt/igneum-floor/home/.sp1/bin/ (the SDK
spawns the server it finds under $HOME/.sp1/bin, so HOME=/opt/igneum-floor/home selects it and the live
/root/.sp1/bin/sp1-gpu-server stays as it is).
What a measurement on PC 2 can and cannot say (C26). The server's allocation pattern is deterministic in the
budget it is given, so a run with SP1_GPU_MEMORY_BUDGET_GB=12 on the 5090 shows the peak a 12 GB card's build
would ask for; it does not show that a 3060 proves it in time, nor what the card's display and driver hold. The
public line keeps "24 GB" until the on-order 12 GB card runs the same fixture. Every row names the arch list and
the card.
What shipping it costs (C26). A patched server means the project signs and distributes its own build of SP1's
prover: the WSL2 package, the DMG's prover inputs, the K1-signed inputs and evidence.md carry it, and every SP1
upgrade repeats the clone, patch, build and measurement. The verifying key and the pinned guest ids do not move
(the patch changes buffer sizes and the shard split, not the circuits), which the verify-segment and --mode compressed VERIFIED lines of the unpatched host show on every row below. The packaging path is a row for the
proving plan before 0.3.12, not this branch.
Step 4 contingency, read not measured: RISC Zero's CUDA prover and its memory per segment
If SP1 could not be brought under 11 GB, the alternative's floor is read from its operators' documentation (not
measured here; a PC 2 run would be the measurement): Boundless' prover guide
(https://docs.boundless.network/provers/performance-optimization) sets the segment size cap by VRAM as 8 GB:
po2 19, 16 GB: po2 20, 20 GB: po2 21, 40 GB: po2 22, with measured peaks po2 20: 13,835 MiB, po2 21: 22,905 MiB,
po2 22: 41,089 MiB; RISC Zero's PR 3761 adds low_vram and pinned_witgen to fit po2 22 on a 24 GB 4090. So
RISC Zero proves a 2^19-cycle segment inside 8 GB and a 2^20 one inside 16 GB, and a shard of 4.7 M cycles is
9 segments at po2 19 plus lift and join steps (times not on the page). Adopting it would cost a second guest (the
chain rule in the RISC Zero zkVM), a second pinned program id, a second verifier in the node and no shared
aggregation between the two formats: docs/analysis/amd-proving.md and the proving plan carry that row already.
The build, as it ran (job floor-build-3, 22:28:24 to 22:32:29Z)
Three runs: floor-build-1 (22:17Z) and floor-build-2 (22:24Z) failed in 2 to 4 minutes on
crates/recursion/gnark-ffi/build.rs:70, "Failed to build Go library: NotFound" (no go on PC 2; the first run's
playbook lost its own log, a bug fixed before the second). floor-build-3 fetched go1.27.1 (tarball sha256
63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445, checked before unpacking under
/opt/igneum-floor/go, on the job's PATH only) and built in 240 s (46 crates on the warm target of run 2, 8
niced jobs, 16 cores). The binary: /opt/igneum-floor/bin/sp1-gpu-server, 166,768,224 bytes, sha256
5568108bf7fb9b0e525d8a08926b7046e51136ffaea53f0ca858631d0e938878, --version 6.8.1, cuobjdump --list-elf
sm_86, sm_89, sm_120 (the stock 251,306,680-byte server lists sm_80, 86, 89, 90, 100, 120 and compute_120 PTX).
The live /root/.sp1/bin/sp1-gpu-server (c2642ad1...) was never touched; the miners mined throughout.
Sweep 1 (job floor-sweep-1, 22:34:56 to 22:37:55Z): the shard term gone, a second floor found
PC 2's RTX 5090 (32,607 MiB, idle 1,755 MiB with the miners stopped and the live prover off), the patched server
5568108b... (sm_86, sm_89, sm_120; the build above), the unpatched pv1 host dae6b006... as client and
verifier, one --mode compressed --shard 0 per point, every server killed and its socket unlinked around every
point, peak = nvidia-smi memory.used at 1 s (the idle 1,755 MiB inside it), time = the compressed proof.
Every proof VERIFIED (1,272,897 bytes, verify 0.037 to 0.040 s), so the unpatched verifier accepts every proof
of the patched server.
| Config (environment to the patched server) | Fixture | Cycles | Peak MiB | Prove s | Verified |
|---|---|---|---|---|---|
control: SP1_GPU_MEMORY_BUDGET_GB=32 (upstream's sizes) |
empty live shard (block 83616) | 280,706 | 13,892 | 2.2 | yes |
| control | v1 shard (fees-v1-shards2 shard 0) | 4,717,439 | 20,516 | 4.2 | yes |
| 12 GB tier: budget 12 (threshold 2^27) | empty | 280,706 | 12,740 | 2.4 | yes |
| 12 GB tier | v1 shard | 4.7 M | 15,396 | 4.1 | yes |
12 GB tier + SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1 |
v1 shard | 4.7 M | 15,428 | 4.0 | yes |
| 16 GB tier: budget 16 (2^27 + 2^26) | v1 shard | 4.7 M | 18,628 | 3.8 | yes |
SP1_GPU_ELEMENT_THRESHOLD=67108864 (2^26) |
empty | 280,706 | 12,772 | 3.1 | yes |
| 2^26 | v1 shard (split into 4 core shards) | 4.7 M | 12,708 | 5.3 | yes |
SP1_GPU_ELEMENT_THRESHOLD=33554432 (2^25) |
v1 shard | 4.7 M | 12,836 | 8.5 | yes |
Reading. The control reproduces the proving agent's curve (13.9 and 20.4 GB), so the patched server behaves as
the stock one at the stock sizes. The shard's term follows the threshold as the model says (20.5 GB at 402 M
elements, 15.4 at 134 M, 12.7 at 67 M), and then stops: 2^26 and 2^25 both sit at 12.7 to 12.8 GB for the empty
shard and the v1 shard alike. The FLOOR memory after setup line names the rest: 9,703 MiB in use before the
first shard (2^26; 11,623 at the stock sizes), and the FLOOR tracegen alloc lines at Setup are five
allocations of 134,217,728 elements (the recursion keys, 0.75 GB each, each using 90,177,536 elements: 35.6 M
preprocessed and 54.5 M main at prove time), one of 33,554,432 (the shrink key, 0.19 GB) and one core key at the
threshold. The NORMALIZE_PROGRAM_CACHE_SIZE knob does not reach them (they are keys built at Setup, not the
program LRU). The time cost of the split: the v1 shard at 2^26 is 4 core shards and 5.3 s against 4.2 s (1.26x);
at 2^27 it is 4.1 s with no split.
So after sweep 1 the binding term is the Setup-time keys allocated at full capacity, and patch v2 sizes every
trace buffer (keys and shards) to its padded need: padded_trace_elements in jagged_tracegen/src/lib.rs
(each phase pads to the next multiple of 2^21 rows, generate_jagged_traces's "final padding"), applied in
setup_tracegen and full_tracegen, one stacking height of slack, SP1_GPU_FLOOR_EXACT=0 restoring upstream.