22 KiB
The prover floor: why a 12 GB card cannot prove on SP1 6.8.1's GPU server, and the patch
5 October 2026, 22:00 UTC on (the project lead: "execute if it will solve the issue"). Branch prover-floor
(worktree igneum-wt-prover-floor). The measured facts this starts from: docs/plans/proving-v1.md and the
bench-log entry "proving v1" (branch proving-v1): the GPU server holds 13.9 GB for an empty shard, 20.4 GB at the
adopted v1 shard, 28.3 GB flat from 20 M to 60 M cycles, and no environment knob moved the floor. Every figure
below is from the source at tag v6.8.1 (cloned to vendor/sp1-6.8.1, gitignored; the fork is the patch
proving/prover-floor/sp1-gpu-6.8.1-floor.patch) or from a PC 2 run named in docs/bench-log.md
("prover floor"). Sizes in GiB are computed from the source constants (4-byte field elements); sizes in MiB are
measured by nvidia-smi at 1 s.
Where the server is built and what it reads
The SDK downloads sp1_gpu_server_v6.8.1_x86_64.tar.gz (133,750,780 bytes) from the SP1 release and runs it from
$HOME/.sp1/bin/sp1-gpu-server (crates/cuda/src/server.rs 19 to 30, 80 to 99). The source is in the same
repository: sp1-gpu/crates/server (the binary), built by .github/workflows/release.yml 234 to 314 on CUDA
12.8.1 with Go and protoc (cargo build --release --bin sp1-gpu-server). The binary takes no options
(sp1-gpu/crates/server/src/main.rs 15 to 18: --version only) and reads CUDA_VISIBLE_DEVICES (32 to 35);
everything else comes from the environment the host process passes it, through SP1CoreOpts::default()
(crates/core/executor/src/opts.rs 99 to 140: SHARD_SIZE, ELEMENT_THRESHOLD, HEIGHT_THRESHOLD,
MINIMAL_TRACE_CHUNK_THRESHOLD, TRACE_CHUNK_SLOTS, FULL_SIZE_SHARDS) and the worker counts
(crates/prover/src/worker/config.rs).
The memory model, term by term
Every device buffer is sized at construction from constants, not from the shard. The server builds the prover at
the first Setup request (sp1-gpu/crates/server/src/server.rs 126 to 137) through
cuda_worker_builder_with_machine (sp1-gpu/crates/prover_components/src/builder.rs 102 to 148):
| Term | Where | Size | On the device | Moves with the shard |
|---|---|---|---|---|
| The gate | builder.rs 35 to 39: gpu_memory_gb = ceil(total / GiB) + 4; panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB") when under 24 |
a 12 GB card reads 16, a 16 GB card 20: both refused before any allocation | no | |
| The core element threshold | builder.rs 41 to 48: ELEMENT_THRESHOLD = 2^28 + 2^27 = 402,653,184 elements (opts.rs 12) on a card reading over 30 (a 32 GB card reads 36); minus 2^26 + 2^25 + 2^24 = 285,212,672 on a card reading 24 to 30 (a 24 GB card). The environment's ELEMENT_THRESHOLD is read at opts.rs 129 and then OVERWRITTEN at builder.rs 48, which is why the sweep's ELEMENT_THRESHOLD rows changed nothing; HEIGHT_THRESHOLD survives (it is not overwritten), which is why the 2^25 + 2^20 row did |
sets the next two terms | no | |
| The core trace area, one per shard in flight | builder.rs 70 to 71: num_elts = element_threshold + 2^21 (CORE_LOG_STACKING_HEIGHT 21, crates/prover/src/components.rs 16) = 404,750,336; allocated on the device at sp1-gpu/crates/jagged_tracegen/src/lib.rs 484 to 500 (allocate_and_initialize_traces: max_trace_size felts + max_trace_size / 2 u32 column index + 2^14 u32) |
6 bytes an element: 2.26 GiB for the full threshold, 1.59 GiB for the 24 GB threshold | yes, in full, whatever the shard holds | no |
| The program's preprocessed traces (the proving key) | sp1-gpu/crates/shard_prover/src/setup.rs 47 to 58 and 107: the same allocate_and_initialize_traces(max_trace_size) at Setup, kept in the key cache for the connection's life (server.rs 139 to 143) |
another 2.26 GiB, held from Setup on |
yes | no |
| The pinned host trace buffers | sp1-gpu/crates/prover_components/src/components.rs 99 to 103: 4 PinnedBuffer of max_trace_size felts per prover (core 4 x 1.51 GiB, recursion 4 x 0.5 GiB, shrink 4 x 0.125 GiB, wrap 4 x 0.32 GiB) |
9.8 GiB of pinned host RAM, not device memory (the WSL2 working set the bench saw) | no | no |
| The recursion trace area | builder.rs 15 and 95: RECURSION_TRACE_ALLOCATION = 2^27 elements, one per recursion tracegen (the recursion program's key at setup and its shard at prove) |
0.75 GiB each | yes | no |
| The shrink and wrap provers | builder.rs 16, 19, 117 to 127: 2^25 and 85,376,340 elements, built at Setup for every proof mode, used only by the Groth16 and PLONK path |
host pinned at build; device only when a wrap runs (never, for a compressed proof) | no | no |
| The codewords (LDE) and the Merkle trees | sp1-gpu/crates/basefold/src/fri.rs 92 to 97: every stacked column of 2^21 rows encoded to 2^(21 + 1) rows (log_blowup 1); kept until the query phase unless drop_ldes (builder.rs 52: only on a 24 GB card with FULL_SIZE_SHARDS); the preprocessed codewords live in the key |
2 x the padded trace, so up to 2 x the term above | yes | yes, with the padded trace |
| The LogUp GKR layers | builder.rs 51: recompute_gkr_trace = false, so the first layer stays materialised (sp1-gpu/crates/logup_gkr/src/tracegen.rs 169 to 224) |
of the order of the interaction count | yes | yes |
| The allocator | sp1-gpu/crates/cuda/src/task.rs 152 and 196: the device's default cudaMallocAsync pool with its release threshold at u64::MAX, so nothing freed is ever returned to the driver: nvidia-smi reads the high-water mark of everything live at once |
So at zero cycles the server already holds the proving key's 2.26 GiB, the shard's 2.26 GiB (both allocated at the threshold, not at the shard's rows), their codewords and trees, and the recursion program's key and traces (2 x 0.75 GiB and their codewords): the 13.9 GB floor. The shard's own content only adds to the codewords, the GKR layers and the working buffers, which is the 13.9 to 20.4 GB step from 0.3 M to 4.7 M cycles, and the flat 28.3 GB from 20 M cycles is the threshold's padded area reached. The witness (5 to 22 KB) never appears.
What the patch does (proving/prover-floor/sp1-gpu-6.8.1-floor.patch, three files)
builder.rs: the panic is gone; the card's memory (orSP1_GPU_MEMORY_BUDGET_GB) picks the element threshold from a tier table (element_threshold_for_budget: over 30 as read, the full 402.6 M; 24 to 30, upstream's 24 GB figure; 18 to 24 (a 16 GB card), 2^27 + 2^26 = 201.3 M; under 18 (a 12 GB card), 2^27 = 134.2 M);SP1_GPU_ELEMENT_THRESHOLDsets it directly andSP1_GPU_RECURSION_TRACE_ALLOCATIONthe recursion buffer. The chosen numbers are printed as aFLOOR optsline. Every other option is as upstream.jagged_tracegen/src/lib.rs: withSP1_GPU_FLOOR_LOGset, every trace allocation prints its capacity and, after the shard's traces are in, the elements actually used and the device memory in use.server.rs: aFLOOR memoryline (device used, free, total) afterSetupand after every proof, with the proof's time.
Nothing in the proof changes: the element threshold only moves where the executor splits shards, exactly what upstream's own 24 GB tier does with the same verifier and the same keys; the recursion program, the verifying key and the pinned guest ids are untouched. The unpatched verifier (the pv1 host's SDK) is the one that verifies every measured proof below.
The build (PC 2, WSL2 Ubuntu-24.04, job floor-toolchain-1 then the build job)
Toolchain found 22:10Z (job floor-toolchain-1, 4 s): nvcc 12.8 at /usr/local/cuda-12.8, cmake 3.28.3, gcc 13.3,
clang 18, protoc 3.21.12, cargo 1.99.0, no Go. The release workflow installs Go for the server's native-gnark
feature (the Groth16 and PLONK wrap through gnark), which a compressed proof never runs, so the build drops that
feature from sp1-gpu/crates/server/Cargo.toml and nothing else. CUDA_ARCHS=86,89,120 (consequences reviewer
C26): the 12 GB tier is sm_86 (RTX 3060) and sm_89 (RTX 4070), the 16 GB tier sm_89 and sm_120 (RTX 5080), PC 2's
5090 is sm_120; the stock server lists sm_80, 86, 89, 90, 100 and 120, which a shipped build repeats. The recipe:
tools/prover-floor/pc2-build-server.ps1 (generated by make-build-playbook.sh from the patch, so the two cannot
drift): clone the tag, git apply the patch, touch the three files, cargo build --release --bin sp1-gpu-server
niced with 8 jobs into /opt/igneum-floor/target, the binary copied to /opt/igneum-floor/home/.sp1/bin/ (the SDK
spawns the server it finds under $HOME/.sp1/bin, so HOME=/opt/igneum-floor/home selects it and the live
/root/.sp1/bin/sp1-gpu-server stays as it is).
What a measurement on PC 2 can and cannot say (C26). The server's allocation pattern is deterministic in the
budget it is given, so a run with SP1_GPU_MEMORY_BUDGET_GB=12 on the 5090 shows the peak a 12 GB card's build
would ask for; it does not show that a 3060 proves it in time, nor what the card's display and driver hold. The
public line keeps "24 GB" until the on-order 12 GB card runs the same fixture. Every row names the arch list and
the card.
What shipping it costs (C26). A patched server means the project signs and distributes its own build of SP1's
prover: the WSL2 package, the DMG's prover inputs, the K1-signed inputs and evidence.md carry it, and every SP1
upgrade repeats the clone, patch, build and measurement. The verifying key and the pinned guest ids do not move
(the patch changes buffer sizes and the shard split, not the circuits), which the verify-segment and --mode compressed VERIFIED lines of the unpatched host show on every row below. The packaging path is a row for the
proving plan before 0.3.12, not this branch.
Step 4 contingency, read not measured: RISC Zero's CUDA prover and its memory per segment
If SP1 could not be brought under 11 GB, the alternative's floor is read from its operators' documentation (not
measured here; a PC 2 run would be the measurement): Boundless' prover guide
(https://docs.boundless.network/provers/performance-optimization) sets the segment size cap by VRAM as 8 GB:
po2 19, 16 GB: po2 20, 20 GB: po2 21, 40 GB: po2 22, with measured peaks po2 20: 13,835 MiB, po2 21: 22,905 MiB,
po2 22: 41,089 MiB; RISC Zero's PR 3761 adds low_vram and pinned_witgen to fit po2 22 on a 24 GB 4090. So
RISC Zero proves a 2^19-cycle segment inside 8 GB and a 2^20 one inside 16 GB, and a shard of 4.7 M cycles is
9 segments at po2 19 plus lift and join steps (times not on the page). Adopting it would cost a second guest (the
chain rule in the RISC Zero zkVM), a second pinned program id, a second verifier in the node and no shared
aggregation between the two formats: docs/analysis/amd-proving.md and the proving plan carry that row already.
The build, as it ran (job floor-build-3, 22:28:24 to 22:32:29Z)
Three runs: floor-build-1 (22:17Z) and floor-build-2 (22:24Z) failed in 2 to 4 minutes on
crates/recursion/gnark-ffi/build.rs:70, "Failed to build Go library: NotFound" (no go on PC 2; the first run's
playbook lost its own log, a bug fixed before the second). floor-build-3 fetched go1.27.1 (tarball sha256
63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445, checked before unpacking under
/opt/igneum-floor/go, on the job's PATH only) and built in 240 s (46 crates on the warm target of run 2, 8
niced jobs, 16 cores). The binary: /opt/igneum-floor/bin/sp1-gpu-server, 166,768,224 bytes, sha256
5568108bf7fb9b0e525d8a08926b7046e51136ffaea53f0ca858631d0e938878, --version 6.8.1, cuobjdump --list-elf
sm_86, sm_89, sm_120 (the stock 251,306,680-byte server lists sm_80, 86, 89, 90, 100, 120 and compute_120 PTX).
The live /root/.sp1/bin/sp1-gpu-server (c2642ad1...) was never touched; the miners mined throughout.
Sweep 1 (job floor-sweep-1, 22:34:56 to 22:37:55Z): the shard term gone, a second floor found
PC 2's RTX 5090 (32,607 MiB, idle 1,755 MiB with the miners stopped and the live prover off), the patched server
5568108b... (sm_86, sm_89, sm_120; the build above), the unpatched pv1 host dae6b006... as client and
verifier, one --mode compressed --shard 0 per point, every server killed and its socket unlinked around every
point, peak = nvidia-smi memory.used at 1 s (the idle 1,755 MiB inside it), time = the compressed proof.
Every proof VERIFIED (1,272,897 bytes, verify 0.037 to 0.040 s), so the unpatched verifier accepts every proof
of the patched server.
| Config (environment to the patched server) | Fixture | Cycles | Peak MiB | Prove s | Verified |
|---|---|---|---|---|---|
control: SP1_GPU_MEMORY_BUDGET_GB=32 (upstream's sizes) |
empty live shard (block 83616) | 280,706 | 13,892 | 2.2 | yes |
| control | v1 shard (fees-v1-shards2 shard 0) | 4,717,439 | 20,516 | 4.2 | yes |
| 12 GB tier: budget 12 (threshold 2^27) | empty | 280,706 | 12,740 | 2.4 | yes |
| 12 GB tier | v1 shard | 4.7 M | 15,396 | 4.1 | yes |
12 GB tier + SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1 |
v1 shard | 4.7 M | 15,428 | 4.0 | yes |
| 16 GB tier: budget 16 (2^27 + 2^26) | v1 shard | 4.7 M | 18,628 | 3.8 | yes |
SP1_GPU_ELEMENT_THRESHOLD=67108864 (2^26) |
empty | 280,706 | 12,772 | 3.1 | yes |
| 2^26 | v1 shard (split into 4 core shards) | 4.7 M | 12,708 | 5.3 | yes |
SP1_GPU_ELEMENT_THRESHOLD=33554432 (2^25) |
v1 shard | 4.7 M | 12,836 | 8.5 | yes |
Reading. The control reproduces the proving agent's curve (13.9 and 20.4 GB), so the patched server behaves as
the stock one at the stock sizes. The shard's term follows the threshold as the model says (20.5 GB at 402 M
elements, 15.4 at 134 M, 12.7 at 67 M), and then stops: 2^26 and 2^25 both sit at 12.7 to 12.8 GB for the empty
shard and the v1 shard alike. The FLOOR memory after setup line names the rest: 9,703 MiB in use before the
first shard (2^26; 11,623 at the stock sizes), and the FLOOR tracegen alloc lines at Setup are five
allocations of 134,217,728 elements (the recursion keys, 0.75 GB each, each using 90,177,536 elements: 35.6 M
preprocessed and 54.5 M main at prove time), one of 33,554,432 (the shrink key, 0.19 GB) and one core key at the
threshold. The NORMALIZE_PROGRAM_CACHE_SIZE knob does not reach them (they are keys built at Setup, not the
program LRU). The time cost of the split: the v1 shard at 2^26 is 4 core shards and 5.3 s against 4.2 s (1.26x);
at 2^27 it is 4.1 s with no split.
So after sweep 1 the binding term is the Setup-time keys allocated at full capacity, and patch v2 sizes every
trace buffer (keys and shards) to its padded need: padded_trace_elements in jagged_tracegen/src/lib.rs
(each phase pads to the next multiple of 2^21 rows, generate_jagged_traces's "final padding"), applied in
setup_tracegen and full_tracegen, one stacking height of slack, SP1_GPU_FLOOR_EXACT=0 restoring upstream.
Sweep 2 (hung), the diagnosis, patch v3, sweep 3: under 11 GB
Sweep 2 (floor-sweep-2, 23:21Z, the v2 server fb3165d8) hung on its first point. The restore job
(floor-restore-1, 23:57Z) read the point's log: the v2 server had panicked in a tokio worker at
jagged_tracegen/src/lib.rs:240 ("range end index 37,428,736 out of range for slice of length 36,700,160") and the
SDK client waited on the socket for 1,740 s. The arithmetic named the bug: 36,700,160 is a recursion key's buffer
as v2 sized it (the padded preprocessed traces, 35,651,584, plus one stacking height), 37,428,736 is that end plus
one main trace of 1,777,152 elements: prove_shard_with_pk (shard_prover/src/prover.rs 332) runs
main_tracegen, which appends the shard's main traces INTO the key's buffer, which upstream sized for a whole shard.
Patch v3 keeps the key preprocessed-sized at Setup and lets main_tracegen grow it on first use (grow_for_main:
a bigger dense buffer and column index, the preprocessed region copied device to device, swapped into the key under
its lock; a loud abort instead of a hang if a bound were ever short). Build 5 (floor-build-5, 120 s):
b37defef9f5de43da06eb99a5aebdb6a0db214f49f485e00b86dd6d6c8b497b4, 166,752,944 bytes, sm_86/89/120.
Sweep 3 (floor-sweep-3, 00:13:46 to 00:16:57Z, idle 2,089 MiB inside every peak, every proof VERIFIED by the
unpatched host): the v1 shard at threshold 2^26 10,291 MiB (8,202 with the idle subtracted), 5.7 s; at 2^27
12,915 MiB 4.3 s; the empty shard 9,939 to 9,971 MiB; one transfer 10,003 MiB 3.5 s; the 60 M-cycle prototype
shard 13,459 MiB 16.8 s at the 12 GB tier (28,295 MiB on the stock server); the 16 GB tier 16,115 MiB; upstream's
threshold 16,851 MiB (20,516 stock); the server after Setup 6,535 MiB (v1 9,703, stock 11,623); the grow line
36,700,160 -> 91,226,112 elements on every recursion key's first use. The full table is in the bench-log entry.
The tier line that follows (for the public copy, once a 12 GB card has run the same fixture)
| Card | Stock SP1 6.8.1 server | The v3 server (this branch), measured on the 5090's allocation | Profile |
|---|---|---|---|
| 8 GB | refused (the 24 GB panic) | not measured; the empty shard alone is 9.9 GB measured (7.9 GB the server's own), so no | mine only |
| 12 GB | refused | MEASURED ON AN RTX 4070 12 GB (PC 1, 6 October): alone 7,553 MiB in 7.7 s at 2^26 (the card held nothing else); beside its own miner (1,449 MiB) 9,034 MiB of 12,282 in 24.1 s; core-only at 2^25 7,242 MiB beside the miner | SP1_GPU_ELEMENT_THRESHOLD=67108864: MINES AND PROVES (3.2 GB spare), no hand-off needed |
| 16 GB | refused | the v1 shard alone 12.9 GB measured (10.95 own) at 2^27, 4.3 s; beside the miner 10.95 own + 1.7, 17.4 s (sweep 4) | 2^27, mine and prove (2.9 GB spare on paper) |
| 24 GB | the v1 shard alone (20.4 GB), never the prototype shard (28.3) | the v1 shard 12.9 GB at 2^27, the prototype shard 13.5 GB (16.8 s) | upstream's 24 GB threshold or 2^27 |
| 32 GB | everything (28.3 GB for the prototype shard beside the miner at 30.1) | 16.9 GB at upstream's threshold | unchanged |
Every row is the 5090's allocation pattern under a budget; the public line keeps "24 GB" until the on-order 12 GB
card runs fees-v1-shards2 shard 0 through this server and its own peak and time are in the bench-log.
Sweep 4, beside the miner (job floor-sweep-4, 00:35 to 00:38Z)
The v3 server with the 0.3.11 worker mining on the same card (95%, 338 W, 3,833 MiB resident before the points): 2^26: the v1 shard 12,066 MiB (8,233 own) 24.4 s, the empty shard 11,586 MiB (7,753 own) 13.0 s; 2^25: the v1 shard 12,066 MiB 43.5 s; 2^27: the v1 shard 14,786 MiB (10,953 own) 17.4 s; every proof VERIFIED. The server's own working set does not move with the miner; the miner costs 4.3x in time; 2^25 is not a lever. The remaining floor is the Setup keys (6,535 MiB in use after Setup with the idle inside, about 4.4 GB own) and the recursion stage's peak (about 3.8 GB): cutting further means fewer keys built at Setup or a smaller recursion program, not a shard knob. The first publish of sweep 4 was refused by the publisher's kit-path-check (the wiped-jobs-folder rule of 21:49Z); the measurement template now tests the chain job's kit before use.
Route 2 measured (jobs floor-core-alone, floor-core-miner, 08:47 to 08:53Z)
Core-only at 2^26 on the v1 shard: 9,874 MiB alone (7,817 own) in 3.1 s and 12,834 MiB beside the miner (7,754 own)
in 12.1 s, against the full compressed run's 8,233 own; the core proof 14,379,043 bytes, CPU verify 0.447 s; the
aggregator's extra 2.5 s alone and 12.6 s beside a miner. The empty shard 6,377 own. The core proof's hand-off is
a prover-protocol change (the pool and igneum_submitProofRecord carry compressed proofs today), not one line.
The 12 GB mine-and-prove line (9.0 GB own) is not met by core-only at 2^26 (7.8 + 1.7 GB): the levers left are the
core threshold at 2^25 and 2^24 for core-only proving and patch v4's SP1_GPU_MEM_RELEASE_THRESHOLD (the pool
returns freed memory between shards; upstream holds the high-water mark for the process's life), both in
tools/prover-floor/pc2-floor-core2-miner.ps1 for the next PC 2 window. The table is in the bench-log entry.
Route 2, second round (jobs floor-build-6, floor-core2-miner, 09:02 to 09:13Z): the verdict
Patch v4 adds SP1_GPU_MEM_RELEASE_THRESHOLD (sp1-gpu/crates/cuda/src/task.rs: upstream's u64::MAX keeps every
freed allocation for the process's life). Core-only beside the miner: 2^25 6,409 MiB own in 19.8 s (core proof
25.6 MB, CPU verify 0.79 s), 2^24 6,025 own in 56.6 s; the pool's release threshold does not move the peak. So a
12 GB card mines and proves as a core-only prover at 2^25 (6.4 + 1.7 GB before the display) under the 9.0 GB line,
with the hand-off to a compressing aggregator as the prover-protocol change; the full table and the per-tier
consequences are in the bench-log entry. The real card's run decides the public line.
Route 1 measured: the RTX 4070 12 GB in PC 1 (jobs card12-alone, card12-miner, 11:53 to 12:17Z)
Alone (the card's idle 0 MiB): the compressed v1 shard 7,553 MiB in 7.7 s at 2^26, 10,177 MiB in 5.5 s at 2^27; core-only 5,761 MiB at 2^25. Beside its own miner (1,449 MiB): 9,034 MiB in 24.1 s at 2^26, 11,754 MiB in 17.3 s at 2^27, core-only 7,242 MiB at 2^25; a 23.6 M-cycle shard 9,066 MiB in 116.6 s at 2^26. Every proof verified by the unpatched verifier. So a 12 GB card mines and proves compressed shards at 2^26 with 3.2 GB spare, and route 2's core-only hand-off is the reserve, not the requirement. The 5090's allocation pattern overstated the card by about 0.65 GB. The full tables and the per-tier consequences are in the bench-log entry; the public line moves to "12 GB mines and proves" when the packaging row ships the server.
The packaging row (6 October 2026, afternoon)
Branch prover-floor carries the whole path: packaging/prover/build-server.sh (the one recipe), the CI workflow
prover-server.yml, push-server.sh (the Mac signs the CI build's manifest with the OTA key and publishes),
fetch-server.sh (verifies and places it for the payload), the make-payload.sh step (wsl2\bin\sp1-gpu-server,
wsl2\prover-server.json, .sig), the Windows build's fetch step, and the app: proverserver.rs (the manifest,
the hash check, the install into ~/.sp1/bin, the fallback), provedefault::profile (the tiers from the VRAM,
Settings' override), /api/state's server_* fields. The tier table and the shipper's steps are in
docs/plans/proving-v1.md, "The packaged GPU server". Tests: 147 in the app, of which the new ones refuse a
modified or truncated binary and an unsigned manifest, choose the tier per VRAM, and fire the fallback only on a
server that did not come up.