igneum/docs/analysis/prover-floor.md

22 KiB

The prover floor: why a 12 GB card cannot prove on SP1 6.8.1's GPU server, and the patch

5 October 2026, 22:00 UTC on (the project lead: "execute if it will solve the issue"). Branch prover-floor (worktree igneum-wt-prover-floor). The measured facts this starts from: docs/plans/proving-v1.md and the bench-log entry "proving v1" (branch proving-v1): the GPU server holds 13.9 GB for an empty shard, 20.4 GB at the adopted v1 shard, 28.3 GB flat from 20 M to 60 M cycles, and no environment knob moved the floor. Every figure below is from the source at tag v6.8.1 (cloned to vendor/sp1-6.8.1, gitignored; the fork is the patch proving/prover-floor/sp1-gpu-6.8.1-floor.patch) or from a PC 2 run named in docs/bench-log.md ("prover floor"). Sizes in GiB are computed from the source constants (4-byte field elements); sizes in MiB are measured by nvidia-smi at 1 s.

Where the server is built and what it reads

The SDK downloads sp1_gpu_server_v6.8.1_x86_64.tar.gz (133,750,780 bytes) from the SP1 release and runs it from $HOME/.sp1/bin/sp1-gpu-server (crates/cuda/src/server.rs 19 to 30, 80 to 99). The source is in the same repository: sp1-gpu/crates/server (the binary), built by .github/workflows/release.yml 234 to 314 on CUDA 12.8.1 with Go and protoc (cargo build --release --bin sp1-gpu-server). The binary takes no options (sp1-gpu/crates/server/src/main.rs 15 to 18: --version only) and reads CUDA_VISIBLE_DEVICES (32 to 35); everything else comes from the environment the host process passes it, through SP1CoreOpts::default() (crates/core/executor/src/opts.rs 99 to 140: SHARD_SIZE, ELEMENT_THRESHOLD, HEIGHT_THRESHOLD, MINIMAL_TRACE_CHUNK_THRESHOLD, TRACE_CHUNK_SLOTS, FULL_SIZE_SHARDS) and the worker counts (crates/prover/src/worker/config.rs).

The memory model, term by term

Every device buffer is sized at construction from constants, not from the shard. The server builds the prover at the first Setup request (sp1-gpu/crates/server/src/server.rs 126 to 137) through cuda_worker_builder_with_machine (sp1-gpu/crates/prover_components/src/builder.rs 102 to 148):

Term Where Size On the device Moves with the shard
The gate builder.rs 35 to 39: gpu_memory_gb = ceil(total / GiB) + 4; panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB") when under 24 a 12 GB card reads 16, a 16 GB card 20: both refused before any allocation no
The core element threshold builder.rs 41 to 48: ELEMENT_THRESHOLD = 2^28 + 2^27 = 402,653,184 elements (opts.rs 12) on a card reading over 30 (a 32 GB card reads 36); minus 2^26 + 2^25 + 2^24 = 285,212,672 on a card reading 24 to 30 (a 24 GB card). The environment's ELEMENT_THRESHOLD is read at opts.rs 129 and then OVERWRITTEN at builder.rs 48, which is why the sweep's ELEMENT_THRESHOLD rows changed nothing; HEIGHT_THRESHOLD survives (it is not overwritten), which is why the 2^25 + 2^20 row did sets the next two terms no
The core trace area, one per shard in flight builder.rs 70 to 71: num_elts = element_threshold + 2^21 (CORE_LOG_STACKING_HEIGHT 21, crates/prover/src/components.rs 16) = 404,750,336; allocated on the device at sp1-gpu/crates/jagged_tracegen/src/lib.rs 484 to 500 (allocate_and_initialize_traces: max_trace_size felts + max_trace_size / 2 u32 column index + 2^14 u32) 6 bytes an element: 2.26 GiB for the full threshold, 1.59 GiB for the 24 GB threshold yes, in full, whatever the shard holds no
The program's preprocessed traces (the proving key) sp1-gpu/crates/shard_prover/src/setup.rs 47 to 58 and 107: the same allocate_and_initialize_traces(max_trace_size) at Setup, kept in the key cache for the connection's life (server.rs 139 to 143) another 2.26 GiB, held from Setup on yes no
The pinned host trace buffers sp1-gpu/crates/prover_components/src/components.rs 99 to 103: 4 PinnedBuffer of max_trace_size felts per prover (core 4 x 1.51 GiB, recursion 4 x 0.5 GiB, shrink 4 x 0.125 GiB, wrap 4 x 0.32 GiB) 9.8 GiB of pinned host RAM, not device memory (the WSL2 working set the bench saw) no no
The recursion trace area builder.rs 15 and 95: RECURSION_TRACE_ALLOCATION = 2^27 elements, one per recursion tracegen (the recursion program's key at setup and its shard at prove) 0.75 GiB each yes no
The shrink and wrap provers builder.rs 16, 19, 117 to 127: 2^25 and 85,376,340 elements, built at Setup for every proof mode, used only by the Groth16 and PLONK path host pinned at build; device only when a wrap runs (never, for a compressed proof) no no
The codewords (LDE) and the Merkle trees sp1-gpu/crates/basefold/src/fri.rs 92 to 97: every stacked column of 2^21 rows encoded to 2^(21 + 1) rows (log_blowup 1); kept until the query phase unless drop_ldes (builder.rs 52: only on a 24 GB card with FULL_SIZE_SHARDS); the preprocessed codewords live in the key 2 x the padded trace, so up to 2 x the term above yes yes, with the padded trace
The LogUp GKR layers builder.rs 51: recompute_gkr_trace = false, so the first layer stays materialised (sp1-gpu/crates/logup_gkr/src/tracegen.rs 169 to 224) of the order of the interaction count yes yes
The allocator sp1-gpu/crates/cuda/src/task.rs 152 and 196: the device's default cudaMallocAsync pool with its release threshold at u64::MAX, so nothing freed is ever returned to the driver: nvidia-smi reads the high-water mark of everything live at once

So at zero cycles the server already holds the proving key's 2.26 GiB, the shard's 2.26 GiB (both allocated at the threshold, not at the shard's rows), their codewords and trees, and the recursion program's key and traces (2 x 0.75 GiB and their codewords): the 13.9 GB floor. The shard's own content only adds to the codewords, the GKR layers and the working buffers, which is the 13.9 to 20.4 GB step from 0.3 M to 4.7 M cycles, and the flat 28.3 GB from 20 M cycles is the threshold's padded area reached. The witness (5 to 22 KB) never appears.

What the patch does (proving/prover-floor/sp1-gpu-6.8.1-floor.patch, three files)

  1. builder.rs: the panic is gone; the card's memory (or SP1_GPU_MEMORY_BUDGET_GB) picks the element threshold from a tier table (element_threshold_for_budget: over 30 as read, the full 402.6 M; 24 to 30, upstream's 24 GB figure; 18 to 24 (a 16 GB card), 2^27 + 2^26 = 201.3 M; under 18 (a 12 GB card), 2^27 = 134.2 M); SP1_GPU_ELEMENT_THRESHOLD sets it directly and SP1_GPU_RECURSION_TRACE_ALLOCATION the recursion buffer. The chosen numbers are printed as a FLOOR opts line. Every other option is as upstream.
  2. jagged_tracegen/src/lib.rs: with SP1_GPU_FLOOR_LOG set, every trace allocation prints its capacity and, after the shard's traces are in, the elements actually used and the device memory in use.
  3. server.rs: a FLOOR memory line (device used, free, total) after Setup and after every proof, with the proof's time.

Nothing in the proof changes: the element threshold only moves where the executor splits shards, exactly what upstream's own 24 GB tier does with the same verifier and the same keys; the recursion program, the verifying key and the pinned guest ids are untouched. The unpatched verifier (the pv1 host's SDK) is the one that verifies every measured proof below.

The build (PC 2, WSL2 Ubuntu-24.04, job floor-toolchain-1 then the build job)

Toolchain found 22:10Z (job floor-toolchain-1, 4 s): nvcc 12.8 at /usr/local/cuda-12.8, cmake 3.28.3, gcc 13.3, clang 18, protoc 3.21.12, cargo 1.99.0, no Go. The release workflow installs Go for the server's native-gnark feature (the Groth16 and PLONK wrap through gnark), which a compressed proof never runs, so the build drops that feature from sp1-gpu/crates/server/Cargo.toml and nothing else. CUDA_ARCHS=86,89,120 (consequences reviewer C26): the 12 GB tier is sm_86 (RTX 3060) and sm_89 (RTX 4070), the 16 GB tier sm_89 and sm_120 (RTX 5080), PC 2's 5090 is sm_120; the stock server lists sm_80, 86, 89, 90, 100 and 120, which a shipped build repeats. The recipe: tools/prover-floor/pc2-build-server.ps1 (generated by make-build-playbook.sh from the patch, so the two cannot drift): clone the tag, git apply the patch, touch the three files, cargo build --release --bin sp1-gpu-server niced with 8 jobs into /opt/igneum-floor/target, the binary copied to /opt/igneum-floor/home/.sp1/bin/ (the SDK spawns the server it finds under $HOME/.sp1/bin, so HOME=/opt/igneum-floor/home selects it and the live /root/.sp1/bin/sp1-gpu-server stays as it is).

What a measurement on PC 2 can and cannot say (C26). The server's allocation pattern is deterministic in the budget it is given, so a run with SP1_GPU_MEMORY_BUDGET_GB=12 on the 5090 shows the peak a 12 GB card's build would ask for; it does not show that a 3060 proves it in time, nor what the card's display and driver hold. The public line keeps "24 GB" until the on-order 12 GB card runs the same fixture. Every row names the arch list and the card.

What shipping it costs (C26). A patched server means the project signs and distributes its own build of SP1's prover: the WSL2 package, the DMG's prover inputs, the K1-signed inputs and evidence.md carry it, and every SP1 upgrade repeats the clone, patch, build and measurement. The verifying key and the pinned guest ids do not move (the patch changes buffer sizes and the shard split, not the circuits), which the verify-segment and --mode compressed VERIFIED lines of the unpatched host show on every row below. The packaging path is a row for the proving plan before 0.3.12, not this branch.

Step 4 contingency, read not measured: RISC Zero's CUDA prover and its memory per segment

If SP1 could not be brought under 11 GB, the alternative's floor is read from its operators' documentation (not measured here; a PC 2 run would be the measurement): Boundless' prover guide (https://docs.boundless.network/provers/performance-optimization) sets the segment size cap by VRAM as 8 GB: po2 19, 16 GB: po2 20, 20 GB: po2 21, 40 GB: po2 22, with measured peaks po2 20: 13,835 MiB, po2 21: 22,905 MiB, po2 22: 41,089 MiB; RISC Zero's PR 3761 adds low_vram and pinned_witgen to fit po2 22 on a 24 GB 4090. So RISC Zero proves a 2^19-cycle segment inside 8 GB and a 2^20 one inside 16 GB, and a shard of 4.7 M cycles is 9 segments at po2 19 plus lift and join steps (times not on the page). Adopting it would cost a second guest (the chain rule in the RISC Zero zkVM), a second pinned program id, a second verifier in the node and no shared aggregation between the two formats: docs/analysis/amd-proving.md and the proving plan carry that row already.

The build, as it ran (job floor-build-3, 22:28:24 to 22:32:29Z)

Three runs: floor-build-1 (22:17Z) and floor-build-2 (22:24Z) failed in 2 to 4 minutes on crates/recursion/gnark-ffi/build.rs:70, "Failed to build Go library: NotFound" (no go on PC 2; the first run's playbook lost its own log, a bug fixed before the second). floor-build-3 fetched go1.27.1 (tarball sha256 63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445, checked before unpacking under /opt/igneum-floor/go, on the job's PATH only) and built in 240 s (46 crates on the warm target of run 2, 8 niced jobs, 16 cores). The binary: /opt/igneum-floor/bin/sp1-gpu-server, 166,768,224 bytes, sha256 5568108bf7fb9b0e525d8a08926b7046e51136ffaea53f0ca858631d0e938878, --version 6.8.1, cuobjdump --list-elf sm_86, sm_89, sm_120 (the stock 251,306,680-byte server lists sm_80, 86, 89, 90, 100, 120 and compute_120 PTX). The live /root/.sp1/bin/sp1-gpu-server (c2642ad1...) was never touched; the miners mined throughout.

Sweep 1 (job floor-sweep-1, 22:34:56 to 22:37:55Z): the shard term gone, a second floor found

PC 2's RTX 5090 (32,607 MiB, idle 1,755 MiB with the miners stopped and the live prover off), the patched server 5568108b... (sm_86, sm_89, sm_120; the build above), the unpatched pv1 host dae6b006... as client and verifier, one --mode compressed --shard 0 per point, every server killed and its socket unlinked around every point, peak = nvidia-smi memory.used at 1 s (the idle 1,755 MiB inside it), time = the compressed proof. Every proof VERIFIED (1,272,897 bytes, verify 0.037 to 0.040 s), so the unpatched verifier accepts every proof of the patched server.

Config (environment to the patched server) Fixture Cycles Peak MiB Prove s Verified
control: SP1_GPU_MEMORY_BUDGET_GB=32 (upstream's sizes) empty live shard (block 83616) 280,706 13,892 2.2 yes
control v1 shard (fees-v1-shards2 shard 0) 4,717,439 20,516 4.2 yes
12 GB tier: budget 12 (threshold 2^27) empty 280,706 12,740 2.4 yes
12 GB tier v1 shard 4.7 M 15,396 4.1 yes
12 GB tier + SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1 v1 shard 4.7 M 15,428 4.0 yes
16 GB tier: budget 16 (2^27 + 2^26) v1 shard 4.7 M 18,628 3.8 yes
SP1_GPU_ELEMENT_THRESHOLD=67108864 (2^26) empty 280,706 12,772 3.1 yes
2^26 v1 shard (split into 4 core shards) 4.7 M 12,708 5.3 yes
SP1_GPU_ELEMENT_THRESHOLD=33554432 (2^25) v1 shard 4.7 M 12,836 8.5 yes

Reading. The control reproduces the proving agent's curve (13.9 and 20.4 GB), so the patched server behaves as the stock one at the stock sizes. The shard's term follows the threshold as the model says (20.5 GB at 402 M elements, 15.4 at 134 M, 12.7 at 67 M), and then stops: 2^26 and 2^25 both sit at 12.7 to 12.8 GB for the empty shard and the v1 shard alike. The FLOOR memory after setup line names the rest: 9,703 MiB in use before the first shard (2^26; 11,623 at the stock sizes), and the FLOOR tracegen alloc lines at Setup are five allocations of 134,217,728 elements (the recursion keys, 0.75 GB each, each using 90,177,536 elements: 35.6 M preprocessed and 54.5 M main at prove time), one of 33,554,432 (the shrink key, 0.19 GB) and one core key at the threshold. The NORMALIZE_PROGRAM_CACHE_SIZE knob does not reach them (they are keys built at Setup, not the program LRU). The time cost of the split: the v1 shard at 2^26 is 4 core shards and 5.3 s against 4.2 s (1.26x); at 2^27 it is 4.1 s with no split.

So after sweep 1 the binding term is the Setup-time keys allocated at full capacity, and patch v2 sizes every trace buffer (keys and shards) to its padded need: padded_trace_elements in jagged_tracegen/src/lib.rs (each phase pads to the next multiple of 2^21 rows, generate_jagged_traces's "final padding"), applied in setup_tracegen and full_tracegen, one stacking height of slack, SP1_GPU_FLOOR_EXACT=0 restoring upstream.

Sweep 2 (hung), the diagnosis, patch v3, sweep 3: under 11 GB

Sweep 2 (floor-sweep-2, 23:21Z, the v2 server fb3165d8) hung on its first point. The restore job (floor-restore-1, 23:57Z) read the point's log: the v2 server had panicked in a tokio worker at jagged_tracegen/src/lib.rs:240 ("range end index 37,428,736 out of range for slice of length 36,700,160") and the SDK client waited on the socket for 1,740 s. The arithmetic named the bug: 36,700,160 is a recursion key's buffer as v2 sized it (the padded preprocessed traces, 35,651,584, plus one stacking height), 37,428,736 is that end plus one main trace of 1,777,152 elements: prove_shard_with_pk (shard_prover/src/prover.rs 332) runs main_tracegen, which appends the shard's main traces INTO the key's buffer, which upstream sized for a whole shard. Patch v3 keeps the key preprocessed-sized at Setup and lets main_tracegen grow it on first use (grow_for_main: a bigger dense buffer and column index, the preprocessed region copied device to device, swapped into the key under its lock; a loud abort instead of a hang if a bound were ever short). Build 5 (floor-build-5, 120 s): b37defef9f5de43da06eb99a5aebdb6a0db214f49f485e00b86dd6d6c8b497b4, 166,752,944 bytes, sm_86/89/120.

Sweep 3 (floor-sweep-3, 00:13:46 to 00:16:57Z, idle 2,089 MiB inside every peak, every proof VERIFIED by the unpatched host): the v1 shard at threshold 2^26 10,291 MiB (8,202 with the idle subtracted), 5.7 s; at 2^27 12,915 MiB 4.3 s; the empty shard 9,939 to 9,971 MiB; one transfer 10,003 MiB 3.5 s; the 60 M-cycle prototype shard 13,459 MiB 16.8 s at the 12 GB tier (28,295 MiB on the stock server); the 16 GB tier 16,115 MiB; upstream's threshold 16,851 MiB (20,516 stock); the server after Setup 6,535 MiB (v1 9,703, stock 11,623); the grow line 36,700,160 -> 91,226,112 elements on every recursion key's first use. The full table is in the bench-log entry.

The tier line that follows (for the public copy, once a 12 GB card has run the same fixture)

Card Stock SP1 6.8.1 server The v3 server (this branch), measured on the 5090's allocation Profile
8 GB refused (the 24 GB panic) not measured; the empty shard alone is 9.9 GB measured (7.9 GB the server's own), so no mine only
12 GB refused MEASURED ON AN RTX 4070 12 GB (PC 1, 6 October): alone 7,553 MiB in 7.7 s at 2^26 (the card held nothing else); beside its own miner (1,449 MiB) 9,034 MiB of 12,282 in 24.1 s; core-only at 2^25 7,242 MiB beside the miner SP1_GPU_ELEMENT_THRESHOLD=67108864: MINES AND PROVES (3.2 GB spare), no hand-off needed
16 GB refused the v1 shard alone 12.9 GB measured (10.95 own) at 2^27, 4.3 s; beside the miner 10.95 own + 1.7, 17.4 s (sweep 4) 2^27, mine and prove (2.9 GB spare on paper)
24 GB the v1 shard alone (20.4 GB), never the prototype shard (28.3) the v1 shard 12.9 GB at 2^27, the prototype shard 13.5 GB (16.8 s) upstream's 24 GB threshold or 2^27
32 GB everything (28.3 GB for the prototype shard beside the miner at 30.1) 16.9 GB at upstream's threshold unchanged

Every row is the 5090's allocation pattern under a budget; the public line keeps "24 GB" until the on-order 12 GB card runs fees-v1-shards2 shard 0 through this server and its own peak and time are in the bench-log.

Sweep 4, beside the miner (job floor-sweep-4, 00:35 to 00:38Z)

The v3 server with the 0.3.11 worker mining on the same card (95%, 338 W, 3,833 MiB resident before the points): 2^26: the v1 shard 12,066 MiB (8,233 own) 24.4 s, the empty shard 11,586 MiB (7,753 own) 13.0 s; 2^25: the v1 shard 12,066 MiB 43.5 s; 2^27: the v1 shard 14,786 MiB (10,953 own) 17.4 s; every proof VERIFIED. The server's own working set does not move with the miner; the miner costs 4.3x in time; 2^25 is not a lever. The remaining floor is the Setup keys (6,535 MiB in use after Setup with the idle inside, about 4.4 GB own) and the recursion stage's peak (about 3.8 GB): cutting further means fewer keys built at Setup or a smaller recursion program, not a shard knob. The first publish of sweep 4 was refused by the publisher's kit-path-check (the wiped-jobs-folder rule of 21:49Z); the measurement template now tests the chain job's kit before use.

Route 2 measured (jobs floor-core-alone, floor-core-miner, 08:47 to 08:53Z)

Core-only at 2^26 on the v1 shard: 9,874 MiB alone (7,817 own) in 3.1 s and 12,834 MiB beside the miner (7,754 own) in 12.1 s, against the full compressed run's 8,233 own; the core proof 14,379,043 bytes, CPU verify 0.447 s; the aggregator's extra 2.5 s alone and 12.6 s beside a miner. The empty shard 6,377 own. The core proof's hand-off is a prover-protocol change (the pool and igneum_submitProofRecord carry compressed proofs today), not one line. The 12 GB mine-and-prove line (9.0 GB own) is not met by core-only at 2^26 (7.8 + 1.7 GB): the levers left are the core threshold at 2^25 and 2^24 for core-only proving and patch v4's SP1_GPU_MEM_RELEASE_THRESHOLD (the pool returns freed memory between shards; upstream holds the high-water mark for the process's life), both in tools/prover-floor/pc2-floor-core2-miner.ps1 for the next PC 2 window. The table is in the bench-log entry.

Route 2, second round (jobs floor-build-6, floor-core2-miner, 09:02 to 09:13Z): the verdict

Patch v4 adds SP1_GPU_MEM_RELEASE_THRESHOLD (sp1-gpu/crates/cuda/src/task.rs: upstream's u64::MAX keeps every freed allocation for the process's life). Core-only beside the miner: 2^25 6,409 MiB own in 19.8 s (core proof 25.6 MB, CPU verify 0.79 s), 2^24 6,025 own in 56.6 s; the pool's release threshold does not move the peak. So a 12 GB card mines and proves as a core-only prover at 2^25 (6.4 + 1.7 GB before the display) under the 9.0 GB line, with the hand-off to a compressing aggregator as the prover-protocol change; the full table and the per-tier consequences are in the bench-log entry. The real card's run decides the public line.

Route 1 measured: the RTX 4070 12 GB in PC 1 (jobs card12-alone, card12-miner, 11:53 to 12:17Z)

Alone (the card's idle 0 MiB): the compressed v1 shard 7,553 MiB in 7.7 s at 2^26, 10,177 MiB in 5.5 s at 2^27; core-only 5,761 MiB at 2^25. Beside its own miner (1,449 MiB): 9,034 MiB in 24.1 s at 2^26, 11,754 MiB in 17.3 s at 2^27, core-only 7,242 MiB at 2^25; a 23.6 M-cycle shard 9,066 MiB in 116.6 s at 2^26. Every proof verified by the unpatched verifier. So a 12 GB card mines and proves compressed shards at 2^26 with 3.2 GB spare, and route 2's core-only hand-off is the reserve, not the requirement. The 5090's allocation pattern overstated the card by about 0.65 GB. The full tables and the per-tier consequences are in the bench-log entry; the public line moves to "12 GB mines and proves" when the packaging row ships the server.

The packaging row (6 October 2026, afternoon)

Branch prover-floor carries the whole path: packaging/prover/build-server.sh (the one recipe), the CI workflow prover-server.yml, push-server.sh (the Mac signs the CI build's manifest with the OTA key and publishes), fetch-server.sh (verifies and places it for the payload), the make-payload.sh step (wsl2\bin\sp1-gpu-server, wsl2\prover-server.json, .sig), the Windows build's fetch step, and the app: proverserver.rs (the manifest, the hash check, the install into ~/.sp1/bin, the fallback), provedefault::profile (the tiers from the VRAM, Settings' override), /api/state's server_* fields. The tier table and the shipper's steps are in docs/plans/proving-v1.md, "The packaged GPU server". Tests: 147 in the app, of which the new ones refuse a modified or truncated binary and an unsigned manifest, choose the tier per VRAM, and fire the fallback only on a server that did not come up.