igneum/docs/analysis/prover-floor.md
igneum-labs 5bc7ce9b55 Prover floor: sweep 1 table and reading (the Setup-time keys bind at 12.7 GB)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 22:42:17 +00:00

15 KiB

The prover floor: why a 12 GB card cannot prove on SP1 6.8.1's GPU server, and the patch

5 October 2026, 22:00 UTC on (the project lead: "execute if it will solve the issue"). Branch prover-floor (worktree igneum-wt-prover-floor). The measured facts this starts from: docs/plans/proving-v1.md and the bench-log entry "proving v1" (branch proving-v1): the GPU server holds 13.9 GB for an empty shard, 20.4 GB at the adopted v1 shard, 28.3 GB flat from 20 M to 60 M cycles, and no environment knob moved the floor. Every figure below is from the source at tag v6.8.1 (cloned to vendor/sp1-6.8.1, gitignored; the fork is the patch proving/prover-floor/sp1-gpu-6.8.1-floor.patch) or from a PC 2 run named in docs/bench-log.md ("prover floor"). Sizes in GiB are computed from the source constants (4-byte field elements); sizes in MiB are measured by nvidia-smi at 1 s.

Where the server is built and what it reads

The SDK downloads sp1_gpu_server_v6.8.1_x86_64.tar.gz (133,750,780 bytes) from the SP1 release and runs it from $HOME/.sp1/bin/sp1-gpu-server (crates/cuda/src/server.rs 19 to 30, 80 to 99). The source is in the same repository: sp1-gpu/crates/server (the binary), built by .github/workflows/release.yml 234 to 314 on CUDA 12.8.1 with Go and protoc (cargo build --release --bin sp1-gpu-server). The binary takes no options (sp1-gpu/crates/server/src/main.rs 15 to 18: --version only) and reads CUDA_VISIBLE_DEVICES (32 to 35); everything else comes from the environment the host process passes it, through SP1CoreOpts::default() (crates/core/executor/src/opts.rs 99 to 140: SHARD_SIZE, ELEMENT_THRESHOLD, HEIGHT_THRESHOLD, MINIMAL_TRACE_CHUNK_THRESHOLD, TRACE_CHUNK_SLOTS, FULL_SIZE_SHARDS) and the worker counts (crates/prover/src/worker/config.rs).

The memory model, term by term

Every device buffer is sized at construction from constants, not from the shard. The server builds the prover at the first Setup request (sp1-gpu/crates/server/src/server.rs 126 to 137) through cuda_worker_builder_with_machine (sp1-gpu/crates/prover_components/src/builder.rs 102 to 148):

Term Where Size On the device Moves with the shard
The gate builder.rs 35 to 39: gpu_memory_gb = ceil(total / GiB) + 4; panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB") when under 24 a 12 GB card reads 16, a 16 GB card 20: both refused before any allocation no
The core element threshold builder.rs 41 to 48: ELEMENT_THRESHOLD = 2^28 + 2^27 = 402,653,184 elements (opts.rs 12) on a card reading over 30 (a 32 GB card reads 36); minus 2^26 + 2^25 + 2^24 = 285,212,672 on a card reading 24 to 30 (a 24 GB card). The environment's ELEMENT_THRESHOLD is read at opts.rs 129 and then OVERWRITTEN at builder.rs 48, which is why the sweep's ELEMENT_THRESHOLD rows changed nothing; HEIGHT_THRESHOLD survives (it is not overwritten), which is why the 2^25 + 2^20 row did sets the next two terms no
The core trace area, one per shard in flight builder.rs 70 to 71: num_elts = element_threshold + 2^21 (CORE_LOG_STACKING_HEIGHT 21, crates/prover/src/components.rs 16) = 404,750,336; allocated on the device at sp1-gpu/crates/jagged_tracegen/src/lib.rs 484 to 500 (allocate_and_initialize_traces: max_trace_size felts + max_trace_size / 2 u32 column index + 2^14 u32) 6 bytes an element: 2.26 GiB for the full threshold, 1.59 GiB for the 24 GB threshold yes, in full, whatever the shard holds no
The program's preprocessed traces (the proving key) sp1-gpu/crates/shard_prover/src/setup.rs 47 to 58 and 107: the same allocate_and_initialize_traces(max_trace_size) at Setup, kept in the key cache for the connection's life (server.rs 139 to 143) another 2.26 GiB, held from Setup on yes no
The pinned host trace buffers sp1-gpu/crates/prover_components/src/components.rs 99 to 103: 4 PinnedBuffer of max_trace_size felts per prover (core 4 x 1.51 GiB, recursion 4 x 0.5 GiB, shrink 4 x 0.125 GiB, wrap 4 x 0.32 GiB) 9.8 GiB of pinned host RAM, not device memory (the WSL2 working set the bench saw) no no
The recursion trace area builder.rs 15 and 95: RECURSION_TRACE_ALLOCATION = 2^27 elements, one per recursion tracegen (the recursion program's key at setup and its shard at prove) 0.75 GiB each yes no
The shrink and wrap provers builder.rs 16, 19, 117 to 127: 2^25 and 85,376,340 elements, built at Setup for every proof mode, used only by the Groth16 and PLONK path host pinned at build; device only when a wrap runs (never, for a compressed proof) no no
The codewords (LDE) and the Merkle trees sp1-gpu/crates/basefold/src/fri.rs 92 to 97: every stacked column of 2^21 rows encoded to 2^(21 + 1) rows (log_blowup 1); kept until the query phase unless drop_ldes (builder.rs 52: only on a 24 GB card with FULL_SIZE_SHARDS); the preprocessed codewords live in the key 2 x the padded trace, so up to 2 x the term above yes yes, with the padded trace
The LogUp GKR layers builder.rs 51: recompute_gkr_trace = false, so the first layer stays materialised (sp1-gpu/crates/logup_gkr/src/tracegen.rs 169 to 224) of the order of the interaction count yes yes
The allocator sp1-gpu/crates/cuda/src/task.rs 152 and 196: the device's default cudaMallocAsync pool with its release threshold at u64::MAX, so nothing freed is ever returned to the driver: nvidia-smi reads the high-water mark of everything live at once

So at zero cycles the server already holds the proving key's 2.26 GiB, the shard's 2.26 GiB (both allocated at the threshold, not at the shard's rows), their codewords and trees, and the recursion program's key and traces (2 x 0.75 GiB and their codewords): the 13.9 GB floor. The shard's own content only adds to the codewords, the GKR layers and the working buffers, which is the 13.9 to 20.4 GB step from 0.3 M to 4.7 M cycles, and the flat 28.3 GB from 20 M cycles is the threshold's padded area reached. The witness (5 to 22 KB) never appears.

What the patch does (proving/prover-floor/sp1-gpu-6.8.1-floor.patch, three files)

  1. builder.rs: the panic is gone; the card's memory (or SP1_GPU_MEMORY_BUDGET_GB) picks the element threshold from a tier table (element_threshold_for_budget: over 30 as read, the full 402.6 M; 24 to 30, upstream's 24 GB figure; 18 to 24 (a 16 GB card), 2^27 + 2^26 = 201.3 M; under 18 (a 12 GB card), 2^27 = 134.2 M); SP1_GPU_ELEMENT_THRESHOLD sets it directly and SP1_GPU_RECURSION_TRACE_ALLOCATION the recursion buffer. The chosen numbers are printed as a FLOOR opts line. Every other option is as upstream.
  2. jagged_tracegen/src/lib.rs: with SP1_GPU_FLOOR_LOG set, every trace allocation prints its capacity and, after the shard's traces are in, the elements actually used and the device memory in use.
  3. server.rs: a FLOOR memory line (device used, free, total) after Setup and after every proof, with the proof's time.

Nothing in the proof changes: the element threshold only moves where the executor splits shards, exactly what upstream's own 24 GB tier does with the same verifier and the same keys; the recursion program, the verifying key and the pinned guest ids are untouched. The unpatched verifier (the pv1 host's SDK) is the one that verifies every measured proof below.

The build (PC 2, WSL2 Ubuntu-24.04, job floor-toolchain-1 then the build job)

Toolchain found 22:10Z (job floor-toolchain-1, 4 s): nvcc 12.8 at /usr/local/cuda-12.8, cmake 3.28.3, gcc 13.3, clang 18, protoc 3.21.12, cargo 1.99.0, no Go. The release workflow installs Go for the server's native-gnark feature (the Groth16 and PLONK wrap through gnark), which a compressed proof never runs, so the build drops that feature from sp1-gpu/crates/server/Cargo.toml and nothing else. CUDA_ARCHS=86,89,120 (consequences reviewer C26): the 12 GB tier is sm_86 (RTX 3060) and sm_89 (RTX 4070), the 16 GB tier sm_89 and sm_120 (RTX 5080), PC 2's 5090 is sm_120; the stock server lists sm_80, 86, 89, 90, 100 and 120, which a shipped build repeats. The recipe: tools/prover-floor/pc2-build-server.ps1 (generated by make-build-playbook.sh from the patch, so the two cannot drift): clone the tag, git apply the patch, touch the three files, cargo build --release --bin sp1-gpu-server niced with 8 jobs into /opt/igneum-floor/target, the binary copied to /opt/igneum-floor/home/.sp1/bin/ (the SDK spawns the server it finds under $HOME/.sp1/bin, so HOME=/opt/igneum-floor/home selects it and the live /root/.sp1/bin/sp1-gpu-server stays as it is).

What a measurement on PC 2 can and cannot say (C26). The server's allocation pattern is deterministic in the budget it is given, so a run with SP1_GPU_MEMORY_BUDGET_GB=12 on the 5090 shows the peak a 12 GB card's build would ask for; it does not show that a 3060 proves it in time, nor what the card's display and driver hold. The public line keeps "24 GB" until the on-order 12 GB card runs the same fixture. Every row names the arch list and the card.

What shipping it costs (C26). A patched server means the project signs and distributes its own build of SP1's prover: the WSL2 package, the DMG's prover inputs, the K1-signed inputs and evidence.md carry it, and every SP1 upgrade repeats the clone, patch, build and measurement. The verifying key and the pinned guest ids do not move (the patch changes buffer sizes and the shard split, not the circuits), which the verify-segment and --mode compressed VERIFIED lines of the unpatched host show on every row below. The packaging path is a row for the proving plan before 0.3.12, not this branch.

Step 4 contingency, read not measured: RISC Zero's CUDA prover and its memory per segment

If SP1 could not be brought under 11 GB, the alternative's floor is read from its operators' documentation (not measured here; a PC 2 run would be the measurement): Boundless' prover guide (https://docs.boundless.network/provers/performance-optimization) sets the segment size cap by VRAM as 8 GB: po2 19, 16 GB: po2 20, 20 GB: po2 21, 40 GB: po2 22, with measured peaks po2 20: 13,835 MiB, po2 21: 22,905 MiB, po2 22: 41,089 MiB; RISC Zero's PR 3761 adds low_vram and pinned_witgen to fit po2 22 on a 24 GB 4090. So RISC Zero proves a 2^19-cycle segment inside 8 GB and a 2^20 one inside 16 GB, and a shard of 4.7 M cycles is 9 segments at po2 19 plus lift and join steps (times not on the page). Adopting it would cost a second guest (the chain rule in the RISC Zero zkVM), a second pinned program id, a second verifier in the node and no shared aggregation between the two formats: docs/analysis/amd-proving.md and the proving plan carry that row already.

The build, as it ran (job floor-build-3, 22:28:24 to 22:32:29Z)

Three runs: floor-build-1 (22:17Z) and floor-build-2 (22:24Z) failed in 2 to 4 minutes on crates/recursion/gnark-ffi/build.rs:70, "Failed to build Go library: NotFound" (no go on PC 2; the first run's playbook lost its own log, a bug fixed before the second). floor-build-3 fetched go1.27.1 (tarball sha256 63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445, checked before unpacking under /opt/igneum-floor/go, on the job's PATH only) and built in 240 s (46 crates on the warm target of run 2, 8 niced jobs, 16 cores). The binary: /opt/igneum-floor/bin/sp1-gpu-server, 166,768,224 bytes, sha256 5568108bf7fb9b0e525d8a08926b7046e51136ffaea53f0ca858631d0e938878, --version 6.8.1, cuobjdump --list-elf sm_86, sm_89, sm_120 (the stock 251,306,680-byte server lists sm_80, 86, 89, 90, 100, 120 and compute_120 PTX). The live /root/.sp1/bin/sp1-gpu-server (c2642ad1...) was never touched; the miners mined throughout.

Sweep 1 (job floor-sweep-1, 22:34:56 to 22:37:55Z): the shard term gone, a second floor found

PC 2's RTX 5090 (32,607 MiB, idle 1,755 MiB with the miners stopped and the live prover off), the patched server 5568108b... (sm_86, sm_89, sm_120; the build above), the unpatched pv1 host dae6b006... as client and verifier, one --mode compressed --shard 0 per point, every server killed and its socket unlinked around every point, peak = nvidia-smi memory.used at 1 s (the idle 1,755 MiB inside it), time = the compressed proof. Every proof VERIFIED (1,272,897 bytes, verify 0.037 to 0.040 s), so the unpatched verifier accepts every proof of the patched server.

Config (environment to the patched server) Fixture Cycles Peak MiB Prove s Verified
control: SP1_GPU_MEMORY_BUDGET_GB=32 (upstream's sizes) empty live shard (block 83616) 280,706 13,892 2.2 yes
control v1 shard (fees-v1-shards2 shard 0) 4,717,439 20,516 4.2 yes
12 GB tier: budget 12 (threshold 2^27) empty 280,706 12,740 2.4 yes
12 GB tier v1 shard 4.7 M 15,396 4.1 yes
12 GB tier + SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1 v1 shard 4.7 M 15,428 4.0 yes
16 GB tier: budget 16 (2^27 + 2^26) v1 shard 4.7 M 18,628 3.8 yes
SP1_GPU_ELEMENT_THRESHOLD=67108864 (2^26) empty 280,706 12,772 3.1 yes
2^26 v1 shard (split into 4 core shards) 4.7 M 12,708 5.3 yes
SP1_GPU_ELEMENT_THRESHOLD=33554432 (2^25) v1 shard 4.7 M 12,836 8.5 yes

Reading. The control reproduces the proving agent's curve (13.9 and 20.4 GB), so the patched server behaves as the stock one at the stock sizes. The shard's term follows the threshold as the model says (20.5 GB at 402 M elements, 15.4 at 134 M, 12.7 at 67 M), and then stops: 2^26 and 2^25 both sit at 12.7 to 12.8 GB for the empty shard and the v1 shard alike. The FLOOR memory after setup line names the rest: 9,703 MiB in use before the first shard (2^26; 11,623 at the stock sizes), and the FLOOR tracegen alloc lines at Setup are five allocations of 134,217,728 elements (the recursion keys, 0.75 GB each, each using 90,177,536 elements: 35.6 M preprocessed and 54.5 M main at prove time), one of 33,554,432 (the shrink key, 0.19 GB) and one core key at the threshold. The NORMALIZE_PROGRAM_CACHE_SIZE knob does not reach them (they are keys built at Setup, not the program LRU). The time cost of the split: the v1 shard at 2^26 is 4 core shards and 5.3 s against 4.2 s (1.26x); at 2^27 it is 4.1 s with no split.

So after sweep 1 the binding term is the Setup-time keys allocated at full capacity, and patch v2 sizes every trace buffer (keys and shards) to its padded need: padded_trace_elements in jagged_tracegen/src/lib.rs (each phase pads to the next multiple of 2^21 rows, generate_jagged_traces's "final padding"), applied in setup_tracegen and full_tracegen, one stacking height of slack, SP1_GPU_FLOOR_EXACT=0 restoring upstream.