the project lead, 5 October 2026 22:05 UTC: if the GPU-server patch does not enable 12 GB cards, research proving and see if there are different methods. docs/analysis/proving-methods.md: (1) the memory anatomy of SP1 6.8.1 from its code (the 20 GB panic in sp1-gpu builder.rs, the maximum-sized trace buffers, the mempool that never releases, the fixed recursion shape), the theoretical floor for the adopted shard (about 10 to 11 GB, approximate); (2) the survey: SP1, RISC Zero, Airbender, ZisK, OpenVM, Pico, Ziren, Stwo, Jolt, Ceno, Nexus, Valida, Powdr, Binius, the 2026 entrants, the sumcheck and GKR family, folding, continuations, distributed proving, AMD and Apple backends, every claim cited; (3) per route the change to the guest, aggregator and node verifier, the cost in agent days, the risk, and whether 12 GB under 60 s is reached; (4) the ranking: re-size SP1's server with S_p as the dial and one server per card (no consensus change), RISC Zero as version 2 for the fallback and Apple, the sumcheck family in five years; (5) the tier consequences and the public line. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
66 KiB
Proving methods: why the prover needs 14 GB, what else exists, and how a 12 GB card gets to prove
5 October 2026, from the project lead at 22:05 UTC: "if this doesn't enable 12 GB cards, then do a full deep research task on proving
and see if there are different methods." "This" is the prover-floor agent's patch of SP1's GPU server (branch
prover-floor), running tonight. This document is research and reading, not measurement: every number of ours is from
docs/bench-log.md with its entry named; every claim about another system cites its repository file, its documentation
page or its paper, or is labelled approximate. Status words follow docs/spec/00-overview.md 0.2. Day estimates follow
the project lead's rule of 3 October 2026: hours of agent time, never weeks.
The facts this starts from (bench-log, "proving v1", 5 October 2026; docs/plans/proving-v1.md; docs/analysis/amd-proving.md):
| Fact | Number |
|---|---|
| SP1 6.8.1's GPU server, an empty shard (280,706 cycles), the card to itself | 13,874 MiB peak, 2.2 s compressed |
The adopted v1 shard (S_p 30,000 pgas, 4,717,439 cycles) |
20,434 MiB, 4.3 s; 22,210 MiB and 13.2 s beside the miner |
| The prototype shard (6.75 M pgas, 60.4 M cycles) | 28,307 MiB, 10.8 s; flat at 28.3 GB from 20 M cycles up |
| Aggregation, chained, per block, on a mining 5090 | 9.6 to 9.7 s; 2.2 to 2.5 s with the card to itself |
| The CPU path (PC 1, 16 cores) | 282 s a shard whatever its size, 29.5 to 30.5 GB RSS |
| AMD and Apple GPUs | no zkVM proves on AMD; RISC Zero has a Metal prover, SP1 does not |
| The promise | site/litepaper.html: "Target: shard size will be set so a 12 GB card proves one shard in about 20 seconds"; the design goal is every block proven within about a minute by the miners' own cards |
1. The memory anatomy of a STARK-based zkVM prover, and why the floor is where it is
1.1 What SP1 6.8.1 is
SP1 6.x is not the univariate FRI STARK of the earlier SP1 releases (Succinct calls Hypercube the "first zkVM built entirely on a multilinear polynomial-based proof system", blog.succinct.xyz, sp1-hypercube, 20 May 2025). The crates it pulls say what it is: slop-multilinear, slop-sumcheck,
slop-jagged, slop-stacked, slop-basefold, slop-whir (the ~/.cargo/registry of this Mac; proving/igneum-prove/Cargo.toml
pins sp1-sdk = "=6.8.1"). The architecture Succinct calls Hypercube: the execution trace is a set of multilinear
polynomials over the 31-bit KoalaBear field (sp1-hypercube-6.8.1/src/verifier/config.rs: SP1BasefoldConfig = Poseidon2KoalaBear16BasefoldConfig), the constraints are checked by a zerocheck sumcheck and the lookups by a LogUp GKR
(sp1-gpu/crates/zerocheck, sp1-gpu/crates/logup_gkr), and the polynomial commitment is "jagged": every table's
columns, whatever their heights, are concatenated into one long vector, stacked into rows of height 2^log_stacking_height
and committed with BaseFold, a FRI-like folding over a Reed-Solomon code (sp1-gpu/crates/basefold/src/fri.rs,
slop-basefold-6.8.1/src/verifier.rs). The proof system parameters, from sp1-primitives-6.8.1/src/fri_params.rs and
sp1-prover-6.8.1/src/components.rs:
| Parameter | Value | Where |
|---|---|---|
| Field | KoalaBear, 31 bits, 4 bytes an element; extension degree 4 (16 bytes) | sp1-primitives |
| Core stage: Reed-Solomon blowup | CORE_LOG_BLOWUP = 2, so the codeword is 4x the data |
fri_params.rs:5 |
| Core stage: stacking height, maximum rows per table | CORE_LOG_STACKING_HEIGHT = 21, CORE_MAX_LOG_ROW_COUNT = 22 |
components.rs:16,17 |
| Core shard limits (the executor's cut) | MAX_SHARD_SIZE = 2^24 cycles, ELEMENT_THRESHOLD = 2^28 + 2^27 = 402,653,184 trace elements, HEIGHT_THRESHOLD = 2^22 rows |
sp1-core-executor-6.8.1/src/opts.rs:9-12 |
| Recursion (compress) stage | blowup 2 (RECURSION_LOG_BLOWUP = 2), stacking height 20, max rows 2^21 |
fri_params.rs:6, sp1-verifier-6.8.1/src/compressed/config.rs:1,2 |
| Shrink and wrap stages | blowup 3 (8x), 22 bits of grinding, stacking 18 and 21 | fri_params.rs:17,18,7, components.rs:37-40 |
| Recursion arity | 4 proofs per compose step (DEFAULT_MAX_COMPOSE_ARITY = 4, DEFAULT_MAX_REDUCE_ARITY = 4) |
sp1-prover-6.8.1/src/worker/config.rs:183,193 |
| Workers | 4 core workers, 8 recursion prover workers, 4 recursion executors, 4 deferred workers, buffers of 4 to 8 | worker/config.rs:188-205 |
The stages a shard goes through (sp1-prover-6.8.1/src/worker/controller/*.rs): execute (the RISC-V executor cuts the
run into core shards at the thresholds above); core (one jagged-PCS proof per core shard, on the GPU); normalize and
compose (each core proof is verified inside a recursion program, then proofs are folded 4 at a time until one remains,
the "compressed" proof, 1,272,897 bytes for every shard we have proven, bench-log 4 and 5 October); deferred (what the
aggregator uses: verify_sp1_proof inside a guest, proving/igneum-prove/aggregator/src/main.rs); shrink and wrap
(to a BN254 STARK, then Groth16 or Plonk; not run here, ledger P3).
1.2 The terms, and which scale with the shard
Every STARK-family prover holds these buffers on the device at its peak, in some order and with some overlap. The
sizes below are from the constants of 1.1 and the allocation code of sp1-gpu; where a buffer's size is the actual
trace rather than the maximum, the row says so.
| Term | What it is | Size rule | SP1 6.8.1 on a 32 GB card | Scales with the shard? |
|---|---|---|---|---|
| Main trace | the witness: one element per cell of every table the shard touched | cells x 4 bytes, where cells = sum over tables of rows x columns; the executor cuts a new core shard at 402,653,184 cells |
the device buffer is allocated at the MAXIMUM, not the actual trace: allocate_and_initialize_traces takes max_trace_size and allocates max_trace_size felts plus max_trace_size / 2 u32 of index (sp1-gpu/crates/jagged_tracegen/src/lib.rs:484-503), 6 bytes a cell; the core prover's max_trace_size is element_threshold + 2^21 (prover_components/src/builder.rs:70-71): 2.26 GiB on a card over 30 GB, 1.61 GiB on a 24 GB card (the threshold drops by 2^26 + 2^25 + 2^24 when memory is 30 GB or under, builder.rs:41-45) |
no: fixed at the maximum shard, whatever the trace |
| Preprocessed trace | the program's own tables (the ELF as a Program AIR, the byte and range tables) |
program size x its columns plus 2 x 2^16-class tables | small for a 2.8 MB guest ELF (elf/manifest.json); not isolated |
with the guest, not the shard |
| Codeword (the LDE) | the stacked polynomial encoded at rate 1/4 for BaseFold | stacked cells x 4 (blowup) x 4 bytes; the stacked length is the actual cell count padded to a multiple of 2^21 |
4.7 M cycles: approximate, the actual trace; 60 M cycles: the shard is 7 to 15 core shards of up to 402 M cells, each encoded to 6 GiB at the blowup, one or more in flight | yes, up to the core-shard cap; past the cap the shard count grows and the per-shard term stays |
| Merkle commitment | Poseidon2 hashes of the codeword rows | rows x 8 elements x 4 bytes x 2, rows = 2^21 x blowup |
about 0.5 GiB at full stacking, approximate | with the stacked rows |
| Zerocheck and GKR | the constraint sumcheck over the extension field, and the LogUp GKR layers | extension elements are 16 bytes; the sumcheck holds a folded copy of the trace in the extension field, which is 4x the base trace at the first round and halves each round | up to about 4x the live trace in the first round, approximate (sp1-gpu/crates/zerocheck/src/primitives.rs:174,287: Buffer<Ext> of new_total_length) |
yes |
| Recursion traces | the normalize and compose programs' own traces, verifying core proofs | fixed-shape programs: RECURSION_TRACE_ALLOCATION = 2^27 cells (builder.rs:15), allocated at 6 bytes a cell: 0.75 GiB per recursion prove, at blowup 4 a 2 GiB codeword plus its own zerocheck |
fixed per recursion step; the number of steps is log4 of the core-shard count | no (per step) |
| Shrink and wrap traces | the two last stages, not run by us | 2^25 and 85,376,340 cells (builder.rs:16,19): 0.19 and 0.48 GiB |
only when wrapping | no |
| Proving keys and program cache | the recursion programs (vk_map.bin, the normalize cache of 5 programs) and the shard program's setup |
DEFAULT_NORMALIZE_PROGRAM_CACHE_SIZE = 5 (worker/config.rs:192); the key setup took 14.6 s on the 5090 (bench-log 4 October) |
not isolated | no |
| Pinned host buffers | the staging copies on the PC side | 4 core workers x max_trace_size x 4 bytes = 6.0 GiB of pinned RAM, plus 4 x 0.5 GiB for recursion (prover_components/src/components.rs:99-103, builder.rs:76,95) |
this is the 7.9 GB WSL2 working set measured on 5 October | no |
| The allocator | CUDA's default memory pool with its release threshold set to u64::MAX (sp1-gpu/crates/cuda/src/task.rs:152,190-199): freed blocks are never returned to the driver |
nvidia-smi therefore reports the high-water mark of everything above, and it stays until the server exits |
this is why the memory curve is flat between shards of different size | no |
Two facts from this table explain the measurements:
- The server refuses small cards by code.
local_gpu_opts()reads the card's total memory, adds 4 GB, and panics under 24:"Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB"(sp1-gpu/crates/prover_components/src/builder.rs:35-38). A 16 GB card (16 + 4 = 20) and a 12 GB card (16) never start; a 20 GB card is the smallest that does. The 13.9 GB floor measured on the 5090 is therefore not the whole story for a 12 GB card: on this build the card is refused before any buffer is allocated. Any route through SP1's GPU server starts by removing this line. - The environment knobs do not reach the floor because the same function overwrites
element_thresholdwith the compile-time constant (builder.rs:41-48); onlyHEIGHT_THRESHOLDpasses through, which is why the sweep'sELEMENT_THRESHOLD 2^26rows changed nothing andHEIGHT_THRESHOLD 2^20took 5.4 GB off the 60 M-cycle shard (bench-log, "the 12 GB requirement", 5 October 2026). The worker counts only slow the proof (11.4 s to 20.8 s) because the device buffers are sized bymax_trace_size, not by the worker count.
1.3 Why the floor is 13.9 GB for an empty shard
With the release threshold at u64::MAX, the peak is the high-water mark over the whole pipeline. For an empty shard
the core trace is small (280,706 cycles), so the fixed-shape terms dominate: the 2.26 GiB main-trace buffer allocated
at the maximum, the recursion step over a fixed-shape normalize program (a 0.75 GiB trace buffer, its 4x codeword in
the extension field for the zerocheck, its Merkle tree), the proving-key and program caches, and the deferred and
compose machinery that a compressed proof always runs once. The decomposition of the 13.9 GB into those terms is
approximate until the prover-floor agent's profile lands (branch prover-floor, tonight): the figure that is not
approximate is that none of it is the witness (5 to 22 KB a shard) and none of it is the shard's cycles (the same
13.9 GB at 280 k cycles and 556 k cycles, bench-log "the S_p curve").
The step from 13.9 GB (empty) to 20.4 GB (4.7 M cycles) is the live trace: the 4.7 M-cycle shard is one core shard
(its trace area is under 402 M cells, so it was not split; approximate from the memory curve, the cell count is not
logged by the host), and its codeword, zerocheck and GKR buffers are sized by its actual cells. The step from 20.4 GB
to 28.3 GB (20 M cycles and up) is the second and later core shards in flight at once: 4 core workers with a buffer of
4 (worker/config.rs:188,189) let several core shards' codewords exist at the same time; past 20 M cycles the pipeline
is full and the peak is flat, which is what the curve shows (28,371 MiB at 20 M cycles, 28,307 at 40 M and 60 M).
1.4 The theoretical floor for our guest at the adopted shard
If every buffer were sized to the shard rather than to the maximum, the adopted v1 shard (4.7 M cycles) would need, approximate, from the rules of 1.2:
| Term | Rule | Approximate bytes |
|---|---|---|
| Main trace, actual | 4.7 M cycles x about 60 cells a cycle (the Add and Addi tables cost 33 and 30 columns a row, a memory access adds 20 and a global interaction 241: sp1-core-executor-6.8.1/src/artifacts/rv64im_costs.json) |
about 280 M cells, 1.1 GB |
| Codeword at blowup 4 | 4x | 4.5 GB |
| Zerocheck first round in the extension field | 4x base, halving each round | 4.5 GB at the peak round, falling |
| Merkle tree | rows x 32 bytes x 2 | 0.3 GB |
| Recursion step, fixed | 2^27 cells x 6 bytes plus its 4x codeword and extension copies | 2 to 3 GB, approximate |
| Keys and caches | under 1 GB, approximate | |
| Peak, if the core stage and the recursion stage do not overlap and the pool releases | about 10 to 11 GB; about 6 GB if the blowup-4 codeword is replaced by a rate the sumcheck does not need (see 2.4) |
So the adopted shard is, on paper, a 12 GB card's shard with the server re-sized and nothing else changed, and it is a
12 GB card's shard with 4 GB to spare if the shard is halved (S_p 15,000 pgas, 2.4 M cycles: the planner cuts at
transaction boundaries to any budget, core/src/plan.rs, and the fee switch of 5 October already moved S_p once).
What the paper figure does not say is the time: a smaller card proves slower, and 60 s with the miner running is the
bound (section 3). The prover-floor agent is measuring the real figure; this section says what it should find and why.
1.5 RISC Zero's anatomy, for comparison
RISC Zero is the FRI STARK the textbooks describe, and its constants make the same table easy to read
(~/.cargo/registry, risc0-zkp-3.0.4/src/lib.rs, risc0-circuit-rv32im-4.0.4/src/zirgen/defs.rs.inc):
| Term | Value | Where |
|---|---|---|
| Field | BabyBear, 31 bits; extension degree 4 | risc0-core |
| Segment size | DEFAULT_SEGMENT_LIMIT_PO2 = 20 (1,048,576 cycles), MIN_CYCLES_PO2 = 13, MAX_CYCLES_PO2 = 24; DEFAULT_MAX_PO2 = 22 for the verifier |
risc0-circuit-rv32im-4.0.4/src/execute/mod.rs:39, risc0-zkp-3.0.4/src/lib.rs:35-38, risc0-zkvm-3.0.4/src/receipt.rs:884 |
| Trace width | data 211 + accum 103 + code 1 = 315 columns; globals 90, mix 36 | defs.rs.inc:7-11 |
| Blowup | INV_RATE = 4; 50 queries; FRI fold 16 |
risc0-zkp-3.0.4/src/lib.rs:41-51 |
| Recursion | lift, join and resolve programs at RECURSION_PO2 = 18 rows |
risc0-zkvm-3.0.4/src/host/recursion/prove/mod.rs:58 |
| The GPU buffers | check + ctrl + data + accum + mix + out elements x 4 bytes, printed by the CUDA HAL at eval_check |
risc0-circuit-rv32im-4.0.4/src/prove/hal/cuda.rs:181-200 |
From those constants the trace of a segment is 2^po2 x 315 x 4 bytes and its LDE 4x that, so, approximate: po2 18 is
0.3 GB of trace and 1.5 GB with the LDE, po2 19 is 3.1 GB, po2 20 is 6.2 GB, po2 21 is 12.3 GB, before the check
polynomial, the extension-field accumulators and the Merkle trees. Two things follow. A RISC Zero segment at the
default 2^20 is in the same class as one SP1 core shard, not smaller. And a RISC Zero segment at 2^18 or 2^19 is a
2 to 4 GB object: the only reason a 12 GB card could not prove one is the fixed overhead of the recursion circuits
(2^18 rows each) and the allocator, which is the measurement the prover-floor agent takes on PC 2 if SP1 cannot go
under 11 GB. Section 2.2 carries the documented numbers.
1.6 Which terms the shard size can move, and which it cannot
| Lever | Moves | Does not move |
|---|---|---|
Our S_p (pgas per shard) |
the live trace, the codeword, the zerocheck: everything in 1.2 marked "yes" | the maximum-sized buffers, the recursion step, the keys, the pool |
SP1's HEIGHT_THRESHOLD (the one knob the server honours) |
the rows per table in one core shard, so the live buffers | the fixed terms (measured: 13,861 MiB on the empty shard with every knob at its minimum) |
A server patch: size max_trace_size to the shard, release the pool, one core worker |
the 2.26 GiB buffer, the high-water mark, the in-flight count | the recursion step's fixed shape and the key caches |
| A different proof system | the blowup (sumcheck-only and linear-code systems have none, 2.4), the recursion shape | the trace itself: a RISC-V cycle costs tens of cells in every zkVM |
2. Every current proving route
Read 5 October 2026, 22:10 to 23:00 UTC, by four research agents and this one; every cell names its page or file. "Not documented" means the project publishes no figure, which for a memory floor is itself the finding.
2.1 The zkVMs with a GPU prover
| Prover | Proof system, field, chunk | GPU support and the documented minimum memory | Throughput, on what | Verification of the recursive proof; wrapper | Licence | State, October 2026 |
|---|---|---|---|---|---|---|
| SP1 6.8.1 (ours) | Hypercube: multilinear, jagged PCS, BaseFold, LogUp GKR; KoalaBear; core shards of up to 2^24 cycles and 402 M cells (section 1) | CUDA only. Docs: "24GB or more VRAM", compute capability 8.0+, Linux x86_64 (docs.succinct.xyz, hardware-acceleration page). Code: panic under 20 GB physical (builder.rs:35-38). Measured here: 13.9 GB floor, 20.4 GB at the adopted shard, 28.3 GB at the prototype shard. Issue #2950: two clients on a 48 GB L40S hold 41 to 43 GB; a single 6 GiB tensor allocation failed |
4.3 s for the adopted shard, 10.8 s for the prototype one on a 5090 (bench-log); Succinct: 99.7% of Ethereum blocks under 12 s on 16 x RTX 5090 (blog.succinct.xyz, 18 Nov 2025) | compressed proof 1,272,897 bytes, verified in 0.032 to 0.040 s here (--mode verify-segment); Groth16 about 260 bytes and about 270 k gas, Plonk about 868 bytes and 300 k gas (docs, proof-types page); the Groth16 wrap needs about 14 GB of host RAM, Plonk about 60 GB (hardware-requirements page) |
Apache-2.0 or MIT for the repository including sp1-gpu/ (LICENSE-APACHE, LICENSE-MIT at the root; no separate licence under sp1-gpu/); sp1-cluster is Business Source 1.1 |
v6.8.1 of 24 Sep 2026 is the latest tag; mainnet for Ethereum proving since 19 Feb 2026 (blog); AMD port PR #2668 closed unmerged 20 Mar 2026; no Metal, Vulkan or WebGPU |
| RISC Zero 3.0.x | FRI STARK (DEEP-ALI), BabyBear, Poseidon2, blowup 4, 50 queries; segments of 2^po2 cycles, default po2 20, allowed 13 to 24 (section 1.5); lift, join, resolve recursion at 2^18 rows; keccak as a separate circuit |
CUDA and Metal (risc0/sys/kernels/zkp/metal/; on Apple silicon the Metal path is on automatically, risc0/zkvm/build.rs). Documented memory per segment: Bento design page, 1 M cycles 9 to 10 GB, 2 M 17 to 18 GB, 4 M 32 to 34 GB; Boundless performance page, the largest SEGMENT_SIZE per card: 8 GB card po2 19, 16 GB po2 20, 20 GB po2 21, 40 GB po2 22; docs: "less than 10 GB available: change the segment size limit" (dev.risczero.com, local proving); env.rs:190-192: "lowering this value by 1 will cut memory consumption by about half". PR #3761 (June 2026): po2 22 did not fit a 24 GB 4090 until the low_vram buffer reuse |
4090: 808 kHz at po2 21, 1,207 kHz at po2 22 with PR #3761 (end to end to a succinct receipt); Apple M2 Pro about 14 kHz on the 2023 datasheet, approximate (the page was unreachable tonight); real-time Ethereum on about 160 x 4090 (blog, approximate) | succinct receipt 222,668 bytes, constant; about 100 ms to verify, approximate (gsr-stark-verifier PR #5, mirrored docs); Groth16 seal 256 bytes, about 200 to 300 k gas, approximate; the Groth16 wrapper is x86 only, not on Apple silicon (docs) |
Apache-2.0 or MIT, CUDA and Metal kernels included (risc0/sys/kernels/zkp/cuda/eltwise.cu:1-13); Bento is BSL 1.1 with a change date already passed |
v3.0.6 of 17 Jul 2026 on the maintained line; main is 5.0.0 with no release body; RISC0_PROVER=actor multi-GPU scheduler experimental since 3.0.1 (r0vm/src/actors/factory.rs:183-195 carries measured per-po2 memory tokens: po2 18 = 8, 19 = 10, 20 = 15, 21 = 24; lift and join = 3) |
| Airbender (Matter Labs) | DEEP STARK, FRI, Mersenne31; chunks of 2^22 cycles; Boojum then FFLONK wrap (docs.zksync.io, airbender page) | CUDA only. "any GPU with 22GB RAM" for production (zksync.io/airbender); the code has memory presets GiB21 (24 GB cards) and GiB30, raised from 29 because Ethereum blocks failed to allocate at 29 GiB (PR #448, gpu/execution_prover/src/prover/config.rs); the final SNARK is CPU with about 150 GB of RAM (docs/gpu.md) |
one H100: 21.8 MHz base layer, 8.5 MHz end to end, about 35 s an Ethereum block (June 2025 post); ethproofs.org today: 4 x 5090 2.3 s average | FFLONK over BN254 on chain; gas not published | MIT or Apache-2.0 | v0.6.0-rc.2; Veridise audit Feb to Apr 2026; live for ZKsync Atlas chains |
| ZisK (Polygon spin-out) | eSTARK over Goldilocks (pil2-stark), Poseidon2, approximate; main instance 2^22 to 2^23 rows, chunks of up to 2^22 steps (PR #1238) | CUDA only, CUDA 12.9+; no VRAM floor documented; workers need about 32 GB of host RAM, the assembly emulator 64 GB (docs, limits and distributed pages); Cysic's Venus fork submits from one RTX 4090 | 4 x 5090: p99 9.62 s on Ethereum blocks (Aug 2026); 24 x 5090 6.56 s average (Nov 2025) | PLONK wrapper verified by Solidity (zisk-contracts); 128-bit claimed |
Apache-2.0 or MIT | v1.3.1-alpha, 30 Sep 2026, "undergoing security and correctness audits" (README) |
| OpenVM 2.0 (Axiom) | SWIRL: sumcheck, zerocheck, LogUp GKR, stacked reduction into WHIR; BabyBear; segments by metered trace height (blog.openvm.dev/2.0) | CUDA only; "at least 24GB of VRAM": L40, 4090, L40S, 5090 (blog.openvm.dev/openvm-gpu) | 11.4 MHz on one 5090, 139 MHz on 16; 2.1 preview: 4 x 5090 p99 9.7 s | STARK proof under 300 KB; Halo2-KZG wrapper, 316 k gas; the Halo2 wrap 8.1 s on a 5090 | MIT or Apache-2.0, GPU prover included | v2.0.2 of 14 Aug 2026; zkSecurity audit of SWIRL; Scroll's prover builds on it |
| Pico (Brevis) | Plonky3 STARK, KoalaBear default; chunk size a parameter with no documented default | CUDA via pico-gpu; no VRAM figure published; every run on 32 GB 5090s |
Prism 2.1: 16 x 5090 over two machines, 4.87 s average on Ethereum blocks | Groth16 via gnark | core MIT or Apache; pico-gpu is BUSL-1.1 and "not recommended for production" (its README) |
v2.1.2, Aug 2026; Sherlock audit |
| Ziren (ZKM, MIPS) | Plonky3-class, KoalaBear, LogUp GKR, WHIR | CUDA 12, compute capability 8.6+, "24 GB VRAM or higher"; the GPU prover is a Docker image pinned by digest, source "planned H1 2026" (docs.zkm.io prover page; an independent evaluation of v1.1.4 says the GPU path is not open) | one 5090: 5.9 MHz on a 288 M-cycle block; 4 GPUs 3.1 to 3.3x | compressed proof 603 KiB; Groth16 or PLONK | core MIT or Apache; GPU image licence unspecified | v1.2.7; no public audit cited |
| Stwo / S-two (StarkWare) | Circle STARK over Mersenne31; blowup 1 (rate 1/2), 70 queries, 26 bits of grinding | CPU SIMD first (AVX2, AVX-512, NEON, WASM); GPU through ICICLE-Stwo (Ingonyama): about 3 GB of trace in GPU memory, out of memory from 2^23 rows; a WebGPU port of the constraint evaluation (zkSecurity blog, April 2025); client-side proving under 1 GB after a spill allocator (third-party PR) | 620 k Poseidon2 a second on an M3 laptop | via a Cairo verifier (proofs of proofs); sizes not published here | Apache-2.0 | live on Starknet mainnet since 3 Nov 2025; no RISC-V guest of its own (Nexus 3.0 is the RISC-V zkVM on it, BUSL-1.1 until 2029) |
| Jolt (a16z) | sumcheck and lookups (Lasso lineage, Twist and Shout memory checking); PCS Dory over BN254 by default, or Akita, a lattice commitment over a 128-bit prime field (Sep 2026, "Lattice Jolt"); RV64IMAC; no continuations (the book's recursion page is "under construction") | No CUDA in the public repository (LayerZero's "Jolt Pro" CUDA port is private); Metal: PR #1938 merged 30 Sep 2026 (the jolt-metal runtime crate), PR #1733 (the full prover on Metal) still a draft. Memory: "about 200 bytes per cycle" with Akita (a16z substack, Sep 2026); the book: "under 2 GB of memory per million cycles"; a streaming prover bounded to "a few GBs" is planned, not shipped |
over 2 M cycles a second on a laptop CPU with Akita, over 10 M with Metal on a MacBook (a16z substack, Sep 2026); PR #1733: M5 Max, 2^25 cycles in 19.8 s, 2^27 in 77 s at an 89.4 GiB footprint | proof about 50 KB (Dory) or 65 to 80 KB (Akita); verify sub-second, approximate; on-chain 1.3 to 2 M gas estimated in 2024; no Groth16 wrapper shipped | MIT or Apache-2.0 | v0.3.0-alpha is the last tag (1 Oct 2025); README: "not suitable for production use"; no audit |
| Ceno (Scroll) | GKR tower prover, BabyBear, WHIR or BaseFold PCS; RV32IM | CUDA, but the real HAL is in a private ceno-gpu repository (the public one is a mock); no memory numbers |
2 GPUs 1.6x over one (PR #1403, Sep 2026) | via OpenVM recursion to Halo2 | Apache-2.0 | README: "under construction and not suitable for use in production" |
| Nexus 3.0 | on Stwo (Circle STARK, M31) | no GPU path documented | none published | not published | BUSL-1.1 until 10 Feb 2029 | last push 6 Jan 2026; folding (Nova family) abandoned June 2025 for the STARK |
| Valida (Lita) | Plonky3 STARK | no GPU; a CUDA port "underway" in July 2025 | none current | not published | Apache or MIT | dormant since Sep 2025; documented soundness issues in its own benchmarks page |
| Powdr | no longer a zkVM: powdrVM archived; powdr is autoprecompiles on OpenVM |
OpenVM's | OpenVM's | OpenVM's | MIT or Apache | tooling layer; "DO NOT USE FOR PRODUCTION" |
| Binius / Binius64 (Irreducible) | binary-field SNARK, BaseFold-style FRI over GF(2^64) words | CPU SIMD only; the FPGA work was dropped 9 Sep 2025 ("FPGAs underperformed GPUs"); no GPU | ECDSA aggregation about 5x over SP1 and R0VM on L40S GPUs, on CPU (the page carries methodology corrections) | hash-based; no recursion shipped | Apache-2.0 or MIT | the company shut down 12 Nov 2025; the only zkVM on it (PetraVM) is archived |
| 2026 entrants | Cysic Venus (a ZisK fork with cudaGraph tuning and an FPGA backend, Apache or MIT, "do not use in production"); Zilkworm (Erigon's C++ guest on Airbender, 2 x 5090 9.3 s); zkDTVM (evmone guest, 4 x 5090 4.7 s, no public docs); Delphinus zkWasm (Halo2 on BN254, a 4090 minimum plus 58 GB of host RAM); Miden (Goldilocks STARK, client-side, Metal via miden-gpu, mainnet alpha planned); Boojum (2023 claim of proving on a 16 GB card, superseded by Airbender) |
none states a floor under 24 GB on a GPU |
The reading of the table. No shipped zkVM documents a GPU floor under 24 GB except RISC Zero, whose memory is a
function of a runtime knob (segment_limit_po2) and is published per card size by Boundless. SP1's 24 GB is a line
of code, not a property of the proof system: Airbender, OpenVM and Pico all pad to the card they tune on, and all
three say 24 or 32 GB because their market is Ethereum blocks on 5090 clusters. The real-time race has collapsed to 2
to 4 consumer cards per block (ethproofs.org, 5 October 2026), which is why nobody is tuning for a 12 GB card: the
customer buys 5090s. Igneum's customer is the miner who already owns the card, so Igneum has to do the tuning itself.
2.2 The sumcheck and GKR family against FRI STARKs, in memory terms
| Family | What it holds at the peak | Blowup | The GPU figure today | Source |
|---|---|---|---|---|
| FRI STARK (RISC Zero, Airbender, ZisK, Pico, Stwo, SP1 3 and 4) | trace, its Reed-Solomon codeword at the blowup, the Merkle trees, the DEEP quotient in the extension field | 4x (RISC Zero, Airbender), 2x (SP1 3 and 4, approximate), 2x (Stwo at rate 1/2) | RISC Zero: 9 to 10 GB per 1 M cycles (Bento) | section 1.5; Boundless pages |
| Sumcheck with a hash-based PCS (SP1 Hypercube, OpenVM SWIRL, Ceno, Ziren) | the trace as multilinears, the extension-field folded copies of the zerocheck and GKR, and the BaseFold or WHIR codeword of the stacked polynomial (still a Reed-Solomon encoding, at 4x in SP1, section 1.1) | 4x of the stacked data in SP1; WHIR's rate is a parameter | SP1: the fixed 13.9 GB plus about 6.5 GB for a 4.7 M-cycle shard (measured) | section 1 |
| Sumcheck with a curve or lattice PCS (Jolt) | the trace and the one-hot columns; no codeword at all: Dory commits by MSM and Akita by lattice hashing, so memory is bytes per cycle with no blowup | none | no GPU figure: 200 bytes a cycle on CPU (Akita), so the adopted 4.7 M-cycle shard is about 0.9 GB of prover RAM, approximate (derived) | a16z substack, Sep 2026; the Jolt book, streaming page |
| GKR (Expander, Ceno) | the circuit witness layer by layer; no codeword for the inner layers | none inside; a PCS for the inputs | Expander: 16 MB per Keccak, approximate | Polyhedra blog (returned 530 tonight) |
| Linear-code PCS (Ligero, Brakedown, Ligerito, Blaze) | one encoded matrix and one Merkle tree; linear time, no FFT | rate 1/2 to 1/4 | no prover memory benchmarks found; Linea's Vortex is the only production use | eprint 2021/1043, 2025/1187, 2024/1609; linea-monorepo/prover/protocol/compiler/vortex |
| Binius (binary field) | words of GF(2^64) and a BaseFold FRI | 2x to 4x | none; CPU only; company closed | irreducible.com posts |
The memory law in one line: a FRI or BaseFold prover holds blowup x trace plus the trace itself plus extension-field
working copies, so 8 to 12 bytes per cell at the peak; a Jolt-class prover holds the trace and its lookups at about 4
bytes per cell and commits without encoding. The figure that matters for us is not the ratio but the absolute: our
adopted shard is small enough (about 280 M cells, section 1.4) that a FRI-class prover sized to it fits a 12 GB card,
and a Jolt-class one fits a phone. The reason SP1 does not fit today is section 1.3, not the proof system.
2.3 Folding schemes
| Scheme | Prover memory per step | The verifier at the end | Field | GPU | Fit for Igneum |
|---|---|---|---|---|---|
| Nova, SuperNova, HyperNova, ProtoStar, Mova; Sonobe as the library | one step's witness plus the running instance: tiny by construction (eprint 2021/370) | an IVC proof of O(F) group elements, compressed by a SNARK: Sonobe's decider is Groth16 over BN254 with KZG, about 11.9 M constraints for a 500 k-constraint step (sonobe.pse.dev, decider page); MicroNova about 2.2 M gas (eprint 2024/2099) | curve cycles (Pasta, BN254 and Grumpkin) | partial: sppark MSM on Pasta, a GPL-3 cuda-nova for BN254; Sonobe lists GPU as a plan |
no: a RISC-V step over a 256-bit curve cycle costs two MSMs per step, the opposite of the hash-based consumer-card design decision (design 5.6), and the verifier changes to pairings |
| Nexus zkVM 1 and 2 | the prover ran "on as little as 1 GB of RAM" (whitepaper, approximate) | a curve SNARK | curve cycle | none | abandoned by its own author: Nexus 3.0 (25 June 2025) moved to a Circle STARK, "proofs are smaller, faster to generate" (StarkWare blog) |
| LatticeFold, LatticeFold+, Neo, SuperNeo; Nightstream as the zkVM | one step plus an accumulator, lattice commitments over 64-bit fields | a Spartan-class decider | Goldilocks named; BabyBear and KoalaBear not | none; Nethermind's LatticeFold is a "proof-of-concept prototype" whose benches take 48 h; Nightstream is "research software, not production-ready" with its RV32IM prototype removed | not before 2027 at the earliest; the first candidate that folds small-field STARK steps |
| Arc, WARP (hash-based accumulation of Reed-Solomon proximity claims) | small: Merkle openings per step (eprint 2024/1731, 2025/753) | a FRI-style accumulator check | any STARK field | none; no public implementation found | the right primitive on paper for folding RISC-V STARK shards with a hash-based verifier; nothing to adopt |
| Mangrove, Nebula | 390 MB peak at 2^24 gates (Mangrove, eprint 2024/416); pay-per-use steps (Nebula) | curve SNARK | curve cycle | none | research |
What folding would mean for a shard: the shard prover would hold one transaction's step at a time and the memory floor would vanish; the price is a curve-based decider at the end of every shard (seconds on a CPU, a different verifier in the node, pairings on the light-client path), and no production code over our field. Today's small-field zkVMs get their bounded memory from segmenting and recursion (2.4), not from folding. Folding is a watch item, not a route.
2.4 Continuations and segment proving at small sizes
| Prover | The segment knob | What a 2^18 or 2^19 segment costs | Can our shard be cut that way inside the guest? |
|---|---|---|---|
| RISC Zero | segment_limit_po2, runtime, 13 to 24 (env.rs:181-186); the recursion lifts every segment at 2^18 rows and joins them in a tree |
po2 19 is the documented fit for an 8 GB card and po2 20 for a 16 GB card (Boundless); the scheduler's measured tokens put po2 18 at about a third of po2 21 and a lift or join at an eighth (factory.rs); the 4.7 M-cycle shard at po2 19 is 9 segments, 9 lifts and 8 joins |
yes, with no guest change: the zkVM cuts at the limit on its own; the shard statement is unchanged and one succinct receipt comes out |
| SP1 | HEIGHT_THRESHOLD (honoured) and ELEMENT_THRESHOLD (overwritten by the server, section 1.2); SHARD_SIZE up to 2^24 cycles |
the live buffers shrink (22.9 GB against 28.3 GB on the prototype shard at HEIGHT_THRESHOLD 2^20, bench-log) and the fixed 13.9 GB does not; the compose tree folds 4 proofs at a time |
yes, the same way; but the floor is the server's, so the cut buys nothing until the server is re-sized (route A) |
| Our own planner | S_p in pgas, a consensus parameter changed by the fee-switch pattern (docs/plans/fee-switch-devnet.md); the cut is at transaction boundaries (core/src/plan.rs) |
halving S_p halves the live trace and doubles the shard count; the aggregator verifies one deferred proof per shard (1.66 M cycles for 4 shards, bench-log 4 October) and the chained aggregation is one per block whatever the count (9.7 s on a mining 5090) |
yes, already implemented; a transaction above S_p stays one shard and the zkVM's own continuations cover it (spec 7.6 item 1) |
| Jolt | none: monolithic; streaming planned | n/a | no |
2.5 Distributed proving across several small cards
| System | How one execution is split | Per-card memory | Several cards on one host | What it means for four 12 GB cards |
|---|---|---|---|---|
SP1 cluster (sp1-cluster, BSL 1.1) |
by core shard: ProveShard, RecursionReduce, RecursionDeferred and ShrinkWrap tasks go to GPU workers, CoreExecute and the Groth16 or Plonk wrap to CPU workers (crates/prover-types/src/lib.rs:31-41); artifacts through Redis and S3 |
"only certain GPUs with >= 24GB RAM are supported" (infra/charts/sp1-cluster/values-example.yaml); one task holds one whole card |
yes: one GPU node process per card (gpu{0..7} services in the docker-compose deployment page); the local server itself supports device 0 only (task.rs:160, "only device 0 is supported at the moment"), one server per CUDA_VISIBLE_DEVICES |
the split is by core shard, and the adopted shard is ONE core shard (section 1.3), so there is nothing to split across cards; the floor per card is unchanged. The cluster is a throughput tool, and its code is BSL |
| RISC Zero Bento (Boundless) | by segment onto Redis; gpu_prove_agent spawns one prove agent per card with CUDA_VISIBLE_DEVICES; the same SEGMENT_SIZE for every card, "the lowest common denominator"; joins form a tree (docs.boundless.network, performance-optimization and bento pages; compose.yml:62-113) |
by SEGMENT_SIZE (2.1): 8 GB po2 19, 16 GB po2 20 |
yes, documented: one 16 GB card 264 kHz, two 431 kHz (sub-linear, "bound by bus bandwidth, memory") | the one documented configuration: four 12 GB cards at po2 19 or 20 take segments off one queue and the joins fold them; the per-card floor is the segment, and the cost of small segments is the lift and join count (9 lifts and 8 joins for the adopted shard at po2 19) |
RISC Zero RISC0_PROVER=actor |
one process, several cards, a token budget per card from measured memory per po2 (r0vm/src/actors/factory.rs) |
per po2 | yes, experimental since 3.0.1 | the same model without Bento's services |
| Pico Prism 2.0 | a global task queue across two machines, 16 x 5090, 100 Gbps between them (Brevis blog, May 2026) | not published | yes | no figure |
| OpenVM | metered execution on the CPU, segments to GPUs, an aggregation tree, "clusters with hundreds of GPUs" (docs, distributed-proving page) | 24 GB | yes | no 12 GB path |
| ZisK | coordinator and stateless workers; "splits the trace into pieces, proves each in parallel on separate machines, and aggregates"; the first worker aggregates a binary tree (docs, distributed execution page) | not documented | yes, --gpu per worker |
no figure |
| Ceno | shards round-robin by shard_id % device_count; a shard never split across cards; one CUDA context per device (PR #1403) |
not stated | yes | the same model |
| Column-split of one trace across cards (FRIttata eprint 2025/1285, HyperFond 2025/1349, deVirgo arXiv 2210.00264, Pianist 2023/1271, Cirrus 2024/1873, SumFold 2025/1653) | the sumcheck or FRI itself is distributed, each worker holding a slice of the columns or rows and exchanging small messages | a slice | research code or CPU clusters only | nothing shipped; the one route that would let four 12 GB cards hold what one 32 GB card holds for a SINGLE core shard, and nobody has it in a zkVM |
The reading. Every shipping system splits by rows (segments, shards, chunks), proves each on one card, and folds
with recursion. So "four 12 GB cards do what one 32 GB card does" is true for throughput (four shards in flight, or
four segments of one shard, then a join tree) and false for a single unit that exceeds one card: that unit must be cut
smaller, by the zkVM's segment knob (RISC Zero) or by our planner (S_p). For Igneum the units are already small and
independent (a shard, assigned by sortition), so the rig's natural mode is one prover process per card, each taking
its own shard. The distributed route therefore costs nothing in protocol and lands as an app change (section 3, route C).
2.6 Proof systems that run on AMD or Apple
| Target | What exists | Status | Source |
|---|---|---|---|
| Apple, RISC Zero Metal | the full STARK prover (rv32im, keccak, recursion) on Metal, automatic on Apple silicon; the Groth16 wrap x86 only | shipped, maintained (PR #3761's June 2026 matrix lists "metal (Mac M-series): build, run"); the only speed published is a 2023 M2 datasheet (14 to 93 kHz, approximate); nothing for M3, M4 or M5 | risc0/sys/kernels/zkp/metal/, dev.risczero.com local-proving page |
| Apple, ICICLE Metal (Ingonyama) | MSM, NTT, sumcheck on Metal since v3.6.0 (Mar 2025); "missing API implementations for Poseidon and Poseidon2 hashes, Merkle tree" at that release; v4.0.0 of 11 Jul 2025 is the latest | a library, closed-source backends under a free research licence (dev.ingonyama.com, install_gpu_backend page); no STARK prover built on it for Metal | ICICLE releases, the Metal blog |
| Apple, Jolt Metal | PR #1938 merged 30 Sep 2026 (the runtime and field kernels); PR #1733, the prover itself, a draft: M5 Max 2^25 cycles in 19.8 s, 3.2x over its CPU | the fastest Apple number anyone has published, in a draft | github.com/a16z/jolt pulls 1733 and 1938 |
| Apple, Stwo | CPU SIMD with NEON; ICICLE-Stwo promises Metal | CPU path shipped; no RISC-V guest of its own | stwo README, Ingonyama blog |
| Apple, Miden | miden-gpu on Metal |
Cairo-class VM, not RISC-V | hackmd (bobbinth) |
| AMD, sppark | "A limited support for AMD's RDNA and CDNA GPUs" (README); SP1's tree carries no HIP build | a library | github.com/supranational/sppark |
| AMD, SP1 PR #2668 | an external port to RDNA3 and RDNA4 with "a caching memory allocator to work around hipMallocAsync leak bug" | closed unmerged 20 Mar 2026 | github.com/succinctlabs/sp1/pull/2668 |
| AMD, OpenVM stark-backend HIP fork | cuda2hip.hpp so the same .cu builds under nvcc and hipcc, native mont32_t.hip, tested on gfx1100, targets MI300X and 7900 XTX |
merged 15 Sep 2026 in a fork (Okm165/stark-backend PR #2), not upstream | the PR |
| AMD, Goldilocks NTT and STARK on ROCm | 19.19 ms NTT at 2^27 on an RX 7900 XTX; a Goldilocks STARK backend on HIP | research posts | ethresear.ch, qingming-g64-ntt and stark-g64 |
| Vulkan and WebGPU | ICICLE's Vulkan build (Jan 2025) with no installable backend; zkSecurity's WebGPU Stwo (5x on constraint evaluation, 2x end to end, no 64-bit integers in WGSL); ZPrize WebGPU MSM | prototypes; nothing proves a RISC-V shard | the pages named |
Said plainly, as docs/analysis/amd-proving.md said it: on 5 October 2026 no zkVM proves on an AMD GPU, and the only
Apple prover that ships is RISC Zero's. The AMD work that exists is two ports of CUDA STARK kernels through a HIP shim,
one closed, one in a fork; both are days of agent work to revive against a given tree, and PC 1's RX 9070 XT (gfx1201)
is the card to measure on.
3. For each route: the change to our guest, the aggregator and the node's verifier; the cost; the risk; 12 GB under 60 s
What the node verifies today: SP1 compressed proofs through igneum-prove-host --mode verify and verify-segment
(vendor/igneum-node-pv1/igneum/exec/src/proving.rs:41-46, 878-926), the pinned ids read at start and named in the
native statement (program_ids, IGNEUM_PROOF_PROGRAM_IDS), the record bound in a BLS-signed ProofRecord (version
- or
SegmentRecord(version 2) with the proof's SHA-256 (spec 7.7 item 1, 7.8 item 3). A different proof system means a newProofSystemversion (design 5.6), a new pinned id, a second verifier command, and the record's version field telling the node which. The swap procedure of design 5.6 (test vectors, 90% signalling, a 3-month overlap with both verifiers, a wrap of the last old proof) is the path for any of the rows below that change the family.
The 60-s test. The litepaper's minute, the launch target of 20 to 60 s behind the tip, and the mine-and-prove
measurement that a shared card proves 3 to 4x slower (bench-log, chain-pc2-pv1c). No 12 GB card has run any
prover in this repository; the 12 GB times below are approximate, scaled from the 5090 by memory bandwidth (an RTX
3060 at 360 GB/s and an RTX 4070 at 504 GB/s against the 5090's 1,792 GB/s, NVIDIA's published figures, approximate),
which is the term a STARK prover is bound by. They are the numbers the first 3060-class run replaces.
| Route | Guest | Aggregator | Node verifier and record | Cost (agent time) | Risk | Reaches 12 GB with a real shard under 60 s? |
|---|---|---|---|---|---|---|
A. Re-size SP1's GPU server (the prover-floor agent's patch, running tonight): remove the 20 GB panic (builder.rs:37), size max_trace_size to the shard (honour ELEMENT_THRESHOLD, or set the core allocation from the shard's measured cells), one core worker and a buffer of 1, a release threshold so the pool returns memory between stages, drop_ldes on; build with CUDA_ARCHS for Ampere, Ada and Blackwell |
none: the same ELF, the same pinned id (the verifying key hashes the program and its preprocessed tables, not the server's buffer sizes; HEIGHT_THRESHOLD only shortens tables below the verifier's 2^22 maximum) |
none: the compressed proof format and the aggregator guest are unchanged | none: the same --mode verify; the record format unchanged |
hours to one day: a fork of sp1-gpu/crates/prover_components and jagged_tracegen (Apache or MIT), the 11-min cross-build, a per-card profile in provedefault.rs, a CI check that the fork's constants match the pinned verifier's |
low on the protocol, medium on the build: the fixed recursion stage may hold the floor near 8 to 9 GB (section 1.3, approximate) and the first measurement says whether 11 GB is reached; a fork of sp1-gpu to carry forward on every SP1 release; the server rejects nothing it cannot hold, so an out-of-memory shard must fail cleanly and be left (the pool's rule today) |
memory: likely for the adopted shard (10 to 11 GB on paper, section 1.4), not for the prototype shard (28 GB of live trace). Time: prove-only yes (4.3 s on the 5090 scales to about 15 to 22 s on a 3060 and 10 to 15 s on a 4070, approximate); mine-and-prove on a 12 GB card: no at S_p (3 to 4x on a shared card puts a 3060 at 45 to 90 s, approximate, and the miner's 1.7 GB on top of 11 GB does not fit), yes at S_p/2 on a 4070 if the floor lands under 9 GB (approximate). The measurement decides; this is the route the gate waits on |
B. Halve S_p (30,000 to 15,000 pgas, the fee-switch pattern): more and smaller shards |
none | none: one deferred proof per shard, so 2x the shards per block; the chained aggregation stays one per block (9.7 s mining, 2.5 s alone) | none | hours: a fee-table change and a rollout plan like fee-switch-devnet.md |
low: more records per block (the coinbase carries at most 8 shard records, spec 7.7 item 2, so B_p / S_p must stay at 8 or under); the assignment window and sortition unchanged |
alone, no: the floor is the server's (13.9 GB at 0 cycles). With A, it is the dial that moves a 12 GB card from prove-only to mine-and-prove, and a 16 GB card to a comfortable fit |
C. One prover process per card on a rig (the distributed route): the app runs one sp1-gpu-server per NVIDIA card (CUDA_VISIBLE_DEVICES, the per-device socket of sp1-cuda/src/client.rs:211, .cuda().with_device_id(n)), one host process per card, each taking its own assigned shard; the rig installer already picks cards (igneum-rig-lib.sh, prover_decision) |
none | none: shards are independent units by design (spec 7.2); the aggregator runs on the biggest card | none | one day: the app's prover loop per card (prover.rs runs one loop today), the Settings and tile per card, the rig installer's prover unit per card, the socket cleanup per device (the root-socket rule of 5 October) |
low; the throughput is per card, the host RAM 6 GB of pinned buffers per server (section 1.2), so a 4-card rig needs 32 GB of RAM or route A's smaller buffers | it does not move the floor: each card still needs A. It is the route that makes four 12 GB cards worth four shards a cycle, and it ships with A, not instead of it. Splitting ONE shard across cards is not a route: the adopted shard is one core shard (2.5), and column-split provers are research |
| D. RISC Zero as proof system version 2 (CUDA and Metal; segments at po2 19 or 20) | a second guest: core/ is plain Rust and ports as is; the precompile patches differ (SP1's sha3 and k256 patches against RISC Zero's sha2, k256 and keccak circuit); the shard statement bytes unchanged; a second pinned ELF and image id in elf/manifest.json |
a RISC Zero aggregator guest using composition (env::verify of the shard receipts, dev.risczero.com composition page); the chain rule (N verifies N-1) inside the family; a block's shards must be one family, and a chain cannot cross families inside the proof: a family switch lands at a segment boundary as a fresh chain (spec 7.8 item 6 already allows one after an unproven segment; the rule gains "or at a proof-system version change") |
a second verifier mode (--mode verify-r0, the risc0-zkvm verifier, pure Rust, about 100 ms, 222 KB receipts); the record's version selects the family; the native statement names the family's pinned id; both verifiers in the node through the overlap of design 5.6 |
3 to 4 days: guest port and pinning 1, aggregator and chain rule 1, node verifier and record version 1, app profile and host modes 0.5, test vectors and the fast-time harness 0.5; plus the measurement day on PC 2 | medium: two proof systems in consensus for the overlap; RISC Zero's Groth16 wrap is x86 only (the light-client path of ledger P3 stays on SP1 or waits); a 222 KB receipt per shard against 1.27 MB today is a gain; the recursion tree per shard (9 lifts and 8 joins at po2 19) is extra time on small cards; main is at 5.0.0 with no release body, so the pin is 3.0.6 |
memory: yes by documentation (po2 19 for an 8 GB card, po2 20 for 16 GB; 9 to 10 GB per 1 M cycles), the first documented sub-12 GB prover. Time: approximate: a 4090 does 808 kHz at po2 21, so the adopted shard is about 6 s on a 4090-class card and about 20 to 30 s on a 3060-class one at po2 19, prove-only; beside the miner over 60 s on a 3060, near it on a 4070. The prover-floor agent's PC 2 run is the first real number |
| E. Airbender, OpenVM, ZisK, Pico, Ziren as version 2 | a new guest each (RISC-V, except Ziren's MIPS); OpenVM's and ZisK's toolchains are the most complete | each has its own recursion; OpenVM's aggregation and Halo2 wrap are the most documented | a new verifier each (STARK under 300 KB for OpenVM; PLONK or FFLONK for ZisK and Airbender) | 4 to 6 days each | the same two-family cost as D with no memory gain: 21 GiB (Airbender), 24 GB (OpenVM, Ziren), undocumented (ZisK, Pico); Pico's and Ziren's GPU code is BUSL or closed | no: none documents a floor under 21 GiB; the race is tuned for 5090 clusters |
| F. Jolt (Lattice Jolt) as version 2: a sumcheck prover with no codeword; CPU and Metal | a new guest (RV64IMAC, Jolt's toolchain; no keccak precompile today, approximate, so the trie hashing costs more cycles than in SP1) | none exists: no recursion or continuation shipped, so the aggregator would verify N shard proofs natively and the chain rule would live in the native statement until Jolt's recursion lands | a Dory verifier (BN254 pairings, about 50 KB, sub-second, approximate) or an Akita verifier (lattice, 65 to 80 KB); no on-chain verifier shipped | 5 to 8 days for the guest, the verifier and the record; the aggregator question has no answer in the code | high: alpha software, no audit, no production user, no recursion; the proof system of the miner's CPU, not of its card | memory: yes by a wide margin (about 0.9 GB for the adopted shard at 200 bytes a cycle, approximate). Time on a CPU: about 2 to 3 s for 4.7 M cycles at over 2 M cycles a second (a16z, Sep 2026, laptop CPU; approximate for our guest), on Metal under 1 s (PR #1733's 2^25 in 19.8 s on an M5 Max, approximate). The numbers are the best in this document and the software is not shippable |
| G. Folding (Nova family, lattice folding) | a step circuit per transaction or per opcode group | a decider per shard | pairing or lattice verifier | weeks of research, no code over our field | the family that Nexus left | no today; the watch item for 2027 |
H. AMD through a HIP port of SP1's kernels (PR #2668 revived against 6.8.1, or the cuda2hip shim of the OpenVM fork) |
none | none | none: the same SP1 proofs | 3 to 5 days plus PC 1's RX 9070 XT to measure; the hipMallocAsync leak needs the caching allocator the PR carried |
medium: a kernel port with no upstream; the sppark NTT has a limited HIP path and cuPQC none | memory as route A (the same buffers); time unmeasured on any AMD card; the one route that gives AMD miners the 20% pool share |
| I. Apple through RISC Zero Metal (route D's Metal half) | as D | as D | as D | inside D's 3 to 4 days | the 2023 M2 figure (14 kHz, approximate) says 5 minutes for the adopted shard; an M5 Max is not measured by anyone | memory: yes (unified memory, 64 GB on the M5 Max). Time: unknown; the Mac measure lock run is the number |
4. The ranked recommendation
| Rank | Route | Why | Gate |
|---|---|---|---|
| 1. Soonest to 12 GB with the least change: A, with B as the dial and C for rigs | re-size SP1's GPU server; keep the guest, the aggregator, the verifier and the pinned ids exactly as they are; set S_p from the first 12 GB measurement; one prover per card on rigs |
nothing in consensus moves; the work is a fork of two Apache crates and an app profile; it is already running tonight; every other route costs days and adds a second verifier | the prover-floor agent's rows: the adopted shard under 11 GB alone and the time on the first 3060-class or 4070-class card, prove-only and beside the miner. If under 11 GB and under 60 s prove-only: ship 0.3.12 with the 12 GB tier as prove-only and S_p/2 measured for mine-and-prove. If not under 11 GB: route D |
| 2. The fallback if A misses 11 GB, and the Apple route either way: D, RISC Zero as version 2 | the only shipped prover with a documented sub-12 GB configuration and a shipped Metal path; Apache or MIT including the kernels; 222 KB receipts | the swappable interface was built for this (design 5.6) and the node already names the pinned id in the statement, so a second family is a version, not a redesign; the cost is 3 to 4 days plus the overlap | PC 2's po2 19 and 20 rows (memory, time per segment, lift and join) tonight; the Mac's Metal row |
| 3. Best in five years: the sumcheck family without a codeword (Jolt-class), or the sumcheck-plus-WHIR family SP1 and OpenVM already converge on | Jolt proves the adopted shard in seconds on a laptop CPU at under 1 GB of memory, which is the only route that gives AMD-only, Apple and 8 GB machines the proving share with their existing hardware; its verifier is small (50 to 80 KB); its licence is MIT or Apache. It is alpha with no recursion, so not before it ships a stable release with continuations and an audit. SP1 Hypercube and OpenVM SWIRL are the same mathematics with a hash-based PCS and a GPU today, which is why staying on SP1 now loses nothing in that direction | do not adopt now; re-read Jolt and the Arc or WARP accumulation line at every 6-month era draw (design 5.6's swap procedure needs 90% signalling and a 3-month overlap, so the lead time is the schedule) | a stable Jolt tag with recursion, an audit, and a CUDA or merged Metal prover |
| The interface question | yes: ProofSystem is versioned (VERSION, program_id, verify_segment), the record carries version, the node reads pinned ids at start and names them in the native statement, and the overlap procedure keeps both verifiers in the node for 3 months with B_p from the stricter table. What is missing for two families at once is small and named: the record version selecting the verifier command, the fresh-chain rule at a version change, and the shard plan carrying the family per block so a block's shards are homogeneous (the aggregator folds one family). Those three items are in route D's day of node work |
so the answer to "ship one now and move to the other later" is yes, and route A ships nothing that has to be undone |
The honest statement of what this ranking does not know: no 12 GB card has run any prover here. Route A's time figures are bandwidth scaling, labelled approximate; route D's are a 4090 figure scaled the same way. The first 3060 or 4070 in this repository replaces both columns, and the plan is to borrow or buy one this week (a 4070 is the common 12 GB card of 2026; a 3060 the common older one; both are the gate's named class, design R2).
5. The tier consequences, and the public line while the change is made
Every number carries its consequences (CLAUDE.md, 5 October 2026). The table says what each tier has today on SP1 6.8.1, what route 1 (A plus B plus C) gives it if the gate is met, what route 2 (D) adds, and what only route 3 would give. "Today" is measured; the rest is the routes' expected outcome, labelled, until the measurement.
| Tier | Today (measured, bench-log 5 October) | Route 1: re-sized SP1 server, S_p as the dial, one server per card |
Route 2: RISC Zero version 2 | Only route 3 (sumcheck without a codeword) |
|---|---|---|---|---|
| Home miner, one 8 GB NVIDIA card | mines; proves nothing (the server panics under 20 GB) | proves nothing at S_p (the floor's fixed terms, 8 to 9 GB approximate, leave no room); perhaps empty shards |
prove-only at po2 19 (Boundless' 8 GB tier), the miner paused per shard; time approximate 30 to 60 s | mines and proves on its CPU |
| Home miner, one 12 GB card (3060, 4070) | mines; proves nothing; the litepaper's gate card | prove-only at S_p if the floor lands under 11 GB (expected, section 1.4): about 15 to 22 s a shard, approximate; mine-and-prove at S_p/2 on a 4070 if the floor is under 9 GB, approximate; on a 3060 the shared card misses 60 s, approximate, so its default is prove-only with the miner paused per shard (the 16 GB rule of provedefault.rs today, moved down a tier) |
prove-only at po2 20 (16 GB tier) or po2 19; mine-and-prove not inside 60 s on a 3060, approximate | mines and proves, CPU |
| Home miner, one 16 GB card (5080, 4080, 4060 Ti 16 GB) | an empty shard alone (13.9 GB); nothing beside the miner | mine-and-prove at S_p (11 GB plus the miner's 1.7 GB), about 7 to 12 s a shard alone and 20 to 40 s beside the miner, approximate |
mine-and-prove at po2 20 | the same |
| Home miner, one 24 GB card (4090, 3090) | the adopted shard alone (20.4 GB) and beside the miner (22.2 GB, approximate for the card); the prototype shard never | mine-and-prove at S_p with 10 GB to spare; the prototype shard (28 GB live) only if the devnet's fee switch has passed, which it has from DAA 210,000 |
the same with Metal irrelevant | the same |
| Home miner, one 32 GB card (5090) | everything, measured | everything, with more shards in flight if the pool releases between stages | the same | the same |
| Rig, several NVIDIA cards | one prover on the biggest card (prover_decision) |
one server per card, each its own shard; the aggregator on the biggest card; host RAM 6 GB pinned per server today, under 2 GB with route A's buffers | the same model (Bento's) | the same |
| Pool user | through the pool; who proves is open (spec 09) | unchanged | unchanged | unchanged |
| AMD-only (RX 9070 XT, 7900 XTX) | mines; proves nothing on the card; the CPU path 282 s a shard at 30 GB | unchanged until route H (a HIP port, 3 to 5 days, measured on PC 1's 9070 XT) | unchanged: RISC Zero is CUDA and Metal only | mines and proves on its CPU |
| Apple silicon (M-series) | mines (26.7 MH/s on the M5 Max); the SP1 CPU prover 41 to 55 s for an empty shard, 272 s for a small one | unchanged | proves on the GPU through Metal (64 GB unified memory on an M5 Max holds any segment); the time is the measurement | proves in seconds on Metal (Jolt's draft PR figure, approximate) |
| Windows under 32 GB of RAM | off (the WSL2 prover held 7.9 GB) | the pinned buffers fall with max_trace_size, so a 16 GB PC likely qualifies, approximate; measure |
RISC Zero's CUDA path also runs in WSL2 |
The deadlines these fit (spec 7.2 item 3, the litepaper, proving-v1.md): the 10-s exclusive window is the 5090's
alone; a 12 GB card at 15 to 22 s proves its assigned shards in the open phase and is paid when no faster card took
them, which on a chain with few 5090s is most of the time; the minute of the litepaper holds for prove-only 12 GB
cards and for mine-and-prove 16 GB cards; the 600-s unproven deadline holds for every tier above the CPU path.
The public line while the change is made
The litepaper's sentence today ("Target: shard size will be set so a 12 GB card proves one shard in about 20
seconds") is a target and says so (fud-ledger P1, overclaim 27). What this document adds, for site/litepaper.html,
site/miner.html and the app's Proving tile, in the copy law:
Proving runs on NVIDIA cards with 24 GB or more today. A build for 12 GB and 16 GB cards is being measured: the memory is the prover's buffers, not the shard, and the fix is a smaller build of the same prover. AMD and Apple cards mine. A second prover with an Apple path exists and is the fallback.
And the rule for the next status line, whichever way the measurement goes: the number, the card it was taken on, and the tier it moves, in one sentence, the day it is taken.
What this document does about it
| Consequence | Action | Owner |
|---|---|---|
| The gate card has never run a prover here | get a 4070 or 3060 into the measurement loop this week; until then every 12 GB figure stays approximate | coordinator; the project lead for the card |
| Route A's gate | the prover-floor agent's rows (asked for by message tonight); if under 11 GB, provedefault.rs gains the 12 GB prove-only and 16 GB mine-and-prove tiers and the rig installer one server per card |
prover-floor agent, then the proving engineer |
| Route D's measurement | RISC Zero 3.0.6 at po2 19 and 20 on PC 2 (CUDA) and on this Mac (Metal), the same shard statement run natively: memory, time per segment, lift and join, receipt size | prover-floor agent (PC 2); a Mac measure job for Metal |
| The two-family node items (record version selects the verifier, fresh chain at a version change, one family per block) | spec 7.8 gains the three rules when route D starts; nothing changes before | execution engineer |
| AMD | route H is a 3-to-5-day job with a measurement on PC 1's 9070 XT; opened as a plan when route A's result is in | execution engineer |
| The public line | the paragraph above to the site and the tile with the next site pass | site-pages owner |
Sources
Our own: docs/bench-log.md entries "proving v1: segment records, the chain rule, the unproven rule" (5 October 2026),
"the SP1 CPU prover on PC 1" (5 October), "shard proving on the RTX 5090" (4 October); docs/plans/proving-v0.md,
proving-v1.md; docs/analysis/amd-proving.md; docs/spec/07-execution.md 7.2, 7.6, 7.7, 7.8; docs/design/execution-layer.md
5.1 to 5.7; proving/igneum-prove (host/src/proof_system.rs, program/src/main.rs, aggregator/src/main.rs,
elf/manifest.json); vendor/igneum-node-pv1/igneum/exec/src/proving.rs; app/igneum-app/src/provedefault.rs, prover.rs.
SP1 6.8.1, read from ~/.cargo/registry/src/index.crates.io-*/ and the vendored tree vendor/sp1-6.8.1 (commit
c84ada1e, 24 Sep 2026) on the prover-floor worktree: sp1-core-executor-6.8.1/src/opts.rs, src/utils.rs,
src/artifacts/rv64im_costs.json; sp1-prover-6.8.1/src/components.rs, src/worker/config.rs, src/shapes.rs;
sp1-primitives-6.8.1/src/fri_params.rs; sp1-verifier-6.8.1/src/compressed/config.rs; sp1-hypercube-6.8.1/src/verifier/config.rs;
sp1-cuda-6.8.1/src/server.rs, src/client.rs; sp1-gpu/README.md, sp1-gpu/crates/prover_components/src/builder.rs,
src/components.rs, sp1-gpu/crates/jagged_tracegen/src/lib.rs, sp1-gpu/crates/shard_prover/src/prover.rs,
sp1-gpu/crates/cuda/src/task.rs, src/device.rs, sp1-gpu/crates/sys/lib/runtime/mem_pool.cu, sp1-gpu/crates/zerocheck/src/primitives.rs.
Web: docs.succinct.xyz (hardware-acceleration, hardware-requirements, proof-types, security-model, provers introduction,
cluster architecture, docker-compose deployment); blog.succinct.xyz (sp1-hypercube, real-time-proving-16-gpus,
sp1-hypercube-is-now-live-on-mainnet); github.com/succinctlabs/sp1 releases v6.0.0 to v6.8.1, issues #2674, #2930,
#2950, #2969, pulls #2631, #2668, #2723, #2917, #2974; github.com/succinctlabs/sp1-cluster (README, LICENSE,
infra/charts/sp1-cluster/values-example.yaml, crates/worker/src/config.rs); eprint 2025/917 (jagged polynomial commitments).
RISC Zero: ~/.cargo/registry crates risc0-zkp-3.0.4/src/lib.rs, risc0-zkvm-3.0.4/src/receipt.rs, src/host/recursion/prove/mod.rs,
risc0-circuit-rv32im-4.0.4/src/execute/mod.rs, src/zirgen/defs.rs.inc, src/prove/hal/cuda.rs; github.com/risc0/risc0
risc0/zkvm/src/host/client/env.rs, risc0/zkvm/Cargo.toml, risc0/zkvm/build.rs, risc0/sys/kernels/zkp/{cuda,metal}/,
risc0/r0vm/src/actors/factory.rs, risc0/circuit/recursion/src/lib.rs, releases v2.0.0, v3.0.1, v3.0.6, pull #3761;
dev.risczero.com (local-proving, composition); docs.boundless.network (bento, performance-optimization, quick-start);
github.com/boundless-xyz/boundless (compose.yml, bento/README.md, bento/LICENSE-BSL); github.com/ekrembal/gsr-stark-verifier pull 5; l2beat.com/zk-catalog/risc0.
Others: zksync.io/airbender, docs.zksync.io airbender and proving pages, github.com/matter-labs/zksync-airbender (README,
docs/gpu.md, pull #448), veridise.com (the Airbender audit); 0xpolygonhermez.github.io/zisk (introduction, limits,
distributed execution, installation), github.com/0xPolygonHermez/zisk (README, pull #1238, zisk-contracts);
blog.openvm.dev (2.0, 2.0-production, 2.1, openvm-gpu, v1), docs.openvm.dev (security-model, distributed-proving, sdk);
pico-docs.brevis.network, github.com/brevis-network/pico and pico-gpu (README, LICENSE), blog.brevis.network (Prism 1.0,
2.0, 2.1); docs.zkm.io (prover, performance), github.com/ProjectZKM/Ziren, zkm.io (the independent evaluation of v1.1.4),
eprint 2026/2330; github.com/starkware-libs/stwo and stwo-cairo (README), ingonyama.com (ICICLE-Stwo, the Starknet
partnership, ICICLE Metal v3.6), dev.ingonyama.com (install_gpu_backend), blog.zksecurity.xyz/posts/webgpu, starkware.co
(S-two 2.0.0, Nexus on S-two), theblock.co (S-two on Starknet); github.com/a16z/jolt (README, book: intro, dory, akita,
streaming, recursion, blindfold; pulls #1733, #1938; tags), a16zcrypto.substack.com ("How to prove software ran
correctly", Sep 2026), a16zcrypto.com (jolt-6x-speedup, 64-bit-proving-jolt, zkvm-jolt-zero-knowledge, faqs-on-jolts-initial-implementation),
eprint 2025/611; github.com/scroll-tech/ceno (README, Cargo.toml, pull #1403), ceno-gpu-mock, scroll.io (Ceno post),
osec.io ("zkVMs' unfaithful claims"); github.com/nexus-xyz/nexus-zkvm (README, LICENSE), blog.nexus.xyz (roadmap);
lita.gitbook.io (Valida architecture, benchmarks); github.com/powdr-labs/powdr; irreducible.com (announcing-binius64,
reinventing-irreducible, irreducible-shutting-down), github.com/binius-zk/binius64, eprint 2026/1656; eprint 2021/1043,
2022/1010, 2024/1609, 2025/1187, 2024/1586, 2024/185 (linear-code commitments, WHIR, Vortex), github.com/Consensys/linea-monorepo;
PolyhedraZK/Expander and blog.polyhedra.network (returned 530 tonight); eprint 2021/370, 2024/2099, 2024/1220, 2024/416,
2024/1605, 2025/247, 2025/294, 2026/242, 2024/1731, 2025/753, 2026/1371 (folding and accumulation), sonobe.pse.dev,
github.com/privacy-scaling-explorations/sonobe, NethermindEth/latticefold, LFDT-Nightstream/Nightstream; eprint 2023/1271,
2024/1208, 2024/1873, 2025/1349, 2025/1653, 2025/1285, 2018/691, arXiv 2210.00264, 2602.16338 (distributed proving);
github.com/supranational/sppark, github.com/Okm165/stark-backend pull 2, ethresear.ch (qingming G64 NTT and STARK on ROCm);
github.com/cysic-labs/venus, erigon.tech (Zilkworm), github.com/DelphinusLab/prover-node-docker, hackmd.io/@bobbinth
(Miden), ethproofs.org/clusters (5 October 2026).