411 lines
66 KiB
Markdown
411 lines
66 KiB
Markdown
# Proving methods: why the prover needs 14 GB, what else exists, and how a 12 GB card gets to prove
|
|
|
|
5 October 2026, from the project lead at 22:05 UTC: "if this doesn't enable 12 GB cards, then do a full deep research task on proving
|
|
and see if there are different methods." "This" is the prover-floor agent's patch of SP1's GPU server (branch
|
|
`prover-floor`), running tonight. This document is research and reading, not measurement: every number of ours is from
|
|
`docs/bench-log.md` with its entry named; every claim about another system cites its repository file, its documentation
|
|
page or its paper, or is labelled approximate. Status words follow `docs/spec/00-overview.md` 0.2. Day estimates follow
|
|
the project lead's rule of 3 October 2026: hours of agent time, never weeks.
|
|
|
|
The facts this starts from (bench-log, "proving v1", 5 October 2026; `docs/plans/proving-v1.md`; `docs/analysis/amd-proving.md`):
|
|
|
|
| Fact | Number |
|
|
|---|---|
|
|
| SP1 6.8.1's GPU server, an empty shard (280,706 cycles), the card to itself | 13,874 MiB peak, 2.2 s compressed |
|
|
| The adopted v1 shard (`S_p` 30,000 pgas, 4,717,439 cycles) | 20,434 MiB, 4.3 s; 22,210 MiB and 13.2 s beside the miner |
|
|
| The prototype shard (6.75 M pgas, 60.4 M cycles) | 28,307 MiB, 10.8 s; flat at 28.3 GB from 20 M cycles up |
|
|
| Aggregation, chained, per block, on a mining 5090 | 9.6 to 9.7 s; 2.2 to 2.5 s with the card to itself |
|
|
| The CPU path (PC 1, 16 cores) | 282 s a shard whatever its size, 29.5 to 30.5 GB RSS |
|
|
| AMD and Apple GPUs | no zkVM proves on AMD; RISC Zero has a Metal prover, SP1 does not |
|
|
| The promise | `site/litepaper.html`: "Target: shard size will be set so a 12 GB card proves one shard in about 20 seconds"; the design goal is every block proven within about a minute by the miners' own cards |
|
|
|
|
## 1. The memory anatomy of a STARK-based zkVM prover, and why the floor is where it is
|
|
|
|
### 1.1 What SP1 6.8.1 is
|
|
|
|
SP1 6.x is not the univariate FRI STARK of the earlier SP1 releases (Succinct calls Hypercube the "first zkVM built entirely on a multilinear polynomial-based proof system", blog.succinct.xyz, sp1-hypercube, 20 May 2025). The crates it pulls say what it is: `slop-multilinear`, `slop-sumcheck`,
|
|
`slop-jagged`, `slop-stacked`, `slop-basefold`, `slop-whir` (the `~/.cargo/registry` of this Mac; `proving/igneum-prove/Cargo.toml`
|
|
pins `sp1-sdk = "=6.8.1"`). The architecture Succinct calls Hypercube: the execution trace is a set of multilinear
|
|
polynomials over the 31-bit KoalaBear field (`sp1-hypercube-6.8.1/src/verifier/config.rs`: `SP1BasefoldConfig =
|
|
Poseidon2KoalaBear16BasefoldConfig`), the constraints are checked by a zerocheck sumcheck and the lookups by a LogUp GKR
|
|
(`sp1-gpu/crates/zerocheck`, `sp1-gpu/crates/logup_gkr`), and the polynomial commitment is "jagged": every table's
|
|
columns, whatever their heights, are concatenated into one long vector, stacked into rows of height `2^log_stacking_height`
|
|
and committed with BaseFold, a FRI-like folding over a Reed-Solomon code (`sp1-gpu/crates/basefold/src/fri.rs`,
|
|
`slop-basefold-6.8.1/src/verifier.rs`). The proof system parameters, from `sp1-primitives-6.8.1/src/fri_params.rs` and
|
|
`sp1-prover-6.8.1/src/components.rs`:
|
|
|
|
| Parameter | Value | Where |
|
|
|---|---|---|
|
|
| Field | KoalaBear, 31 bits, 4 bytes an element; extension degree 4 (16 bytes) | `sp1-primitives` |
|
|
| Core stage: Reed-Solomon blowup | `CORE_LOG_BLOWUP = 2`, so the codeword is 4x the data | `fri_params.rs:5` |
|
|
| Core stage: stacking height, maximum rows per table | `CORE_LOG_STACKING_HEIGHT = 21`, `CORE_MAX_LOG_ROW_COUNT = 22` | `components.rs:16,17` |
|
|
| Core shard limits (the executor's cut) | `MAX_SHARD_SIZE = 2^24` cycles, `ELEMENT_THRESHOLD = 2^28 + 2^27 = 402,653,184` trace elements, `HEIGHT_THRESHOLD = 2^22` rows | `sp1-core-executor-6.8.1/src/opts.rs:9-12` |
|
|
| Recursion (compress) stage | blowup 2 (`RECURSION_LOG_BLOWUP = 2`), stacking height 20, max rows 2^21 | `fri_params.rs:6`, `sp1-verifier-6.8.1/src/compressed/config.rs:1,2` |
|
|
| Shrink and wrap stages | blowup 3 (8x), 22 bits of grinding, stacking 18 and 21 | `fri_params.rs:17,18,7`, `components.rs:37-40` |
|
|
| Recursion arity | 4 proofs per compose step (`DEFAULT_MAX_COMPOSE_ARITY = 4`, `DEFAULT_MAX_REDUCE_ARITY = 4`) | `sp1-prover-6.8.1/src/worker/config.rs:183,193` |
|
|
| Workers | 4 core workers, 8 recursion prover workers, 4 recursion executors, 4 deferred workers, buffers of 4 to 8 | `worker/config.rs:188-205` |
|
|
|
|
The stages a shard goes through (`sp1-prover-6.8.1/src/worker/controller/*.rs`): execute (the RISC-V executor cuts the
|
|
run into core shards at the thresholds above); core (one jagged-PCS proof per core shard, on the GPU); normalize and
|
|
compose (each core proof is verified inside a recursion program, then proofs are folded 4 at a time until one remains,
|
|
the "compressed" proof, 1,272,897 bytes for every shard we have proven, bench-log 4 and 5 October); deferred (what the
|
|
aggregator uses: `verify_sp1_proof` inside a guest, `proving/igneum-prove/aggregator/src/main.rs`); shrink and wrap
|
|
(to a BN254 STARK, then Groth16 or Plonk; not run here, ledger P3).
|
|
|
|
### 1.2 The terms, and which scale with the shard
|
|
|
|
Every STARK-family prover holds these buffers on the device at its peak, in some order and with some overlap. The
|
|
sizes below are from the constants of 1.1 and the allocation code of `sp1-gpu`; where a buffer's size is the actual
|
|
trace rather than the maximum, the row says so.
|
|
|
|
| Term | What it is | Size rule | SP1 6.8.1 on a 32 GB card | Scales with the shard? |
|
|
|---|---|---|---|---|
|
|
| Main trace | the witness: one element per cell of every table the shard touched | `cells x 4 bytes`, where cells = sum over tables of rows x columns; the executor cuts a new core shard at 402,653,184 cells | the device buffer is allocated at the MAXIMUM, not the actual trace: `allocate_and_initialize_traces` takes `max_trace_size` and allocates `max_trace_size` felts plus `max_trace_size / 2` u32 of index (`sp1-gpu/crates/jagged_tracegen/src/lib.rs:484-503`), 6 bytes a cell; the core prover's `max_trace_size` is `element_threshold + 2^21` (`prover_components/src/builder.rs:70-71`): **2.26 GiB** on a card over 30 GB, 1.61 GiB on a 24 GB card (the threshold drops by 2^26 + 2^25 + 2^24 when memory is 30 GB or under, `builder.rs:41-45`) | no: fixed at the maximum shard, whatever the trace |
|
|
| Preprocessed trace | the program's own tables (the ELF as a `Program` AIR, the byte and range tables) | program size x its columns plus 2 x 2^16-class tables | small for a 2.8 MB guest ELF (`elf/manifest.json`); not isolated | with the guest, not the shard |
|
|
| Codeword (the LDE) | the stacked polynomial encoded at rate 1/4 for BaseFold | `stacked cells x 4 (blowup) x 4 bytes`; the stacked length is the actual cell count padded to a multiple of 2^21 | 4.7 M cycles: approximate, the actual trace; 60 M cycles: the shard is 7 to 15 core shards of up to 402 M cells, each encoded to 6 GiB at the blowup, one or more in flight | yes, up to the core-shard cap; past the cap the shard count grows and the per-shard term stays |
|
|
| Merkle commitment | Poseidon2 hashes of the codeword rows | `rows x 8 elements x 4 bytes x 2`, rows = 2^21 x blowup | about 0.5 GiB at full stacking, approximate | with the stacked rows |
|
|
| Zerocheck and GKR | the constraint sumcheck over the extension field, and the LogUp GKR layers | extension elements are 16 bytes; the sumcheck holds a folded copy of the trace in the extension field, which is 4x the base trace at the first round and halves each round | up to about 4x the live trace in the first round, approximate (`sp1-gpu/crates/zerocheck/src/primitives.rs:174,287`: `Buffer<Ext>` of `new_total_length`) | yes |
|
|
| Recursion traces | the normalize and compose programs' own traces, verifying core proofs | fixed-shape programs: `RECURSION_TRACE_ALLOCATION = 2^27` cells (`builder.rs:15`), allocated at 6 bytes a cell: **0.75 GiB** per recursion prove, at blowup 4 a 2 GiB codeword plus its own zerocheck | fixed per recursion step; the number of steps is log4 of the core-shard count | no (per step) |
|
|
| Shrink and wrap traces | the two last stages, not run by us | 2^25 and 85,376,340 cells (`builder.rs:16,19`): 0.19 and 0.48 GiB | only when wrapping | no |
|
|
| Proving keys and program cache | the recursion programs (`vk_map.bin`, the normalize cache of 5 programs) and the shard program's setup | `DEFAULT_NORMALIZE_PROGRAM_CACHE_SIZE = 5` (`worker/config.rs:192`); the key setup took 14.6 s on the 5090 (bench-log 4 October) | not isolated | no |
|
|
| Pinned host buffers | the staging copies on the PC side | 4 core workers x `max_trace_size` x 4 bytes = **6.0 GiB** of pinned RAM, plus 4 x 0.5 GiB for recursion (`prover_components/src/components.rs:99-103`, `builder.rs:76,95`) | this is the 7.9 GB WSL2 working set measured on 5 October | no |
|
|
| The allocator | CUDA's default memory pool with its release threshold set to `u64::MAX` (`sp1-gpu/crates/cuda/src/task.rs:152,190-199`): freed blocks are never returned to the driver | `nvidia-smi` therefore reports the high-water mark of everything above, and it stays until the server exits | this is why the memory curve is flat between shards of different size | no |
|
|
|
|
Two facts from this table explain the measurements:
|
|
|
|
1. **The server refuses small cards by code.** `local_gpu_opts()` reads the card's total memory, adds 4 GB, and panics
|
|
under 24: `"Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB"` (`sp1-gpu/crates/prover_components/src/builder.rs:35-38`).
|
|
A 16 GB card (16 + 4 = 20) and a 12 GB card (16) never start; a 20 GB card is the smallest that does. The 13.9 GB
|
|
floor measured on the 5090 is therefore not the whole story for a 12 GB card: on this build the card is refused
|
|
before any buffer is allocated. Any route through SP1's GPU server starts by removing this line.
|
|
2. **The environment knobs do not reach the floor** because the same function overwrites `element_threshold` with the
|
|
compile-time constant (`builder.rs:41-48`); only `HEIGHT_THRESHOLD` passes through, which is why the sweep's
|
|
`ELEMENT_THRESHOLD 2^26` rows changed nothing and `HEIGHT_THRESHOLD 2^20` took 5.4 GB off the 60 M-cycle shard
|
|
(bench-log, "the 12 GB requirement", 5 October 2026). The worker counts only slow the proof (11.4 s to 20.8 s)
|
|
because the device buffers are sized by `max_trace_size`, not by the worker count.
|
|
|
|
### 1.3 Why the floor is 13.9 GB for an empty shard
|
|
|
|
With the release threshold at `u64::MAX`, the peak is the high-water mark over the whole pipeline. For an empty shard
|
|
the core trace is small (280,706 cycles), so the fixed-shape terms dominate: the 2.26 GiB main-trace buffer allocated
|
|
at the maximum, the recursion step over a fixed-shape normalize program (a 0.75 GiB trace buffer, its 4x codeword in
|
|
the extension field for the zerocheck, its Merkle tree), the proving-key and program caches, and the deferred and
|
|
compose machinery that a compressed proof always runs once. The decomposition of the 13.9 GB into those terms is
|
|
approximate until the prover-floor agent's profile lands (branch `prover-floor`, tonight): the figure that is not
|
|
approximate is that none of it is the witness (5 to 22 KB a shard) and none of it is the shard's cycles (the same
|
|
13.9 GB at 280 k cycles and 556 k cycles, bench-log "the S_p curve").
|
|
|
|
The step from 13.9 GB (empty) to 20.4 GB (4.7 M cycles) is the live trace: the 4.7 M-cycle shard is one core shard
|
|
(its trace area is under 402 M cells, so it was not split; approximate from the memory curve, the cell count is not
|
|
logged by the host), and its codeword, zerocheck and GKR buffers are sized by its actual cells. The step from 20.4 GB
|
|
to 28.3 GB (20 M cycles and up) is the second and later core shards in flight at once: 4 core workers with a buffer of
|
|
4 (`worker/config.rs:188,189`) let several core shards' codewords exist at the same time; past 20 M cycles the pipeline
|
|
is full and the peak is flat, which is what the curve shows (28,371 MiB at 20 M cycles, 28,307 at 40 M and 60 M).
|
|
|
|
### 1.4 The theoretical floor for our guest at the adopted shard
|
|
|
|
If every buffer were sized to the shard rather than to the maximum, the adopted v1 shard (4.7 M cycles) would need,
|
|
approximate, from the rules of 1.2:
|
|
|
|
| Term | Rule | Approximate bytes |
|
|
|---|---|---|
|
|
| Main trace, actual | 4.7 M cycles x about 60 cells a cycle (the `Add` and `Addi` tables cost 33 and 30 columns a row, a memory access adds 20 and a global interaction 241: `sp1-core-executor-6.8.1/src/artifacts/rv64im_costs.json`) | about 280 M cells, 1.1 GB |
|
|
| Codeword at blowup 4 | 4x | 4.5 GB |
|
|
| Zerocheck first round in the extension field | 4x base, halving each round | 4.5 GB at the peak round, falling |
|
|
| Merkle tree | rows x 32 bytes x 2 | 0.3 GB |
|
|
| Recursion step, fixed | 2^27 cells x 6 bytes plus its 4x codeword and extension copies | 2 to 3 GB, approximate |
|
|
| Keys and caches | | under 1 GB, approximate |
|
|
| Peak, if the core stage and the recursion stage do not overlap and the pool releases | | **about 10 to 11 GB**; about 6 GB if the blowup-4 codeword is replaced by a rate the sumcheck does not need (see 2.4) |
|
|
|
|
So the adopted shard is, on paper, a 12 GB card's shard with the server re-sized and nothing else changed, and it is a
|
|
12 GB card's shard with 4 GB to spare if the shard is halved (`S_p` 15,000 pgas, 2.4 M cycles: the planner cuts at
|
|
transaction boundaries to any budget, `core/src/plan.rs`, and the fee switch of 5 October already moved `S_p` once).
|
|
What the paper figure does not say is the time: a smaller card proves slower, and 60 s with the miner running is the
|
|
bound (section 3). The prover-floor agent is measuring the real figure; this section says what it should find and why.
|
|
|
|
### 1.5 RISC Zero's anatomy, for comparison
|
|
|
|
RISC Zero is the FRI STARK the textbooks describe, and its constants make the same table easy to read
|
|
(`~/.cargo/registry`, `risc0-zkp-3.0.4/src/lib.rs`, `risc0-circuit-rv32im-4.0.4/src/zirgen/defs.rs.inc`):
|
|
|
|
| Term | Value | Where |
|
|
|---|---|---|
|
|
| Field | BabyBear, 31 bits; extension degree 4 | `risc0-core` |
|
|
| Segment size | `DEFAULT_SEGMENT_LIMIT_PO2 = 20` (1,048,576 cycles), `MIN_CYCLES_PO2 = 13`, `MAX_CYCLES_PO2 = 24`; `DEFAULT_MAX_PO2 = 22` for the verifier | `risc0-circuit-rv32im-4.0.4/src/execute/mod.rs:39`, `risc0-zkp-3.0.4/src/lib.rs:35-38`, `risc0-zkvm-3.0.4/src/receipt.rs:884` |
|
|
| Trace width | data 211 + accum 103 + code 1 = 315 columns; globals 90, mix 36 | `defs.rs.inc:7-11` |
|
|
| Blowup | `INV_RATE = 4`; 50 queries; FRI fold 16 | `risc0-zkp-3.0.4/src/lib.rs:41-51` |
|
|
| Recursion | lift, join and resolve programs at `RECURSION_PO2 = 18` rows | `risc0-zkvm-3.0.4/src/host/recursion/prove/mod.rs:58` |
|
|
| The GPU buffers | `check + ctrl + data + accum + mix + out` elements x 4 bytes, printed by the CUDA HAL at `eval_check` | `risc0-circuit-rv32im-4.0.4/src/prove/hal/cuda.rs:181-200` |
|
|
|
|
From those constants the trace of a segment is `2^po2 x 315 x 4` bytes and its LDE 4x that, so, approximate: po2 18 is
|
|
0.3 GB of trace and 1.5 GB with the LDE, po2 19 is 3.1 GB, po2 20 is 6.2 GB, po2 21 is 12.3 GB, before the check
|
|
polynomial, the extension-field accumulators and the Merkle trees. Two things follow. A RISC Zero segment at the
|
|
default 2^20 is in the same class as one SP1 core shard, not smaller. And a RISC Zero segment at 2^18 or 2^19 is a
|
|
2 to 4 GB object: the only reason a 12 GB card could not prove one is the fixed overhead of the recursion circuits
|
|
(2^18 rows each) and the allocator, which is the measurement the prover-floor agent takes on PC 2 if SP1 cannot go
|
|
under 11 GB. Section 2.2 carries the documented numbers.
|
|
|
|
### 1.6 Which terms the shard size can move, and which it cannot
|
|
|
|
| Lever | Moves | Does not move |
|
|
|---|---|---|
|
|
| Our `S_p` (pgas per shard) | the live trace, the codeword, the zerocheck: everything in 1.2 marked "yes" | the maximum-sized buffers, the recursion step, the keys, the pool |
|
|
| SP1's `HEIGHT_THRESHOLD` (the one knob the server honours) | the rows per table in one core shard, so the live buffers | the fixed terms (measured: 13,861 MiB on the empty shard with every knob at its minimum) |
|
|
| A server patch: size `max_trace_size` to the shard, release the pool, one core worker | the 2.26 GiB buffer, the high-water mark, the in-flight count | the recursion step's fixed shape and the key caches |
|
|
| A different proof system | the blowup (sumcheck-only and linear-code systems have none, 2.4), the recursion shape | the trace itself: a RISC-V cycle costs tens of cells in every zkVM |
|
|
|
|
## 2. Every current proving route
|
|
|
|
Read 5 October 2026, 22:10 to 23:00 UTC, by four research agents and this one; every cell names its page or file.
|
|
"Not documented" means the project publishes no figure, which for a memory floor is itself the finding.
|
|
|
|
### 2.1 The zkVMs with a GPU prover
|
|
|
|
| Prover | Proof system, field, chunk | GPU support and the documented minimum memory | Throughput, on what | Verification of the recursive proof; wrapper | Licence | State, October 2026 |
|
|
|---|---|---|---|---|---|---|
|
|
| **SP1 6.8.1** (ours) | Hypercube: multilinear, jagged PCS, BaseFold, LogUp GKR; KoalaBear; core shards of up to 2^24 cycles and 402 M cells (section 1) | CUDA only. Docs: "24GB or more VRAM", compute capability 8.0+, Linux x86_64 (docs.succinct.xyz, hardware-acceleration page). Code: panic under 20 GB physical (`builder.rs:35-38`). Measured here: 13.9 GB floor, 20.4 GB at the adopted shard, 28.3 GB at the prototype shard. Issue #2950: two clients on a 48 GB L40S hold 41 to 43 GB; a single 6 GiB tensor allocation failed | 4.3 s for the adopted shard, 10.8 s for the prototype one on a 5090 (bench-log); Succinct: 99.7% of Ethereum blocks under 12 s on 16 x RTX 5090 (blog.succinct.xyz, 18 Nov 2025) | compressed proof 1,272,897 bytes, verified in 0.032 to 0.040 s here (`--mode verify-segment`); Groth16 about 260 bytes and about 270 k gas, Plonk about 868 bytes and 300 k gas (docs, proof-types page); the Groth16 wrap needs about 14 GB of host RAM, Plonk about 60 GB (hardware-requirements page) | Apache-2.0 or MIT for the repository including `sp1-gpu/` (`LICENSE-APACHE`, `LICENSE-MIT` at the root; no separate licence under `sp1-gpu/`); `sp1-cluster` is Business Source 1.1 | v6.8.1 of 24 Sep 2026 is the latest tag; mainnet for Ethereum proving since 19 Feb 2026 (blog); AMD port PR #2668 closed unmerged 20 Mar 2026; no Metal, Vulkan or WebGPU |
|
|
| **RISC Zero 3.0.x** | FRI STARK (DEEP-ALI), BabyBear, Poseidon2, blowup 4, 50 queries; segments of `2^po2` cycles, default po2 20, allowed 13 to 24 (section 1.5); lift, join, resolve recursion at 2^18 rows; keccak as a separate circuit | CUDA and **Metal** (`risc0/sys/kernels/zkp/metal/`; on Apple silicon the Metal path is on automatically, `risc0/zkvm/build.rs`). Documented memory per segment: Bento design page, 1 M cycles 9 to 10 GB, 2 M 17 to 18 GB, 4 M 32 to 34 GB; Boundless performance page, the largest `SEGMENT_SIZE` per card: 8 GB card po2 19, 16 GB po2 20, 20 GB po2 21, 40 GB po2 22; docs: "less than 10 GB available: change the segment size limit" (dev.risczero.com, local proving); `env.rs:190-192`: "lowering this value by 1 will cut memory consumption by about half". PR #3761 (June 2026): po2 22 did not fit a 24 GB 4090 until the `low_vram` buffer reuse | 4090: 808 kHz at po2 21, 1,207 kHz at po2 22 with PR #3761 (end to end to a succinct receipt); Apple M2 Pro about 14 kHz on the 2023 datasheet, approximate (the page was unreachable tonight); real-time Ethereum on about 160 x 4090 (blog, approximate) | succinct receipt 222,668 bytes, constant; about 100 ms to verify, approximate (`gsr-stark-verifier` PR #5, mirrored docs); Groth16 seal 256 bytes, about 200 to 300 k gas, approximate; the Groth16 wrapper is x86 only, not on Apple silicon (docs) | Apache-2.0 or MIT, CUDA and Metal kernels included (`risc0/sys/kernels/zkp/cuda/eltwise.cu:1-13`); Bento is BSL 1.1 with a change date already passed | v3.0.6 of 17 Jul 2026 on the maintained line; `main` is 5.0.0 with no release body; `RISC0_PROVER=actor` multi-GPU scheduler experimental since 3.0.1 (`r0vm/src/actors/factory.rs:183-195` carries measured per-po2 memory tokens: po2 18 = 8, 19 = 10, 20 = 15, 21 = 24; lift and join = 3) |
|
|
| **Airbender** (Matter Labs) | DEEP STARK, FRI, Mersenne31; chunks of 2^22 cycles; Boojum then FFLONK wrap (docs.zksync.io, airbender page) | CUDA only. "any GPU with 22GB RAM" for production (zksync.io/airbender); the code has memory presets `GiB21` (24 GB cards) and `GiB30`, raised from 29 because Ethereum blocks failed to allocate at 29 GiB (PR #448, `gpu/execution_prover/src/prover/config.rs`); the final SNARK is CPU with about 150 GB of RAM (`docs/gpu.md`) | one H100: 21.8 MHz base layer, 8.5 MHz end to end, about 35 s an Ethereum block (June 2025 post); ethproofs.org today: 4 x 5090 2.3 s average | FFLONK over BN254 on chain; gas not published | MIT or Apache-2.0 | v0.6.0-rc.2; Veridise audit Feb to Apr 2026; live for ZKsync Atlas chains |
|
|
| **ZisK** (Polygon spin-out) | eSTARK over Goldilocks (pil2-stark), Poseidon2, approximate; main instance 2^22 to 2^23 rows, chunks of up to 2^22 steps (PR #1238) | CUDA only, CUDA 12.9+; **no VRAM floor documented**; workers need about 32 GB of host RAM, the assembly emulator 64 GB (docs, limits and distributed pages); Cysic's Venus fork submits from one RTX 4090 | 4 x 5090: p99 9.62 s on Ethereum blocks (Aug 2026); 24 x 5090 6.56 s average (Nov 2025) | PLONK wrapper verified by Solidity (`zisk-contracts`); 128-bit claimed | Apache-2.0 or MIT | v1.3.1-alpha, 30 Sep 2026, "undergoing security and correctness audits" (README) |
|
|
| **OpenVM 2.0** (Axiom) | SWIRL: sumcheck, zerocheck, LogUp GKR, stacked reduction into WHIR; BabyBear; segments by metered trace height (blog.openvm.dev/2.0) | CUDA only; "at least 24GB of VRAM": L40, 4090, L40S, 5090 (blog.openvm.dev/openvm-gpu) | 11.4 MHz on one 5090, 139 MHz on 16; 2.1 preview: 4 x 5090 p99 9.7 s | STARK proof under 300 KB; Halo2-KZG wrapper, 316 k gas; the Halo2 wrap 8.1 s on a 5090 | MIT or Apache-2.0, GPU prover included | v2.0.2 of 14 Aug 2026; zkSecurity audit of SWIRL; Scroll's prover builds on it |
|
|
| **Pico** (Brevis) | Plonky3 STARK, KoalaBear default; chunk size a parameter with no documented default | CUDA via `pico-gpu`; **no VRAM figure published**; every run on 32 GB 5090s | Prism 2.1: 16 x 5090 over two machines, 4.87 s average on Ethereum blocks | Groth16 via gnark | core MIT or Apache; **`pico-gpu` is BUSL-1.1** and "not recommended for production" (its README) | v2.1.2, Aug 2026; Sherlock audit |
|
|
| **Ziren** (ZKM, MIPS) | Plonky3-class, KoalaBear, LogUp GKR, WHIR | CUDA 12, compute capability 8.6+, "24 GB VRAM or higher"; the GPU prover is a Docker image pinned by digest, source "planned H1 2026" (docs.zkm.io prover page; an independent evaluation of v1.1.4 says the GPU path is not open) | one 5090: 5.9 MHz on a 288 M-cycle block; 4 GPUs 3.1 to 3.3x | compressed proof 603 KiB; Groth16 or PLONK | core MIT or Apache; GPU image licence unspecified | v1.2.7; no public audit cited |
|
|
| **Stwo / S-two** (StarkWare) | Circle STARK over Mersenne31; blowup 1 (rate 1/2), 70 queries, 26 bits of grinding | **CPU SIMD first** (AVX2, AVX-512, NEON, WASM); GPU through ICICLE-Stwo (Ingonyama): about 3 GB of trace in GPU memory, out of memory from 2^23 rows; a WebGPU port of the constraint evaluation (zkSecurity blog, April 2025); client-side proving under 1 GB after a spill allocator (third-party PR) | 620 k Poseidon2 a second on an M3 laptop | via a Cairo verifier (proofs of proofs); sizes not published here | Apache-2.0 | live on Starknet mainnet since 3 Nov 2025; no RISC-V guest of its own (Nexus 3.0 is the RISC-V zkVM on it, BUSL-1.1 until 2029) |
|
|
| **Jolt** (a16z) | sumcheck and lookups (Lasso lineage, Twist and Shout memory checking); PCS Dory over BN254 by default, or **Akita**, a lattice commitment over a 128-bit prime field (Sep 2026, "Lattice Jolt"); RV64IMAC; no continuations (the book's recursion page is "under construction") | **No CUDA in the public repository** (LayerZero's "Jolt Pro" CUDA port is private); **Metal**: PR #1938 merged 30 Sep 2026 (the `jolt-metal` runtime crate), PR #1733 (the full prover on Metal) still a draft. Memory: "about 200 bytes per cycle" with Akita (a16z substack, Sep 2026); the book: "under 2 GB of memory per million cycles"; a streaming prover bounded to "a few GBs" is planned, not shipped | over 2 M cycles a second on a laptop CPU with Akita, over 10 M with Metal on a Apple laptop (a16z substack, Sep 2026); PR #1733: M5 Max, 2^25 cycles in 19.8 s, 2^27 in 77 s at an 89.4 GiB footprint | proof about 50 KB (Dory) or 65 to 80 KB (Akita); verify sub-second, approximate; on-chain 1.3 to 2 M gas estimated in 2024; no Groth16 wrapper shipped | MIT or Apache-2.0 | `v0.3.0-alpha` is the last tag (1 Oct 2025); README: "not suitable for production use"; no audit |
|
|
| **Ceno** (Scroll) | GKR tower prover, BabyBear, WHIR or BaseFold PCS; RV32IM | CUDA, but the real HAL is in a **private** `ceno-gpu` repository (the public one is a mock); no memory numbers | 2 GPUs 1.6x over one (PR #1403, Sep 2026) | via OpenVM recursion to Halo2 | Apache-2.0 | README: "under construction and not suitable for use in production" |
|
|
| **Nexus 3.0** | on Stwo (Circle STARK, M31) | no GPU path documented | none published | not published | **BUSL-1.1** until 10 Feb 2029 | last push 6 Jan 2026; folding (Nova family) abandoned June 2025 for the STARK |
|
|
| **Valida** (Lita) | Plonky3 STARK | no GPU; a CUDA port "underway" in July 2025 | none current | not published | Apache or MIT | dormant since Sep 2025; documented soundness issues in its own benchmarks page |
|
|
| **Powdr** | no longer a zkVM: `powdrVM` archived; powdr is autoprecompiles on OpenVM | OpenVM's | OpenVM's | OpenVM's | MIT or Apache | tooling layer; "DO NOT USE FOR PRODUCTION" |
|
|
| **Binius / Binius64** (Irreducible) | binary-field SNARK, BaseFold-style FRI over GF(2^64) words | CPU SIMD only; the FPGA work was dropped 9 Sep 2025 ("FPGAs underperformed GPUs"); no GPU | ECDSA aggregation about 5x over SP1 and R0VM on L40S GPUs, on CPU (the page carries methodology corrections) | hash-based; no recursion shipped | Apache-2.0 or MIT | **the company shut down 12 Nov 2025**; the only zkVM on it (PetraVM) is archived |
|
|
| 2026 entrants | Cysic Venus (a ZisK fork with cudaGraph tuning and an FPGA backend, Apache or MIT, "do not use in production"); Zilkworm (Erigon's C++ guest on Airbender, 2 x 5090 9.3 s); zkDTVM (evmone guest, 4 x 5090 4.7 s, no public docs); Delphinus zkWasm (Halo2 on BN254, a 4090 minimum plus 58 GB of host RAM); Miden (Goldilocks STARK, client-side, Metal via `miden-gpu`, mainnet alpha planned); Boojum (2023 claim of proving on a 16 GB card, superseded by Airbender) | none states a floor under 24 GB on a GPU | | | | |
|
|
|
|
The reading of the table. No shipped zkVM documents a GPU floor under 24 GB except RISC Zero, whose memory is a
|
|
function of a runtime knob (`segment_limit_po2`) and is published per card size by Boundless. SP1's 24 GB is a line
|
|
of code, not a property of the proof system: Airbender, OpenVM and Pico all pad to the card they tune on, and all
|
|
three say 24 or 32 GB because their market is Ethereum blocks on 5090 clusters. The real-time race has collapsed to 2
|
|
to 4 consumer cards per block (ethproofs.org, 5 October 2026), which is why nobody is tuning for a 12 GB card: the
|
|
customer buys 5090s. Igneum's customer is the miner who already owns the card, so Igneum has to do the tuning itself.
|
|
|
|
### 2.2 The sumcheck and GKR family against FRI STARKs, in memory terms
|
|
|
|
| Family | What it holds at the peak | Blowup | The GPU figure today | Source |
|
|
|---|---|---|---|---|
|
|
| FRI STARK (RISC Zero, Airbender, ZisK, Pico, Stwo, SP1 3 and 4) | trace, its Reed-Solomon codeword at the blowup, the Merkle trees, the DEEP quotient in the extension field | 4x (RISC Zero, Airbender), 2x (SP1 3 and 4, approximate), 2x (Stwo at rate 1/2) | RISC Zero: 9 to 10 GB per 1 M cycles (Bento) | section 1.5; Boundless pages |
|
|
| Sumcheck with a hash-based PCS (SP1 Hypercube, OpenVM SWIRL, Ceno, Ziren) | the trace as multilinears, the extension-field folded copies of the zerocheck and GKR, and the BaseFold or WHIR codeword of the stacked polynomial (still a Reed-Solomon encoding, at 4x in SP1, section 1.1) | 4x of the stacked data in SP1; WHIR's rate is a parameter | SP1: the fixed 13.9 GB plus about 6.5 GB for a 4.7 M-cycle shard (measured) | section 1 |
|
|
| Sumcheck with a curve or lattice PCS (Jolt) | the trace and the one-hot columns; **no codeword at all**: Dory commits by MSM and Akita by lattice hashing, so memory is bytes per cycle with no blowup | none | no GPU figure: 200 bytes a cycle on CPU (Akita), so the adopted 4.7 M-cycle shard is about 0.9 GB of prover RAM, approximate (derived) | a16z substack, Sep 2026; the Jolt book, streaming page |
|
|
| GKR (Expander, Ceno) | the circuit witness layer by layer; no codeword for the inner layers | none inside; a PCS for the inputs | Expander: 16 MB per Keccak, approximate | Polyhedra blog (returned 530 tonight) |
|
|
| Linear-code PCS (Ligero, Brakedown, Ligerito, Blaze) | one encoded matrix and one Merkle tree; linear time, no FFT | rate 1/2 to 1/4 | no prover memory benchmarks found; Linea's Vortex is the only production use | eprint 2021/1043, 2025/1187, 2024/1609; `linea-monorepo/prover/protocol/compiler/vortex` |
|
|
| Binius (binary field) | words of GF(2^64) and a BaseFold FRI | 2x to 4x | none; CPU only; company closed | irreducible.com posts |
|
|
|
|
The memory law in one line: a FRI or BaseFold prover holds `blowup x trace` plus the trace itself plus extension-field
|
|
working copies, so 8 to 12 bytes per cell at the peak; a Jolt-class prover holds the trace and its lookups at about 4
|
|
bytes per cell and commits without encoding. The figure that matters for us is not the ratio but the absolute: our
|
|
adopted shard is small enough (about 280 M cells, section 1.4) that a FRI-class prover sized to it fits a 12 GB card,
|
|
and a Jolt-class one fits a phone. The reason SP1 does not fit today is section 1.3, not the proof system.
|
|
|
|
### 2.3 Folding schemes
|
|
|
|
| Scheme | Prover memory per step | The verifier at the end | Field | GPU | Fit for Igneum |
|
|
|---|---|---|---|---|---|
|
|
| Nova, SuperNova, HyperNova, ProtoStar, Mova; Sonobe as the library | one step's witness plus the running instance: tiny by construction (eprint 2021/370) | an IVC proof of O(F) group elements, compressed by a SNARK: Sonobe's decider is Groth16 over BN254 with KZG, about 11.9 M constraints for a 500 k-constraint step (sonobe.pse.dev, decider page); MicroNova about 2.2 M gas (eprint 2024/2099) | curve cycles (Pasta, BN254 and Grumpkin) | partial: sppark MSM on Pasta, a GPL-3 `cuda-nova` for BN254; Sonobe lists GPU as a plan | **no**: a RISC-V step over a 256-bit curve cycle costs two MSMs per step, the opposite of the hash-based consumer-card design decision (design 5.6), and the verifier changes to pairings |
|
|
| Nexus zkVM 1 and 2 | the prover ran "on as little as 1 GB of RAM" (whitepaper, approximate) | a curve SNARK | curve cycle | none | **abandoned by its own author**: Nexus 3.0 (25 June 2025) moved to a Circle STARK, "proofs are smaller, faster to generate" (StarkWare blog) |
|
|
| LatticeFold, LatticeFold+, Neo, SuperNeo; Nightstream as the zkVM | one step plus an accumulator, lattice commitments over 64-bit fields | a Spartan-class decider | Goldilocks named; BabyBear and KoalaBear not | none; Nethermind's LatticeFold is a "proof-of-concept prototype" whose benches take 48 h; Nightstream is "research software, not production-ready" with its RV32IM prototype removed | **not before 2027 at the earliest**; the first candidate that folds small-field STARK steps |
|
|
| Arc, WARP (hash-based accumulation of Reed-Solomon proximity claims) | small: Merkle openings per step (eprint 2024/1731, 2025/753) | a FRI-style accumulator check | any STARK field | none; no public implementation found | the right primitive on paper for folding RISC-V STARK shards with a hash-based verifier; nothing to adopt |
|
|
| Mangrove, Nebula | 390 MB peak at 2^24 gates (Mangrove, eprint 2024/416); pay-per-use steps (Nebula) | curve SNARK | curve cycle | none | research |
|
|
|
|
What folding would mean for a shard: the shard prover would hold one transaction's step at a time and the memory
|
|
floor would vanish; the price is a curve-based decider at the end of every shard (seconds on a CPU, a different
|
|
verifier in the node, pairings on the light-client path), and no production code over our field. Today's small-field
|
|
zkVMs get their bounded memory from segmenting and recursion (2.4), not from folding. Folding is a watch item, not a
|
|
route.
|
|
|
|
### 2.4 Continuations and segment proving at small sizes
|
|
|
|
| Prover | The segment knob | What a 2^18 or 2^19 segment costs | Can our shard be cut that way inside the guest? |
|
|
|---|---|---|---|
|
|
| RISC Zero | `segment_limit_po2`, runtime, 13 to 24 (`env.rs:181-186`); the recursion lifts every segment at 2^18 rows and joins them in a tree | po2 19 is the documented fit for an 8 GB card and po2 20 for a 16 GB card (Boundless); the scheduler's measured tokens put po2 18 at about a third of po2 21 and a lift or join at an eighth (`factory.rs`); the 4.7 M-cycle shard at po2 19 is 9 segments, 9 lifts and 8 joins | yes, with no guest change: the zkVM cuts at the limit on its own; the shard statement is unchanged and one succinct receipt comes out |
|
|
| SP1 | `HEIGHT_THRESHOLD` (honoured) and `ELEMENT_THRESHOLD` (overwritten by the server, section 1.2); `SHARD_SIZE` up to 2^24 cycles | the live buffers shrink (22.9 GB against 28.3 GB on the prototype shard at `HEIGHT_THRESHOLD 2^20`, bench-log) and the fixed 13.9 GB does not; the compose tree folds 4 proofs at a time | yes, the same way; but the floor is the server's, so the cut buys nothing until the server is re-sized (route A) |
|
|
| Our own planner | `S_p` in pgas, a consensus parameter changed by the fee-switch pattern (`docs/plans/fee-switch-devnet.md`); the cut is at transaction boundaries (`core/src/plan.rs`) | halving `S_p` halves the live trace and doubles the shard count; the aggregator verifies one deferred proof per shard (1.66 M cycles for 4 shards, bench-log 4 October) and the chained aggregation is one per block whatever the count (9.7 s on a mining 5090) | yes, already implemented; a transaction above `S_p` stays one shard and the zkVM's own continuations cover it (spec 7.6 item 1) |
|
|
| Jolt | none: monolithic; streaming planned | n/a | no |
|
|
|
|
### 2.5 Distributed proving across several small cards
|
|
|
|
| System | How one execution is split | Per-card memory | Several cards on one host | What it means for four 12 GB cards |
|
|
|---|---|---|---|---|
|
|
| SP1 cluster (`sp1-cluster`, BSL 1.1) | by core shard: `ProveShard`, `RecursionReduce`, `RecursionDeferred` and `ShrinkWrap` tasks go to GPU workers, `CoreExecute` and the Groth16 or Plonk wrap to CPU workers (`crates/prover-types/src/lib.rs:31-41`); artifacts through Redis and S3 | "only certain GPUs with >= 24GB RAM are supported" (`infra/charts/sp1-cluster/values-example.yaml`); one task holds one whole card | yes: one GPU node process per card (`gpu{0..7}` services in the docker-compose deployment page); the local server itself supports device 0 only (`task.rs:160`, "only device 0 is supported at the moment"), one server per `CUDA_VISIBLE_DEVICES` | the split is by core shard, and the adopted shard is ONE core shard (section 1.3), so there is nothing to split across cards; the floor per card is unchanged. The cluster is a throughput tool, and its code is BSL |
|
|
| RISC Zero Bento (Boundless) | by segment onto Redis; `gpu_prove_agent` spawns one prove agent per card with `CUDA_VISIBLE_DEVICES`; the same `SEGMENT_SIZE` for every card, "the lowest common denominator"; joins form a tree (docs.boundless.network, performance-optimization and bento pages; `compose.yml:62-113`) | by `SEGMENT_SIZE` (2.1): 8 GB po2 19, 16 GB po2 20 | yes, documented: one 16 GB card 264 kHz, two 431 kHz (sub-linear, "bound by bus bandwidth, memory") | **the one documented configuration**: four 12 GB cards at po2 19 or 20 take segments off one queue and the joins fold them; the per-card floor is the segment, and the cost of small segments is the lift and join count (9 lifts and 8 joins for the adopted shard at po2 19) |
|
|
| RISC Zero `RISC0_PROVER=actor` | one process, several cards, a token budget per card from measured memory per po2 (`r0vm/src/actors/factory.rs`) | per po2 | yes, experimental since 3.0.1 | the same model without Bento's services |
|
|
| Pico Prism 2.0 | a global task queue across two machines, 16 x 5090, 100 Gbps between them (Brevis blog, May 2026) | not published | yes | no figure |
|
|
| OpenVM | metered execution on the CPU, segments to GPUs, an aggregation tree, "clusters with hundreds of GPUs" (docs, distributed-proving page) | 24 GB | yes | no 12 GB path |
|
|
| ZisK | coordinator and stateless workers; "splits the trace into pieces, proves each in parallel on separate machines, and aggregates"; the first worker aggregates a binary tree (docs, distributed execution page) | not documented | yes, `--gpu` per worker | no figure |
|
|
| Ceno | shards round-robin by `shard_id % device_count`; a shard never split across cards; one CUDA context per device (PR #1403) | not stated | yes | the same model |
|
|
| Column-split of one trace across cards (FRIttata eprint 2025/1285, HyperFond 2025/1349, deVirgo arXiv 2210.00264, Pianist 2023/1271, Cirrus 2024/1873, SumFold 2025/1653) | the sumcheck or FRI itself is distributed, each worker holding a slice of the columns or rows and exchanging small messages | a slice | research code or CPU clusters only | nothing shipped; the one route that would let four 12 GB cards hold what one 32 GB card holds for a SINGLE core shard, and nobody has it in a zkVM |
|
|
|
|
The reading. Every shipping system splits by rows (segments, shards, chunks), proves each on one card, and folds
|
|
with recursion. So "four 12 GB cards do what one 32 GB card does" is true for throughput (four shards in flight, or
|
|
four segments of one shard, then a join tree) and false for a single unit that exceeds one card: that unit must be cut
|
|
smaller, by the zkVM's segment knob (RISC Zero) or by our planner (`S_p`). For Igneum the units are already small and
|
|
independent (a shard, assigned by sortition), so the rig's natural mode is one prover process per card, each taking
|
|
its own shard. The distributed route therefore costs nothing in protocol and lands as an app change (section 3, route C).
|
|
|
|
### 2.6 Proof systems that run on AMD or Apple
|
|
|
|
| Target | What exists | Status | Source |
|
|
|---|---|---|---|
|
|
| Apple, RISC Zero Metal | the full STARK prover (rv32im, keccak, recursion) on Metal, automatic on Apple silicon; the Groth16 wrap x86 only | shipped, maintained (PR #3761's June 2026 matrix lists "metal (Mac M-series): build, run"); the only speed published is a 2023 M2 datasheet (14 to 93 kHz, approximate); nothing for M3, M4 or M5 | `risc0/sys/kernels/zkp/metal/`, dev.risczero.com local-proving page |
|
|
| Apple, ICICLE Metal (Ingonyama) | MSM, NTT, sumcheck on Metal since v3.6.0 (Mar 2025); "missing API implementations for Poseidon and Poseidon2 hashes, Merkle tree" at that release; v4.0.0 of 11 Jul 2025 is the latest | a library, closed-source backends under a free research licence (dev.ingonyama.com, install_gpu_backend page); no STARK prover built on it for Metal | ICICLE releases, the Metal blog |
|
|
| Apple, Jolt Metal | PR #1938 merged 30 Sep 2026 (the runtime and field kernels); PR #1733, the prover itself, a draft: M5 Max 2^25 cycles in 19.8 s, 3.2x over its CPU | the fastest Apple number anyone has published, in a draft | github.com/a16z/jolt pulls 1733 and 1938 |
|
|
| Apple, Stwo | CPU SIMD with NEON; ICICLE-Stwo promises Metal | CPU path shipped; no RISC-V guest of its own | stwo README, Ingonyama blog |
|
|
| Apple, Miden | `miden-gpu` on Metal | Cairo-class VM, not RISC-V | hackmd (bobbinth) |
|
|
| AMD, sppark | "A limited support for AMD's RDNA and CDNA GPUs" (README); SP1's tree carries no HIP build | a library | github.com/supranational/sppark |
|
|
| AMD, SP1 PR #2668 | an external port to RDNA3 and RDNA4 with "a caching memory allocator to work around hipMallocAsync leak bug" | **closed unmerged 20 Mar 2026** | github.com/succinctlabs/sp1/pull/2668 |
|
|
| AMD, OpenVM stark-backend HIP fork | `cuda2hip.hpp` so the same `.cu` builds under nvcc and hipcc, native `mont32_t.hip`, tested on gfx1100, targets MI300X and 7900 XTX | merged 15 Sep 2026 in a fork (Okm165/stark-backend PR #2), not upstream | the PR |
|
|
| AMD, Goldilocks NTT and STARK on ROCm | 19.19 ms NTT at 2^27 on an RX 7900 XTX; a Goldilocks STARK backend on HIP | research posts | ethresear.ch, qingming-g64-ntt and stark-g64 |
|
|
| Vulkan and WebGPU | ICICLE's Vulkan build (Jan 2025) with no installable backend; zkSecurity's WebGPU Stwo (5x on constraint evaluation, 2x end to end, no 64-bit integers in WGSL); ZPrize WebGPU MSM | prototypes; nothing proves a RISC-V shard | the pages named |
|
|
|
|
Said plainly, as `docs/analysis/amd-proving.md` said it: on 5 October 2026 no zkVM proves on an AMD GPU, and the only
|
|
Apple prover that ships is RISC Zero's. The AMD work that exists is two ports of CUDA STARK kernels through a HIP shim,
|
|
one closed, one in a fork; both are days of agent work to revive against a given tree, and PC 1's RX 9070 XT (gfx1201)
|
|
is the card to measure on.
|
|
|
|
## 3. For each route: the change to our guest, the aggregator and the node's verifier; the cost; the risk; 12 GB under 60 s
|
|
|
|
What the node verifies today: SP1 compressed proofs through `igneum-prove-host --mode verify` and `verify-segment`
|
|
(`vendor/igneum-node-pv1/igneum/exec/src/proving.rs:41-46, 878-926`), the pinned ids read at start and named in the
|
|
native statement (`program_ids`, `IGNEUM_PROOF_PROGRAM_IDS`), the record bound in a BLS-signed `ProofRecord` (version
|
|
1) or `SegmentRecord` (version 2) with the proof's SHA-256 (spec 7.7 item 1, 7.8 item 3). A different proof system
|
|
means a new `ProofSystem` version (design 5.6), a new pinned id, a second verifier command, and the record's version
|
|
field telling the node which. The swap procedure of design 5.6 (test vectors, 90% signalling, a 3-month overlap with
|
|
both verifiers, a wrap of the last old proof) is the path for any of the rows below that change the family.
|
|
|
|
The 60-s test. The litepaper's minute, the launch target of 20 to 60 s behind the tip, and the mine-and-prove
|
|
measurement that a shared card proves 3 to 4x slower (bench-log, `chain-pc2-pv1c`). No 12 GB card has run any
|
|
prover in this repository; the 12 GB times below are approximate, scaled from the 5090 by memory bandwidth (an RTX
|
|
3060 at 360 GB/s and an RTX 4070 at 504 GB/s against the 5090's 1,792 GB/s, NVIDIA's published figures, approximate),
|
|
which is the term a STARK prover is bound by. They are the numbers the first 3060-class run replaces.
|
|
|
|
| Route | Guest | Aggregator | Node verifier and record | Cost (agent time) | Risk | Reaches 12 GB with a real shard under 60 s? |
|
|
|---|---|---|---|---|---|---|
|
|
| **A. Re-size SP1's GPU server** (the prover-floor agent's patch, running tonight): remove the 20 GB panic (`builder.rs:37`), size `max_trace_size` to the shard (honour `ELEMENT_THRESHOLD`, or set the core allocation from the shard's measured cells), one core worker and a buffer of 1, a release threshold so the pool returns memory between stages, `drop_ldes` on; build with `CUDA_ARCHS` for Ampere, Ada and Blackwell | none: the same ELF, the same pinned id (the verifying key hashes the program and its preprocessed tables, not the server's buffer sizes; `HEIGHT_THRESHOLD` only shortens tables below the verifier's 2^22 maximum) | none: the compressed proof format and the aggregator guest are unchanged | none: the same `--mode verify`; the record format unchanged | hours to one day: a fork of `sp1-gpu/crates/prover_components` and `jagged_tracegen` (Apache or MIT), the 11-min cross-build, a per-card profile in `provedefault.rs`, a CI check that the fork's constants match the pinned verifier's | low on the protocol, medium on the build: the fixed recursion stage may hold the floor near 8 to 9 GB (section 1.3, approximate) and the first measurement says whether 11 GB is reached; a fork of `sp1-gpu` to carry forward on every SP1 release; the server rejects nothing it cannot hold, so an out-of-memory shard must fail cleanly and be left (the pool's rule today) | **memory: likely for the adopted shard** (10 to 11 GB on paper, section 1.4), **not** for the prototype shard (28 GB of live trace). **Time: prove-only yes** (4.3 s on the 5090 scales to about 15 to 22 s on a 3060 and 10 to 15 s on a 4070, approximate); **mine-and-prove on a 12 GB card: no at `S_p`** (3 to 4x on a shared card puts a 3060 at 45 to 90 s, approximate, and the miner's 1.7 GB on top of 11 GB does not fit), yes at `S_p/2` on a 4070 if the floor lands under 9 GB (approximate). The measurement decides; this is the route the gate waits on |
|
|
| **B. Halve `S_p`** (30,000 to 15,000 pgas, the fee-switch pattern): more and smaller shards | none | none: one deferred proof per shard, so 2x the shards per block; the chained aggregation stays one per block (9.7 s mining, 2.5 s alone) | none | hours: a fee-table change and a rollout plan like `fee-switch-devnet.md` | low: more records per block (the coinbase carries at most 8 shard records, spec 7.7 item 2, so `B_p / S_p` must stay at 8 or under); the assignment window and sortition unchanged | **alone, no**: the floor is the server's (13.9 GB at 0 cycles). **With A, it is the dial** that moves a 12 GB card from prove-only to mine-and-prove, and a 16 GB card to a comfortable fit |
|
|
| **C. One prover process per card on a rig** (the distributed route): the app runs one `sp1-gpu-server` per NVIDIA card (`CUDA_VISIBLE_DEVICES`, the per-device socket of `sp1-cuda/src/client.rs:211`, `.cuda().with_device_id(n)`), one host process per card, each taking its own assigned shard; the rig installer already picks cards (`igneum-rig-lib.sh`, `prover_decision`) | none | none: shards are independent units by design (spec 7.2); the aggregator runs on the biggest card | none | one day: the app's prover loop per card (`prover.rs` runs one loop today), the Settings and tile per card, the rig installer's prover unit per card, the socket cleanup per device (the root-socket rule of 5 October) | low; the throughput is per card, the host RAM 6 GB of pinned buffers per server (section 1.2), so a 4-card rig needs 32 GB of RAM or route A's smaller buffers | **it does not move the floor**: each card still needs A. It is the route that makes four 12 GB cards worth four shards a cycle, and it ships with A, not instead of it. Splitting ONE shard across cards is not a route: the adopted shard is one core shard (2.5), and column-split provers are research |
|
|
| **D. RISC Zero as proof system version 2** (CUDA and Metal; segments at po2 19 or 20) | a second guest: `core/` is plain Rust and ports as is; the precompile patches differ (SP1's `sha3` and `k256` patches against RISC Zero's `sha2`, `k256` and keccak circuit); the shard statement bytes unchanged; a second pinned ELF and image id in `elf/manifest.json` | a RISC Zero aggregator guest using composition (`env::verify` of the shard receipts, dev.risczero.com composition page); the chain rule (N verifies N-1) inside the family; **a block's shards must be one family**, and a chain cannot cross families inside the proof: a family switch lands at a segment boundary as a fresh chain (spec 7.8 item 6 already allows one after an unproven segment; the rule gains "or at a proof-system version change") | a second verifier mode (`--mode verify-r0`, the `risc0-zkvm` verifier, pure Rust, about 100 ms, 222 KB receipts); the record's `version` selects the family; the native statement names the family's pinned id; both verifiers in the node through the overlap of design 5.6 | 3 to 4 days: guest port and pinning 1, aggregator and chain rule 1, node verifier and record version 1, app profile and host modes 0.5, test vectors and the fast-time harness 0.5; plus the measurement day on PC 2 | medium: two proof systems in consensus for the overlap; RISC Zero's Groth16 wrap is x86 only (the light-client path of ledger P3 stays on SP1 or waits); a 222 KB receipt per shard against 1.27 MB today is a gain; the recursion tree per shard (9 lifts and 8 joins at po2 19) is extra time on small cards; `main` is at 5.0.0 with no release body, so the pin is 3.0.6 | **memory: yes by documentation** (po2 19 for an 8 GB card, po2 20 for 16 GB; 9 to 10 GB per 1 M cycles), the first documented sub-12 GB prover. **Time: approximate**: a 4090 does 808 kHz at po2 21, so the adopted shard is about 6 s on a 4090-class card and about 20 to 30 s on a 3060-class one at po2 19, prove-only; beside the miner over 60 s on a 3060, near it on a 4070. The prover-floor agent's PC 2 run is the first real number |
|
|
| **E. Airbender, OpenVM, ZisK, Pico, Ziren** as version 2 | a new guest each (RISC-V, except Ziren's MIPS); OpenVM's and ZisK's toolchains are the most complete | each has its own recursion; OpenVM's aggregation and Halo2 wrap are the most documented | a new verifier each (STARK under 300 KB for OpenVM; PLONK or FFLONK for ZisK and Airbender) | 4 to 6 days each | the same two-family cost as D with no memory gain: 21 GiB (Airbender), 24 GB (OpenVM, Ziren), undocumented (ZisK, Pico); Pico's and Ziren's GPU code is BUSL or closed | **no**: none documents a floor under 21 GiB; the race is tuned for 5090 clusters |
|
|
| **F. Jolt (Lattice Jolt) as version 2**: a sumcheck prover with no codeword; CPU and Metal | a new guest (RV64IMAC, Jolt's toolchain; no keccak precompile today, approximate, so the trie hashing costs more cycles than in SP1) | **none exists**: no recursion or continuation shipped, so the aggregator would verify N shard proofs natively and the chain rule would live in the native statement until Jolt's recursion lands | a Dory verifier (BN254 pairings, about 50 KB, sub-second, approximate) or an Akita verifier (lattice, 65 to 80 KB); no on-chain verifier shipped | 5 to 8 days for the guest, the verifier and the record; the aggregator question has no answer in the code | high: alpha software, no audit, no production user, no recursion; the proof system of the miner's CPU, not of its card | **memory: yes by a wide margin** (about 0.9 GB for the adopted shard at 200 bytes a cycle, approximate). **Time on a CPU: about 2 to 3 s** for 4.7 M cycles at over 2 M cycles a second (a16z, Sep 2026, laptop CPU; approximate for our guest), **on Metal under 1 s** (PR #1733's 2^25 in 19.8 s on an M5 Max, approximate). The numbers are the best in this document and the software is not shippable |
|
|
| **G. Folding** (Nova family, lattice folding) | a step circuit per transaction or per opcode group | a decider per shard | pairing or lattice verifier | weeks of research, no code over our field | the family that Nexus left | **no** today; the watch item for 2027 |
|
|
| **H. AMD through a HIP port of SP1's kernels** (PR #2668 revived against 6.8.1, or the `cuda2hip` shim of the OpenVM fork) | none | none | none: the same SP1 proofs | 3 to 5 days plus PC 1's RX 9070 XT to measure; the `hipMallocAsync` leak needs the caching allocator the PR carried | medium: a kernel port with no upstream; the sppark NTT has a limited HIP path and cuPQC none | memory as route A (the same buffers); **time unmeasured on any AMD card**; the one route that gives AMD miners the 20% pool share |
|
|
| **I. Apple through RISC Zero Metal** (route D's Metal half) | as D | as D | as D | inside D's 3 to 4 days | the 2023 M2 figure (14 kHz, approximate) says 5 minutes for the adopted shard; an M5 Max is not measured by anyone | **memory: yes** (unified memory, 64 GB on the M5 Max). **Time: unknown**; the Mac measure lock run is the number |
|
|
|
|
## 4. The ranked recommendation
|
|
|
|
| Rank | Route | Why | Gate |
|
|
|---|---|---|---|
|
|
| **1. Soonest to 12 GB with the least change: A, with B as the dial and C for rigs** | re-size SP1's GPU server; keep the guest, the aggregator, the verifier and the pinned ids exactly as they are; set `S_p` from the first 12 GB measurement; one prover per card on rigs | nothing in consensus moves; the work is a fork of two Apache crates and an app profile; it is already running tonight; every other route costs days and adds a second verifier | the prover-floor agent's rows: the adopted shard under 11 GB alone and the time on the first 3060-class or 4070-class card, prove-only and beside the miner. If under 11 GB and under 60 s prove-only: ship 0.3.12 with the 12 GB tier as prove-only and `S_p/2` measured for mine-and-prove. If not under 11 GB: route D |
|
|
| **2. The fallback if A misses 11 GB, and the Apple route either way: D, RISC Zero as version 2** | the only shipped prover with a documented sub-12 GB configuration and a shipped Metal path; Apache or MIT including the kernels; 222 KB receipts | the swappable interface was built for this (design 5.6) and the node already names the pinned id in the statement, so a second family is a version, not a redesign; the cost is 3 to 4 days plus the overlap | PC 2's po2 19 and 20 rows (memory, time per segment, lift and join) tonight; the Mac's Metal row |
|
|
| **3. Best in five years: the sumcheck family without a codeword (Jolt-class), or the sumcheck-plus-WHIR family SP1 and OpenVM already converge on** | Jolt proves the adopted shard in seconds on a laptop CPU at under 1 GB of memory, which is the only route that gives AMD-only, Apple and 8 GB machines the proving share with their existing hardware; its verifier is small (50 to 80 KB); its licence is MIT or Apache. It is alpha with no recursion, so not before it ships a stable release with continuations and an audit. SP1 Hypercube and OpenVM SWIRL are the same mathematics with a hash-based PCS and a GPU today, which is why staying on SP1 now loses nothing in that direction | do not adopt now; re-read Jolt and the Arc or WARP accumulation line at every 6-month era draw (design 5.6's swap procedure needs 90% signalling and a 3-month overlap, so the lead time is the schedule) | a stable Jolt tag with recursion, an audit, and a CUDA or merged Metal prover |
|
|
| **The interface question** | yes: `ProofSystem` is versioned (`VERSION`, `program_id`, `verify_segment`), the record carries `version`, the node reads pinned ids at start and names them in the native statement, and the overlap procedure keeps both verifiers in the node for 3 months with `B_p` from the stricter table. What is missing for two families at once is small and named: the record version selecting the verifier command, the fresh-chain rule at a version change, and the shard plan carrying the family per block so a block's shards are homogeneous (the aggregator folds one family). Those three items are in route D's day of node work | so the answer to "ship one now and move to the other later" is yes, and route A ships nothing that has to be undone | |
|
|
|
|
The honest statement of what this ranking does not know: no 12 GB card has run any prover here. Route A's time
|
|
figures are bandwidth scaling, labelled approximate; route D's are a 4090 figure scaled the same way. The first 3060
|
|
or 4070 in this repository replaces both columns, and the plan is to borrow or buy one this week (a 4070 is the
|
|
common 12 GB card of 2026; a 3060 the common older one; both are the gate's named class, design R2).
|
|
|
|
## 5. The tier consequences, and the public line while the change is made
|
|
|
|
Every number carries its consequences (CLAUDE.md, 5 October 2026). The table says what each tier has today on SP1
|
|
6.8.1, what route 1 (A plus B plus C) gives it if the gate is met, what route 2 (D) adds, and what only route 3 would
|
|
give. "Today" is measured; the rest is the routes' expected outcome, labelled, until the measurement.
|
|
|
|
| Tier | Today (measured, bench-log 5 October) | Route 1: re-sized SP1 server, `S_p` as the dial, one server per card | Route 2: RISC Zero version 2 | Only route 3 (sumcheck without a codeword) |
|
|
|---|---|---|---|---|
|
|
| Home miner, one 8 GB NVIDIA card | mines; proves nothing (the server panics under 20 GB) | proves nothing at `S_p` (the floor's fixed terms, 8 to 9 GB approximate, leave no room); perhaps empty shards | prove-only at po2 19 (Boundless' 8 GB tier), the miner paused per shard; time approximate 30 to 60 s | mines and proves on its CPU |
|
|
| Home miner, one 12 GB card (3060, 4070) | mines; proves nothing; the litepaper's gate card | **prove-only at `S_p`** if the floor lands under 11 GB (expected, section 1.4): about 15 to 22 s a shard, approximate; **mine-and-prove at `S_p/2`** on a 4070 if the floor is under 9 GB, approximate; on a 3060 the shared card misses 60 s, approximate, so its default is prove-only with the miner paused per shard (the 16 GB rule of `provedefault.rs` today, moved down a tier) | prove-only at po2 20 (16 GB tier) or po2 19; mine-and-prove not inside 60 s on a 3060, approximate | mines and proves, CPU |
|
|
| Home miner, one 16 GB card (5080, 4080, 4060 Ti 16 GB) | an empty shard alone (13.9 GB); nothing beside the miner | **mine-and-prove at `S_p`** (11 GB plus the miner's 1.7 GB), about 7 to 12 s a shard alone and 20 to 40 s beside the miner, approximate | mine-and-prove at po2 20 | the same |
|
|
| Home miner, one 24 GB card (4090, 3090) | the adopted shard alone (20.4 GB) and beside the miner (22.2 GB, approximate for the card); the prototype shard never | mine-and-prove at `S_p` with 10 GB to spare; the prototype shard (28 GB live) only if the devnet's fee switch has passed, which it has from DAA 210,000 | the same with Metal irrelevant | the same |
|
|
| Home miner, one 32 GB card (5090) | everything, measured | everything, with more shards in flight if the pool releases between stages | the same | the same |
|
|
| Rig, several NVIDIA cards | one prover on the biggest card (`prover_decision`) | **one server per card**, each its own shard; the aggregator on the biggest card; host RAM 6 GB pinned per server today, under 2 GB with route A's buffers | the same model (Bento's) | the same |
|
|
| Pool user | through the pool; who proves is open (spec 09) | unchanged | unchanged | unchanged |
|
|
| AMD-only (RX 9070 XT, 7900 XTX) | mines; proves nothing on the card; the CPU path 282 s a shard at 30 GB | unchanged until route H (a HIP port, 3 to 5 days, measured on PC 1's 9070 XT) | unchanged: RISC Zero is CUDA and Metal only | mines and proves on its CPU |
|
|
| Apple silicon (M-series) | mines (26.7 MH/s on the M5 Max); the SP1 CPU prover 41 to 55 s for an empty shard, 272 s for a small one | unchanged | **proves on the GPU through Metal** (64 GB unified memory on an M5 Max holds any segment); the time is the measurement | proves in seconds on Metal (Jolt's draft PR figure, approximate) |
|
|
| Windows under 32 GB of RAM | off (the WSL2 prover held 7.9 GB) | the pinned buffers fall with `max_trace_size`, so a 16 GB PC likely qualifies, approximate; measure | RISC Zero's CUDA path also runs in WSL2 | |
|
|
|
|
The deadlines these fit (spec 7.2 item 3, the litepaper, `proving-v1.md`): the 10-s exclusive window is the 5090's
|
|
alone; a 12 GB card at 15 to 22 s proves its assigned shards in the open phase and is paid when no faster card took
|
|
them, which on a chain with few 5090s is most of the time; the minute of the litepaper holds for prove-only 12 GB
|
|
cards and for mine-and-prove 16 GB cards; the 600-s unproven deadline holds for every tier above the CPU path.
|
|
|
|
### The public line while the change is made
|
|
|
|
The litepaper's sentence today ("Target: shard size will be set so a 12 GB card proves one shard in about 20
|
|
seconds") is a target and says so (fud-ledger P1, overclaim 27). What this document adds, for `site/litepaper.html`,
|
|
`site/miner.html` and the app's Proving tile, in the copy law:
|
|
|
|
> Proving runs on NVIDIA cards with 24 GB or more today. A build for 12 GB and 16 GB cards is being measured: the
|
|
> memory is the prover's buffers, not the shard, and the fix is a smaller build of the same prover. AMD and Apple
|
|
> cards mine. A second prover with an Apple path exists and is the fallback.
|
|
|
|
And the rule for the next status line, whichever way the measurement goes: the number, the card it was taken on, and
|
|
the tier it moves, in one sentence, the day it is taken.
|
|
|
|
### What this document does about it
|
|
|
|
| Consequence | Action | Owner |
|
|
|---|---|---|
|
|
| The gate card has never run a prover here | get a 4070 or 3060 into the measurement loop this week; until then every 12 GB figure stays approximate | coordinator; the project lead for the card |
|
|
| Route A's gate | the prover-floor agent's rows (asked for by message tonight); if under 11 GB, `provedefault.rs` gains the 12 GB prove-only and 16 GB mine-and-prove tiers and the rig installer one server per card | prover-floor agent, then the proving engineer |
|
|
| Route D's measurement | RISC Zero 3.0.6 at po2 19 and 20 on PC 2 (CUDA) and on this Mac (Metal), the same shard statement run natively: memory, time per segment, lift and join, receipt size | prover-floor agent (PC 2); a Mac measure job for Metal |
|
|
| The two-family node items (record version selects the verifier, fresh chain at a version change, one family per block) | spec 7.8 gains the three rules when route D starts; nothing changes before | execution engineer |
|
|
| AMD | route H is a 3-to-5-day job with a measurement on PC 1's 9070 XT; opened as a plan when route A's result is in | execution engineer |
|
|
| The public line | the paragraph above to the site and the tile with the next site pass | site-pages owner |
|
|
|
|
## Sources
|
|
|
|
Our own: `docs/bench-log.md` entries "proving v1: segment records, the chain rule, the unproven rule" (5 October 2026),
|
|
"the SP1 CPU prover on PC 1" (5 October), "shard proving on the RTX 5090" (4 October); `docs/plans/proving-v0.md`,
|
|
`proving-v1.md`; `docs/analysis/amd-proving.md`; `docs/spec/07-execution.md` 7.2, 7.6, 7.7, 7.8; `docs/design/execution-layer.md`
|
|
5.1 to 5.7; `proving/igneum-prove` (`host/src/proof_system.rs`, `program/src/main.rs`, `aggregator/src/main.rs`,
|
|
`elf/manifest.json`); `vendor/igneum-node-pv1/igneum/exec/src/proving.rs`; `app/igneum-app/src/provedefault.rs`, `prover.rs`.
|
|
|
|
SP1 6.8.1, read from `~/.cargo/registry/src/index.crates.io-*/` and the vendored tree `vendor/sp1-6.8.1` (commit
|
|
c84ada1e, 24 Sep 2026) on the `prover-floor` worktree: `sp1-core-executor-6.8.1/src/opts.rs`, `src/utils.rs`,
|
|
`src/artifacts/rv64im_costs.json`; `sp1-prover-6.8.1/src/components.rs`, `src/worker/config.rs`, `src/shapes.rs`;
|
|
`sp1-primitives-6.8.1/src/fri_params.rs`; `sp1-verifier-6.8.1/src/compressed/config.rs`; `sp1-hypercube-6.8.1/src/verifier/config.rs`;
|
|
`sp1-cuda-6.8.1/src/server.rs`, `src/client.rs`; `sp1-gpu/README.md`, `sp1-gpu/crates/prover_components/src/builder.rs`,
|
|
`src/components.rs`, `sp1-gpu/crates/jagged_tracegen/src/lib.rs`, `sp1-gpu/crates/shard_prover/src/prover.rs`,
|
|
`sp1-gpu/crates/cuda/src/task.rs`, `src/device.rs`, `sp1-gpu/crates/sys/lib/runtime/mem_pool.cu`, `sp1-gpu/crates/zerocheck/src/primitives.rs`.
|
|
Web: docs.succinct.xyz (hardware-acceleration, hardware-requirements, proof-types, security-model, provers introduction,
|
|
cluster architecture, docker-compose deployment); blog.succinct.xyz (sp1-hypercube, real-time-proving-16-gpus,
|
|
sp1-hypercube-is-now-live-on-mainnet); github.com/succinctlabs/sp1 releases v6.0.0 to v6.8.1, issues #2674, #2930,
|
|
#2950, #2969, pulls #2631, #2668, #2723, #2917, #2974; github.com/succinctlabs/sp1-cluster (README, LICENSE,
|
|
`infra/charts/sp1-cluster/values-example.yaml`, `crates/worker/src/config.rs`); eprint 2025/917 (jagged polynomial commitments).
|
|
|
|
RISC Zero: `~/.cargo/registry` crates `risc0-zkp-3.0.4/src/lib.rs`, `risc0-zkvm-3.0.4/src/receipt.rs`, `src/host/recursion/prove/mod.rs`,
|
|
`risc0-circuit-rv32im-4.0.4/src/execute/mod.rs`, `src/zirgen/defs.rs.inc`, `src/prove/hal/cuda.rs`; github.com/risc0/risc0
|
|
`risc0/zkvm/src/host/client/env.rs`, `risc0/zkvm/Cargo.toml`, `risc0/zkvm/build.rs`, `risc0/sys/kernels/zkp/{cuda,metal}/`,
|
|
`risc0/r0vm/src/actors/factory.rs`, `risc0/circuit/recursion/src/lib.rs`, releases v2.0.0, v3.0.1, v3.0.6, pull #3761;
|
|
dev.risczero.com (local-proving, composition); docs.boundless.network (bento, performance-optimization, quick-start);
|
|
github.com/boundless-xyz/boundless (`compose.yml`, `bento/README.md`, `bento/LICENSE-BSL`); github.com/ekrembal/gsr-stark-verifier pull 5; l2beat.com/zk-catalog/risc0.
|
|
|
|
Others: zksync.io/airbender, docs.zksync.io airbender and proving pages, github.com/matter-labs/zksync-airbender (README,
|
|
`docs/gpu.md`, pull #448), veridise.com (the Airbender audit); 0xpolygonhermez.github.io/zisk (introduction, limits,
|
|
distributed execution, installation), github.com/0xPolygonHermez/zisk (README, pull #1238, `zisk-contracts`);
|
|
blog.openvm.dev (2.0, 2.0-production, 2.1, openvm-gpu, v1), docs.openvm.dev (security-model, distributed-proving, sdk);
|
|
pico-docs.brevis.network, github.com/brevis-network/pico and pico-gpu (README, LICENSE), blog.brevis.network (Prism 1.0,
|
|
2.0, 2.1); docs.zkm.io (prover, performance), github.com/ProjectZKM/Ziren, zkm.io (the independent evaluation of v1.1.4),
|
|
eprint 2026/2330; github.com/starkware-libs/stwo and stwo-cairo (README), ingonyama.com (ICICLE-Stwo, the Starknet
|
|
partnership, ICICLE Metal v3.6), dev.ingonyama.com (install_gpu_backend), blog.zksecurity.xyz/posts/webgpu, starkware.co
|
|
(S-two 2.0.0, Nexus on S-two), theblock.co (S-two on Starknet); github.com/a16z/jolt (README, book: intro, dory, akita,
|
|
streaming, recursion, blindfold; pulls #1733, #1938; tags), a16zcrypto.substack.com ("How to prove software ran
|
|
correctly", Sep 2026), a16zcrypto.com (jolt-6x-speedup, 64-bit-proving-jolt, zkvm-jolt-zero-knowledge, faqs-on-jolts-initial-implementation),
|
|
eprint 2025/611; github.com/scroll-tech/ceno (README, Cargo.toml, pull #1403), ceno-gpu-mock, scroll.io (Ceno post),
|
|
osec.io ("zkVMs' unfaithful claims"); github.com/nexus-xyz/nexus-zkvm (README, LICENSE), blog.nexus.xyz (roadmap);
|
|
lita.gitbook.io (Valida architecture, benchmarks); github.com/powdr-labs/powdr; irreducible.com (announcing-binius64,
|
|
reinventing-irreducible, irreducible-shutting-down), github.com/binius-zk/binius64, eprint 2026/1656; eprint 2021/1043,
|
|
2022/1010, 2024/1609, 2025/1187, 2024/1586, 2024/185 (linear-code commitments, WHIR, Vortex), github.com/Consensys/linea-monorepo;
|
|
PolyhedraZK/Expander and blog.polyhedra.network (returned 530 tonight); eprint 2021/370, 2024/2099, 2024/1220, 2024/416,
|
|
2024/1605, 2025/247, 2025/294, 2026/242, 2024/1731, 2025/753, 2026/1371 (folding and accumulation), sonobe.pse.dev,
|
|
github.com/privacy-scaling-explorations/sonobe, NethermindEth/latticefold, LFDT-Nightstream/Nightstream; eprint 2023/1271,
|
|
2024/1208, 2024/1873, 2025/1349, 2025/1653, 2025/1285, 2018/691, arXiv 2210.00264, 2602.16338 (distributed proving);
|
|
github.com/supranational/sppark, github.com/Okm165/stark-backend pull 2, ethresear.ch (qingming G64 NTT and STARK on ROCm);
|
|
github.com/cysic-labs/venus, erigon.tech (Zilkworm), github.com/DelphinusLab/prover-node-docker, hackmd.io/@bobbinth
|
|
(Miden), ethproofs.org/clusters (5 October 2026).
|