131 lines
20 KiB
Markdown
131 lines
20 KiB
Markdown
# Proving on AMD and Apple cards: what exists, what the CPU can do, what to tell the public
|
|
|
|
5 October 2026, from the project lead's two questions that evening: "test proving on the amd card?" and "can we test proving on
|
|
mac?". PC 1 holds an RTX 5090 and an RX 9070 XT (gfx1201, 16 GB) in an eGPU; this Mac is an M5 Max. The prover is
|
|
SP1 (`proving/igneum-prove`, `docs/plans/proving-v0.md`, `proving-v1.md`), run on the GPU only through SP1's CUDA
|
|
server. Every figure below is measured (with its bench-log entry or job id) or cited (with its file or page); the
|
|
rest is labelled approximate. Status words follow `docs/spec/00-overview.md` 0.2.
|
|
|
|
**The answer in three lines.** No zkVM proves on an AMD GPU on 5 October 2026: not SP1, not RISC Zero, not Jolt, not
|
|
OpenVM, and the ICICLE library underneath them has no AMD backend either. Apple silicon has a shipped Metal prover in
|
|
RISC Zero and a Metal backend in ICICLE, but SP1, the prover Igneum runs, is CPU-only on a Mac. So an AMD-only or
|
|
Apple-only machine mines and does not prove on its card; it can prove on its CPU, at the times measured in section 2.
|
|
|
|
## 1. The backends (read 5 October 2026, 20:30 to 20:50 UTC)
|
|
|
|
| Prover | Version read | CPU | NVIDIA (CUDA) | AMD (ROCm or HIP) | Apple (Metal) | Vulkan or WebGPU | Where it says so |
|
|
|---|---|---|---|---|---|---|---|
|
|
| SP1 (ours) | v6.8.1, 24 Sep 2026 (pinned); `dev` head 318dd530, 28 Sep 2026 | yes; AVX2 and AVX-512 on x86 through Plonky3 | yes: `sp1-gpu-server`, "Compute Capability 8.0 or higher", "24GB or more VRAM", "the CUDA 12 runtime and a compatible NVIDIA driver", Linux x86_64 | **no** | **no** | **no** | docs.succinct.xyz, SP1 docs "Hardware acceleration" page; `crates/sdk/src/lib.rs` (`pub mod cpu`, `mock`, `light`, `#[cfg(feature = "cuda")] pub mod cuda`, `#[cfg(feature = "network")] pub mod network`: no other backend module); `sp1-gpu/README.md` (`CUDA_ARCHS` 89, 90, 100, 120; NTT by NVIDIA cuPQC or sppark); release notes v6.2.3 to v6.8.1 (the only backend line: "add optional cuPQC NTT backend", v6.8.0); a code search of the repository on 5 October: "rocm" 0 files, "metal" 0, "vulkan" 0, "webgpu" 0; `cuobjdump` of sp1-gpu-server 6.8.1: sm_80, 86, 89, 90, 100, 120 and compute_120 PTX, nothing else (`docs/bench-log.md`, "proving v1", 5 October 2026) |
|
|
| sppark (SP1's NTT fallback, vendored at `sp1-gpu/crates/sys/sppark`) | `main` README, read 5 October 2026 | | yes: "x86_64 with Nvidia's Volta+ GPU hardware platforms on Linux and Windows" | "A limited support for AMD's RDNA and CDNA GPUs is provided" (upstream README). SP1's tree carries no HIP build: the 0 "rocm" files above, and `sp1-gpu/crates/sys/sppark/util/gpu_t.cuh` is CUDA only | no | no | github.com/supranational/sppark README; the SP1 files named |
|
|
| RISC Zero | latest release v3.0.6, 17 Jul 2026 (a v5.0.0-rc.1 of 15 Jan 2026 is also on the releases page) | yes, "nearly any modern CPU (x86 or ARM)" | yes, "RISC Zero targets NVIDIA GPUs using the CUDA framework" | **no** ("rocm", "vulkan": 0 files in the repository) | **yes**: `metal = ["prove"]` in `risc0/zkvm/Cargo.toml`; kernels in `risc0/sys/kernels/zkp/metal/*.metal` (zk, fri, mix, sha); docs: "RISC Zero will use the integrated Metal compute cores" on Apple silicon. The Groth16 wrapper "only works on x86 architecture, and so Apple Silicon is currently unsupported (even via Docker)" | no | dev.risczero.com "Local proving"; `risc0/zkvm/Cargo.toml` features `cuda = [... risc0-zkp/cuda ...]`, `metal = ["prove"]` |
|
|
| Jolt (a16z) | v0.3.0-alpha, 1 Oct 2025; "Jolt is in alpha and is not suitable for production use" | yes, "state-of-the-art performance on CPU" | no | **no** | a **draft** PR #1733 (opened 3 Aug 2026, not merged): titled "feat: Metal GPU backend (Apple Silicon)" and marked experimental: 92 Metal kernels, "2.12x speedup at 2^20 scale" on an M4 mini and "3.20x vs same-binary CPU" on an M5 Max, "Apple Silicon + macOS only. No CI coverage" | no | github.com/a16z/jolt README and book (jolt.a16zcrypto.com); PR #1733 |
|
|
| OpenVM | v2.0.2, 14 Aug 2026 | yes | yes: `cuda-backend` (v1.4.2 notes), "Improves the Halo2 GPU prover" (v2.0.2) | **no** | **no** | no | github.com/openvm-org/openvm releases |
|
|
| ICICLE (Ingonyama; the GPU library behind several provers, not SP1) | v4.0.0, 11 Jul 2025 | yes (MIT) | yes, "CUDA (for NVIDIA GPUs)" | **no** backend listed | yes, "Metal (for Apple Silicon GPUs)"; both under a special licence with a free research licence | Vulkan in the build system (PR #735, merged Jan 2025) and a draft "Vulkan NTT" PR #1019 (Jul 2025, "still wip"); nothing installable | dev.ingonyama.com "Install GPU backend"; the releases page; the README ("backends ... are distributed under a special license") |
|
|
|
|
Said plainly: **on 5 October 2026 no zkVM proves on an AMD GPU.** The only AMD code in the whole chain is sppark's
|
|
limited HIP path, which SP1 does not build. For Apple silicon the answer is split: RISC Zero ships a Metal prover
|
|
and ICICLE a Metal backend; SP1, Jolt and OpenVM do not. Nothing read tonight names an AMD plan with a date.
|
|
|
|
## 2. The CPU fallback, measured
|
|
|
|
SP1's CPU prover is the path an AMD-only or Apple-only machine has today. Three machines, the same pinned guests
|
|
(shard program id `0x2b1a81cb...`, aggregator `0x474678f3...`, pinned 2026-10-05T16:20:38Z), `SP1_PROVER=cpu`,
|
|
`--mode shard --shard 0` (execute, core proof, compressed proof, each verified). The RTX 5090 rows are the reference.
|
|
|
|
| Fixture (SP1 cycles) | Stage | PC 1 CPU, miner running on both cards (job `cpu-prove-pc1-small2`) | Apple M5 Max CPU (4 October, loaded; bench-log) | RTX 5090 (bench-log) |
|
|
|---|---|---|---|---|
|
|
| block-56-transfers-3shards shard 0, 200 pgas (315 k) | core | 82.5 s, 7,310,257 B, verify 0.210 s | 83.1 s, 7,310,257 B | not run on the 5090; the nearest rows are block-78 below and an empty live shard: 7.0 to 7.7 s compressed with the miner on the card (5 Oct, `chain-pc2-pv1b`, `pv1c`) |
|
|
| | compressed | 199.2 s, 1,272,897 B, verify 0.035 s; 312 s wall for setup 22.8 s, execute 0.14 s, core, compressed | 272.3 s, 1,272,897 B | |
|
|
| | peak RSS, CPU | 29.5 GB peak RSS; 978% CPU (9.8 of 16 cores), user 2,516 s, system 537 s | not recorded | |
|
|
| block-78-increment, 2 transactions (626 k) | core | 87.0 s, 7,317,857 B, verify 0.209 s | 22.0 s, 7.3 MB (3 October, v0 guest) | 1.4 s (4 October, mining paused) |
|
|
| | compressed | 202.3 s, 1,272,897 B, verify 0.034 s; 322 s wall (setup 21.8 s) | 55.7 s, 1.27 MB | 2.7 s |
|
|
| | peak RSS, CPU | 30.5 GB peak RSS; 979% CPU, user 2,616 s, system 541 s | not recorded | |
|
|
| block-338-shard1, one shard at `S_p` (60.8 M) | core | **not run**, by the PC 1 scheduler's decision at 21:05Z (PC 1's time tonight belongs to the Counter ASIC 2.0 gates; the job `cpu-prove-pc1-sp`, script `tools/amd-prove/pc1-cpu-prove-sp.ps1`, is written and unpublished). Extrapolation, approximate: 60.8 M cycles is about 29 SP1 shards of 2^21 cycles where the small fixtures are one, so the core proof alone is about 29 x 80 s, 40 min, and the compressed recursion over 29 shard proofs adds hours; the floor from the 5090's own ratios (6x on core, 4x on compressed between block-78 and `S_p`) is 9 min core and 13 min compressed. Either way far outside every deadline | not run on the CPU (execute alone 6.9 s) | 8.3 s |
|
|
| | compressed | not run (see the core cell) | not run | 10.9 s with the card to itself (4 Oct); 33.0 s with the miner running (5 Oct, `memminer-pc2-pv1`); 7.3 to 7.7 s per EMPTY shard with the miner running (`chain-pc2-pv1c`) |
|
|
| | peak RSS, CPU | not run; at least the 30 GB of the small rows | | GPU peak 28,295 MiB alone, 30,039 MiB beside the miner |
|
|
|
|
PC 1: Windows 11, WSL2 Ubuntu 24.04 as root, 16 cores and 46,994 MB visible to the VM, the Igneum Miner app 0.3.9 mining on the RTX 5090 (89% mean utilisation through both runs, 59 to 70% minimum: the miner, untouched) and on the RX 9070 XT (not visible to nvidia-smi, mining through the app's OpenCL worker). Job `cpu-prove-pc1-small2`, 20:49:00Z to 20:59:49Z, 649 s wall including a 6-s warm build; the host built without the `cuda` feature from the hosted package `igneum-prove-wsl2-pv1b.zip`, `--mode id` the pinned pair. The first job, `cpu-prove-pc1-small` (20:44 to 20:46Z), built the host cold in 126 s and proved nothing: an apostrophe inside a single-quoted awk program ended the quote, bash refused the whole loop and the job reported exit 0. The class fix: `tools/amd-prove/check-job-bash.sh` runs `bash -n` on the bash body of a PowerShell job before it is published, and the job itself runs `bash -n` inside the distro before the run; both were shown to fire on the bad body and pass the fixed one. Host RAM in the VM: 968 MB used before, 2,351 MB after; the prover's own peak 29.5 to 30.5 GB.
|
|
|
|
Mac, fresh run tonight: not taken. The Mac measure lock was held from 20:31Z (a read-width `packbench` under `measure`, three build slots, then a 1,500-s proving-v1 network under `run`) and did not free inside the 10-minute window the coordinator set, so the Apple column is the 4 October rows (M5 Max, 18 cores, 64 GB, load 38 to 47, `nice -n 19`): the same host modes on the same fixture, under heavier load than PC 1 tonight. The Mac's RAM peak was not recorded on 4 October; PC 1's 30 GB says a Mac needs more than 32 GB for the CPU prover, which a 64 GB M5 Max has and a 16 or 24 GB Mac does not.
|
|
|
|
The deadlines a CPU proof has to fit (all in the spec and the v1 plan): the exclusive window of an assigned shard is
|
|
10 s of DAA time (spec 7.2 item 3; after it anyone may prove and be paid first); the litepaper promises the block's
|
|
proof "within about a minute"; the launch target is 20 to 60 s behind the tip; from proving v1 a segment nobody has
|
|
proven in `T` = 600 DAA s (10 min) pays nothing (`docs/plans/proving-v1.md`, decisions). So a CPU shard proof is
|
|
useful only if it lands inside 10 min and competitive only if it lands inside about a minute.
|
|
|
|
Reading. On PC 1 the CPU proof of the smallest shard (315 k cycles) and of the two-transaction block (631 k cycles) cost the same: 82.5 and 87.0 s core, 199.2 and 202.3 s compressed. Doubling the cycles added 4.5 s to the core proof and 3.1 s to the compressed one, so about 280 s of every CPU proof is fixed cost (the recursion that turns the core proof into the 1.27 MB compressed proof the chain carries), and no shard size removes it. Against the deadlines: 282 s a shard (core plus compressed, the client already set up, as the app's loop runs it) is 28x the 10-s assignment window, 4.7x the minute the litepaper promises, and inside the 600-s unproven deadline of v1 with 5 min to spare; but an NVIDIA card proves the same shard in 2.7 to 7.7 s, so a CPU prover only ever wins a shard that no card has taken in 10 minutes. The Mac's 83.1 and 272.3 s of 4 October have the same shape. The RAM peak of 29.5 to 30.5 GB is the second finding: the SP1 CPU prover does not fit a 16 GB machine at all, and WSL2 gives a Windows VM half the host's RAM by default, so the CPU path needs a 64 GB Windows PC or a 32 GB Linux or Mac machine. The `S_p` shard on the CPU can only be slower (the 5090 takes 6x longer at `S_p` than on block-78: 8.3 s against 1.4 s core); it was not run tonight (the scheduler kept PC 1 for the Counter ASIC 2.0 gates) and could not change the conclusion.
|
|
|
|
## 3. What this means for each tier (the every-number rule, CLAUDE.md 5 October 2026)
|
|
|
|
| Tier | Mines | Proves on the card | The 20% proving-pool share (spec 2.5) | What the software does today |
|
|
|---|---|---|---|---|
|
|
| AMD-only home miner, one card of 8, 12 or 16 GB (an RX 9070 XT is 16 GB), Windows or Linux | yes (OpenCL worker, `proto-opencl`; PC 1's 9070 XT mines on the devnet) | **no**: no prover exists for the card | **lost**, unless CPU proving at a small shard size becomes a tier (section 4a) | the rig installer: `prover_decision` in `packaging/linux/bin/igneum-rig-lib.sh` (branch `rig-install`) skips every non-NVIDIA card (`[[ "$vendor" == nvidia ]] \|\| continue`) and prints "proving off by default: no NVIDIA card (no CUDA prover for AMD or Intel yet)"; the app: `provedefault.rs` (branch `proving-v1`) considers NVIDIA cards only. Both already right; neither offers the CPU path |
|
|
| Apple silicon (M-series, unified memory) | yes: the M5 Max at 26.7 MH/s (bench-log 4 October, "first hourly program swap", Metal `prepare 1` row) | **no** with SP1; RISC Zero and ICICLE have Metal, SP1 does not | **lost** today; a Metal prover behind the swappable interface would restore it (section 4b) | `provedefault.rs`: "proving stays off on Apple silicon: the M5 Max CPU took 41 to 55 s for an empty shard and minutes for a full one; Settings switches it on (CPU, slow)". Right |
|
|
| Mixed rig (NVIDIA and AMD cards in one box) | every card | the NVIDIA cards prove for the box; the AMD cards mine | kept, earned by the NVIDIA cards | the rig installer picks the biggest NVIDIA card (`prover_decision`, `PROVER_CARD` overrides), pauses its miner under 20 GB, keeps it mining at 20 GB or more; the AMD cards get a miner unit each. **The prover unit must never select an AMD card**: it does not (the vendor filter above), and that filter is now a stated requirement, not an accident |
|
|
| NVIDIA home miner, 8 or 12 GB | yes | no on this SP1 build (13.9 GB floor on an empty shard, `memsweep-pc2-pv1`) | lost unless the shard size moves | unchanged from `proving-v1.md` |
|
|
| NVIDIA 16 GB | yes | prove-only, miner paused per shard | kept | unchanged |
|
|
| NVIDIA 24 or 32 GB | yes | mines and proves (peak 16.8 GB on empty shards, 30.0 GB on a full prototype shard beside the miner) | kept | unchanged |
|
|
| Pool user | through the pool | the pool's own NVIDIA cards prove the shards assigned to the pool's keys (approximate: the pool protocol, spec 09, does not yet say who proves) | by the pool's rules | open, spec 09 |
|
|
|
|
## 4. The options
|
|
|
|
### 4a. CPU proving at a small shard size, as a tier
|
|
|
|
What it is: an AMD-only or Apple machine proves shards cut at a smaller budget than `S_p` on its CPU, through the
|
|
same host (`SP1_PROVER=cpu`; the host's `--budget` re-plan from branch `proving-v1`, commit c2544be, cuts a fixture at
|
|
any budget). The miner keeps the card; the prover takes the CPU.
|
|
|
|
What the numbers say: the fixed cost kills it. 282 s a shard on a 16-core PC and 355 s on the loaded M5 Max, with 30 GB of RAM, at the smallest shard there is; the time sits in the compressed-proof recursion, not in the cycles, so cutting shards smaller does not help, and the launch deadline (20 to 60 s behind the tip) is missed by 5x. It fits only the v1 unproven deadline (600 s), which pays a CPU prover only when no card has proven the shard in 10 minutes: on a chain with one NVIDIA prover that never happens. Recommendation: **no CPU tier**. Settings may still switch the CPU prover on (it does on macOS today), and the Proving tile must then say the proof takes about five minutes and is paid only when no card proves first.
|
|
|
|
What it costs the chain: a block cut into more, smaller shards costs more aggregation work (the aggregator guest
|
|
verifies one deferred proof per shard; 1.66 M cycles for four shards on the executor, bench-log 4 October; the
|
|
chained aggregation is 9.6 to 9.7 s per block on a mining 5090, `chain-pc2-pv1c`) and more records; the assignment
|
|
rule (8 assignees, 10 s window, spec 7.2) would need a CPU class with a longer window or the CPU provers only ever
|
|
win the open phase. None of that is measured. Status: Designed, nothing implemented.
|
|
|
|
### 4b. A second prover backend behind the swappable interface
|
|
|
|
The seam exists: `proving/igneum-prove/host/src/proof_system.rs` (`ProofSystem` trait, `Sp1ProofSystem`,
|
|
`StubProofSystem`), versioned per the design. The candidates:
|
|
|
|
| Target | Most likely backend | What exists | What adopting it costs |
|
|
|---|---|---|---|
|
|
| Apple silicon | RISC Zero's Metal prover (`metal` feature, shipped) | a shipped feature with kernels in the tree; ICICLE's Metal backend as the other library | a second guest program (the shard statement, `core/` is plain Rust and ports; the precompile patches for keccak and secp256k1 differ), a second pinned program id and verifying key in `elf/manifest.json`, the node's verifier for both proof formats (RISC Zero receipt and SP1 compressed proof) in `--mode verify` and the proof pool, and an aggregation problem: SP1's aggregator folds SP1 proofs by deferred verification; it cannot fold a RISC Zero receipt, so a block with shards from both families needs two aggregations or a wrapper. Approximate: weeks of a person's time, no measurement of a Metal shard time exists; RISC Zero's Groth16 wrapper for light clients does not run on Apple silicon at all |
|
|
| AMD | nothing | sppark's limited HIP path (not in SP1's tree); ICICLE's and Jolt's Vulkan and Metal work are not AMD | no backend to adopt. The honest statement is that it lands when a zkVM ships one |
|
|
|
|
### 4c. The public line
|
|
|
|
The site says today (read 5 October 2026 from `site/litepaper.html`, `site/miner.html`, `site/index.html`): "The same
|
|
card proves every block", "The card mines and proves", "Ember finds your GPU, makes a wallet for you and runs the
|
|
node, the miner and the prover as one app", "Target: shard size will be set so a 12 GB card proves one shard in about
|
|
20 seconds". Every one of those is true of an NVIDIA card with enough memory and false of an AMD or Apple card, and the
|
|
litepaper's own rule is "If consumer GPUs cannot prove shards fast enough, Igneum says so and does not launch on promises".
|
|
|
|
The recommended line, for the litepaper's proving section, the miner page and the app's Proving tile (copy law):
|
|
|
|
> Proving needs an NVIDIA card with 16 GB or more today (20 GB to mine and prove on the same card). AMD and Apple
|
|
> cards mine. A prover for them lands when a zkVM ships one. A CPU can prove a small shard in about five minutes with 32 GB of RAM free; the chain pays the first proof, which a card delivers in seconds, so CPU proving is for testing, not income.
|
|
|
|
Where the numbers come from: 16 GB and 20 GB are the measured gates of `proving-v1.md` (13.8 GB prover-alone peak,
|
|
16.8 GB mine-and-prove peak); "when a zkVM ships one" is section 1. The line changes when the memory sweep moves the
|
|
gates or a backend ships; it is reviewed with every prover release.
|
|
|
|
## 5. What this analysis does about it (the consequences, before anyone asks)
|
|
|
|
| Consequence | Action | Owner |
|
|
|---|---|---|
|
|
| An AMD-only miner loses the proving share | the CPU tier of 4a is measured here (section 2); whether it becomes a tier is a decision for the project lead on those numbers | this analysis; the project lead |
|
|
| The rig's prover unit must select NVIDIA cards only | already true in `prover_decision`; told the rig-installer agent to keep it as a stated rule and to print the CPU-fallback line for AMD-only rigs | rig-installer agent |
|
|
| The app's Proving tile on an AMD-only or Apple machine should say why it is off and name the CPU path | the `provedefault.rs` lines already say so for Apple; AMD-only Windows machines get "no NVIDIA card ..." | proving agent (told) |
|
|
| The site and litepaper over-promise for AMD and Apple | the line of 4c, to land with the next site pass (copy law; `node site/build.mjs`; link-check) | site-pages owner; not changed here |
|
|
| A Metal prover is the only non-NVIDIA path with a shipped backend | 4b names RISC Zero's Metal path and its cost; no work started | proving agent (told) |
|
|
|
|
## 6. Commands, jobs and sources
|
|
|
|
| What | Where |
|
|
|---|---|
|
|
| The PC 1 jobs (signed `run` jobs, PowerShell, not elevated, miners untouched, SP1_PROVER=cpu, host built without the `cuda` feature) | `tools/amd-prove/pc1-cpu-prove.ps1` (small fixtures), `pc1-cpu-prove-sp.ps1` (the `S_p` shard); published as `cpu-prove-pc1-small` (built, proved nothing: the quote bug) and `cpu-prove-pc1-small2` (the numbers) by `packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --shell powershell --timeout-minutes 60`; `cpu-prove-pc1-sp` written, not published; `check-job-bash.sh` gates the bash body of every job script here; the package `igneum-prove-wsl2-pv1b.zip` (sha256 df50dee5...), the URL and hash filled at publish time, never committed |
|
|
| The Mac run | `tools/lock/with-lock.sh measure /usr/bin/time -l igneum-prove-host block-56-transfers-3shards.json --mode shard --shard 0` with the `proving-v1` worktree's host (pinned ids checked with `--mode id`) |
|
|
| Results | `node tools/jobs.mjs cpu-prove-pc1-small2 --all`, `docs/bench-log.md` entry "5 October 2026, the CPU prover on PC 1 and the backend survey" |
|
|
| Pages read | SP1: docs.succinct.xyz hardware-acceleration page, github.com/succinctlabs/sp1 (releases, `crates/sdk/src/lib.rs`, `sp1-gpu/README.md`, code search); RISC Zero: dev.risczero.com local-proving, `risc0/zkvm/Cargo.toml`; Jolt: README, book, PR #1733; OpenVM: releases; ICICLE: install_gpu_backend page, releases, README, PRs #735 and #1019; sppark README |
|