igneum/docs/analysis/amd-proving.md

20 KiB

Proving on AMD and Apple cards: what exists, what the CPU can do, what to tell the public

5 October 2026, from the project lead's two questions that evening: "test proving on the amd card?" and "can we test proving on mac?". PC 1 holds an RTX 5090 and an RX 9070 XT (gfx1201, 16 GB) in an eGPU; this Mac is an M5 Max. The prover is SP1 (proving/igneum-prove, docs/plans/proving-v0.md, proving-v1.md), run on the GPU only through SP1's CUDA server. Every figure below is measured (with its bench-log entry or job id) or cited (with its file or page); the rest is labelled approximate. Status words follow docs/spec/00-overview.md 0.2.

The answer in three lines. No zkVM proves on an AMD GPU on 5 October 2026: not SP1, not RISC Zero, not Jolt, not OpenVM, and the ICICLE library underneath them has no AMD backend either. Apple silicon has a shipped Metal prover in RISC Zero and a Metal backend in ICICLE, but SP1, the prover Igneum runs, is CPU-only on a Mac. So an AMD-only or Apple-only machine mines and does not prove on its card; it can prove on its CPU, at the times measured in section 2.

1. The backends (read 5 October 2026, 20:30 to 20:50 UTC)

Prover Version read CPU NVIDIA (CUDA) AMD (ROCm or HIP) Apple (Metal) Vulkan or WebGPU Where it says so
SP1 (ours) v6.8.1, 24 Sep 2026 (pinned); dev head 318dd530, 28 Sep 2026 yes; AVX2 and AVX-512 on x86 through Plonky3 yes: sp1-gpu-server, "Compute Capability 8.0 or higher", "24GB or more VRAM", "the CUDA 12 runtime and a compatible NVIDIA driver", Linux x86_64 no no no docs.succinct.xyz, SP1 docs "Hardware acceleration" page; crates/sdk/src/lib.rs (pub mod cpu, mock, light, #[cfg(feature = "cuda")] pub mod cuda, #[cfg(feature = "network")] pub mod network: no other backend module); sp1-gpu/README.md (CUDA_ARCHS 89, 90, 100, 120; NTT by NVIDIA cuPQC or sppark); release notes v6.2.3 to v6.8.1 (the only backend line: "add optional cuPQC NTT backend", v6.8.0); a code search of the repository on 5 October: "rocm" 0 files, "metal" 0, "vulkan" 0, "webgpu" 0; cuobjdump of sp1-gpu-server 6.8.1: sm_80, 86, 89, 90, 100, 120 and compute_120 PTX, nothing else (docs/bench-log.md, "proving v1", 5 October 2026)
sppark (SP1's NTT fallback, vendored at sp1-gpu/crates/sys/sppark) main README, read 5 October 2026 yes: "x86_64 with Nvidia's Volta+ GPU hardware platforms on Linux and Windows" "A limited support for AMD's RDNA and CDNA GPUs is provided" (upstream README). SP1's tree carries no HIP build: the 0 "rocm" files above, and sp1-gpu/crates/sys/sppark/util/gpu_t.cuh is CUDA only no no github.com/supranational/sppark README; the SP1 files named
RISC Zero latest release v3.0.6, 17 Jul 2026 (a v5.0.0-rc.1 of 15 Jan 2026 is also on the releases page) yes, "nearly any modern CPU (x86 or ARM)" yes, "RISC Zero targets NVIDIA GPUs using the CUDA framework" no ("rocm", "vulkan": 0 files in the repository) yes: metal = ["prove"] in risc0/zkvm/Cargo.toml; kernels in risc0/sys/kernels/zkp/metal/*.metal (zk, fri, mix, sha); docs: "RISC Zero will use the integrated Metal compute cores" on Apple silicon. The Groth16 wrapper "only works on x86 architecture, and so Apple Silicon is currently unsupported (even via Docker)" no dev.risczero.com "Local proving"; risc0/zkvm/Cargo.toml features cuda = [... risc0-zkp/cuda ...], metal = ["prove"]
Jolt (a16z) v0.3.0-alpha, 1 Oct 2025; "Jolt is in alpha and is not suitable for production use" yes, "state-of-the-art performance on CPU" no no a draft PR #1733 (opened 3 Aug 2026, not merged): titled "feat: Metal GPU backend (Apple Silicon)" and marked experimental: 92 Metal kernels, "2.12x speedup at 2^20 scale" on an M4 mini and "3.20x vs same-binary CPU" on an M5 Max, "Apple Silicon + macOS only. No CI coverage" no github.com/a16z/jolt README and book (jolt.a16zcrypto.com); PR #1733
OpenVM v2.0.2, 14 Aug 2026 yes yes: cuda-backend (v1.4.2 notes), "Improves the Halo2 GPU prover" (v2.0.2) no no no github.com/openvm-org/openvm releases
ICICLE (Ingonyama; the GPU library behind several provers, not SP1) v4.0.0, 11 Jul 2025 yes (MIT) yes, "CUDA (for NVIDIA GPUs)" no backend listed yes, "Metal (for Apple Silicon GPUs)"; both under a special licence with a free research licence Vulkan in the build system (PR #735, merged Jan 2025) and a draft "Vulkan NTT" PR #1019 (Jul 2025, "still wip"); nothing installable dev.ingonyama.com "Install GPU backend"; the releases page; the README ("backends ... are distributed under a special license")

Said plainly: on 5 October 2026 no zkVM proves on an AMD GPU. The only AMD code in the whole chain is sppark's limited HIP path, which SP1 does not build. For Apple silicon the answer is split: RISC Zero ships a Metal prover and ICICLE a Metal backend; SP1, Jolt and OpenVM do not. Nothing read tonight names an AMD plan with a date.

2. The CPU fallback, measured

SP1's CPU prover is the path an AMD-only or Apple-only machine has today. Three machines, the same pinned guests (shard program id 0x2b1a81cb..., aggregator 0x474678f3..., pinned 2026-10-05T16:20:38Z), SP1_PROVER=cpu, --mode shard --shard 0 (execute, core proof, compressed proof, each verified). The RTX 5090 rows are the reference.

Fixture (SP1 cycles) Stage PC 1 CPU, miner running on both cards (job cpu-prove-pc1-small2) Apple M5 Max CPU (4 October, loaded; bench-log) RTX 5090 (bench-log)
block-56-transfers-3shards shard 0, 200 pgas (315 k) core 82.5 s, 7,310,257 B, verify 0.210 s 83.1 s, 7,310,257 B not run on the 5090; the nearest rows are block-78 below and an empty live shard: 7.0 to 7.7 s compressed with the miner on the card (5 Oct, chain-pc2-pv1b, pv1c)
compressed 199.2 s, 1,272,897 B, verify 0.035 s; 312 s wall for setup 22.8 s, execute 0.14 s, core, compressed 272.3 s, 1,272,897 B
peak RSS, CPU 29.5 GB peak RSS; 978% CPU (9.8 of 16 cores), user 2,516 s, system 537 s not recorded
block-78-increment, 2 transactions (626 k) core 87.0 s, 7,317,857 B, verify 0.209 s 22.0 s, 7.3 MB (3 October, v0 guest) 1.4 s (4 October, mining paused)
compressed 202.3 s, 1,272,897 B, verify 0.034 s; 322 s wall (setup 21.8 s) 55.7 s, 1.27 MB 2.7 s
peak RSS, CPU 30.5 GB peak RSS; 979% CPU, user 2,616 s, system 541 s not recorded
block-338-shard1, one shard at S_p (60.8 M) core not run, by the PC 1 scheduler's decision at 21:05Z (PC 1's time tonight belongs to the Counter ASIC 2.0 gates; the job cpu-prove-pc1-sp, script tools/amd-prove/pc1-cpu-prove-sp.ps1, is written and unpublished). Extrapolation, approximate: 60.8 M cycles is about 29 SP1 shards of 2^21 cycles where the small fixtures are one, so the core proof alone is about 29 x 80 s, 40 min, and the compressed recursion over 29 shard proofs adds hours; the floor from the 5090's own ratios (6x on core, 4x on compressed between block-78 and S_p) is 9 min core and 13 min compressed. Either way far outside every deadline not run on the CPU (execute alone 6.9 s) 8.3 s
compressed not run (see the core cell) not run 10.9 s with the card to itself (4 Oct); 33.0 s with the miner running (5 Oct, memminer-pc2-pv1); 7.3 to 7.7 s per EMPTY shard with the miner running (chain-pc2-pv1c)
peak RSS, CPU not run; at least the 30 GB of the small rows GPU peak 28,295 MiB alone, 30,039 MiB beside the miner

PC 1: Windows 11, WSL2 Ubuntu 24.04 as root, 16 cores and 46,994 MB visible to the VM, the Igneum Miner app 0.3.9 mining on the RTX 5090 (89% mean utilisation through both runs, 59 to 70% minimum: the miner, untouched) and on the RX 9070 XT (not visible to nvidia-smi, mining through the app's OpenCL worker). Job cpu-prove-pc1-small2, 20:49:00Z to 20:59:49Z, 649 s wall including a 6-s warm build; the host built without the cuda feature from the hosted package igneum-prove-wsl2-pv1b.zip, --mode id the pinned pair. The first job, cpu-prove-pc1-small (20:44 to 20:46Z), built the host cold in 126 s and proved nothing: an apostrophe inside a single-quoted awk program ended the quote, bash refused the whole loop and the job reported exit 0. The class fix: tools/amd-prove/check-job-bash.sh runs bash -n on the bash body of a PowerShell job before it is published, and the job itself runs bash -n inside the distro before the run; both were shown to fire on the bad body and pass the fixed one. Host RAM in the VM: 968 MB used before, 2,351 MB after; the prover's own peak 29.5 to 30.5 GB.

Mac, fresh run tonight: not taken. The Mac measure lock was held from 20:31Z (a read-width packbench under measure, three build slots, then a 1,500-s proving-v1 network under run) and did not free inside the 10-minute window the coordinator set, so the Apple column is the 4 October rows (M5 Max, 18 cores, 64 GB, load 38 to 47, nice -n 19): the same host modes on the same fixture, under heavier load than PC 1 tonight. The Mac's RAM peak was not recorded on 4 October; PC 1's 30 GB says a Mac needs more than 32 GB for the CPU prover, which a 64 GB M5 Max has and a 16 or 24 GB Mac does not.

The deadlines a CPU proof has to fit (all in the spec and the v1 plan): the exclusive window of an assigned shard is 10 s of DAA time (spec 7.2 item 3; after it anyone may prove and be paid first); the litepaper promises the block's proof "within about a minute"; the launch target is 20 to 60 s behind the tip; from proving v1 a segment nobody has proven in T = 600 DAA s (10 min) pays nothing (docs/plans/proving-v1.md, decisions). So a CPU shard proof is useful only if it lands inside 10 min and competitive only if it lands inside about a minute.

Reading. On PC 1 the CPU proof of the smallest shard (315 k cycles) and of the two-transaction block (631 k cycles) cost the same: 82.5 and 87.0 s core, 199.2 and 202.3 s compressed. Doubling the cycles added 4.5 s to the core proof and 3.1 s to the compressed one, so about 280 s of every CPU proof is fixed cost (the recursion that turns the core proof into the 1.27 MB compressed proof the chain carries), and no shard size removes it. Against the deadlines: 282 s a shard (core plus compressed, the client already set up, as the app's loop runs it) is 28x the 10-s assignment window, 4.7x the minute the litepaper promises, and inside the 600-s unproven deadline of v1 with 5 min to spare; but an NVIDIA card proves the same shard in 2.7 to 7.7 s, so a CPU prover only ever wins a shard that no card has taken in 10 minutes. The Mac's 83.1 and 272.3 s of 4 October have the same shape. The RAM peak of 29.5 to 30.5 GB is the second finding: the SP1 CPU prover does not fit a 16 GB machine at all, and WSL2 gives a Windows VM half the host's RAM by default, so the CPU path needs a 64 GB Windows PC or a 32 GB Linux or Mac machine. The S_p shard on the CPU can only be slower (the 5090 takes 6x longer at S_p than on block-78: 8.3 s against 1.4 s core); it was not run tonight (the scheduler kept PC 1 for the Counter ASIC 2.0 gates) and could not change the conclusion.

3. What this means for each tier (the every-number rule, CLAUDE.md 5 October 2026)

Tier Mines Proves on the card The 20% proving-pool share (spec 2.5) What the software does today
AMD-only home miner, one card of 8, 12 or 16 GB (an RX 9070 XT is 16 GB), Windows or Linux yes (OpenCL worker, proto-opencl; PC 1's 9070 XT mines on the devnet) no: no prover exists for the card lost, unless CPU proving at a small shard size becomes a tier (section 4a) the rig installer: prover_decision in packaging/linux/bin/igneum-rig-lib.sh (branch rig-install) skips every non-NVIDIA card ([[ "$vendor" == nvidia ]] || continue) and prints "proving off by default: no NVIDIA card (no CUDA prover for AMD or Intel yet)"; the app: provedefault.rs (branch proving-v1) considers NVIDIA cards only. Both already right; neither offers the CPU path
Apple silicon (M-series, unified memory) yes: the M5 Max at 26.7 MH/s (bench-log 4 October, "first hourly program swap", Metal prepare 1 row) no with SP1; RISC Zero and ICICLE have Metal, SP1 does not lost today; a Metal prover behind the swappable interface would restore it (section 4b) provedefault.rs: "proving stays off on Apple silicon: the M5 Max CPU took 41 to 55 s for an empty shard and minutes for a full one; Settings switches it on (CPU, slow)". Right
Mixed rig (NVIDIA and AMD cards in one box) every card the NVIDIA cards prove for the box; the AMD cards mine kept, earned by the NVIDIA cards the rig installer picks the biggest NVIDIA card (prover_decision, PROVER_CARD overrides), pauses its miner under 20 GB, keeps it mining at 20 GB or more; the AMD cards get a miner unit each. The prover unit must never select an AMD card: it does not (the vendor filter above), and that filter is now a stated requirement, not an accident
NVIDIA home miner, 8 or 12 GB yes no on this SP1 build (13.9 GB floor on an empty shard, memsweep-pc2-pv1) lost unless the shard size moves unchanged from proving-v1.md
NVIDIA 16 GB yes prove-only, miner paused per shard kept unchanged
NVIDIA 24 or 32 GB yes mines and proves (peak 16.8 GB on empty shards, 30.0 GB on a full prototype shard beside the miner) kept unchanged
Pool user through the pool the pool's own NVIDIA cards prove the shards assigned to the pool's keys (approximate: the pool protocol, spec 09, does not yet say who proves) by the pool's rules open, spec 09

4. The options

4a. CPU proving at a small shard size, as a tier

What it is: an AMD-only or Apple machine proves shards cut at a smaller budget than S_p on its CPU, through the same host (SP1_PROVER=cpu; the host's --budget re-plan from branch proving-v1, commit c2544be, cuts a fixture at any budget). The miner keeps the card; the prover takes the CPU.

What the numbers say: the fixed cost kills it. 282 s a shard on a 16-core PC and 355 s on the loaded M5 Max, with 30 GB of RAM, at the smallest shard there is; the time sits in the compressed-proof recursion, not in the cycles, so cutting shards smaller does not help, and the launch deadline (20 to 60 s behind the tip) is missed by 5x. It fits only the v1 unproven deadline (600 s), which pays a CPU prover only when no card has proven the shard in 10 minutes: on a chain with one NVIDIA prover that never happens. Recommendation: no CPU tier. Settings may still switch the CPU prover on (it does on macOS today), and the Proving tile must then say the proof takes about five minutes and is paid only when no card proves first.

What it costs the chain: a block cut into more, smaller shards costs more aggregation work (the aggregator guest verifies one deferred proof per shard; 1.66 M cycles for four shards on the executor, bench-log 4 October; the chained aggregation is 9.6 to 9.7 s per block on a mining 5090, chain-pc2-pv1c) and more records; the assignment rule (8 assignees, 10 s window, spec 7.2) would need a CPU class with a longer window or the CPU provers only ever win the open phase. None of that is measured. Status: Designed, nothing implemented.

4b. A second prover backend behind the swappable interface

The seam exists: proving/igneum-prove/host/src/proof_system.rs (ProofSystem trait, Sp1ProofSystem, StubProofSystem), versioned per the design. The candidates:

Target Most likely backend What exists What adopting it costs
Apple silicon RISC Zero's Metal prover (metal feature, shipped) a shipped feature with kernels in the tree; ICICLE's Metal backend as the other library a second guest program (the shard statement, core/ is plain Rust and ports; the precompile patches for keccak and secp256k1 differ), a second pinned program id and verifying key in elf/manifest.json, the node's verifier for both proof formats (RISC Zero receipt and SP1 compressed proof) in --mode verify and the proof pool, and an aggregation problem: SP1's aggregator folds SP1 proofs by deferred verification; it cannot fold a RISC Zero receipt, so a block with shards from both families needs two aggregations or a wrapper. Approximate: weeks of a person's time, no measurement of a Metal shard time exists; RISC Zero's Groth16 wrapper for light clients does not run on Apple silicon at all
AMD nothing sppark's limited HIP path (not in SP1's tree); ICICLE's and Jolt's Vulkan and Metal work are not AMD no backend to adopt. The honest statement is that it lands when a zkVM ships one

4c. The public line

The site says today (read 5 October 2026 from site/litepaper.html, site/miner.html, site/index.html): "The same card proves every block", "The card mines and proves", "Ember finds your GPU, makes a wallet for you and runs the node, the miner and the prover as one app", "Target: shard size will be set so a 12 GB card proves one shard in about 20 seconds". Every one of those is true of an NVIDIA card with enough memory and false of an AMD or Apple card, and the litepaper's own rule is "If consumer GPUs cannot prove shards fast enough, Igneum says so and does not launch on promises".

The recommended line, for the litepaper's proving section, the miner page and the app's Proving tile (copy law):

Proving needs an NVIDIA card with 16 GB or more today (20 GB to mine and prove on the same card). AMD and Apple cards mine. A prover for them lands when a zkVM ships one. A CPU can prove a small shard in about five minutes with 32 GB of RAM free; the chain pays the first proof, which a card delivers in seconds, so CPU proving is for testing, not income.

Where the numbers come from: 16 GB and 20 GB are the measured gates of proving-v1.md (13.8 GB prover-alone peak, 16.8 GB mine-and-prove peak); "when a zkVM ships one" is section 1. The line changes when the memory sweep moves the gates or a backend ships; it is reviewed with every prover release.

5. What this analysis does about it (the consequences, before anyone asks)

Consequence Action Owner
An AMD-only miner loses the proving share the CPU tier of 4a is measured here (section 2); whether it becomes a tier is a decision for the project lead on those numbers this analysis; the project lead
The rig's prover unit must select NVIDIA cards only already true in prover_decision; told the rig-installer agent to keep it as a stated rule and to print the CPU-fallback line for AMD-only rigs rig-installer agent
The app's Proving tile on an AMD-only or Apple machine should say why it is off and name the CPU path the provedefault.rs lines already say so for Apple; AMD-only Windows machines get "no NVIDIA card ..." proving agent (told)
The site and litepaper over-promise for AMD and Apple the line of 4c, to land with the next site pass (copy law; node site/build.mjs; link-check) site-pages owner; not changed here
A Metal prover is the only non-NVIDIA path with a shipped backend 4b names RISC Zero's Metal path and its cost; no work started proving agent (told)

6. Commands, jobs and sources

What Where
The PC 1 jobs (signed run jobs, PowerShell, not elevated, miners untouched, SP1_PROVER=cpu, host built without the cuda feature) tools/amd-prove/pc1-cpu-prove.ps1 (small fixtures), pc1-cpu-prove-sp.ps1 (the S_p shard); published as cpu-prove-pc1-small (built, proved nothing: the quote bug) and cpu-prove-pc1-small2 (the numbers) by packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --shell powershell --timeout-minutes 60; cpu-prove-pc1-sp written, not published; check-job-bash.sh gates the bash body of every job script here; the package igneum-prove-wsl2-pv1b.zip (sha256 df50dee5...), the URL and hash filled at publish time, never committed
The Mac run tools/lock/with-lock.sh measure /usr/bin/time -l igneum-prove-host block-56-transfers-3shards.json --mode shard --shard 0 with the proving-v1 worktree's host (pinned ids checked with --mode id)
Results node tools/jobs.mjs cpu-prove-pc1-small2 --all, docs/bench-log.md entry "5 October 2026, the CPU prover on PC 1 and the backend survey"
Pages read SP1: docs.succinct.xyz hardware-acceleration page, github.com/succinctlabs/sp1 (releases, crates/sdk/src/lib.rs, sp1-gpu/README.md, code search); RISC Zero: dev.risczero.com local-proving, risc0/zkvm/Cargo.toml; Jolt: README, book, PR #1733; OpenVM: releases; ICICLE: install_gpu_backend page, releases, README, PRs #735 and #1019; sppark README