proving v2 plan: the CUDA rows (4090, 3060, 4060, PC 2) and the per-tier table; the verdict: SP1 patched 4 to 6x faster on every card, RISC Zero's edge the 5.7x smaller proof

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 15:11:45 +00:00
parent 95e3ccea12
commit bd322e9db1

View file

@ -91,3 +91,39 @@ verify-mode runs).
RISC Zero stays in the version 2 slot for two things, neither of them Apple: the 8 GB CUDA tier (po2 19 takes 8.4 GB of RSS on the Mac's CPU for the v1 shard; the rented RTX 4060 8 GB measures it on a card) and its 5.7x smaller proof (223,882 B against SP1's 1,272,897 B, 0.21 s to verify). The Apple route is SP1 on the CPU: 71.7 s for the v1 shard on the M5 Max (bench/proof-systems), inside the 600 DAA-s window with no toolchain work, about one shard a minute, mining paused or not; the base Mac mini row is owed when the project lead's mini arrives. The analysis (docs/analysis/proving-methods.md) carried "RISC Zero has a Metal prover"; that sentence, route I and the ranking are corrected from this measurement.
## The CUDA rows and the per-tier table (6 October 2026, rented cards and PC 2)
Every row: one shard proved on the card alone (no miner) unless marked, every proof verified by the pinned
verifier, `bench/proof-systems/rows` (the watch), the v1 shard = `fees-v1-shards2.json` shard 0 (4,717,439 SP1
cycles, 10,049,917 RISC Zero cycles), the empty shard = `block-58927-empty-reward.json`, the prototype shard =
`block-338-shard1.json` (60 M SP1 cycles). Peak = nvidia-smi memory.used; the RISC Zero po2 is the segment limit.
| Card | SP1 stock | SP1 patched, the profile that fits | RISC Zero po2 20 | RISC Zero po2 19 |
|---|---|---|---|---|
| RTX 4090 24 GB, v1 shard | 6.4 s at 17,800 MiB | 2^27: 6.2 s at 10,850; 2^26: 6.8 s at 8,192 | 23.1 s at 13,280 MiB | 32.7 s at 8,500 MiB |
| RTX 4090, prototype shard | 19.1 s at 20,680 | 2^26: 29.2 s at 8,192 | 178.5 s at 13,390 | 281.4 s at 8,846 |
| RTX 3060 12 GB, v1 shard | refused (the stock server needs 24 GB) | 2^27: 9.5 s at 10,181; 2^26: 13.1 s at 7,587 | 53.0 s at 10,071 | 68.5 s at 7,339 |
| RTX 3060, empty shard | refused | 2^26: 7.3 s at 7,203 | 5.6 s at 7,345 | 5.7 s at 7,345 |
| RTX 3060, prototype shard | refused | 2^26: 58.3 s at 7,555 | n/a (not run) | 578.2 s at 7,285 |
| RTX 4060 8 GB, v1 shard | refused | 2^27: the server panics (does not fit); 2^26: 7.9 s at 7,692; 2^25: 12.8 s at 7,660 | out of memory | 47.1 s at 6,720 MiB |
| RTX 4060, empty shard | refused | (running) | 3.8 s at 6,812 | 3.8 s at 6,876 |
| RTX 5090 32 GB beside its miner (PC 2) | 14.8 s at 22,477 (the miner inside) | 2^26: 27.0 s at 12,012 | not run (the r0 host was not on PC 2) | |
What the rows say, per tier (every line a measured row above, alone on the card unless the PC 2 line):
| Tier | SP1 (the patched server, the profile for the card) | RISC Zero (proof system 2) | Verdict for the slot |
|---|---|---|---|
| 8 GB (4060) | mines? not measured here; proves alone the v1 shard in 7.9 s at 2^26 (7.7 GB) and 12.8 s at 2^25 (7.7 GB); the stock server refuses the card | proves alone the v1 shard at po2 19 in 47.1 s at 6.7 GB; po2 20 does not fit | SP1 patched is 6x faster at the same memory; RISC Zero's one edge is the 5.7x smaller proof (223,882 B against 1,272,897) |
| 12 GB (3060) | the v1 shard 13.1 s at 2^26 (7.6 GB), 9.5 s at 2^27 (10.2 GB); the prototype shard 58.3 s at 7.6 GB | the v1 shard 68.5 s at po2 19 (7.3 GB), 53.0 s at po2 20 (10.1 GB); the prototype shard 578 s | SP1 patched 5x faster; the 12 GB tier stays SP1 |
| 16 GB | no 16 GB card measured today; by the rows the 2^27 profile (10.2 to 11.2 GB) fits with the display | po2 20 (10.1 to 13.4 GB) fits | SP1 |
| 24 GB (4090) | stock 6.4 s at 17.8 GB; patched 2^27 6.2 s at 10.9 GB | 23.1 s at po2 20 | SP1 3.7x faster |
| 32 GB (5090, beside its miner) | 14.8 s stock, 27.0 s at 2^26 beside the miner | not measured beside a miner | SP1 |
| Apple (M5 Max) | the CPU: 71.7 s | the CPU: 574 s (no Metal path in 3.0.6) | SP1 on the CPU |
The verdict: RISC Zero 3.0.6 proves our shard on every card SP1's patched server proves it on, with no card it
alone reaches (the 8 GB card takes both at 2^25 or po2 19), at 4 to 6x the time; its proof is 5.7x smaller and
verifies in 0.22 to 0.46 s against 0.32 to 0.55 s. So the version 2 slot is built and measured, and SP1 with the
patched server stays the default on every tier; RISC Zero earns its place when proof size matters (a light client,
a bridge) or when SP1's server cannot be patched for a card. The slot's switch (`proving_v2_activation_daa`) stays
never until the project lead wants the second system live; the quarterly watch re-measures both.