Compare commits

...

2 commits

View file

@ -133,7 +133,7 @@ The GPU server of SP1 6.8.1 sets the memory, not the shard: a floor of 13.9 GB f
| 32 GB (RTX 5090) | the prototype shard, 28.3 GB, 10.8 s; the v1 shard 20.4 GB, 4.3 s | the prototype shard 30.1 GB, 33 s; the v1 shard 22.2 GB, 13.2 s | on, mine and prove, today |
| 24 GB (RTX 4090, 3090) | the v1 shard 20.4 GB; the prototype shard does NOT fit (28.3 GB) | the v1 shard 22.2 GB measured on the 5090's allocation (2.3 GB spare on a 24 GB card; approximate for the card itself) | on, mine and prove, with the line "until the devnet's fee switch its shards are the prototype size, which needs 32 GB, so this card proves from the switch on" |
| 16 GB (RTX 5080, 4080) | the shipped server: an empty shard only (13.9 GB); the patched server v3 b37defef at threshold 2^27: the v1 shard 12,915 MiB and 4.3 s, the prototype shard 13,459 MiB and 16.8 s (measured by the prover-floor agent on the 5090's allocation, job `floor-sweep-3`, 00:13 to 00:17Z 6 October; 2^27 + 2^26 gives 16,115 MiB, over the card) | the shipped server: nothing (15.7 GB for an empty shard); the patched server at 2^27 beside the miner (the 5090 mining at 95%, 338 W, same card; `floor-sweep-4`, 00:35 to 00:38Z): the v1 shard 14,786 MiB total with the miner's 3,833 MiB resident inside it, the server's own 10,953 MiB, 17.4 s a shard; on a 16 GB card that is 10.95 GB server + 1.7 GB miner = 12.7 GB plus the display, under the 15.0 GB line | off on the shipped server, with the line; on (mines and proves, 17 s a v1 shard, 4.3x the alone time) once the patched server ships (the packaging row below) |
| 12 GB (RTX 3060, 4070) | the shipped server: nothing (the floor is 13.9 GB, and the server refuses the card outright); the patched server v3 b37defef at threshold 2^26 (`SP1_GPU_ELEMENT_THRESHOLD=67108864`, the 12 GB profile): **the v1 shard 10,291 MiB and 5.7 s, an empty shard 9,971 MiB and 3.3 s**, the card's 2,089 MiB idle inside the peak and the server's own working set about 8.2 GB (6,535 MiB after Setup), so a 12 GB card proves alone with about 3 GB over it (measured by the prover-floor agent on the 5090's allocation, `floor-sweep-3`; the on-order RTX 3060 run is pending) | measured beside the miner (`floor-sweep-4`): at 2^26 the v1 shard 12,066 MiB total with the miner's 3,833 MiB inside, the server's own 8,233 MiB, 24.4 s; the empty shard 11,586 MiB, 13.0 s; 2^25 gains nothing (12,066 MiB, 43.5 s). On a 12 GB card that is 8.2 GB server + 1.7 GB miner = 9.9 GB before the display, over the 9.0 GB line the project lead set, so mine-and-prove on 12 GB is NOT claimed | off on the shipped server; "proves alone" on the patched one once it ships (the packaging row below); mine-and-prove stays off on 12 GB (9.9 GB plus the display, over the 9.0 GB line). The public gate stays "12 GB proves; 16 GB mines and proves", both on the patched server, both pending a run on the card itself. the project lead's "make sure we can prove on 12 GB cards" is answered on the 5090's allocation and OPEN on the card itself: the prover-floor agent (branch prover-floor, 5 October night) read SP1 v6.8.1's GPU server source (`sp1-gpu/crates/prover_components/src/builder.rs` lines 35 to 39): it reads the card's memory, adds 4 and panics under 24 ("Unsupported GPU memory ... must be at least 24GB"), and builds its core (ELEMENT_THRESHOLD 2^28 + 2^27 elements + 2^21), recursion (2^27), shrink (2^25) and wrap (85 M element) provers at Setup whatever the mode, which is the 13.9 GB floor; no knob reaches them, so the fix is a server rebuilt from source on PC 2 (WSL2, nvcc 12.8, CUDA_ARCHS=120) with those sizes cut, measured on the same fixtures and recipe as the curve above (D2 carries the curve) |
| 12 GB (RTX 3060, 4070) | the shipped server: nothing (the floor is 13.9 GB, and the server refuses the card outright); the patched server v3 b37defef at threshold 2^26 (`SP1_GPU_ELEMENT_THRESHOLD=67108864`, the 12 GB profile): **the v1 shard 10,291 MiB and 5.7 s, an empty shard 9,971 MiB and 3.3 s**, the card's 2,089 MiB idle inside the peak and the server's own working set about 8.2 GB (6,535 MiB after Setup), so a 12 GB card proves alone with about 3 GB over it (measured by the prover-floor agent on the 5090's allocation, `floor-sweep-3`; the on-order RTX 3060 run is pending) | measured beside the miner (`floor-sweep-4`): at 2^26 the v1 shard 12,066 MiB total with the miner's 3,833 MiB inside, the server's own 8,233 MiB, 24.4 s; the empty shard 11,586 MiB, 13.0 s; 2^25 gains nothing (12,066 MiB, 43.5 s). On a 12 GB card that is 8.2 GB server + 1.7 GB miner = 9.9 GB before the display, over the 9.0 GB line the project lead set, so mine-and-prove on 12 GB is NOT claimed | off on the shipped server; "proves alone" on the patched one once it ships (the packaging row below); mine-and-prove stays off on 12 GB (9.9 GB plus the display, over the 9.0 GB line). Route 2, measured by the prover-floor agent on 6 October (floor-core-alone 08:48Z, floor-core-miner 08:53Z, the v3 server, every proof verified): a prover that stops at the CORE proof and leaves the compression to the aggregator has a working set of 7.8 GB for the v1 shard at 2^26 beside the miner (12,834 MiB peak with the miner's 5,080 MiB resident), 6.4 GB for an empty shard; the core proof is 14.4 MB (6.9 MB empty), its CPU verify 0.45 s, and the aggregator pays 2.5 s a shard alone (12.6 s beside a miner) to compress it. Still 7.8 + 1.7 GB = 9.5 GB over the 9.0 GB line, because the server's Setup builds the recursion, shrink and wrap keys (4.4 GB) it never uses; the core-only server v4 68b2f512 (floor-core2-miner, 09:05 to 09:13Z, the 5090 mining at 95%, 3,801 MiB resident inside every peak, every proof verified by the unpatched host) brings the v1 shard at threshold 2^25 to 6,409 MiB own (10,210 MiB peak with the miner) in 19.8 s, the empty shard to 6,121 MiB own in 8.7 s; 2^24 gains 0.4 GB for 56.6 s and is not worth it; the pool returning freed memory does not lower the peak. So the 12 GB mine-and-prove line is MET on the 5090's allocation by a core-only prover at 2^25: 6.4 + 1.7 = 8.1 GB before the display, under the 9.0 GB line, 19.8 s a v1 shard, with a 25.6 MB core proof (CPU verify 0.79 s) handed to an aggregator that compresses it. That hand-off is a prover-protocol change (the pool and the aggregator take compressed proofs today), so the row ships only with it; compressed proving on the card itself stays "proves alone". Still "measured on the 5090's allocation, the 12 GB card pending". Its rows: docs/analysis/prover-floor.md, route 2, and the bench-log entry "prover floor: the SP1 6.8.1 GPU server rebuilt for small cards" (its branch). The public gate stays "12 GB proves; 16 GB mines and proves", both on the patched server, both pending a run on the card itself. the project lead's "make sure we can prove on 12 GB cards" is answered on the 5090's allocation and OPEN on the card itself: the prover-floor agent (branch prover-floor, 5 October night) read SP1 v6.8.1's GPU server source (`sp1-gpu/crates/prover_components/src/builder.rs` lines 35 to 39): it reads the card's memory, adds 4 and panics under 24 ("Unsupported GPU memory ... must be at least 24GB"), and builds its core (ELEMENT_THRESHOLD 2^28 + 2^27 elements + 2^21), recursion (2^27), shrink (2^25) and wrap (85 M element) provers at Setup whatever the mode, which is the 13.9 GB floor; no knob reaches them, so the fix is a server rebuilt from source on PC 2 (WSL2, nvcc 12.8, CUDA_ARCHS=120) with those sizes cut, measured on the same fixtures and recipe as the curve above (D2 carries the curve) |
| under 12 GB | nothing | nothing | off, mine only |
| AMD-only and Apple machines | nothing on the GPU: no zkVM proves on an AMD GPU today (`docs/analysis/amd-proving.md`, branch amd-prove); the CPU prover is about 5 minutes a shard at a 30 GB RSS whatever the shard size (PC 1, bench-log "the SP1 CPU prover on PC 1") | | off, "mines and does not prove"; the only non-NVIDIA path with a shipped backend is RISC Zero's Metal prover behind the `ProofSystem` seam (a second guest and pinned id, a verifier for both formats, no shared aggregation): an open item, not 0.3.11 |