29 KiB
Proving v1: segment records, the chain rule, the unproven rule; the 0.3.11 rollout
5 October 2026, from 18:55 UTC (the project lead: "open the proving round asap"). Branches proving-v1 in the main repository
(worktree /Users/joshm/Projects/igneum-wt-proving-v1, from master a93199a) and in the fork
(vendor/igneum-node-pv1, from release-0.3.6 a24ab01a; to be rebased onto the 0.3.10 tip when it lands on
release-0.3.6). Status words follow docs/spec/00-overview.md 0.2. Every number here is in docs/bench-log.md
with its command. Nothing ships from this plan: it delivers branches, numbers and the rollout for 0.3.11.
The gap this closes
The litepaper says every block is proven within about a minute. Proving v0 (proving-v0.md, spec 7.7) proves
some shards: one prover (PC 2's RTX 5090) takes the newest shard assigned to it, about one shard every 30 s, so
under a tenth of blocks carry a proof; consensus does not require one; the aggregator guest (design 5.3) runs
on fixtures only. Proving v1 adds the aggregated segment record on chain (spec 7.8), the chain rule (segment N's
record verifies N-1, inside the proof by recursion), the unproven rule (a segment nobody proves in T seconds pays
nothing and may be skipped), the prover on by default on every machine that can prove, and the measurements
that say how many cards cover the chain.
The round, step by step
| Step | What | State |
|---|---|---|
| 1 | Prover on by default (app/igneum-app/src/provedefault.rs, engine.rs apply_prove_default): on at install when the machine can prove (NVIDIA card with 12 GB or more; WSL2 answering on Windows; Linux native; Apple silicon off until measured), never switching an explicit on back off; the Settings switch line and the tile line say why. Unit tests (5). The cost of proving on a mining machine: PC 2 job prover-cost-pc2-pv1 (5 min mining alone, 5 min with the prover, the GPU memory peak and the host RAM peak, the sp1-gpu-server's compiled SM targets) |
Implemented; the measurement is HELD (coordinator, 19:00Z): PC 2's RTX 5090 worker has been exiting on a pack seed mismatch since 18:35Z, so the first run's "mining alone" is 0 MH/s and void; re-run after the go |
| 2 | Segment aggregation: SegmentRecord (586 bytes, the aggregator guest's 340-byte statement inline), section IGNS before the shard section, p2p message 75 at protocol 15 (14 went to the EVM transaction relay in 0.3.10), the native block statement and the veto, the credit split (split_pool_credit), the payout at the carrier, igneum-miner sign-segment-record, the RPCs; host modes chain (consecutive fixtures), aggregate (live shard proofs from the pool, a run of blocks in one process) and verify-segment (the node's verifier, pinned aggregator key); the app's aggregator step (prover.rs aggregate_once) |
Implemented, unit-tested (consensus core 2 new tests, exec 2, params 1); the GPU measurement (N = 2, 4, 8 blocks on PC 2) is HELD with step 1; the Mac CPU run of --mode chain over 2 live blocks is the known-finished case |
| 3 | Coverage: tools/proving-v1/coverage.mjs (the proven-block share and the on-chain proof latency over a window from one node's RPC, the live page beside it) |
Implemented and run for 3 min (below); the 30-min window waits for the fleet |
| 4 | The chain rule and the unproven rule in consensus behind proving_v1_activation_daa (spec 7.8 items 2, 6, 7); unit tests; the fast-time 3-node harness tools/proving-v1/net.mjs (ports 29950+, suffix 956, trust mode) with the known-finished and known-failed cases |
Implemented; the harness run waits for the Mac build of the fork (vendor/igneum-node/target-pv1) |
| 5 | This plan: the rollout for 0.3.11 and the project lead's decisions | Written below |
Numbers (every one from docs/bench-log.md, "proving v1: segment records ...", 5 October 2026 evening)
| What | Number |
|---|---|
| The prover's cost to a mining 5090 (the re-run with the fleet mining, hash rate from the miner's own STATUS lines) | 124.7 MH/s alone, 119.7 MH/s with the prover on: 5.0 MH/s, 4.0%, on empty shards at 1.4 a minute |
| GPU memory on the 5090: the prover alone (empty shards) / the miner and the prover together | max 13,816 MiB / max 15,590 MiB (the miner holds 3,396 MiB); a 24 GB 4090 has 8.4 GB of headroom, a 16 GB card 0.4 GB, a 12 GB card cannot do both on this build; the full-shard peak is the chain job's row |
| Coverage, 30-min window with the fleet mining, one prover | 2.4% of blocks, latency p50 44 s, p99 52 s |
| Host RAM | host used 25.6 GB of 63 GB; the WSL2 VM 7.9 GB working set |
| sp1-gpu-server 6.8.1 compiled targets (cuobjdump) | sm_80, sm_86, sm_89, sm_90, sm_100, sm_120 and compute_120 PTX: Ada (4090) is native, no JIT; nothing for AMD |
| Shards a minute, one 5090 through the app's loop (empty shards) | 1.6 |
Chain of 2 live blocks on the Mac CPU (--mode chain) |
shard 55.4 and 41.3 s, aggregate 52.0 s then 59.1 s with the previous proof, chain_len 2, final proof 1,272,909 bytes, verify-segment 0.032 s |
| Unit tests | consensus core 13, exec 8, app 5, all passing on the Mac |
| The harness (3 nodes, fast time, trust mode) | PASSED, 21 checks in 197 s on b177718e (N = 4) and 244 s on the final tree ece42979 (N = 8): paid 1.0 s after submit, every node agreeing; the fresh chain refused after a proven segment; the unproven segment skipped after its deadline; shards at 90% |
| Coverage, 3-min window, the degraded fleet (one card, the Mac verifier down) | 4.7% of blocks proven, on-chain latency p50 39 s |
The chain of 8 live blocks on the 5090, the card also mining (chain-pc2-pv1c) |
shard 7.3 to 7.7 s, first aggregation 7.9 s, every chained one 9.6 to 9.7 s; N = 2 in 32.6 s, N = 4 in 66.8 s, N = 8 in 135.6 s (17.0 s a block); the final proof 1,272,909 bytes whatever N, the record 586 bytes, verify-segment 0.037 to 0.040 s; GPU peak 16,751 MiB with the miner resident. Against 4 October with the miner stopped (aggregate 2.2 s): the miner slows the prover 3 to 4x |
| 5090-class cards for 100% at 1 block/s, measured rows | 47 with the loop as it is, 18 through the chain mode on mining cards, 6 (approximate) on proving-only cards, at empty blocks; 14 proving-only at one full shard a block; 45 at B_p: the table in the bench log |
The rule, in one paragraph (spec 7.8)
From the first chain block A at or above proving_v1_activation_daa, chain blocks form fixed segments of N
(proving_v1_segment_blocks). A segment record carries the aggregated proof of the segment's last block, whose
chain_len says how many consecutive blocks the recursion attests. Every node checks the record's statement
against its own native block statement (every field but the provers commitment and chain_len), the chain rule
(a proof that does not chain to the previous segment, chain_len = N, is valid only for the first segment or
after an unproven one) and the deadline (T = proving_v1_unproven_daa DAA seconds after the segment's last block;
a record carried later pays nothing). The pool credit of every attested block splits: proving_v1_aggregator_share_bps
to the aggregator, the rest to the shards as v0. A block is never invalid for lack of a proof; the mandatory rule
(spec 7.8 item 10) is Designed and off, with no switch yet.
Rollout for 0.3.11 (the digest handshake pattern of 0.3.9 and tonight's switches)
The switch moves the consensus digest only once it is set (consensus_digest: the four v1 fields enter the hash
when proving_v1_activation_daa != never), so a 0.3.11 node on the unswitched devnet keeps the 0.3.10 digest and
the rolling upgrade does not partition the network. The order, each step with its check:
- Rebase and build. Fork
proving-v1rebased onto the 0.3.10 tip onrelease-0.3.6; the six node suites and the app tests as PC 2 build jobs; the Mac node and the Windows exes by the Mac cross-build; the Linux node by PC 1; the HiveOS package republished from the same fork commit (rule: the HiveOS package carries the node of the release commit and the same override object,infra/hive, as 0.3.9'shive-sync-039ochecked it). - Pinned guests. The guest ids do not change in this round (shard
0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a, aggregator0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896, pinned 2026-10-05T16:20:38Z): the aggregator guest already carried the chain rule and only the HOST gained modes. So no provers-off drain is needed for the guests;--mode idon every machine after the update must print the same two ids, and the node now reads them at start (program_ids:IGNEUM_PROOF_PROGRAM_IDSor the verifier's--mode id) and names them in the native statement. - Hand nodes and the seed first, with the UNCHANGED override object (the digest stays): observer, node 1, the
seed on the 0.3.11 node; peers back within 20 s; the
proving v1: segment records from DAA score neverline in each log. - Manifest and apps.
publish-manifest.sh --version 0.3.11with the unchanged object;update-nowto every app; every machine on 0.3.11 with a DAA score and a hash rate (the watcher takes the commit as an argument). The app's prover default applies at the first start on 0.3.11: every NVIDIA machine with WSL2 goes on; the log lineprover default: ...on each. - The switch. When every node runs 0.3.11: publish the object with
proving_v1_activation_daa= H (24 h ahead, the rule offee-switch-devnet.md) and the three parameters; read the expected digest on a scratch node first;update-now; the hand nodes and the seed with the same object; the digest sweep; the first paid segment record (igneum_getSegmentRecords) andigneum_getProvingStatus.v1.segmentsInWindowafter H. - The mandatory rule stays off: no switch exists for it yet; it gets one when the measured share is one.
The decisions (Decided 5 October 2026, delegated: the project lead, "I have no idea for most of this stuff so do a lot of research and deploy what is absolute best")
Each with its rule, its number and its evidence. The deploy is the DEVNET through 0.3.11 (not the public testnet).
What the other networks do (read 5 October 2026, 20:05 to 20:15 UTC; every figure from the page named, else labelled approximate)
| Network | Unit proven | Deadline | What a miss costs | Who is paid what | Measured latency |
|---|---|---|---|---|---|
| Taiko Alethia (L2BEAT page, protocol v2.1.0 notes) | a batch of blocks | proving window 2 h, cooldown 2 h (v2.1.0, February 2025); the Shasta inbox targets a 4-h proof submission cadence | the proposer's liveness bond is credited back in full when the batch is proved inside the window, half when outside; mainnet currently sets minBond and livenessBond to 0 | the prover earns the proving fee; two of four proofs needed (SGX Geth, SGX Reth, SP1, RISC0, at least one ZK) | 100% ZK coverage of mainnet blocks reached December 2025 (blockchain.news); preconfirmations 2 s |
| Boundless (docs.boundless.network, proof lifecycle) | one request | the requester's timeout (example 3,600 s) and a lock timeout (example 2,700 s); a reverse Dutch auction ramps the price from the minimum to the maximum over a ramp-up (example 300 s) | the locked collateral (example 5 ZKC) is slashed and used to pay another prover who fulfils the request | the prover's fee = the bid minus the market fee | not stated on the page |
| Succinct Prover Network (docs.succinct.xyz, SPN architecture and quickstart) | one request | the requester's deadline (the quickstart example: 10 minutes, 50 PROVE staked to bid, 100 PROVE maximum fee) | part or all of the winning prover's collateral slashed "according to protocol rules" | a reverse auction: the lowest bidder is assigned | "real-time", no number on the page |
| Aztec (docs.aztec.network economics; L2BEAT; forum) | an epoch of 32 blocks (38 min 24 s), a proof may cover one checkpoint (1 min 12 s) up to one epoch; maximum proof window 1 h 16 min | the epoch is declared failed only when its submission window expires | an unproven epoch is reorged out (no reward); proposals under discussion remove bonds and pay every prover that delivers on time | 400 AZTEC a slot: 70% sequencers, 30% provers (120 AZTEC), provers' share by an activity score | the public testnet proved by community provers (zkcloud blog), no page number |
| zkSync Era, Linea, Scroll (eco.com comparisons) | a batch | none on chain (the operator proves) | none | the operator | proof latency about 30 min (zkSync Era), 75 min (Linea), 90 min (Scroll), approximate |
Reading. Nobody pays an aggregator as a separate role: Aztec's 30% goes to whoever delivers the epoch proof, Taiko's fee to whoever proves the batch, the markets to the request's winner. Deadlines run from 10 minutes (Succinct's example) through 1 h (Boundless' example) to 2 h (Taiko) and 1 h 16 min (Aztec's maximum window); a miss forfeits the reward or part of a bond, and the slashed value goes to the prover who steps in (Boundless). Igneum has no bond on shards by decision (spec 7.2 item 4), so the forfeit here is the reward only.
The decisions
| Decision | Decided | Rule and number | Evidence |
|---|---|---|---|
proving_v1_segment_blocks (N) |
8 | Aggregation is a fixed cost per block, not per segment: 9.6 to 9.7 s for every chained block on a mining 5090, 7.9 s unchained (chain-pc2-pv1c), so N buys nothing in card time; it sets the record cadence and the forfeit. At N = 8 and 1 block/s a record every 8 s, a 1.27 MB proof gossiped every 8 s (159 KB/s per path, half of N = 4's 318 KB/s), and a missed segment forfeits 8 blocks' aggregator share (8 x 0.088 IGN at today's credit). The chain for 8 blocks cost 135.6 s cold on a mining card (66.8 s for 4), a fifth of T; pipelined per block it is 17 s after the last block. Aztec proves 32 blocks (38 min) as one; 8 blocks at 1 block/s is 8 s of chain, so the record lands well inside the minute the litepaper promises |
bench-log "proving v1" chain rows; Aztec economics page |
proving_v1_unproven_daa (T) |
600 DAA s (10 min) | T = p99 x 10: the measured block-to-carried-record latency of a shard record is p99 52 to 62 s (two 30-min windows), a cold chain of 8 adds 136 s and relay plus inclusion 10 to 40 s, about 240 s worst case; 600 leaves 2.5x on that and equals the 600-block record window of v0, so nothing is payable past it either way. Succinct's example deadline is the same 10 minutes; Boundless' example 1 h, Taiko 2 h, Aztec up to 1 h 16 min: Igneum's blocks are 1 s and its proofs seconds, so the shortest of the field. The forfeited aggregator share of an unproven segment STAYS IN THE POOL ESCROW (it is never paid, as an unproven shard's part today): no burn and no roll-over, the rule the pool already has, and the escrow is what later proofs are paid from | coverage rows; chain-pc2-pv1c; the table above |
proving_v1_aggregator_share_bps |
1,000 (a tenth) | The aggregator's card time per block is 9.7 s on a mining card against 4 x 10.6 s of shard proofs at B_p (19% of the card time) and 2.5 s against 42.5 s with the card to itself (6%); on tonight's empty blocks it is half the card time. A tenth of every attested block's pool credit sits between the two full-block ratios, pays a role no other network pays separately (Aztec pays its 30% to whoever delivers the epoch; the markets pay the winner), and leaves the shard provers 90%, which the fast-time harness showed paid exactly (shardWei 90% of the credit). The pool's 20% emission share itself is unchanged (spec 2.5, 5.3) |
chain-pc2-pv1c; bench-log 4 October 5090 rows; the harness |
proving_v1_activation_daa (H) |
the devnet tip + 14,400 at publish (4 h at 1 block/s), set by the 0.3.11 publisher in the same override object as program_class_v3 |
tonight's rule for consensus switches (the coordinator, 5 October 2026); the digest moves only once H is set, so the rolling update does not partition | |
| Aggregator sortition | none in v1: the first valid record carried wins | design 5.3's VRF draw (O-7.3) with one or two aggregators on the devnet changes nothing; the segment grid and the deadline already bound the race; revisit when a second aggregator exists | |
| Apple silicon default | off | the gate was "a shard under 60 s with the miner running": the M5 Max CPU took 41.3 and 55.4 s for EMPTY shards under tonight's load and 272 s for a 200-pgas shard on 4 October; a full shard at S_p was never under 60 s. Settings switches it on |
bench-log "proving v1" CPU chain row; 4 October CPU rows |
| The prover profile per card and the 12 GB and 16 GB gates (the project lead: "make sure we can prove on 12gb cards"; "is there any way we can make 12gb cards mine and prove?") | measured on PC 2, the rows below | the SP1 6.8.1 GPU server reads ELEMENT_THRESHOLD, HEIGHT_THRESHOLD, SHARD_SIZE and the SP1_WORKER_NUM_*/BUFFER_SIZE knobs from the environment it inherits (sp1-core-executor-6.8.1/src/opts.rs, sp1-prover-6.8.1/src/worker/config.rs); the app passes a profile per card (provedefault.rs) and the host forwards it |
the sweep job memsweep-pc2-pv1 and the miner-on run |
The resume path (5 October 2026, the 0.3.11 app): POST /api/resume on 0.3.9 re-armed only FAULTED cards (stop_miners("paused") clears every slot's restart_at), so a healthy paused card stayed "off" at 0 MH/s until the app was relaunched: PC 2 at 21:25:11Z (the aggregation-cost job's pause and resume; [ok] mining resumed then 0.00 MH/s, waiting for 20 minutes), the Mac that afternoon. Now every slot without a live worker is re-armed and its pack exported again before the start, and 90 s later resume_check logs resume: <card> is not mining 90 s after resume (state ..., pid ...) for every enabled card without a hash rate (engine.rs, three unit tests: the state machine, the 21:25:11Z case against the old rule, the check).
The prover-floor agent's first sweep (job floor-sweep-1, 22:34 to 22:38Z, PC 2's 5090, the miners stopped, this plan's per-point recipe, its patched sp1-gpu-server 5568108b built for sm_86, sm_89 and sm_120, every proof VERIFIED by the unpatched pv1 host): the control at upstream's sizes reproduces the curve above (empty shard 13,892 MiB and 2.2 s; the v1 shard 20,516 MiB and 4.2 s); with the core element threshold at 2^26 the v1 shard proves as four core shards in 5.3 s at 12,708 MiB and the empty shard at 12,772 MiB; 2^25 gives 12,836 MiB at 8.5 s; 2^27 gives 15,396 MiB. The 12.7 GB left is the server's Setup (five recursion keys pre-built at a fixed 2^27 capacity plus the shrink and core keys: 9.7 GB before the first shard), which its patch v2 sizes to the need. Decided for the 12 GB profile: the split that lands under 11 GB wins (5.3 s a shard is inside the loop's own 25 to 30 s of carriage and 100x inside T); 2^27 is the second profile only if v2 leaves it under 11 GB with the miner's 1.8 GB beside it. The 12 GB row stays OPEN until the final pair (alone and beside the miner) lands and the on-order 3060 runs it.
A self-built CUDA server (the 12 GB path), before 0.3.12 (consequences C26)
If the prover-floor agent's rebuilt sp1-gpu-server (the Setup sizes cut, built on PC 2 under WSL2) proves a shard under 11 GB, it becomes a shipped artefact and needs its own row of rules before 0.3.12: it is built from a pinned SP1 source tag with CUDA_ARCHS covering sm_86, sm_89 and sm_120 (the 12 and 16 GB tiers are Ampere and Ada, not only the 5090's Blackwell; one card family per measured row), by the packaging path that builds the Windows payload (PC 1's build job for the Linux binary, the Mac signs the manifest as it does the DMG), lands in the DMG and the WSL2 package beside the host as wsl2/bin/sp1-gpu-server with its sha256 in payload-inputs.json, is named in evidence.md beside the prover rows ("prover built from SP1 at "), is rebuilt and re-measured at every SP1 upgrade, and ships only after --mode verify-segment and --mode verify on proofs it made show the pinned verifying keys unchanged (the server changes allocation, not the circuit; the ids 0x2b1a81cb... and 0x474678f3... must still verify them). The 12 GB claim itself waits for the on-order RTX 3060 to run that server on the same fixtures and recipe as the curve; until then the public line stays at 24 GB.
The root-socket class on PC 2, the two times: 20:00:56Z (my chain job's root run; the live prover failed with Connect(PermissionDenied) until the socket was gone) and 21:25:24Z (the aggregation-cost job's root run; the prover stayed dark through the 0.3.10 restart at 21:49:41Z until socketfix-pc2-pv1 removed the root-owned /tmp/sp1-cuda-0.sock at 22:01:16Z; the next shard, block 89011, was proven at 22:02:13Z and paid, and every shard since). The permanent fix in the 0.3.11 app tree: every committed playbook that runs a prove mode as root carries pkill -f sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock at its start and end, tools/ci/prover-socket-check.sh (in ci.yml) fails a playbook without them, and the app's prover names the cause in its log line when the host reports PermissionDenied. The app itself cannot remove a socket another user owns, so a job written outside the tree must still follow the rule.
A prover job on a shared card runs as the app's user or cleans its socket (pkill -f sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock at the start and the end; tools/ci/prover-socket-check.sh): the root-socket fault of 20:00Z, bench-log.
The prover profiles: the tiers from the S_p curve (bench-log, "proving v1", the sweep, the miner-on pair and the curve)
The GPU server of SP1 6.8.1 sets the memory, not the shard: a floor of 13.9 GB for an empty shard, 20.4 GB for a full shard at the adopted v1 budget (30,000 pgas, 4.7 M cycles), 28.3 GB for the prototype shard (6.75 M pgas, 60 M cycles); the miner adds 1.7 GB when it shares the card; no environment knob moves the floor and the server has no options of its own; the witness is 5 to 22 KB a shard and never binds.
| Card | Alone | Beside the miner | Default (provedefault.rs) |
|---|---|---|---|
| 32 GB (RTX 5090) | the prototype shard, 28.3 GB, 10.8 s; the v1 shard 20.4 GB, 4.3 s | the prototype shard 30.1 GB, 33 s; the v1 shard 22.2 GB, 13.2 s | on, mine and prove, today |
| 24 GB (RTX 4090, 3090) | the v1 shard 20.4 GB; the prototype shard does NOT fit (28.3 GB) | the v1 shard 22.2 GB measured on the 5090's allocation (2.3 GB spare on a 24 GB card; approximate for the card itself) | on, mine and prove, with the line "until the devnet's fee switch its shards are the prototype size, which needs 32 GB, so this card proves from the switch on" |
| 16 GB (RTX 5080, 4080) | the shipped server: an empty shard only (13.9 GB); the patched server v3 b37defef at threshold 2^27: the v1 shard 12,915 MiB and 4.3 s, the prototype shard 13,459 MiB and 16.8 s (measured by the prover-floor agent on the 5090's allocation, job floor-sweep-3, 00:13 to 00:17Z 6 October; 2^27 + 2^26 gives 16,115 MiB, over the card) |
the shipped server: nothing (15.7 GB for an empty shard); the patched server at 2^27 beside the miner (the 5090 mining at 95%, 338 W, same card; floor-sweep-4, 00:35 to 00:38Z): the v1 shard 14,786 MiB total with the miner's 3,833 MiB resident inside it, the server's own 10,953 MiB, 17.4 s a shard; on a 16 GB card that is 10.95 GB server + 1.7 GB miner = 12.7 GB plus the display, under the 15.0 GB line |
off on the shipped server, with the line; on (mines and proves, 17 s a v1 shard, 4.3x the alone time) once the patched server ships (the packaging row below) |
| 12 GB (RTX 3060, 4070) | the shipped server: nothing (the floor is 13.9 GB, and the server refuses the card outright); the patched server v3 b37defef at threshold 2^26 (SP1_GPU_ELEMENT_THRESHOLD=67108864, the 12 GB profile): the v1 shard 10,291 MiB and 5.7 s, an empty shard 9,971 MiB and 3.3 s, the card's 2,089 MiB idle inside the peak and the server's own working set about 8.2 GB (6,535 MiB after Setup), so a 12 GB card proves alone with about 3 GB over it (measured by the prover-floor agent on the 5090's allocation, floor-sweep-3; the on-order RTX 3060 run is pending) |
measured beside the miner (floor-sweep-4): at 2^26 the v1 shard 12,066 MiB total with the miner's 3,833 MiB inside, the server's own 8,233 MiB, 24.4 s; the empty shard 11,586 MiB, 13.0 s; 2^25 gains nothing (12,066 MiB, 43.5 s). On a 12 GB card that is 8.2 GB server + 1.7 GB miner = 9.9 GB before the display, over the 9.0 GB line the project lead set, so mine-and-prove on 12 GB is NOT claimed |
off on the shipped server; "proves alone" on the patched one once it ships (the packaging row below); mine-and-prove stays off on 12 GB (9.9 GB plus the display, over the 9.0 GB line). The public gate stays "12 GB proves; 16 GB mines and proves", both on the patched server, both pending a run on the card itself. the project lead's "make sure we can prove on 12 GB cards" is answered on the 5090's allocation and OPEN on the card itself: the prover-floor agent (branch prover-floor, 5 October night) read SP1 v6.8.1's GPU server source (sp1-gpu/crates/prover_components/src/builder.rs lines 35 to 39): it reads the card's memory, adds 4 and panics under 24 ("Unsupported GPU memory ... must be at least 24GB"), and builds its core (ELEMENT_THRESHOLD 2^28 + 2^27 elements + 2^21), recursion (2^27), shrink (2^25) and wrap (85 M element) provers at Setup whatever the mode, which is the 13.9 GB floor; no knob reaches them, so the fix is a server rebuilt from source on PC 2 (WSL2, nvcc 12.8, CUDA_ARCHS=120) with those sizes cut, measured on the same fixtures and recipe as the curve above (D2 carries the curve) |
| under 12 GB | nothing | nothing | off, mine only |
| AMD-only and Apple machines | nothing on the GPU: no zkVM proves on an AMD GPU today (docs/analysis/amd-proving.md, branch amd-prove); the CPU prover is about 5 minutes a shard at a 30 GB RSS whatever the shard size (PC 1, bench-log "the SP1 CPU prover on PC 1") |
off, "mines and does not prove"; the only non-NVIDIA path with a shipped backend is RISC Zero's Metal prover behind the ProofSystem seam (a second guest and pinned id, a verifier for both formats, no shared aggregation): an open item, not 0.3.11 |
The three profile numbers the coordinator asked for, as measured: under 9.0 GB does not exist on this build (floor 13.9); under 15.0 GB mine-and-prove does not exist for any full shard (the v1 shard alone is 20.4); the full profile is the 32 GB card. The fleet table's "proving-only" rows therefore read 24 GB cards at the v1 budget and 32 GB cards at the prototype budget. Shards per block at the v1 budget: 1 on tonight's empty chain, 2 to 4 on blocks with transactions (B_p 120,000 = 4 x S_p); the aggregation count is one per block whatever the shard count (the chained recursion), so the aggregation-cost agent's target is per block.
The aggregation-cost agent's first rows (branch agg-cost, 5 October 2026 night, the same four live blocks on PC 2): with the miners paused an empty shard proves in 1.9 to 2.2 s and an aggregation in 1.7 to 2.2 s with the card 15.8% busy; mining, 7.4 to 7.8 s and 7.9 to 9.8 s at 93.9% busy, so the miner's kernels take the card and the prover runs 3.6x (shards) to 4.5x (aggregations) slower beside them; its batch-size curve is still open. That puts a proving-only card at about 4 s per empty block (one shard and one aggregation), 4 cards for an empty-block chain at 1 block/s, against 18 mining cards.
The re-plans of block 344 at 2.25 M and 4.5 M pgas peak at 28.3 to 28.4 GB alone (the server's buffers step up between 4.7 M and 20 M cycles and are flat to 60 M), so no shard size between the v1 budget and the prototype one changes a tier; with the miner the adopted shard proves 3.1x slower (13.2 s against 4.2 s) and the chained aggregation 9.7 s against 2.5 s: a mining 24 GB card delivers one adopted-size shard plus one aggregation in about 23 s, inside T by 25x.
The empty /api/state reply (6 October 2026)
The aggregation-cost agent's jobs read the two bytes {} from /api/state on PC 2 at 22:22Z, 22:41Z and 00:18Z (0.3.10 and 0.3.11); the 21:01Z reply was full. Cause, from the app source and node 1's RPC: ProvingState.paid_wei is a u128, and serde_json's to_value refuses a u128 over u64::MAX (18,446,744,073,709,551,615 wei, 18.45 IGN); state_json() turned that refusal into json!({}) with no log line. A paid shard is 1.23 IGN on average (node 1, igneum_getProvingStatus: 814.64 IGN over 663 shards at 00:3xZ), so the fifteenth paid shard after an app start empties the reply. PC 2's prover was blind to the root-owned socket from 20:00:56Z to 22:01Z (paid_wei stayed 0, hence the full reply at 21:01Z), proved from 22:02:13Z, and crossed 18.45 IGN inside its first 15 paid shards, before 22:22Z. Every app restart resets the counter, so the reply comes back for about 15 shards and goes again.
What it means: the dashboard on a proving machine shows nothing within about 12 minutes of its prover's first payout; every PC playbook that reads a card from /api/state fails the same way (the agent's job 5 reads settings.json instead). Mining, proving and payouts are untouched; it is the status page only.
Fix on the app branch: paid_wei serialises as a decimal string (the dashboard already reads it with Number()), state_json logs [error] state_json: ... once instead of answering {}, and the reply on any future serialisation error carries error and version rather than nothing; unit test a_paid_total_over_u64_max_still_serialises_the_whole_state. Not in 0.3.11 (that tree closed at 22c2363, master 630da6b, published); 6714a45 heads 0.3.12, the morning's first cut, app only, before PC 1's relaunch (coordinator, counter-asic-2-rollout.md 8a); until then the workaround is settings.json for the card keys.