igneum/docs/plans/proving-v1.md

20 KiB

Proving v1: segment records, the chain rule, the unproven rule; the 0.3.11 rollout

5 October 2026, from 18:55 UTC (the project lead: "open the proving round asap"). Branches proving-v1 in the main repository (worktree /Users/joshm/Projects/igneum-wt-proving-v1, from master a93199a) and in the fork (vendor/igneum-node-pv1, from release-0.3.6 a24ab01a; to be rebased onto the 0.3.10 tip when it lands on release-0.3.6). Status words follow docs/spec/00-overview.md 0.2. Every number here is in docs/bench-log.md with its command. Nothing ships from this plan: it delivers branches, numbers and the rollout for 0.3.11.

The gap this closes

The litepaper says every block is proven within about a minute. Proving v0 (proving-v0.md, spec 7.7) proves some shards: one prover (PC 2's RTX 5090) takes the newest shard assigned to it, about one shard every 30 s, so under a tenth of blocks carry a proof; consensus does not require one; the aggregator guest (design 5.3) runs on fixtures only. Proving v1 adds the aggregated segment record on chain (spec 7.8), the chain rule (segment N's record verifies N-1, inside the proof by recursion), the unproven rule (a segment nobody proves in T seconds pays nothing and may be skipped), the prover on by default on every machine that can prove, and the measurements that say how many cards cover the chain.

The round, step by step

Step What State
1 Prover on by default (app/igneum-app/src/provedefault.rs, engine.rs apply_prove_default): on at install when the machine can prove (NVIDIA card with 12 GB or more; WSL2 answering on Windows; Linux native; Apple silicon off until measured), never switching an explicit on back off; the Settings switch line and the tile line say why. Unit tests (5). The cost of proving on a mining machine: PC 2 job prover-cost-pc2-pv1 (5 min mining alone, 5 min with the prover, the GPU memory peak and the host RAM peak, the sp1-gpu-server's compiled SM targets) Implemented; the measurement is HELD (coordinator, 19:00Z): PC 2's RTX 5090 worker has been exiting on a pack seed mismatch since 18:35Z, so the first run's "mining alone" is 0 MH/s and void; re-run after the go
2 Segment aggregation: SegmentRecord (586 bytes, the aggregator guest's 340-byte statement inline), section IGNS before the shard section, p2p message 75 at protocol 15 (14 went to the EVM transaction relay in 0.3.10), the native block statement and the veto, the credit split (split_pool_credit), the payout at the carrier, igneum-miner sign-segment-record, the RPCs; host modes chain (consecutive fixtures), aggregate (live shard proofs from the pool, a run of blocks in one process) and verify-segment (the node's verifier, pinned aggregator key); the app's aggregator step (prover.rs aggregate_once) Implemented, unit-tested (consensus core 2 new tests, exec 2, params 1); the GPU measurement (N = 2, 4, 8 blocks on PC 2) is HELD with step 1; the Mac CPU run of --mode chain over 2 live blocks is the known-finished case
3 Coverage: tools/proving-v1/coverage.mjs (the proven-block share and the on-chain proof latency over a window from one node's RPC, the live page beside it) Implemented and run for 3 min (below); the 30-min window waits for the fleet
4 The chain rule and the unproven rule in consensus behind proving_v1_activation_daa (spec 7.8 items 2, 6, 7); unit tests; the fast-time 3-node harness tools/proving-v1/net.mjs (ports 29950+, suffix 956, trust mode) with the known-finished and known-failed cases Implemented; the harness run waits for the Mac build of the fork (vendor/igneum-node/target-pv1)
5 This plan: the rollout for 0.3.11 and the project lead's decisions Written below

Numbers (every one from docs/bench-log.md, "proving v1: segment records ...", 5 October 2026 evening)

What Number
The prover's cost to a mining 5090 (the re-run with the fleet mining, hash rate from the miner's own STATUS lines) 124.7 MH/s alone, 119.7 MH/s with the prover on: 5.0 MH/s, 4.0%, on empty shards at 1.4 a minute
GPU memory on the 5090: the prover alone (empty shards) / the miner and the prover together max 13,816 MiB / max 15,590 MiB (the miner holds 3,396 MiB); a 24 GB 4090 has 8.4 GB of headroom, a 16 GB card 0.4 GB, a 12 GB card cannot do both on this build; the full-shard peak is the chain job's row
Coverage, 30-min window with the fleet mining, one prover 2.4% of blocks, latency p50 44 s, p99 52 s
Host RAM host used 25.6 GB of 63 GB; the WSL2 VM 7.9 GB working set
sp1-gpu-server 6.8.1 compiled targets (cuobjdump) sm_80, sm_86, sm_89, sm_90, sm_100, sm_120 and compute_120 PTX: Ada (4090) is native, no JIT; nothing for AMD
Shards a minute, one 5090 through the app's loop (empty shards) 1.6
Chain of 2 live blocks on the Mac CPU (--mode chain) shard 55.4 and 41.3 s, aggregate 52.0 s then 59.1 s with the previous proof, chain_len 2, final proof 1,272,909 bytes, verify-segment 0.032 s
Unit tests consensus core 13, exec 8, app 5, all passing on the Mac
The harness (3 nodes, fast time, trust mode) PASSED, 21 checks in 197 s on b177718e (N = 4) and 244 s on the final tree ece42979 (N = 8): paid 1.0 s after submit, every node agreeing; the fresh chain refused after a proven segment; the unproven segment skipped after its deadline; shards at 90%
Coverage, 3-min window, the degraded fleet (one card, the Mac verifier down) 4.7% of blocks proven, on-chain latency p50 39 s
The chain of 8 live blocks on the 5090, the card also mining (chain-pc2-pv1c) shard 7.3 to 7.7 s, first aggregation 7.9 s, every chained one 9.6 to 9.7 s; N = 2 in 32.6 s, N = 4 in 66.8 s, N = 8 in 135.6 s (17.0 s a block); the final proof 1,272,909 bytes whatever N, the record 586 bytes, verify-segment 0.037 to 0.040 s; GPU peak 16,751 MiB with the miner resident. Against 4 October with the miner stopped (aggregate 2.2 s): the miner slows the prover 3 to 4x
5090-class cards for 100% at 1 block/s, measured rows 47 with the loop as it is, 18 through the chain mode on mining cards, 6 (approximate) on proving-only cards, at empty blocks; 14 proving-only at one full shard a block; 45 at B_p: the table in the bench log

The rule, in one paragraph (spec 7.8)

From the first chain block A at or above proving_v1_activation_daa, chain blocks form fixed segments of N (proving_v1_segment_blocks). A segment record carries the aggregated proof of the segment's last block, whose chain_len says how many consecutive blocks the recursion attests. Every node checks the record's statement against its own native block statement (every field but the provers commitment and chain_len), the chain rule (a proof that does not chain to the previous segment, chain_len = N, is valid only for the first segment or after an unproven one) and the deadline (T = proving_v1_unproven_daa DAA seconds after the segment's last block; a record carried later pays nothing). The pool credit of every attested block splits: proving_v1_aggregator_share_bps to the aggregator, the rest to the shards as v0. A block is never invalid for lack of a proof; the mandatory rule (spec 7.8 item 10) is Designed and off, with no switch yet.

Rollout for 0.3.11 (the digest handshake pattern of 0.3.9 and tonight's switches)

The switch moves the consensus digest only once it is set (consensus_digest: the four v1 fields enter the hash when proving_v1_activation_daa != never), so a 0.3.11 node on the unswitched devnet keeps the 0.3.10 digest and the rolling upgrade does not partition the network. The order, each step with its check:

  1. Rebase and build. Fork proving-v1 rebased onto the 0.3.10 tip on release-0.3.6; the six node suites and the app tests as PC 2 build jobs; the Mac node and the Windows exes by the Mac cross-build; the Linux node by PC 1; the HiveOS package republished from the same fork commit (rule: the HiveOS package carries the node of the release commit and the same override object, infra/hive, as 0.3.9's hive-sync-039o checked it).
  2. Pinned guests. The guest ids do not change in this round (shard 0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a, aggregator 0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896, pinned 2026-10-05T16:20:38Z): the aggregator guest already carried the chain rule and only the HOST gained modes. So no provers-off drain is needed for the guests; --mode id on every machine after the update must print the same two ids, and the node now reads them at start (program_ids: IGNEUM_PROOF_PROGRAM_IDS or the verifier's --mode id) and names them in the native statement.
  3. Hand nodes and the seed first, with the UNCHANGED override object (the digest stays): observer, node 1, the seed on the 0.3.11 node; peers back within 20 s; the proving v1: segment records from DAA score never line in each log.
  4. Manifest and apps. publish-manifest.sh --version 0.3.11 with the unchanged object; update-now to every app; every machine on 0.3.11 with a DAA score and a hash rate (the watcher takes the commit as an argument). The app's prover default applies at the first start on 0.3.11: every NVIDIA machine with WSL2 goes on; the log line prover default: ... on each.
  5. The switch. When every node runs 0.3.11: publish the object with proving_v1_activation_daa = H (24 h ahead, the rule of fee-switch-devnet.md) and the three parameters; read the expected digest on a scratch node first; update-now; the hand nodes and the seed with the same object; the digest sweep; the first paid segment record (igneum_getSegmentRecords) and igneum_getProvingStatus.v1.segmentsInWindow after H.
  6. The mandatory rule stays off: no switch exists for it yet; it gets one when the measured share is one.

The decisions (Decided 5 October 2026, delegated: the project lead, "I have no idea for most of this stuff so do a lot of research and deploy what is absolute best")

Each with its rule, its number and its evidence. The deploy is the DEVNET through 0.3.11 (not the public testnet).

What the other networks do (read 5 October 2026, 20:05 to 20:15 UTC; every figure from the page named, else labelled approximate)

Network Unit proven Deadline What a miss costs Who is paid what Measured latency
Taiko Alethia (L2BEAT page, protocol v2.1.0 notes) a batch of blocks proving window 2 h, cooldown 2 h (v2.1.0, February 2025); the Shasta inbox targets a 4-h proof submission cadence the proposer's liveness bond is credited back in full when the batch is proved inside the window, half when outside; mainnet currently sets minBond and livenessBond to 0 the prover earns the proving fee; two of four proofs needed (SGX Geth, SGX Reth, SP1, RISC0, at least one ZK) 100% ZK coverage of mainnet blocks reached December 2025 (blockchain.news); preconfirmations 2 s
Boundless (docs.boundless.network, proof lifecycle) one request the requester's timeout (example 3,600 s) and a lock timeout (example 2,700 s); a reverse Dutch auction ramps the price from the minimum to the maximum over a ramp-up (example 300 s) the locked collateral (example 5 ZKC) is slashed and used to pay another prover who fulfils the request the prover's fee = the bid minus the market fee not stated on the page
Succinct Prover Network (docs.succinct.xyz, SPN architecture and quickstart) one request the requester's deadline (the quickstart example: 10 minutes, 50 PROVE staked to bid, 100 PROVE maximum fee) part or all of the winning prover's collateral slashed "according to protocol rules" a reverse auction: the lowest bidder is assigned "real-time", no number on the page
Aztec (docs.aztec.network economics; L2BEAT; forum) an epoch of 32 blocks (38 min 24 s), a proof may cover one checkpoint (1 min 12 s) up to one epoch; maximum proof window 1 h 16 min the epoch is declared failed only when its submission window expires an unproven epoch is reorged out (no reward); proposals under discussion remove bonds and pay every prover that delivers on time 400 AZTEC a slot: 70% sequencers, 30% provers (120 AZTEC), provers' share by an activity score the public testnet proved by community provers (zkcloud blog), no page number
zkSync Era, Linea, Scroll (eco.com comparisons) a batch none on chain (the operator proves) none the operator proof latency about 30 min (zkSync Era), 75 min (Linea), 90 min (Scroll), approximate

Reading. Nobody pays an aggregator as a separate role: Aztec's 30% goes to whoever delivers the epoch proof, Taiko's fee to whoever proves the batch, the markets to the request's winner. Deadlines run from 10 minutes (Succinct's example) through 1 h (Boundless' example) to 2 h (Taiko) and 1 h 16 min (Aztec's maximum window); a miss forfeits the reward or part of a bond, and the slashed value goes to the prover who steps in (Boundless). Igneum has no bond on shards by decision (spec 7.2 item 4), so the forfeit here is the reward only.

The decisions

Decision Decided Rule and number Evidence
proving_v1_segment_blocks (N) 8 Aggregation is a fixed cost per block, not per segment: 9.6 to 9.7 s for every chained block on a mining 5090, 7.9 s unchained (chain-pc2-pv1c), so N buys nothing in card time; it sets the record cadence and the forfeit. At N = 8 and 1 block/s a record every 8 s, a 1.27 MB proof gossiped every 8 s (159 KB/s per path, half of N = 4's 318 KB/s), and a missed segment forfeits 8 blocks' aggregator share (8 x 0.088 IGN at today's credit). The chain for 8 blocks cost 135.6 s cold on a mining card (66.8 s for 4), a fifth of T; pipelined per block it is 17 s after the last block. Aztec proves 32 blocks (38 min) as one; 8 blocks at 1 block/s is 8 s of chain, so the record lands well inside the minute the litepaper promises bench-log "proving v1" chain rows; Aztec economics page
proving_v1_unproven_daa (T) 600 DAA s (10 min) T = p99 x 10: the measured block-to-carried-record latency of a shard record is p99 52 to 62 s (two 30-min windows), a cold chain of 8 adds 136 s and relay plus inclusion 10 to 40 s, about 240 s worst case; 600 leaves 2.5x on that and equals the 600-block record window of v0, so nothing is payable past it either way. Succinct's example deadline is the same 10 minutes; Boundless' example 1 h, Taiko 2 h, Aztec up to 1 h 16 min: Igneum's blocks are 1 s and its proofs seconds, so the shortest of the field. The forfeited aggregator share of an unproven segment STAYS IN THE POOL ESCROW (it is never paid, as an unproven shard's part today): no burn and no roll-over, the rule the pool already has, and the escrow is what later proofs are paid from coverage rows; chain-pc2-pv1c; the table above
proving_v1_aggregator_share_bps 1,000 (a tenth) The aggregator's card time per block is 9.7 s on a mining card against 4 x 10.6 s of shard proofs at B_p (19% of the card time) and 2.5 s against 42.5 s with the card to itself (6%); on tonight's empty blocks it is half the card time. A tenth of every attested block's pool credit sits between the two full-block ratios, pays a role no other network pays separately (Aztec pays its 30% to whoever delivers the epoch; the markets pay the winner), and leaves the shard provers 90%, which the fast-time harness showed paid exactly (shardWei 90% of the credit). The pool's 20% emission share itself is unchanged (spec 2.5, 5.3) chain-pc2-pv1c; bench-log 4 October 5090 rows; the harness
proving_v1_activation_daa (H) the devnet tip + 14,400 at publish (4 h at 1 block/s), set by the 0.3.11 publisher in the same override object as program_class_v3 tonight's rule for consensus switches (the coordinator, 5 October 2026); the digest moves only once H is set, so the rolling update does not partition
Aggregator sortition none in v1: the first valid record carried wins design 5.3's VRF draw (O-7.3) with one or two aggregators on the devnet changes nothing; the segment grid and the deadline already bound the race; revisit when a second aggregator exists
Apple silicon default off the gate was "a shard under 60 s with the miner running": the M5 Max CPU took 41.3 and 55.4 s for EMPTY shards under tonight's load and 272 s for a 200-pgas shard on 4 October; a full shard at S_p was never under 60 s. Settings switches it on bench-log "proving v1" CPU chain row; 4 October CPU rows
The prover profile per card and the 12 GB and 16 GB gates (the project lead: "make sure we can prove on 12gb cards"; "is there any way we can make 12gb cards mine and prove?") measured on PC 2, the rows below the SP1 6.8.1 GPU server reads ELEMENT_THRESHOLD, HEIGHT_THRESHOLD, SHARD_SIZE and the SP1_WORKER_NUM_*/BUFFER_SIZE knobs from the environment it inherits (sp1-core-executor-6.8.1/src/opts.rs, sp1-prover-6.8.1/src/worker/config.rs); the app passes a profile per card (provedefault.rs) and the host forwards it the sweep job memsweep-pc2-pv1 and the miner-on run

A prover job on a shared card runs as the app's user or cleans its socket (pkill -f sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock at the start and the end; tools/ci/prover-socket-check.sh): the root-socket fault of 20:00Z, bench-log.

The prover profiles: the tiers from the S_p curve (bench-log, "proving v1", the sweep, the miner-on pair and the curve)

The GPU server of SP1 6.8.1 sets the memory, not the shard: a floor of 13.9 GB for an empty shard, 20.4 GB for a full shard at the adopted v1 budget (30,000 pgas, 4.7 M cycles), 28.3 GB for the prototype shard (6.75 M pgas, 60 M cycles); the miner adds 1.7 GB when it shares the card; no environment knob moves the floor and the server has no options of its own; the witness is 5 to 22 KB a shard and never binds.

Card Alone Beside the miner Default (provedefault.rs)
32 GB (RTX 5090) the prototype shard, 28.3 GB, 10.8 s; the v1 shard 20.4 GB, 4.3 s the prototype shard 30.1 GB, 33 s; the v1 shard 22.2 GB, 13.2 s on, mine and prove, today
24 GB (RTX 4090, 3090) the v1 shard 20.4 GB; the prototype shard does NOT fit (28.3 GB) the v1 shard 22.2 GB measured on the 5090's allocation (2.3 GB spare on a 24 GB card; approximate for the card itself) on, mine and prove, with the line "until the devnet's fee switch its shards are the prototype size, which needs 32 GB, so this card proves from the switch on"
16 GB (RTX 5080, 4080) an empty shard only (13.9 GB) nothing (15.7 GB for an empty shard, no room for the display) off, with the line
12 GB (RTX 3060, 4070) nothing: the floor is 13.9 GB nothing off; the project lead's "make sure we can prove on 12 GB cards" is OPEN: it needs a prover build with a smaller floor (an SP1 release or a fork of its GPU server), measured on a 12 GB card; the curve above is the evidence for that ask (D2)
under 12 GB nothing nothing off, mine only
AMD-only and Apple machines nothing on the GPU: no zkVM proves on an AMD GPU today (docs/analysis/amd-proving.md, branch amd-prove); the CPU prover is about 5 minutes a shard at a 30 GB RSS whatever the shard size (PC 1, bench-log "the SP1 CPU prover on PC 1") off, "mines and does not prove"; the only non-NVIDIA path with a shipped backend is RISC Zero's Metal prover behind the ProofSystem seam (a second guest and pinned id, a verifier for both formats, no shared aggregation): an open item, not 0.3.11

The three profile numbers the coordinator asked for, as measured: under 9.0 GB does not exist on this build (floor 13.9); under 15.0 GB mine-and-prove does not exist for any full shard (the v1 shard alone is 20.4); the full profile is the 32 GB card. The fleet table's "proving-only" rows therefore read 24 GB cards at the v1 budget and 32 GB cards at the prototype budget. Shards per block at the v1 budget: 1 on tonight's empty chain, 2 to 4 on blocks with transactions (B_p 120,000 = 4 x S_p); the aggregation count is one per block whatever the shard count (the chained recursion), so the aggregation-cost agent's target is per block.

The re-plans of block 344 at 2.25 M and 4.5 M pgas peak at 28.3 to 28.4 GB alone (the server's buffers step up between 4.7 M and 20 M cycles and are flat to 60 M), so no shard size between the v1 budget and the prototype one changes a tier; with the miner the adopted shard proves 3.1x slower (13.2 s against 4.2 s) and the chained aggregation 9.7 s against 2.5 s: a mining 24 GB card delivers one adopted-size shard plus one aggregation in about 23 s, inside T by 25x.