igneum/docs/plans/proving-v0.md
igneum-labs 326953e0ac Proving v0: SP1 guest and host for one Igneum block, real-block fixtures, versioned ProofSystem trait, WSL2 package for the RTX 5090 run
proving/igneum-prove: core (port of igneum-exec at fb33069 as the block statement), program (SP1 v6.8.1 guest),
host (execute, core, compressed; ProofSystem trait with the stub and the SP1 implementation), export (cuts a block
out of igneum_exportSegments and checks every state root against the node's). Fixtures block-78-increment and
block-56-transfers from the 3-node simnet. proving/windows-wsl2: SETUP-PROVER.bat, setup-wsl.sh, PROVE-BLOCK.bat,
make-package.sh. docs/plans/proving-v0.md: the devnet v4 shard plan, what tonight's proof shows and does not, the
morning acceptance line. Mac CPU baseline (block 78: 626 k cycles, core 22.0 s and 7.3 MB, compressed 55.7 s and
1.27 MB, both verified) appended to docs/bench-log.md, left uncommitted because that file carries another agent's
pending changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 22:30:12 +00:00

8.2 KiB

Proving v0 and the devnet v4 shard plan

3 October 2026. Code in proving/ (SP1 v6.8.1 guest and host, the versioned ProofSystem trait, two real-block fixtures, the Windows WSL2 package). Status words follow docs/spec/00-overview.md 0.2. Nothing here is a mainnet number.

What tonight's single-block proof is

The guest (proving/igneum-prove/program) re-executes one Igneum chain block with the execution layer's own rules: the transaction decoder is the node's crate (igneum-evm-types), the executor is a port of igneum-exec at commit fb33069 (rewards by rule, the nonce-rule skip, two-dimensional gas with the prototype pgas table, fee flows with the 80/20 developer split, the registry population on CREATE, the state root over every non-empty account through alloy-trie). It commits the chain id, the block number and hash, a commitment to the ordered transactions, the pre-state root, the post-state root, the receipts root, gas and pgas used, and the executed and skipped counts (igneum_prove_core::ProveOutput, 200 bytes, fixed layout).

The fixtures are blocks that happened: proving/fixtures/block-78-increment.json (the increment(5) call on the Counter contract, one executed transaction and one duplicate skipped by the nonce rule, 10 accounts) and block-56-transfers.json (the three funding transfers, 5 accounts), cut by igneum-prove-export from tools/evm-smoke/seq.json, the igneum_exportSegments dump of the 3-node simnet run. The exporter replays all 79 segments from genesis through the port and refuses to write unless every segment's state root equals the node's; it did (Measured, 3 October 2026, final root 0x5b18b3a5...). So the proof statement is "the node's executor, run twice", which is what design 2.1 asks for, up to the duplicate of the executor that the port is until the executor becomes the no-network library of design 2.1.

The host runs the statement natively first and refuses to prove a fixture whose expected roots it does not reproduce, then runs SP1 in three modes behind the trait (proof_system.rs): execute (cycle count, no proof), core (prove_shard), compressed (aggregate over one shard). Every proof is verified and its public values are compared with the native run. SP1_PROVER=cuda selects the GPU prover through the same binary (feature cuda).

What it does not show

Not shown Why Where it goes
Shard time at S_p pgas on a 12 GB card (the phase 2 gate, ledger P1, overclaim 27) the fixture blocks carry 600 and 1,488 pgas, not a shard's budget devnet v4 shard plan below, 3060-class card
MPT witnesses the whole in-memory devnet state (5 to 10 accounts) is the input; a real shard gets the touched accounts, slots and trie nodes (design 5.1), and the trie-proof share of pgas (R3) is unmeasured witness generator in the node, v4
Recursion over shards and over the chain (segment N verifies N-1) v0 aggregates one shard: SP1's compress stage on the block v4 aggregator
The Groth16 or Plonk wrapper for light clients (ledger P3, overclaim 25) not run; needs SP1's circuit artifacts and a measurement on consumer hardware R4, phase 2 benchmark
The prover's payout key in the statement (ledger P12) ShardWitness.prover is carried, not committed guest change, v4
pgas calibration (R1) the table is the prototype; tonight's cycles per EVM gas are one data point calibrate per opcode in SP1 with three input sizes

The shard plan for devnet v4 (Designed)

  1. Cut. The native executor emits per segment the transaction boundaries with cumulative gas and pgas and the state root at each boundary (design 5.1). The planner cuts at boundaries so each shard's pgas is at most S_p; one transaction above S_p is one shard proven with continuations (SP1 checkpoints its own execution; approximate). The plan is a pure function of the trace, so every node computes the same shard list and the same ids (segment hash, shard index).
  2. Statement per shard. From pre-root r_i and the committed transaction list of shard i, running the executor yields r_{i+1} and receipts root h_i, with the executed set and the prover's payout key as public outputs (P12). The witness (touched accounts, slots, MPT nodes, block environment) is produced by any full node from its own execution and served over the proving gossip; it is not consensus data. Tonight's guest is this statement with the whole state as the witness and one shard per block.
  3. Claim and sortition (spec 7.2). Eligible provers are the vote keys above dust (100 blue blocks in the 30-day window). For each shard, draws r_n = H("igneum-shard/" || epoch_seed || shard_id || n) mod N over the window's blue blocks pick 8 distinct keys by weight. They hold the shard for 10 DAA seconds from the moment the segment is executed: a valid proof by one of them, included in a block in the window, earns the shard's part of the pool; after the window anyone's first valid proof earns it. No claim and no bond (ledger P8, P9). Parameters 8 and 10 s are set on the devnet from the shard-time distribution across three prover speeds (R7, fastest prover under 25% of shards; one key versus 1,000 keys of equal weight win the same number, F17).
  4. Aggregation by recursion (design 5.3). Shard proofs of a segment fold into one segment proof (SP1 compress, recursion over the shard proofs, not over the witness as v0 does); the segment proof for N also verifies the segment proof for N-1 so one proof attests the chain of state and a client keeps only the latest. The aggregator is anyone, paid the aggregator share. The wrapper to bn254 for the bridge verifier and phones is measured, not assumed (R4).
  5. Proof records and the veto (design 5.4, 5.5, spec 7.2 item 5). A block carries zero or more ProofRecord {version, segment, pre_root, post_root, receipts, provers, aggregator, proof}; full nodes verify the proof natively (a compressed proof, tonight's verify time is the first number for that) and run the native-execution veto: the block is invalid if the record names a chain block off the carrying block's own selected-parent chain or if post_root or receipts differ from the native execution of that segment along that chain. The test is relative to the carrying block's past, never to the validating node's current chain (ledger P11). A forged proof from a soundness bug is therefore rejected by every full node; the damage is bounded to light clients.
  6. Devnet v4 run. Nodes emit shard plans; miners run the host as a prover service that takes gossiped witnesses and returns shard proofs; a block producer includes records. Acceptance for the execution layer's A5 and A6 (docs/design/execution-layer.md 9.2): records within 60 s of 95% of segments, a wrong post_root makes the block invalid on every other node.

The Mac CPU baseline (Measured, 3 October 2026, docs/bench-log.md)

block-78-increment on the Apple M5 Max CPU under load, SP1 v6.8.1: 626,246 cycles (14 per EVM gas), core proof 22.0 s and 7.3 MB (verify 0.16 s), compressed proof 55.7 s and 1.27 MB (verify 0.03 s), both verified and both reproducing the node's post-state root 0x5b18b3a5.... block-56-transfers: 549,469 cycles. These are fixed-overhead numbers for a block far below one SP1 shard, not throughput.

The morning run

On the Windows PC through WSL2 (proving/windows-wsl2/): SETUP-PROVER.bat (reboot once), SETUP-PROVER.bat again, pause mining, PROVE-BLOCK.bat. It proves block-78-increment on the RTX 5090 (execute, core, compressed, each verified) and then a core proof on the CPU, prints the RESULT lines and uploads the log to the intake. Acceptance line: one Igneum devnet block proven on the RTX 5090 and verified, time and size recorded (the RESULT lines go to docs/bench-log.md next to the Mac CPU baseline).

Two uncertainties the run settles. First, whether SP1's GPU prover runs under WSL2 at all with the 5090 (Blackwell, compute capability 12.0; the documentation lists 8.0 or higher and 24 GB of VRAM, so it should; the server binary is prebuilt for Linux x86_64 and the SDK downloads it, and the WSL CUDA driver path is the usual failure point). Second, how long the compressed proof takes on one consumer card for a 1,488-pgas block: this is the first point on the curve that sets S_p, and no shard has been proven on any card before this (ledger P1).