igneum/docs/plans/proving-v0.md
2026-10-04 10:26:52 +00:00

7.8 KiB

Proving v0 and the devnet v4 shards

3 October 2026 (v0: one proof per block) and 4 October 2026 (devnet v4: shards implemented, measured on the Mac CPU; the GPU run is pending). Code in proving/ (SP1 v6.8.1, two guests and a host, the versioned ProofSystem trait, six fixtures, the Windows WSL2 package). Status words follow docs/spec/00-overview.md 0.2. Nothing here is a mainnet number.

Shards implemented, measured: what

Implemented 4 October 2026 (proving/igneum-prove, commits "Proving: shard cutter..." and after), checked on this Mac's CPU (docs/bench-log.md, "proving: devnet v4 shards"):

Step Status Where
Cutter Implemented, Measured. The executor runs any contiguous range of a segment from a carried-in position and records a boundary per transaction (cumulative gas and pgas, the carry link, natively the state root); the planner cuts at transaction boundaries to at most S_p pgas, a pure function of the trace. The host checks that the shards chain (roots, links) and sum (gas, pgas) to the whole block on every fixture core/src/executor.rs, core/src/plan.rs
Witnesses Implemented, Measured. Per shard: the touched accounts, slots and code, and partial Merkle Patricia tries (the touched leaves in full, every untouched subtree as a hash, collapse-safe siblings carried) for the account trie and each touched storage trie. The guest rebuilds the pre-root from them, refuses any read they do not cover (a dropped account makes the proof impossible), and rebuilds the post-root after execution. 60 randomised rounds of inserts, updates and deletes against the full trie in the unit test; three tamper variants rejected on every fixture (a re-encoded balance moves the pre-root away from the node's; a flipped storage value or code byte breaks the storage-root or code-hash check; a dropped account is refused) core/src/trie.rs, core/src/witness.rs
Shard statement Implemented, Measured. Public values: chain id, block number and hash, shard index, transaction range, carry links in and out, transaction accumulators in and out, pre-root, post-root, the shard's receipts root, gas, pgas, executed and skipped counts, and the prover's payout address (ledger P12). 328 bytes, fixed layout core/src/shard.rs, program/
Aggregator Implemented; Measured in SP1's executor and (CPU) end to end on the three-shard test fixture. The aggregator guest verifies every shard proof against the shard program's key (SP1 deferred proofs), checks the chain (indices contiguous, link_in = previous link_out, pre-root = previous post-root, accumulators), sums gas and pgas, and commits the block statement with keccak over the shard receipts roots and over the provers, plus the shard program id. The chain rule (segment N verifies N-1, design 5.3) is in the guest and the proof system (AggInput.prev, BlockOutput.agg_vk, chain_len) but has no host mode and no measurement yet: next step core/src/agg.rs, aggregator/, host/src/proof_system.rs
Fixtures of real size Implemented. Blocks 338, 341 and 344 of a private one-node simnet (4 October 2026): 0.90, 1.80 and 3.60 S_p of proving gas, one, two and four shards; plus the v0 blocks re-cut (one shard each) and block 56 cut at a test budget of 200 pgas into three shards for the Mac CPU check proving/fixtures/, tools/prove-fixtures/
Host modes Implemented. native, execute, shard (execute, core, compressed, verified), block (compressed proof per shard, aggregation, verified against the shard program id and the claim), all; --shard, --prover, --out. A STAGE line and a RESULT line with a UTC timestamp per stage; a Tokio runtime held for the whole run with the proof system dropped inside it (ledger P20) host/src/main.rs
WSL2 package Implemented, not yet run. PROVE-SHARD.bat: the shard at S_p in all three stages, then the two- and four-shard blocks end to end, on the GPU; PROVE-BLOCK.bat kept for the small block. Zip at ~/Desktop/igneum-prove-wsl2.zip proving/windows-wsl2/

S_p. The specification had no number. Provisional since 4 October 2026: S_p = B_p / 4 = 7,500,000 pgas, so a block at its proving budget is exactly four shards (spec 7.4, 7.6). It is set from the shard-time measurement on the card, never from the fixtures (benchmark standard 2.3).

Measured on the Mac CPU (4 October 2026, loaded machine, docs/bench-log.md)

What Number
One shard at S_p (block 344, 6.75 M pgas, three modexp calls plus transfers): SP1 cycles 59.6 M to 60.4 M per shard, 9 cycles per pgas, 46 to 50 per EVM gas
Shard input (witness and transactions) 17 to 21 KB per shard; 7 to 19 accounts, 1 to 3 slots, 7 to 24 trie leaves, 10 to 21 subtree hashes
Aggregator statement over four shards (no proofs, executor only) 1.66 M cycles
A 200-pgas shard (block 56 test cut, one transfer): execute, core, compressed 315 k cycles; core 83.1 s, 7.31 MB, verify 0.37 s; compressed 272.3 s, 1.27 MB, verify 0.08 s; all verified
Three 200-pgas shards plus aggregation on the CPU (--mode block) compressed shard proofs 336.9, 305.4 and 245.3 s (1.27 MB each, verify 0.06 to 0.07 s); aggregation by recursion over the three 244.5 s, block proof 1,272,909 bytes, verify 0.084 s, VERIFIED with the shard program id and the claim checked; 19 minutes end to end from the first shard proof to the verified block proof

Reading. The modexp-heavy shard costs 9 SP1 cycles per pgas against the unit's 1,000: the prototype table's modexp entry (1,000 + 10 per input byte) is two orders of magnitude above its SP1 cost, which is the R1 calibration in one number; at the current table a shard of S_p is about 60 M cycles, far below what the unit would imply. The witness is small because the devnet state is small; the share of the shard's cycles spent on the trie is not isolated yet (R3). The CPU times are on a machine at load 40 shared with the live devnet and other agents' builds; they are correctness runs, not throughput.

What waits for the GPU

PROVE-SHARD.bat on the RTX 5090: the shard at S_p in the three stages and the two- and four-shard blocks end to end. The RESULT lines fill the GPU row of the bench-log entry and give the first point for S_p. The P20 Drop fix and the gap before the first compressed stage are confirmed or not by the same run.

What it still does not show

Not shown Why Where it goes
Shard time on a 12 GB card (the phase 2 gate, ledger P1, overclaim 27) the 5090 run is pending; no 3060-class card has run it PROVE-SHARD.bat, then the 3060-class card
The chain rule measured (segment N verifies N-1) in the guest and the proof system, no host mode for two consecutive fixtures yet next step: --mode chain over consecutive blocks
The Groth16 or Plonk wrapper for light clients (ledger P3, overclaim 25) not run; needs SP1's circuit artifacts and a measurement on consumer hardware R4, phase 2 benchmark
Sortition, proof records, the native-execution veto in the node the node has no proof records on devnet v4 design 5.4, 5.5; the stub ProofSystem path exists in the host
pgas calibration (R1) the table is the prototype; 9 cycles per pgas on modexp, 1,400 to 1,600 on a plain transfer shard calibrate per opcode in SP1 with three input sizes
Trie share of pgas (R3) not isolated instrument the guest

The v0 statement (3 October 2026), for the record

The v0 guest re-executed one whole block with the whole in-memory state as input and committed the pre-root, post-root, receipts root, gas and pgas (ProveOutput, 200 bytes). Mac CPU, loaded: block-78-increment 626,246 cycles, core 22.0 s, compressed 55.7 s; RTX 5090 through WSL2 (4 October morning): core 1.4 s, compressed 2.7 s, with the two P20 defects. The v0 fixtures are re-cut in the v1 format (one shard each); the v0 host modes core and compressed are the shard mode now.