igneum/docs/plans/proving-v0.md
igneum-labs 64009a676a Proving: pinned guest programs, the verifier on SP1's light verifier
On 5 October 2026 the Mac's host (shard program id 0x0559759b...) rejected every
proof from PC 2's host (0x05db1aca...). Both were built from the same guest
sources: host/build.rs compiled the guests on each machine and the ELF depends
on where it is built (cargo's -C metadata for a path crate includes the checkout
path; a worktree on the same Mac gave a third id, 0x0dfade07...). The node's
verifier also spent 114 s to 138 s per proof in the prover client and both key
setups before a 0.1 s to 0.4 s verify.

- elf/: both guest ELFs, their verifying keys and manifest.json (sha256, ids);
  host/src/pinned.rs embeds and checks them at every start; the prove modes
  refuse when SP1's setup does not derive the manifest's id
- --mode verify: LightProver with the pinned key, no prover client, no key
  setup; prints the proof's own program id next to ours ("IS NOT OURS")
- --mode id; igneum-prove-pin and pin-guests.sh to re-pin; build.rs builds a
  guest only under IGNEUM_BUILD_GUESTS=1
- tools/ci/pinned-guests-check.sh: elf/ must match its manifest, no script
  builds a guest outside pin-guests.sh; make-package.sh and build-dmg.sh print
  the pinned ids
- unit tests on the pinned set; bench-log entry with the three ids, the cause
  and the timing: 127.0 s wall per verify before, 1.8 s to 2.4 s after
- rollout order in proving/README.md: every prover and verifier moves together

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 12:54:31 +00:00

72 lines
13 KiB
Markdown

# Proving v0 and the devnet v4 shards
3 October 2026 (v0: one proof per block) and 4 October 2026 (devnet v4: shards implemented, measured on the Mac CPU; the GPU run is pending). Code in `proving/` (SP1 v6.8.1, two guests and a host, the versioned `ProofSystem` trait, six fixtures, the Windows WSL2 package). Status words follow `docs/spec/00-overview.md` 0.2. Nothing here is a mainnet number.
## Shards implemented, measured: what
Implemented 4 October 2026 (`proving/igneum-prove`, commits "Proving: shard cutter..." and after), checked on this Mac's CPU (`docs/bench-log.md`, "proving: devnet v4 shards"):
| Step | Status | Where |
|---|---|---|
| Cutter | Implemented, Measured. The executor runs any contiguous range of a segment from a carried-in position and records a boundary per transaction (cumulative gas and pgas, the carry link, natively the state root); the planner cuts at transaction boundaries to at most `S_p` pgas, a pure function of the trace. The host checks that the shards chain (roots, links) and sum (gas, pgas) to the whole block on every fixture | `core/src/executor.rs`, `core/src/plan.rs` |
| Witnesses | Implemented, Measured. Per shard: the touched accounts, slots and code, and partial Merkle Patricia tries (the touched leaves in full, every untouched subtree as a hash, collapse-safe siblings carried) for the account trie and each touched storage trie. The guest rebuilds the pre-root from them, refuses any read they do not cover (a dropped account makes the proof impossible), and rebuilds the post-root after execution. 60 randomised rounds of inserts, updates and deletes against the full trie in the unit test; three tamper variants rejected on every fixture (a re-encoded balance moves the pre-root away from the node's; a flipped storage value or code byte breaks the storage-root or code-hash check; a dropped account is refused) | `core/src/trie.rs`, `core/src/witness.rs` |
| Shard statement | Implemented, Measured. Public values: chain id, block number and hash, shard index, transaction range, carry links in and out, transaction accumulators in and out, pre-root, post-root, the shard's receipts root, gas, pgas, executed and skipped counts, and the prover's payout address (ledger P12). 328 bytes, fixed layout | `core/src/shard.rs`, `program/` |
| Aggregator | Implemented; Measured in SP1's executor and (CPU) end to end on the three-shard test fixture. The aggregator guest verifies every shard proof against the shard program's key (SP1 deferred proofs), checks the chain (indices contiguous, `link_in` = previous `link_out`, pre-root = previous post-root, accumulators), sums gas and pgas, and commits the block statement with keccak over the shard receipts roots and over the provers, plus the shard program id. The chain rule (segment N verifies N-1, design 5.3) is in the guest and the proof system (`AggInput.prev`, `BlockOutput.agg_vk`, `chain_len`) but has no host mode and no measurement yet: next step | `core/src/agg.rs`, `aggregator/`, `host/src/proof_system.rs` |
| Fixtures of real size | Implemented. Blocks 338, 341 and 344 of a private one-node simnet (4 October 2026): 0.90, 1.80 and 3.60 `S_p` of proving gas, one, two and four shards; plus the v0 blocks re-cut (one shard each) and block 56 cut at a test budget of 200 pgas into three shards for the Mac CPU check | `proving/fixtures/`, `tools/prove-fixtures/` |
| Host modes | Implemented. `native`, `execute`, `shard` (execute, core, compressed, verified), `block` (compressed proof per shard, aggregation, verified against the shard program id and the claim), `all`; `--shard`, `--prover`, `--out`. A `STAGE` line and a `RESULT` line with a UTC timestamp per stage; a Tokio runtime held for the whole run with the proof system dropped inside it (ledger P20) | `host/src/main.rs` |
| WSL2 package | Implemented, not yet run. `PROVE-SHARD.bat`: the shard at `S_p` in all three stages, then the two- and four-shard blocks end to end, on the GPU; `PROVE-BLOCK.bat` kept for the small block. Zip at `~/Desktop/igneum-prove-wsl2.zip` | `proving/windows-wsl2/` |
`S_p`. The specification had no number. Provisional since 4 October 2026: `S_p = B_p / 4 = 7,500,000` pgas, so a block at its proving budget is exactly four shards (spec 7.4, 7.6). It is set from the shard-time measurement on the card, never from the fixtures (benchmark standard 2.3).
## Measured on the Mac CPU (4 October 2026, loaded machine, `docs/bench-log.md`)
| What | Number |
|---|---|
| One shard at `S_p` (block 344, 6.75 M pgas, three modexp calls plus transfers): SP1 cycles | 59.6 M to 60.4 M per shard, 9 cycles per pgas, 46 to 50 per EVM gas |
| Shard input (witness and transactions) | 17 to 21 KB per shard; 7 to 19 accounts, 1 to 3 slots, 7 to 24 trie leaves, 10 to 21 subtree hashes |
| Aggregator statement over four shards (no proofs, executor only) | 1.66 M cycles |
| A 200-pgas shard (block 56 test cut, one transfer): execute, core, compressed | 315 k cycles; core 83.1 s, 7.31 MB, verify 0.37 s; compressed 272.3 s, 1.27 MB, verify 0.08 s; all verified |
| Three 200-pgas shards plus aggregation on the CPU (`--mode block`) | compressed shard proofs 336.9, 305.4 and 245.3 s (1.27 MB each, verify 0.06 to 0.07 s); aggregation by recursion over the three 244.5 s, block proof 1,272,909 bytes, verify 0.084 s, VERIFIED with the shard program id and the claim checked; 19 minutes end to end from the first shard proof to the verified block proof |
Reading. The modexp-heavy shard costs 9 SP1 cycles per pgas against the unit's 1,000: the prototype table's modexp entry (1,000 + 10 per input byte) is two orders of magnitude above its SP1 cost, which is the R1 calibration in one number; at the current table a shard of `S_p` is about 60 M cycles, far below what the unit would imply. The witness is small because the devnet state is small; the share of the shard's cycles spent on the trie is not isolated yet (R3). The CPU times are on a machine at load 40 shared with the live devnet and other agents' builds; they are correctness runs, not throughput.
## Proving layer v0 in the node and the app (4 October 2026, afternoon)
Implemented on the fork branch `proving` (`vendor/igneum-node-proving`, from devnet-v4 3bfe346f), in `proving/igneum-prove` and in `app/igneum-app/src/prover.rs`; the rules are spec 7.7. Status words as above.
| Piece | Status | Where |
|---|---|---|
| Proof record (per shard, BLS-signed by the vote key), coinbase record section `IGNP` before the finality section, p2p message 71 at protocol 13, shard id, sortition draw, payout arithmetic | Implemented, unit-tested (record round trip and signature, nested sections, sortition determinism and weighting, dust, payouts) | `consensus/core/src/proving.rs`, `protocol/flows/src/v10/proving.rs`, `v13/` |
| `proving_v0_activation_daa` height switch (default never on every network; the override file carries it) | Implemented, unit-tested | `consensus/core/src/config/params.rs`, `infra/fast-time/override-60x.json` |
| The node's shard plan: per-transaction boundaries with the carry of 7.6 and the state root, the cut of `igneum_prove_core::plan`, the sortition window from the finality parameters, assignees per shard | Implemented, unit-tested against the core planner's test vector; the exporter's plan checked equal to the node's on the test network | `igneum/exec/src/executor.rs` (boundaries), `igneum/exec/src/proving.rs` |
| Native shard statement (the 328-byte public values recomputed for any payout address), the record checks (chain membership, record window, plan, signature, assignment inside the exclusive window, native-execution veto) | Implemented, unit-tested | `igneum/exec/src/proving.rs` |
| Proof pool with the external verifier (`igneum-prove-host --mode verify`), trust mode for test networks, template section of verified records, payout at the carrying segment, reorg unwinding | Implemented; pool unit-tested; verifier and payout exercised on the test network | `igneum/exec/src/proving.rs`, `service.rs` |
| RPCs `igneum_getShardPlan`, `igneum_getProofRecords`, `igneum_submitProofRecord`, `igneum_getAssignedShards`, `igneum_getProvingStatus`; `igneum_exportSegments` carries payouts | Implemented | `igneum/exec/src/rpc.rs` |
| Payouts in the shard statement (fixture, `ShardInput`, `execute_range`, the guest), the empty-segment plan fix, host modes `compressed` (execute + compressed, statement and proof hash in the results) and `verify` | Implemented; guest rebuilt (new shard vk) | `proving/igneum-prove` |
| Pinned guest programs (5 October 2026): the two guest ELFs and their verifying keys committed under `elf/` with a manifest of hashes and program ids, embedded by the host and checked at every start; `--mode verify` on SP1's light verifier with the pinned key (no prover client, no key setup); `--mode id`; `igneum-prove-pin` and `pin-guests.sh` to re-pin; CI check `tools/ci/pinned-guests-check.sh` | Implemented after the Mac (shard program id `0x0559759b...`) rejected every proof of PC 2 (`0x05db1aca...`): the same guest sources built on two machines gave two ELFs. Not yet rolled out: every prover and verifier moves together (proving/README.md, "Pinned guest programs") | `proving/igneum-prove/elf/`, `host/src/pinned.rs`, `host/src/bin/pin.rs` |
| `igneum-miner vmine` (voting producer for PoW-less test networks, with an EVM payout address), `sign-record`, `key-hash` | Implemented | `igneum/miner/src/proving.rs` |
| The app's prover service: `prove` setting (default off), the loop (work list, export, cut, prove, sign, submit, open-shard fallback), the Proving tile (assigned, proving, submitted, paid), WSL2 detection and Set up on Windows, CPU on macOS | Implemented; run end to end inside the engine on the Mac with the packaged binaries against the 3-node test network (bench-log, 4 October 2026 afternoon, "the app's prover loop"): proving, submitted, paid in 71 s for the smallest shard; the Windows/WSL2 path has not run | `app/igneum-app/src/prover.rs` |
| Packaging: the Mac DMG carries the host and the exporter; the Windows payload carries the Linux x86_64 host and exporter (cargo-zigbuild, glibc 2.36, sp1 `cuda` feature, 11 min 18 s on the Mac) under `wsl2/bin/` with the WSL2 scripts and the fixtures | Implemented; dry-run staged; no payload cut (0.3.3 is the coordinator's) | `packaging/mac/build-dmg.sh`, `packaging/windows/make-payload.sh` |
| 3-node test network on 29800+ at 60x with activation 60 | Implemented, run; see the bench-log entry of 4 October 2026 (afternoon) for what it showed | `tools/proving-v0/run.mjs` |
Decisions the implementation forced, all in spec 7.7: per-shard records instead of the aggregated record of design 5.4 (the aggregated record is the next step); the SP1 proof verified off the consensus path, the native statement being the consensus check; payouts as data in the shard statement beside the rewards; an empty segment is one shard ending at the root after rewards and payouts.
## What waits for the GPU
`PROVE-SHARD.bat` on the RTX 5090: the shard at `S_p` in the three stages and the two- and four-shard blocks end to end. The RESULT lines fill the GPU row of the bench-log entry and give the first point for `S_p`. The P20 Drop fix and the gap before the first compressed stage are confirmed or not by the same run.
## What it still does not show
| Not shown | Why | Where it goes |
|---|---|---|
| Shard time on a 12 GB card (the phase 2 gate, ledger P1, overclaim 27) | the 5090 run is pending; no 3060-class card has run it | PROVE-SHARD.bat, then the 3060-class card |
| The chain rule measured (segment N verifies N-1) | in the guest and the proof system, no host mode for two consecutive fixtures yet | next step: `--mode chain` over consecutive blocks |
| The Groth16 or Plonk wrapper for light clients (ledger P3, overclaim 25) | not run; needs SP1's circuit artifacts and a measurement on consumer hardware | R4, phase 2 benchmark |
| The aggregated segment record and the chain rule in consensus | proving v0 carries per-shard records and verifies the SP1 proof off the consensus path (spec 7.7 items 1 and 4) | the aggregator's record on top of the shard records; verification of the aggregated proof in consensus |
| Proving on the GPU through the app | the app's prover loop is written and unit-tested; the WSL2 path has never run | PC 2 after the Windows exe ships (OTA) |
| pgas calibration (R1) | the table is the prototype; 9 cycles per pgas on modexp, 1,400 to 1,600 on a plain transfer shard | calibrate per opcode in SP1 with three input sizes |
| Trie share of pgas (R3) | not isolated | instrument the guest |
## The v0 statement (3 October 2026), for the record
The v0 guest re-executed one whole block with the whole in-memory state as input and committed the pre-root, post-root, receipts root, gas and pgas (`ProveOutput`, 200 bytes). Mac CPU, loaded: block-78-increment 626,246 cycles, core 22.0 s, compressed 55.7 s; RTX 5090 through WSL2 (4 October morning): core 1.4 s, compressed 2.7 s, with the two P20 defects. The v0 fixtures are re-cut in the v1 format (one shard each); the v0 host modes `core` and `compressed` are the `shard` mode now.