diff --git a/docs/bench-log.md b/docs/bench-log.md index ab0465f6b..d56780f01 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -646,3 +646,47 @@ next program in the background while the current one mined. Reading: the compile-ahead rule (3 October 2026 decision: never restart every miner at once) holds on real values with three vendors; the hash rate is unbroken through the boundary and the launcher's exit-42 rebuild path was not used (`rebuilds 0`). The next boundary is DAA 7,200, which is also the first finality lock (weight window 7,200). + +## 4 October 2026, proving: devnet v4 shards on the Apple M5 Max CPU, loaded machine (execution-engineer, proving) + +Machine: Apple M5 Max (18 cores, 64 GB), macOS Darwin 25.6.0, load average 38 to 47 during the runs (the live devnet node and Metal miner, another agent's cross-builds, this agent's exports), everything under `nice -n 19`. Toolchain: SP1 v6.8.1 (cargo-prove c84ada1, succinct rustc 1.96.0-dev, circuit v6.1.0), sp1-sdk 6.8.1 CPU prover, revm 43.0.3, alloy-trie 0.9.8. Code: `proving/igneum-prove` at the "Proving fixtures: blocks of one, two and four shards" commit: two guests (shard program id `0x7b274fc9...`, aggregator id `0x36e952f0...`), port of igneum-exec b7fca5a0. Fixtures: `proving/fixtures/block-{338,341,344}` cut by `igneum-prove-export` from `tools/prove-fixtures/seq.json` (a private one-node simnet, execution-layer b7fca5a0 binaries whose exec crate is byte-identical on devnet-v4; the devnet-v4 worktree had another agent's uncommitted edits and was not built), which replayed all 346 segments from genesis and matched every node state root (final root `0xe0f269dc...`); `block-56-transfers-3shards` is the v0 block 56 cut at a test budget of 200 pgas. `S_p` = 7,500,000 pgas (provisional, `B_p / 4`). + +Statement per shard: from the carried-in position and a witness of the touched accounts, slots and trie nodes (checked against the pre-root), execute the shard's transactions and commit the post-root, the shard's receipts root, gas, pgas, the carry links and the prover's payout address; the aggregator verifies the shard proofs in order and commits the block. The host checks the native cut (shards chain and sum to the block) and rejects three tampered witnesses before any proof, on every fixture. + +| Fixture | Txs | EVM gas | pgas | Shards (pgas each) | Witness per shard: accounts / slots / leaves / hashes | Input bytes per shard | +|---|---|---|---|---|---|---| +| block-338-shard1 (3 modexp calls of 652 iterations, 6 transfers, 2 Counter increments) | 11 | 1,390,773 | 6,751,568 | 1 (6.75 M) | 19 / 3 / 24 / 13 | 21,447 | +| block-341-shards2 | 14 | 2,564,838 | 13,499,360 | 2 (6.75 M, 6.75 M) | 11 / 1 / 12 / 17; 16 / 3 / 22 / 12 | 18,372; 20,335 | +| block-344-shards4 (near `B_p`) | 20 | 4,947,168 | 26,994,944 | 4 (6.75 M each) | 11 / 1 / 12 / 17; 7 / 1 / 8 / 18; 7 / 1 / 7 / 21; 16 / 3 / 23 / 10 | 18,371; 17,125; 17,150; 20,362 | +| block-56-transfers-3shards (test cut at 200 pgas) | 3 | 63,000 | 600 | 3 (200 each) | 4 / 0 / 5 / 0; 2 / 0 / 1 / 4; 2 / 0 / 1 / 5 | 4,298; 3,671; 3,720 | + +SP1 executor (`--mode execute`, no proof): + +| Shard | Cycles | Prover gas | Cycles per EVM gas | Cycles per pgas | Execute s | +|---|---|---|---|---|---| +| block-344 shard 0 (6 txs) | 60,015,755 | 49,669,617 | 48 | 9 | 6.9 | +| block-344 shard 1 (3 modexp calls) | 59,619,778 | 49,236,458 | 50 | 9 | 7.7 | +| block-344 shard 2 (3 modexp calls) | 59,628,334 | 49,243,900 | 50 | 9 | 4.6 | +| block-344 shard 3 (8 txs) | 60,417,382 | 50,172,279 | 46 | 9 | 6.5 | +| block-344 aggregator over 4 shards (deferred verification off) | 1,664,255 | | | | 0.1 | +| block-56 test shards 0, 1, 2 (one transfer each) | 315,235; 271,190; 274,954 | | 13 to 15 | 1,356 to 1,576 | 0.1 | +| block-56 aggregator over 3 shards | 1,544,681 | | | | 0.1 | + +CPU proofs (`SP1_PROVER=cpu`), block-56-transfers-3shards: + +| Stage | Prove s | Proof bytes | Verify s | Verified | +|---|---|---|---|---| +| setup (prover client plus two key setups) | 39.4 to 60.6 (keys 3.3 + 3.2 of it; the rest is the client) | | | | +| shard 0: core | 83.1 | 7,310,257 | 0.368 | yes | +| shard 0: compressed | 272.3 | 1,272,897 | 0.075 | yes | +| block mode, shard 0: compressed | 336.9 | 1,272,897 | 0.068 | yes | +| block mode, shard 1: compressed | 305.4 | 1,272,897 | 0.066 | yes | +| block mode, shard 2: compressed | 245.3 | 1,272,897 | 0.064 | yes | +| block mode, aggregation over the 3 shard proofs (recursion, deferred proofs) | 244.5 | 1,272,909 | 0.084 | yes, shard program id and claim checked | +| block mode, end to end (first shard proof to the verified block proof, 10:07:22 to 10:26:21 UTC) | 1,139 | | | | + +| RTX 5090 (PROVE-SHARD.bat) | shard at `S_p`: execute, core, compressed | two-shard block end to end | four-shard block end to end | +|---|---|---|---| +| pending | | | | + +Reading: the shard at `S_p` is 60 M cycles on the prototype table, 9 cycles per pgas against the unit's 1,000 (the modexp entry is about 100x its SP1 cost: R1, one number); a plain transfer shard is 1,400 to 1,600 cycles per pgas because the 200-pgas intrinsic charge carries the fixed cost of the witness check and the two root computations. The aggregator statement is 1.5 to 1.7 M cycles (bincode, an unpatched sha256 of each shard's public values, the keccaks), small next to a shard. The CPU proof times are 3.8x and 4.9x the 3 October v0 numbers on a comparable statement, on a machine three times as loaded; the GPU row stays empty until the PC runs. The devnet was not touched; the simnet ran on ports 29300, 29301 and 29390 and was stopped. diff --git a/docs/plans/proving-v0.md b/docs/plans/proving-v0.md index 34b3b5d85..519e72557 100644 --- a/docs/plans/proving-v0.md +++ b/docs/plans/proving-v0.md @@ -1,41 +1,50 @@ -# Proving v0 and the devnet v4 shard plan +# Proving v0 and the devnet v4 shards -3 October 2026. Code in `proving/` (SP1 v6.8.1 guest and host, the versioned `ProofSystem` trait, two real-block fixtures, the Windows WSL2 package). Status words follow `docs/spec/00-overview.md` 0.2. Nothing here is a mainnet number. +3 October 2026 (v0: one proof per block) and 4 October 2026 (devnet v4: shards implemented, measured on the Mac CPU; the GPU run is pending). Code in `proving/` (SP1 v6.8.1, two guests and a host, the versioned `ProofSystem` trait, six fixtures, the Windows WSL2 package). Status words follow `docs/spec/00-overview.md` 0.2. Nothing here is a mainnet number. -## What tonight's single-block proof is +## Shards implemented, measured: what -The guest (`proving/igneum-prove/program`) re-executes one Igneum chain block with the execution layer's own rules: the transaction decoder is the node's crate (`igneum-evm-types`), the executor is a port of `igneum-exec` at commit fb33069 (rewards by rule, the nonce-rule skip, two-dimensional gas with the prototype pgas table, fee flows with the 80/20 developer split, the registry population on CREATE, the state root over every non-empty account through alloy-trie). It commits the chain id, the block number and hash, a commitment to the ordered transactions, the pre-state root, the post-state root, the receipts root, gas and pgas used, and the executed and skipped counts (`igneum_prove_core::ProveOutput`, 200 bytes, fixed layout). +Implemented 4 October 2026 (`proving/igneum-prove`, commits "Proving: shard cutter..." and after), checked on this Mac's CPU (`docs/bench-log.md`, "proving: devnet v4 shards"): -The fixtures are blocks that happened: `proving/fixtures/block-78-increment.json` (the `increment(5)` call on the Counter contract, one executed transaction and one duplicate skipped by the nonce rule, 10 accounts) and `block-56-transfers.json` (the three funding transfers, 5 accounts), cut by `igneum-prove-export` from `tools/evm-smoke/seq.json`, the `igneum_exportSegments` dump of the 3-node simnet run. The exporter replays all 79 segments from genesis through the port and refuses to write unless every segment's state root equals the node's; it did (Measured, 3 October 2026, final root `0x5b18b3a5...`). So the proof statement is "the node's executor, run twice", which is what design 2.1 asks for, up to the duplicate of the executor that the port is until the executor becomes the no-network library of design 2.1. +| Step | Status | Where | +|---|---|---| +| Cutter | Implemented, Measured. The executor runs any contiguous range of a segment from a carried-in position and records a boundary per transaction (cumulative gas and pgas, the carry link, natively the state root); the planner cuts at transaction boundaries to at most `S_p` pgas, a pure function of the trace. The host checks that the shards chain (roots, links) and sum (gas, pgas) to the whole block on every fixture | `core/src/executor.rs`, `core/src/plan.rs` | +| Witnesses | Implemented, Measured. Per shard: the touched accounts, slots and code, and partial Merkle Patricia tries (the touched leaves in full, every untouched subtree as a hash, collapse-safe siblings carried) for the account trie and each touched storage trie. The guest rebuilds the pre-root from them, refuses any read they do not cover (a dropped account makes the proof impossible), and rebuilds the post-root after execution. 60 randomised rounds of inserts, updates and deletes against the full trie in the unit test; three tamper variants rejected on every fixture (a re-encoded balance moves the pre-root away from the node's; a flipped storage value or code byte breaks the storage-root or code-hash check; a dropped account is refused) | `core/src/trie.rs`, `core/src/witness.rs` | +| Shard statement | Implemented, Measured. Public values: chain id, block number and hash, shard index, transaction range, carry links in and out, transaction accumulators in and out, pre-root, post-root, the shard's receipts root, gas, pgas, executed and skipped counts, and the prover's payout address (ledger P12). 328 bytes, fixed layout | `core/src/shard.rs`, `program/` | +| Aggregator | Implemented; Measured in SP1's executor and (CPU) end to end on the three-shard test fixture. The aggregator guest verifies every shard proof against the shard program's key (SP1 deferred proofs), checks the chain (indices contiguous, `link_in` = previous `link_out`, pre-root = previous post-root, accumulators), sums gas and pgas, and commits the block statement with keccak over the shard receipts roots and over the provers, plus the shard program id. The chain rule (segment N verifies N-1, design 5.3) is in the guest and the proof system (`AggInput.prev`, `BlockOutput.agg_vk`, `chain_len`) but has no host mode and no measurement yet: next step | `core/src/agg.rs`, `aggregator/`, `host/src/proof_system.rs` | +| Fixtures of real size | Implemented. Blocks 338, 341 and 344 of a private one-node simnet (4 October 2026): 0.90, 1.80 and 3.60 `S_p` of proving gas, one, two and four shards; plus the v0 blocks re-cut (one shard each) and block 56 cut at a test budget of 200 pgas into three shards for the Mac CPU check | `proving/fixtures/`, `tools/prove-fixtures/` | +| Host modes | Implemented. `native`, `execute`, `shard` (execute, core, compressed, verified), `block` (compressed proof per shard, aggregation, verified against the shard program id and the claim), `all`; `--shard`, `--prover`, `--out`. A `STAGE` line and a `RESULT` line with a UTC timestamp per stage; a Tokio runtime held for the whole run with the proof system dropped inside it (ledger P20) | `host/src/main.rs` | +| WSL2 package | Implemented, not yet run. `PROVE-SHARD.bat`: the shard at `S_p` in all three stages, then the two- and four-shard blocks end to end, on the GPU; `PROVE-BLOCK.bat` kept for the small block. Zip at `~/Desktop/igneum-prove-wsl2.zip` | `proving/windows-wsl2/` | -The host runs the statement natively first and refuses to prove a fixture whose expected roots it does not reproduce, then runs SP1 in three modes behind the trait (`proof_system.rs`): execute (cycle count, no proof), core (`prove_shard`), compressed (`aggregate` over one shard). Every proof is verified and its public values are compared with the native run. `SP1_PROVER=cuda` selects the GPU prover through the same binary (feature `cuda`). +`S_p`. The specification had no number. Provisional since 4 October 2026: `S_p = B_p / 4 = 7,500,000` pgas, so a block at its proving budget is exactly four shards (spec 7.4, 7.6). It is set from the shard-time measurement on the card, never from the fixtures (benchmark standard 2.3). -## What it does not show +## Measured on the Mac CPU (4 October 2026, loaded machine, `docs/bench-log.md`) + +| What | Number | +|---|---| +| One shard at `S_p` (block 344, 6.75 M pgas, three modexp calls plus transfers): SP1 cycles | 59.6 M to 60.4 M per shard, 9 cycles per pgas, 46 to 50 per EVM gas | +| Shard input (witness and transactions) | 17 to 21 KB per shard; 7 to 19 accounts, 1 to 3 slots, 7 to 24 trie leaves, 10 to 21 subtree hashes | +| Aggregator statement over four shards (no proofs, executor only) | 1.66 M cycles | +| A 200-pgas shard (block 56 test cut, one transfer): execute, core, compressed | 315 k cycles; core 83.1 s, 7.31 MB, verify 0.37 s; compressed 272.3 s, 1.27 MB, verify 0.08 s; all verified | +| Three 200-pgas shards plus aggregation on the CPU (`--mode block`) | compressed shard proofs 336.9, 305.4 and 245.3 s (1.27 MB each, verify 0.06 to 0.07 s); aggregation by recursion over the three 244.5 s, block proof 1,272,909 bytes, verify 0.084 s, VERIFIED with the shard program id and the claim checked; 19 minutes end to end from the first shard proof to the verified block proof | + +Reading. The modexp-heavy shard costs 9 SP1 cycles per pgas against the unit's 1,000: the prototype table's modexp entry (1,000 + 10 per input byte) is two orders of magnitude above its SP1 cost, which is the R1 calibration in one number; at the current table a shard of `S_p` is about 60 M cycles, far below what the unit would imply. The witness is small because the devnet state is small; the share of the shard's cycles spent on the trie is not isolated yet (R3). The CPU times are on a machine at load 40 shared with the live devnet and other agents' builds; they are correctness runs, not throughput. + +## What waits for the GPU + +`PROVE-SHARD.bat` on the RTX 5090: the shard at `S_p` in the three stages and the two- and four-shard blocks end to end. The RESULT lines fill the GPU row of the bench-log entry and give the first point for `S_p`. The P20 Drop fix and the gap before the first compressed stage are confirmed or not by the same run. + +## What it still does not show | Not shown | Why | Where it goes | |---|---|---| -| Shard time at `S_p` pgas on a 12 GB card (the phase 2 gate, ledger P1, overclaim 27) | the fixture blocks carry 600 and 1,488 pgas, not a shard's budget | devnet v4 shard plan below, 3060-class card | -| MPT witnesses | the whole in-memory devnet state (5 to 10 accounts) is the input; a real shard gets the touched accounts, slots and trie nodes (design 5.1), and the trie-proof share of pgas (R3) is unmeasured | witness generator in the node, v4 | -| Recursion over shards and over the chain (segment N verifies N-1) | v0 aggregates one shard: SP1's compress stage on the block | v4 aggregator | +| Shard time on a 12 GB card (the phase 2 gate, ledger P1, overclaim 27) | the 5090 run is pending; no 3060-class card has run it | PROVE-SHARD.bat, then the 3060-class card | +| The chain rule measured (segment N verifies N-1) | in the guest and the proof system, no host mode for two consecutive fixtures yet | next step: `--mode chain` over consecutive blocks | | The Groth16 or Plonk wrapper for light clients (ledger P3, overclaim 25) | not run; needs SP1's circuit artifacts and a measurement on consumer hardware | R4, phase 2 benchmark | -| The prover's payout key in the statement (ledger P12) | `ShardWitness.prover` is carried, not committed | guest change, v4 | -| pgas calibration (R1) | the table is the prototype; tonight's cycles per EVM gas are one data point | calibrate per opcode in SP1 with three input sizes | +| Sortition, proof records, the native-execution veto in the node | the node has no proof records on devnet v4 | design 5.4, 5.5; the stub `ProofSystem` path exists in the host | +| pgas calibration (R1) | the table is the prototype; 9 cycles per pgas on modexp, 1,400 to 1,600 on a plain transfer shard | calibrate per opcode in SP1 with three input sizes | +| Trie share of pgas (R3) | not isolated | instrument the guest | -## The shard plan for devnet v4 (Designed) +## The v0 statement (3 October 2026), for the record -1. **Cut.** The native executor emits per segment the transaction boundaries with cumulative gas and pgas and the state root at each boundary (design 5.1). The planner cuts at boundaries so each shard's pgas is at most `S_p`; one transaction above `S_p` is one shard proven with continuations (SP1 checkpoints its own execution; approximate). The plan is a pure function of the trace, so every node computes the same shard list and the same ids `(segment hash, shard index)`. -2. **Statement per shard.** From pre-root `r_i` and the committed transaction list of shard i, running the executor yields `r_{i+1}` and receipts root `h_i`, with the executed set and the prover's payout key as public outputs (P12). The witness (touched accounts, slots, MPT nodes, block environment) is produced by any full node from its own execution and served over the proving gossip; it is not consensus data. Tonight's guest is this statement with the whole state as the witness and one shard per block. -3. **Claim and sortition (spec 7.2).** Eligible provers are the vote keys above dust (100 blue blocks in the 30-day window). For each shard, draws `r_n = H("igneum-shard/" || epoch_seed || shard_id || n) mod N` over the window's blue blocks pick 8 distinct keys by weight. They hold the shard for 10 DAA seconds from the moment the segment is executed: a valid proof by one of them, included in a block in the window, earns the shard's part of the pool; after the window anyone's first valid proof earns it. No claim and no bond (ledger P8, P9). Parameters 8 and 10 s are set on the devnet from the shard-time distribution across three prover speeds (R7, fastest prover under 25% of shards; one key versus 1,000 keys of equal weight win the same number, F17). -4. **Aggregation by recursion (design 5.3).** Shard proofs of a segment fold into one segment proof (SP1 compress, recursion over the shard proofs, not over the witness as v0 does); the segment proof for N also verifies the segment proof for N-1 so one proof attests the chain of state and a client keeps only the latest. The aggregator is anyone, paid the aggregator share. The wrapper to bn254 for the bridge verifier and phones is measured, not assumed (R4). -5. **Proof records and the veto (design 5.4, 5.5, spec 7.2 item 5).** A block carries zero or more `ProofRecord {version, segment, pre_root, post_root, receipts, provers, aggregator, proof}`; full nodes verify the proof natively (a compressed proof, tonight's verify time is the first number for that) and run the native-execution veto: the block is invalid if the record names a chain block off the carrying block's own selected-parent chain or if `post_root` or `receipts` differ from the native execution of that segment along that chain. The test is relative to the carrying block's past, never to the validating node's current chain (ledger P11). A forged proof from a soundness bug is therefore rejected by every full node; the damage is bounded to light clients. -6. **Devnet v4 run.** Nodes emit shard plans; miners run the host as a prover service that takes gossiped witnesses and returns shard proofs; a block producer includes records. Acceptance for the execution layer's A5 and A6 (`docs/design/execution-layer.md` 9.2): records within 60 s of 95% of segments, a wrong `post_root` makes the block invalid on every other node. - -## The Mac CPU baseline (Measured, 3 October 2026, `docs/bench-log.md`) - -block-78-increment on the Apple M5 Max CPU under load, SP1 v6.8.1: 626,246 cycles (14 per EVM gas), core proof 22.0 s and 7.3 MB (verify 0.16 s), compressed proof 55.7 s and 1.27 MB (verify 0.03 s), both verified and both reproducing the node's post-state root `0x5b18b3a5...`. block-56-transfers: 549,469 cycles. These are fixed-overhead numbers for a block far below one SP1 shard, not throughput. - -## The morning run - -On the Windows PC through WSL2 (`proving/windows-wsl2/`): SETUP-PROVER.bat (reboot once), SETUP-PROVER.bat again, pause mining, PROVE-BLOCK.bat. It proves `block-78-increment` on the RTX 5090 (execute, core, compressed, each verified) and then a core proof on the CPU, prints the RESULT lines and uploads the log to the intake. Acceptance line: **one Igneum devnet block proven on the RTX 5090 and verified, time and size recorded** (the RESULT lines go to `docs/bench-log.md` next to the Mac CPU baseline). - -Two uncertainties the run settles. First, whether SP1's GPU prover runs under WSL2 at all with the 5090 (Blackwell, compute capability 12.0; the documentation lists 8.0 or higher and 24 GB of VRAM, so it should; the server binary is prebuilt for Linux x86_64 and the SDK downloads it, and the WSL CUDA driver path is the usual failure point). Second, how long the compressed proof takes on one consumer card for a 1,488-pgas block: this is the first point on the curve that sets `S_p`, and no shard has been proven on any card before this (ledger P1). +The v0 guest re-executed one whole block with the whole in-memory state as input and committed the pre-root, post-root, receipts root, gas and pgas (`ProveOutput`, 200 bytes). Mac CPU, loaded: block-78-increment 626,246 cycles, core 22.0 s, compressed 55.7 s; RTX 5090 through WSL2 (4 October morning): core 1.4 s, compressed 2.7 s, with the two P20 defects. The v0 fixtures are re-cut in the v1 format (one shard each); the v0 host modes `core` and `compressed` are the `shard` mode now.