Proving v1: the chain of 8 on the 5090 measured (N = 2, 4, 8: 32.6, 66.8, 135.6 s; chained aggregation 9.7 s a block on a mining card), the fleet table re-cut on the measured rows

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-05 20:10:50 +00:00
parent e9508926dc
commit d4848453ae
3 changed files with 170 additions and 10 deletions

View file

@ -1612,6 +1612,10 @@ Branches `proving-v1` (main repository, worktree `igneum-wt-proving-v1`; fork `v
| Chain of 2 on the Mac CPU (the known-finished case of `--mode chain` before the GPU; M5 Max under the live nodes, the harness and two builds) | `SP1_PROVER=cpu igneum-prove-host --mode chain --chain block-81046.json,block-81047.json --out results.json` under the run lock, 19:07:48Z to 19:11:28Z: setup 12.2 s; block 81046: shard 0 compressed 55.4 s (1,272,897 bytes, verify 0.036 s), aggregate 52.0 s (1,272,909 bytes, verify 0.031 s), chain_len 1, agg_vk zero; block 81047: shard 41.3 s, aggregate WITH the previous block proof 59.1 s, chain_len 2, agg_vk = the pinned aggregator id; end to end 207.9 s; final proof 1,272,909 bytes, statement 0x232276f4... The recursion over the previous proof cost 7 s more than the first aggregation on this CPU |
| `--mode verify-segment` on that proof (the node's path: SP1 light verifier, pinned aggregator key) | VERIFIED in 0.032 s (0.27 s wall, three runs: 0.033, 0.032, 0.032); known-failed: a wrong statement NOT VERIFIED (0.032 s); the shard verifier (`--mode verify`) on the segment proof NOT VERIFIED, "program id 0x474678f3... IS NOT OURS 0x2b1a81cb..." |
| Chain of 8 on the RTX 5090 (N = 2, 4, 8), job `chain-pc2-pv1b` (`tools/proving-v1/pc2-chain.ps1`; the package `igneum-prove-wsl2-pv1.zip` eb6dccf8..., 1.5 MB, fetched by `fetch-prove-pv1` 19:51Z; the first try `chain-pc2-pv1` died in its own export step, fixed) | Ran 19:58:37Z: the export from PC 2's node (72,901,414 bytes, 1.4 s), the host built in WSL2 against the live build's warm target dir in 6 s and installed to `/opt/igneum-pv1` (the live `/opt/igneum` host untouched, sha 29cc4768...), `--mode id` the pinned pair; eight consecutive fixtures 83346..83353 cut, every one MATCHES natively. The chain on the GPU (SP1_PROVER=cuda, the miner mining on the same card at 119 MH/s): setup 12.7 s; block 83346: shard 7.4 s, aggregate 7.6 s (chain_len 1), 15.1 s; block 83347: shard 7.2 s, aggregate WITH the previous proof 9.5 s (chain_len 2, agg_vk the pinned aggregator id), 16.8 s, cumulative 31.8 s over 2 blocks; block 83348: shard 7.0 s, then at 20:01:09Z the app quit for the 0.3.10 update and aborted the job ("aborted (the app is quitting)"). So N = 2 measured: 31.8 s of GPU time for two empty blocks, the chained aggregation 1.9 s dearer than the first; N = 4 and 8 are the re-run `chain-pc2-pv1c` after the restart. An empty shard's compressed proof on the 5090 is 7.0 to 7.4 s (the 200-pgas shard of 4 October took 2.7 s with the card to itself; tonight the miner held it at 92% utilisation) |
| The chain of 8, the third run `chain-pc2-pv1c` (20:05:21Z to 20:08:33Z, after the app restart; blocks 83616..83623 from PC 2's node at tip 83646, the same script; results `tools/proving-v1/chain-pc2-2026-10-05.json`) | Build 5 s (warm), eight fixtures cut and MATCHING natively, setup 15.7 s, then on the GPU with the miner mining on the same card: shard proofs 7.3 to 7.7 s each (8 x, 59.5 s), aggregations 7.9 s for the first block and 9.6 to 9.7 s for every chained one (75.5 s), every proof VERIFIED, end to end 135.6 s for 8 blocks (17.0 s a block from the second on). Cumulative: N = 2 at 32.6 s, N = 4 at 66.8 s, N = 8 at 135.6 s. The final proof is 1,272,909 bytes whatever N (chain_len 8, agg_vk the pinned aggregator id), the record 586 bytes; `--mode verify-segment` on it: VERIFIED in 0.039, 0.037, 0.040 s after a 0.26-s light-verifier setup, the same three runs each time. GPU memory over the chain (152 one-second samples): max 16,751 MiB with the miner's 3.4 GB resident, so the chained aggregation holds about 13.4 GB, 1.2 GB over the shard-only peak; WSL used 2,456 MB |
Reading the chain numbers. Aggregation is a fixed cost per block (9.7 s here), not per segment: the recursion verifies one more proof whatever `chain_len`, so the record for N blocks costs N aggregations and the verifier one. Against 4 October with the miner stopped (aggregate 2.2 to 2.5 s, a 200-pgas shard 2.7 s), tonight's 9.7 s and 7.3 s say the miner's 92% utilisation slows the prover about 3 to 4x while the prover slows the miner 4%: the card is shared, and the lottery wins the arbitration. A machine that mines and proves at once delivers one empty block's proof and aggregation in 17 s; one that only proves, about 5 s (approximate, from the 4 October stages).
### Step 3, coverage
@ -1628,13 +1632,15 @@ Inputs, all RTX 5090 (PC 2), SP1 6.8.1 cuda: a full shard at the provisional `S_
| Block content at 1 block/s | Shard proofs a second (fleet) | Card-seconds a second for shards | Aggregations a second | Card-seconds a second for aggregation | 5090-class cards for 100% | Rule |
|---|---|---|---|---|---|---|
| empty blocks (tonight's devnet), the app's loop as it is | 1 | 37 | 1 | 2.5 (approximate) | 40 | one shard per block, the loop's 37 s each, one card does 1.6 a minute |
| empty blocks, the loop with one key setup per process and the export cached (the chain mode's shape: 12 s setup once, then proofs back to back) | 1 | about 5 (approximate: an empty shard's compressed proof on the 5090 is not measured; the 200-pgas shard took 2.7 s on 4 October) | 1 | 2.5 | 8 (approximate) | the held PC 2 chain job gives the empty-shard proof time |
| one full shard a block (`S_p`, 6.75 M pgas) | 1 | 10.9 | 1 | 2.5 | 14 | 10.9 + 2.5 card-seconds per block-second |
| blocks at `B_p` (four full shards) | 4 | 42.5 | 1 | 2.5 | 45 | 4 x 10.6 + 2.5 |
| at the adopted v1 budgets (`B_p` 120,000 pgas, `S_p` 30,000, from DAA 210,000 on the devnet): a v1 shard of transfers ran at 213 to 236 cycles per pgas (bench-log 5 October, "the prover carries both fee tables"), 7 M cycles a shard against 60 M for the prototype shard | 4 | under 42.5 (the 5090 time for a 7 M-cycle shard is not measured; scaling 10.9 s by cycles gives about 1.3 s, approximate) | 1 | 2.5 | about 8 (approximate) | measure before the switch lands |
| empty blocks (tonight's devnet), the app's loop as it is, the card also mining | 1 | 37 | 1 | 9.7 (measured, `chain-pc2-pv1c`) | 47 | one shard per block, the loop's 37 s each plus a chained aggregation |
| empty blocks, the chain mode's shape (one key setup per process, proofs back to back), the card also mining | 1 | 7.4 (measured) | 1 | 9.7 (measured) | 18 | 17.1 card-seconds a block, `chain-pc2-pv1c` |
| empty blocks, cards that only prove | 1 | 2.7 (4 October, a 200-pgas shard) | 1 | 2.5 (4 October) | 6 (approximate) | the miner's 92% utilisation costs the prover 3 to 4x |
| one full shard a block (`S_p`, 6.75 M pgas), cards that only prove | 1 | 10.9 | 1 | 2.5 | 14 | 4 October's stages |
| one full shard a block, the card also mining | 1 | about 35 (approximate: 10.9 x 3.2, tonight's ratio) | 1 | 9.7 | about 45 (approximate) | the full-shard proof with the miner on the card is not measured |
| blocks at `B_p` (four full shards), cards that only prove | 4 | 42.5 | 1 | 2.5 | 45 | 4 x 10.6 + 2.5 |
| at the adopted v1 budgets (`B_p` 120,000 pgas, `S_p` 30,000, from DAA 210,000 on the devnet): a v1 shard of transfers ran at 213 to 236 cycles per pgas (bench-log 5 October, "the prover carries both fee tables"), 7 M cycles a shard against 60 M for the prototype shard | 4 | under 42.5 (the 5090 time for a 7 M-cycle shard is not measured; scaling 10.9 s by cycles gives about 1.3 s, approximate) | 1 | 2.5 to 9.7 | 8 to 15 (approximate) | measure before the switch lands |
Reading. The card count is the sum of card-seconds of work per block-second, rounded up, with no slack for the exclusive window, the relay or a card's idle gaps; the devnet's own numbers tonight (one card, 1.6 shards a minute, 4.7% of blocks) are the first row. The lever is the loop, not the proof: a shard's proof on the 5090 is 3 to 11 s and its carriage through export, cut and a 12-s key setup is 25 s more. The host's `--mode aggregate` and `--mode chain` already hold one key setup per process; the prover loop should do the same (one host process per segment, the 0.3.11 item in the plan).
Reading. The card count is the sum of card-seconds of work per block-second, rounded up, with no slack for the exclusive window, the relay or a card's idle gaps; the devnet's own numbers tonight (one card, 1.4 to 1.6 shards a minute, 2.4 to 4.7% of blocks) are the first row. Two levers, both measured tonight: the loop (a shard's carriage through export, cut and a 12-s key setup is 25 s on top of a 7-s proof; the host's `--mode aggregate` and `--mode chain` hold one key setup per process and the prover loop should do the same, the 0.3.11 item in the plan) and the card's other job (a mining card proves 3 to 4x slower than an idle one, `chain-pc2-pv1c` against 4 October; the prover's cost to mining is 4%). A fleet of 18 mining 5090s, or 6 proving-only ones, covers an empty-block chain at 1 block/s through the chain mode; the mandatory rule waits for the measured share to reach one, not for these rows.
### Step 4, the rule

View file

@ -40,8 +40,8 @@ that say how many cards cover the chain.
| Unit tests | consensus core 13, exec 8, app 5, all passing on the Mac |
| The harness (3 nodes, fast time, trust mode) | PASSED, 21 checks in 197 s: paid 1.0 s after submit, every node agreeing; the fresh chain refused after a proven segment; the unproven segment skipped after its deadline; shards at 90% |
| Coverage, 3-min window, the degraded fleet (one card, the Mac verifier down) | 4.7% of blocks proven, on-chain latency p50 39 s |
| 5090-class cards for 100% at 1 block/s | 40 with the loop as it is (empty blocks), 14 at one full shard a block, 45 at `B_p` (four full shards), about 8 at the adopted v1 budgets (approximate): the table in the bench log |
| The chain of 8 on the 5090 (N = 2, 4, 8), job `chain-pc2-pv1` (published 19:52Z after the go) | (the bench log row when it ends) |
| The chain of 8 live blocks on the 5090, the card also mining (`chain-pc2-pv1c`) | shard 7.3 to 7.7 s, first aggregation 7.9 s, every chained one 9.6 to 9.7 s; N = 2 in 32.6 s, N = 4 in 66.8 s, N = 8 in 135.6 s (17.0 s a block); the final proof 1,272,909 bytes whatever N, the record 586 bytes, `verify-segment` 0.037 to 0.040 s; GPU peak 16,751 MiB with the miner resident. Against 4 October with the miner stopped (aggregate 2.2 s): the miner slows the prover 3 to 4x |
| 5090-class cards for 100% at 1 block/s, measured rows | 47 with the loop as it is, 18 through the chain mode on mining cards, 6 (approximate) on proving-only cards, at empty blocks; 14 proving-only at one full shard a block; 45 at `B_p`: the table in the bench log |
## The rule, in one paragraph (spec 7.8)
@ -88,8 +88,8 @@ the rolling upgrade does not partition the network. The order, each step with it
| Decision | Proposed | Why |
|---|---|---|
| `proving_v1_segment_blocks` (N) | 4 | the N = 2, 4, 8 measurement on the 5090 decides; 4 keeps a record every 4 s at 1 block/s and the recursion cost per block constant |
| `proving_v1_unproven_daa` (T) | 600 | equals the 600-block record window of v0: nothing is payable for a segment after it either way; the fleet's measured latency (p99) must sit well inside it |
| `proving_v1_segment_blocks` (N) | 4 | measured: aggregation is a fixed 9.7 s per block on a mining 5090 whatever N, so N only sets how often a record is carried (every 4 s at 1 block/s) and how much a missed deadline forfeits (4 blocks' aggregator share); 8 halves the record traffic for the same card time |
| `proving_v1_unproven_daa` (T) | 600 | equals the 600-block record window of v0: nothing is payable for a segment after it either way; tonight's measured latency from block to carried shard record is p99 52 to 62 s, so 600 leaves 10x |
| `proving_v1_aggregator_share_bps` | 1,000 (a tenth) | the aggregation is one recursion per block, far cheaper than the shards; a tenth pays a second role without starving the shard provers; it is a consensus parameter in the digest |
| `proving_v1_activation_daa` (H) | 24 h after the 0.3.11 publish | the fee-switch rule |
| Aggregator sortition | none in v1 (first valid record wins) | design 5.3's VRF draw is O-7.3; with one or two aggregators on the devnet a draw changes nothing yet |

View file

@ -0,0 +1,154 @@
{
"aggregate_prove_seconds_total": 75.492162475,
"aggregator_id": "0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896",
"blocks": [
{
"aggregate_prove_seconds": 7.903165861,
"aggregate_verify_seconds": 0.037313475,
"block_seconds": 15.53678578,
"chain_len": 1,
"cumulative_seconds": 15.53678733,
"number": 83616,
"post_root": "0x85e67dcea1453afbc17cc0c83ddee0ae12e43d000a4821b4f1c6b2351d677fbc",
"proof_bytes": 1272909,
"shard_prove_seconds": [
7.557706469
],
"shards": 1,
"statement": "0x1f3c209362cc0e6a8b0f4c988728e073331f00d4a2ebdc58cc5cfe078cdf42df"
},
{
"aggregate_prove_seconds": 9.625608758,
"aggregate_verify_seconds": 0.037767114,
"block_seconds": 17.043759121,
"chain_len": 2,
"cumulative_seconds": 32.580617039,
"number": 83617,
"post_root": "0x6e85cac1c9c7973c2cb4d0ff38f96edc054c092fa414f26bf53b237e08e81b77",
"proof_bytes": 1272909,
"shard_prove_seconds": [
7.344317298
],
"shards": 1,
"statement": "0x903b6bf10915403c9ed8a7bc87951c6bc975ba992f947c3b02af48c7e9331306"
},
{
"aggregate_prove_seconds": 9.554146543,
"aggregate_verify_seconds": 0.037548738,
"block_seconds": 16.93764002,
"chain_len": 3,
"cumulative_seconds": 49.518392437,
"number": 83618,
"post_root": "0x4a386002c4df5c9a0ea0494442d252c02860bf2c44bc626321c8b4bb11c5b7d5",
"proof_bytes": 1272909,
"shard_prove_seconds": [
7.307433102
],
"shards": 1,
"statement": "0x0913be6ca9c2671c53e202a144e9a57ecb9b45224ad764031b5e01487dccfa94"
},
{
"aggregate_prove_seconds": 9.725408301,
"aggregate_verify_seconds": 0.038074529,
"block_seconds": 17.316632646,
"chain_len": 4,
"cumulative_seconds": 66.835133749,
"number": 83619,
"post_root": "0x2d46b11d9f257e3e6a6d6e3ebd947a3498c3c5d9a80ef2651caf72f5d2af47ce",
"proof_bytes": 1272909,
"shard_prove_seconds": [
7.516456343
],
"shards": 1,
"statement": "0x00e0a2c9f1e16e06491ce8d26f457b35c4ac60cd3b609e8a87134a31a06a1fec"
},
{
"aggregate_prove_seconds": 9.694312635,
"aggregate_verify_seconds": 0.038877759,
"block_seconds": 17.159775091,
"chain_len": 5,
"cumulative_seconds": 83.995023181,
"number": 83620,
"post_root": "0x209d61c1fa5d410a252a829f186b22b01840729619b69177647ad6d73c87af5d",
"proof_bytes": 1272909,
"shard_prove_seconds": [
7.389889021
],
"shards": 1,
"statement": "0xf26202b09c544d042d78d4844e6589d467f8621d9c4f13d61685fa36d58444d9"
},
{
"aggregate_prove_seconds": 9.664895146,
"aggregate_verify_seconds": 0.035052656,
"block_seconds": 17.453177945,
"chain_len": 6,
"cumulative_seconds": 101.448329353,
"number": 83621,
"post_root": "0x78034f6b455e956da74f6e685b89ead732a45b665ff4706be6deea7cbb402c53",
"proof_bytes": 1272909,
"shard_prove_seconds": [
7.71706568
],
"shards": 1,
"statement": "0x53b07bfc69d17ca7504590bfe52f1d0915e39b2e3e497c57117d871ab275d2c9"
},
{
"aggregate_prove_seconds": 9.634209172,
"aggregate_verify_seconds": 0.037378761,
"block_seconds": 17.047352194,
"chain_len": 7,
"cumulative_seconds": 118.49578211,
"number": 83622,
"post_root": "0x258fbe0b700833451b4bde8139cc6eb60da4fd959dafdd2784572ead6e39d908",
"proof_bytes": 1272909,
"shard_prove_seconds": [
7.338077294
],
"shards": 1,
"statement": "0xc425c86b6569c80481691e6c89abfbd1481cc555030b791178ef536b1b3264a7"
},
{
"aggregate_prove_seconds": 9.690416059,
"aggregate_verify_seconds": 0.036976436,
"block_seconds": 17.115775928,
"chain_len": 8,
"cumulative_seconds": 135.611679366,
"number": 83623,
"post_root": "0x8840c082c4406252bcd65a31ed0102a2eba1b59130582bbba0f87d64ebbc80ed",
"proof_bytes": 1272909,
"shard_prove_seconds": [
7.347491181
],
"shards": 1,
"statement": "0xa99aba5397bae3fd8323dcbc4a93105ae531356fff92feaccca0c3284168ca49"
}
],
"chain_seconds": 135.611788663,
"first": 83616,
"fixtures": [
"/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83616.json",
"/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83617.json",
"/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83618.json",
"/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83619.json",
"/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83620.json",
"/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83621.json",
"/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83622.json",
"/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/block-83623.json"
],
"last": 83623,
"mode": "chain",
"prover": "cuda",
"segment_block_hash": "0x6a3295e7481cb3d5db67758d95280debb72b000a3a0c8a99fe30bf4f3d66aed1",
"segment_chain_len": 8,
"segment_number": 83623,
"segment_proof_bytes": 1272909,
"segment_proof_file": "/mnt/c/Users/Admin/AppData/Local/igneum/app/jobs/chain-pc2-pv1c/segment-83623-aggregated.bin",
"segment_proof_sha256": "0xc28c2c1e8054523c1cff5e38cef8d86498bc004b9beeff010a1d2acc7ef1435c",
"segment_provers": "0x1805afcd50a68f5237f8a5e2e41a4270457be031d1158f5247af6f68808b0d11",
"segment_public_values": "0x000000000000116f00000000000146a76a3295e7481cb3d5db67758d95280debb72b000a3a0c8a99fe30bf4f3d66aed16cad584f6a86085d410c13728d11fa5c4daf583ebbc443512999174b9f7d3b68000000010000000000000000000000000000000000000000000000000000000000000000258fbe0b700833451b4bde8139cc6eb60da4fd959dafdd2784572ead6e39d9088840c082c4406252bcd65a31ed0102a2eba1b59130582bbba0f87d64ebbc80ed6fcd3e8e97da273711ccefb79abdd246c5663c7d61057f82bae745ceac5dcc750000000000000000000000000000000000000000000000001805afcd50a68f5237f8a5e2e41a4270457be031d1158f5247af6f68808b0d112b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b091238960000000000000008",
"segment_statement": "0xa99aba5397bae3fd8323dcbc4a93105ae531356fff92feaccca0c3284168ca49",
"setup_seconds": 15.734044172,
"shard_program_id": "0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a",
"shard_prove_seconds_total": 59.518436388000005,
"shards": 8
}