diff --git a/docs/bench-log.md b/docs/bench-log.md index ec821672e..c0999a0dd 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1618,6 +1618,21 @@ Branches `proving-v1` (main repository, worktree `igneum-wt-proving-v1`; fork `v | A 3-minute window at 18:57Z on node 1 (`node tools/proving-v1/coverage.mjs --minutes 3`, chain blocks 80754..80839, 86 blocks) | 4 blocks with a paid shard (4.7%), 4 fully proven, 4 of 86 shards; on-chain latency (carrier timestamp minus block timestamp) n 4: min 36 s, p50 39 s, max 44 s; 0 content blocks. One prover (PC 2), the Mac verifier node down, PC 2 producing few blocks (3.4 MH/s): the degraded state above, not the fleet's number | | The 30-minute window | waits for the fleet (the coordinator's go) | + +### Step 3, the fleet size (arithmetic from measured inputs; every input names its entry) + +Inputs, all RTX 5090 (PC 2), SP1 6.8.1 cuda: a full shard at the provisional `S_p` (6.75 M pgas) compressed in 10.9 s and the four shards of a near-`B_p` block in 10.2 to 10.7 s each (bench-log 4 October 2026, "shard proving on the RTX 5090", runs run-20261004-173115 and run-20261004-r3-shards); one aggregation 2.2 s (two shards) to 2.5 s (four shards), the same entry; tonight's chain of 2 on the Mac CPU shows the recursion over the previous block proof costs the same order as a first aggregation (52.0 s against 59.1 s), so the GPU figure for a chained aggregation is taken as 2.5 s, approximate, until the held PC 2 chain job measures it; the app's live loop tonight: 1.6 shards a minute per card on empty shards (export, cut, key setup, prove, sign, submit: about 37 s a shard, of which the proof is a few seconds), bench-log step 1 above. A 5090 proves one thing at a time. + +| Block content at 1 block/s | Shard proofs a second (fleet) | Card-seconds a second for shards | Aggregations a second | Card-seconds a second for aggregation | 5090-class cards for 100% | Rule | +|---|---|---|---|---|---|---| +| empty blocks (tonight's devnet), the app's loop as it is | 1 | 37 | 1 | 2.5 (approximate) | 40 | one shard per block, the loop's 37 s each, one card does 1.6 a minute | +| empty blocks, the loop with one key setup per process and the export cached (the chain mode's shape: 12 s setup once, then proofs back to back) | 1 | about 5 (approximate: an empty shard's compressed proof on the 5090 is not measured; the 200-pgas shard took 2.7 s on 4 October) | 1 | 2.5 | 8 (approximate) | the held PC 2 chain job gives the empty-shard proof time | +| one full shard a block (`S_p`, 6.75 M pgas) | 1 | 10.9 | 1 | 2.5 | 14 | 10.9 + 2.5 card-seconds per block-second | +| blocks at `B_p` (four full shards) | 4 | 42.5 | 1 | 2.5 | 45 | 4 x 10.6 + 2.5 | +| at the adopted v1 budgets (`B_p` 120,000 pgas, `S_p` 30,000, from DAA 210,000 on the devnet): a v1 shard of transfers ran at 213 to 236 cycles per pgas (bench-log 5 October, "the prover carries both fee tables"), 7 M cycles a shard against 60 M for the prototype shard | 4 | under 42.5 (the 5090 time for a 7 M-cycle shard is not measured; scaling 10.9 s by cycles gives about 1.3 s, approximate) | 1 | 2.5 | about 8 (approximate) | measure before the switch lands | + +Reading. The card count is the sum of card-seconds of work per block-second, rounded up, with no slack for the exclusive window, the relay or a card's idle gaps; the devnet's own numbers tonight (one card, 1.6 shards a minute, 4.7% of blocks) are the first row. The lever is the loop, not the proof: a shard's proof on the 5090 is 3 to 11 s and its carriage through export, cut and a 12-s key setup is 25 s more. The host's `--mode aggregate` and `--mode chain` already hold one key setup per process; the prover loop should do the same (one host process per segment, the 0.3.11 item in the plan). + ### Step 4, the rule | What | Measured | diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md index b9927add7..cac888026 100644 --- a/docs/plans/proving-v1.md +++ b/docs/plans/proving-v1.md @@ -26,9 +26,20 @@ that say how many cards cover the chain. | 4 | The chain rule and the unproven rule in consensus behind `proving_v1_activation_daa` (spec 7.8 items 2, 6, 7); unit tests; the fast-time 3-node harness `tools/proving-v1/net.mjs` (ports 29950+, suffix 956, trust mode) with the known-finished and known-failed cases | Implemented; the harness run waits for the Mac build of the fork (`vendor/igneum-node/target-pv1`) | | 5 | This plan: the rollout for 0.3.11 and the project lead's decisions | Written below | -## Numbers (every one from `docs/bench-log.md`, "proving v1") +## Numbers (every one from `docs/bench-log.md`, "proving v1: segment records ...", 5 October 2026 evening) -(filled as the measurements land; see the bench log entries of 5 October 2026 named "proving v1 ...") +| What | Number | +|---|---| +| GPU memory during proving on the 5090 (empty shards, 298 one-second samples) | min 1,654 MiB, max 13,816 MiB; so a 16 GB card clears it and a 12 GB card does not; the full-shard peak is the held chain job's row | +| Host RAM | host used 25.6 GB of 63 GB; the WSL2 VM 7.9 GB working set | +| sp1-gpu-server 6.8.1 compiled targets (cuobjdump) | sm_80, sm_86, sm_89, sm_90, sm_100, sm_120 and compute_120 PTX: Ada (4090) is native, no JIT; nothing for AMD | +| Shards a minute, one 5090 through the app's loop (empty shards) | 1.6 | +| Chain of 2 live blocks on the Mac CPU (`--mode chain`) | shard 55.4 and 41.3 s, aggregate 52.0 s then 59.1 s with the previous proof, chain_len 2, final proof 1,272,909 bytes, `verify-segment` 0.032 s | +| Unit tests | consensus core 13, exec 8, app 5, all passing on the Mac | +| The harness (3 nodes, fast time, trust mode) | PASSED, 21 checks in 197 s: paid 1.0 s after submit, every node agreeing; the fresh chain refused after a proven segment; the unproven segment skipped after its deadline; shards at 90% | +| Coverage, 3-min window, the degraded fleet (one card, the Mac verifier down) | 4.7% of blocks proven, on-chain latency p50 39 s | +| 5090-class cards for 100% at 1 block/s | 40 with the loop as it is (empty blocks), 14 at one full shard a block, 45 at `B_p` (four full shards), about 8 at the adopted v1 budgets (approximate): the table in the bench log | +| HELD (coordinator, 19:00Z): mining alone against mining with the prover; the chain of 8 on the 5090 (N = 2, 4, 8); the 30-min coverage with the fleet mining | re-run after the go: jobs `prover-cost-pc2-pv1` (script fixed to print its state error) and `pc2-chain.ps1` with the package `igneum-prove-wsl2-pv1.zip` | ## The rule, in one paragraph (spec 7.8)