igneum/docs/plans/proving-v1.md

223 lines
44 KiB
Markdown

# Proving v1: segment records, the chain rule, the unproven rule; the 0.3.11 rollout
5 October 2026, from 18:55 UTC (the project lead: "open the proving round asap"). Branches `proving-v1` in the main repository
(worktree `/Users/joshm/Projects/igneum-wt-proving-v1`, from master a93199a) and in the fork
(`vendor/igneum-node-pv1`, from release-0.3.6 a24ab01a; to be rebased onto the 0.3.10 tip when it lands on
release-0.3.6). Status words follow `docs/spec/00-overview.md` 0.2. Every number here is in `docs/bench-log.md`
with its command. Nothing ships from this plan: it delivers branches, numbers and the rollout for 0.3.11.
## The gap this closes
The litepaper says every block is proven within about a minute. Proving v0 (`proving-v0.md`, spec 7.7) proves
some shards: one prover (PC 2's RTX 5090) takes the newest shard assigned to it, about one shard every 30 s, so
under a tenth of blocks carry a proof; consensus does not require one; the aggregator guest (design 5.3) runs
on fixtures only. Proving v1 adds the aggregated segment record on chain (spec 7.8), the chain rule (segment N's
record verifies N-1, inside the proof by recursion), the unproven rule (a segment nobody proves in T seconds pays
nothing and may be skipped), the prover on by default on every machine that can prove, and the measurements
that say how many cards cover the chain.
## The round, step by step
| Step | What | State |
|---|---|---|
| 1 | Prover on by default (`app/igneum-app/src/provedefault.rs`, `engine.rs apply_prove_default`): on at install when the machine can prove (NVIDIA card with 12 GB or more; WSL2 answering on Windows; Linux native; Apple silicon off until measured), never switching an explicit on back off; the Settings switch line and the tile line say why. Unit tests (5). The cost of proving on a mining machine: PC 2 job `prover-cost-pc2-pv1` (5 min mining alone, 5 min with the prover, the GPU memory peak and the host RAM peak, the sp1-gpu-server's compiled SM targets) | Implemented; the measurement is HELD (coordinator, 19:00Z): PC 2's RTX 5090 worker has been exiting on a pack seed mismatch since 18:35Z, so the first run's "mining alone" is 0 MH/s and void; re-run after the go |
| 2 | Segment aggregation: `SegmentRecord` (586 bytes, the aggregator guest's 340-byte statement inline), section `IGNS` before the shard section, p2p message 75 at protocol 15 (14 went to the EVM transaction relay in 0.3.10), the native block statement and the veto, the credit split (`split_pool_credit`), the payout at the carrier, `igneum-miner sign-segment-record`, the RPCs; host modes `chain` (consecutive fixtures), `aggregate` (live shard proofs from the pool, a run of blocks in one process) and `verify-segment` (the node's verifier, pinned aggregator key); the app's aggregator step (`prover.rs aggregate_once`) | Implemented, unit-tested (consensus core 2 new tests, exec 2, params 1); the GPU measurement (N = 2, 4, 8 blocks on PC 2) is HELD with step 1; the Mac CPU run of `--mode chain` over 2 live blocks is the known-finished case |
| 3 | Coverage: `tools/proving-v1/coverage.mjs` (the proven-block share and the on-chain proof latency over a window from one node's RPC, the live page beside it) | Implemented and run for 3 min (below); the 30-min window waits for the fleet |
| 4 | The chain rule and the unproven rule in consensus behind `proving_v1_activation_daa` (spec 7.8 items 2, 6, 7); unit tests; the fast-time 3-node harness `tools/proving-v1/net.mjs` (ports 29950+, suffix 956, trust mode) with the known-finished and known-failed cases | Implemented; the harness run waits for the Mac build of the fork (`vendor/igneum-node/target-pv1`) |
| 5 | This plan: the rollout for 0.3.11 and the project lead's decisions | Written below |
## Numbers (every one from `docs/bench-log.md`, "proving v1: segment records ...", 5 October 2026 evening)
| What | Number |
|---|---|
| The prover's cost to a mining 5090 (the re-run with the fleet mining, hash rate from the miner's own STATUS lines) | 124.7 MH/s alone, 119.7 MH/s with the prover on: 5.0 MH/s, 4.0%, on empty shards at 1.4 a minute |
| GPU memory on the 5090: the prover alone (empty shards) / the miner and the prover together | max 13,816 MiB / max 15,590 MiB (the miner holds 3,396 MiB); a 24 GB 4090 has 8.4 GB of headroom, a 16 GB card 0.4 GB, a 12 GB card cannot do both on this build; the full-shard peak is the chain job's row |
| Coverage, 30-min window with the fleet mining, one prover | 2.4% of blocks, latency p50 44 s, p99 52 s |
| Host RAM | host used 25.6 GB of 63 GB; the WSL2 VM 7.9 GB working set |
| sp1-gpu-server 6.8.1 compiled targets (cuobjdump) | sm_80, sm_86, sm_89, sm_90, sm_100, sm_120 and compute_120 PTX: Ada (4090) is native, no JIT; nothing for AMD |
| Shards a minute, one 5090 through the app's loop (empty shards) | 1.6 |
| Chain of 2 live blocks on the Mac CPU (`--mode chain`) | shard 55.4 and 41.3 s, aggregate 52.0 s then 59.1 s with the previous proof, chain_len 2, final proof 1,272,909 bytes, `verify-segment` 0.032 s |
| Unit tests | consensus core 13, exec 8, app 5, all passing on the Mac |
| The harness (3 nodes, fast time, trust mode) | PASSED, 21 checks in 197 s on b177718e (N = 4) and 244 s on the final tree ece42979 (N = 8): paid 1.0 s after submit, every node agreeing; the fresh chain refused after a proven segment; the unproven segment skipped after its deadline; shards at 90% |
| Coverage, 3-min window, the degraded fleet (one card, the Mac verifier down) | 4.7% of blocks proven, on-chain latency p50 39 s |
| The chain of 8 live blocks on the 5090, the card also mining (`chain-pc2-pv1c`) | shard 7.3 to 7.7 s, first aggregation 7.9 s, every chained one 9.6 to 9.7 s; N = 2 in 32.6 s, N = 4 in 66.8 s, N = 8 in 135.6 s (17.0 s a block); the final proof 1,272,909 bytes whatever N, the record 586 bytes, `verify-segment` 0.037 to 0.040 s; GPU peak 16,751 MiB with the miner resident. Against 4 October with the miner stopped (aggregate 2.2 s): the miner slows the prover 3 to 4x |
| 5090-class cards for 100% at 1 block/s, measured rows | 47 with the loop as it is, 18 through the chain mode on mining cards, 6 (approximate) on proving-only cards, at empty blocks; 14 proving-only at one full shard a block; 45 at `B_p`: the table in the bench log |
## The rule, in one paragraph (spec 7.8)
From the first chain block `A` at or above `proving_v1_activation_daa`, chain blocks form fixed segments of `N`
(`proving_v1_segment_blocks`). A segment record carries the aggregated proof of the segment's last block, whose
`chain_len` says how many consecutive blocks the recursion attests. Every node checks the record's statement
against its own native block statement (every field but the provers commitment and `chain_len`), the chain rule
(a proof that does not chain to the previous segment, `chain_len = N`, is valid only for the first segment or
after an unproven one) and the deadline (`T = proving_v1_unproven_daa` DAA seconds after the segment's last block;
a record carried later pays nothing). The pool credit of every attested block splits: `proving_v1_aggregator_share_bps`
to the aggregator, the rest to the shards as v0. A block is never invalid for lack of a proof; the mandatory rule
(spec 7.8 item 10) is Designed and off, with no switch yet.
## Rollout for 0.3.11 (the digest handshake pattern of 0.3.9 and tonight's switches)
The switch moves the consensus digest only once it is set (`consensus_digest`: the four v1 fields enter the hash
when `proving_v1_activation_daa != never`), so a 0.3.11 node on the unswitched devnet keeps the 0.3.10 digest and
the rolling upgrade does not partition the network. The order, each step with its check:
1. **Rebase and build.** Fork `proving-v1` rebased onto the 0.3.10 tip on `release-0.3.6`; the six node suites and
the app tests as PC 2 build jobs; the Mac node and the Windows exes by the Mac cross-build; the Linux node by
PC 1; the HiveOS package republished from the same fork commit (rule: the HiveOS package carries the node of
the release commit and the same override object, `infra/hive`, as 0.3.9's `hive-sync-039o` checked it).
2. **Pinned guests.** The guest ids do not change in this round (shard `0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a`,
aggregator `0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896`, pinned 2026-10-05T16:20:38Z):
the aggregator guest already carried the chain rule and only the HOST gained modes. So no provers-off drain is
needed for the guests; `--mode id` on every machine after the update must print the same two ids, and the node
now reads them at start (`program_ids`: `IGNEUM_PROOF_PROGRAM_IDS` or the verifier's `--mode id`) and names them
in the native statement.
3. **Hand nodes and the seed first**, with the UNCHANGED override object (the digest stays): observer, node 1, the
seed on the 0.3.11 node; peers back within 20 s; the `proving v1: segment records from DAA score never` line in
each log.
4. **Manifest and apps.** `publish-manifest.sh --version 0.3.11` with the unchanged object; `update-now` to every
app; every machine on 0.3.11 with a DAA score and a hash rate (the watcher takes the commit as an argument).
The app's prover default applies at the first start on 0.3.11: every NVIDIA machine with WSL2 goes on; the
log line `prover default: ...` on each.
5. **The switch.** When every node runs 0.3.11: publish the object with `proving_v1_activation_daa` = H (24 h
ahead, the rule of `fee-switch-devnet.md`) and the three parameters; read the expected digest on a scratch node
first; `update-now`; the hand nodes and the seed with the same object; the digest sweep; the first paid segment
record (`igneum_getSegmentRecords`) and `igneum_getProvingStatus.v1.segmentsInWindow` after H.
6. **The mandatory rule** stays off: no switch exists for it yet; it gets one when the measured share is one.
## The decisions (Decided 5 October 2026, delegated: the project lead, "I have no idea for most of this stuff so do a lot of research and deploy what is absolute best")
Each with its rule, its number and its evidence. The deploy is the DEVNET through 0.3.11 (not the public testnet).
### What the other networks do (read 5 October 2026, 20:05 to 20:15 UTC; every figure from the page named, else labelled approximate)
| Network | Unit proven | Deadline | What a miss costs | Who is paid what | Measured latency |
|---|---|---|---|---|---|
| Taiko Alethia (L2BEAT page, protocol v2.1.0 notes) | a batch of blocks | proving window 2 h, cooldown 2 h (v2.1.0, February 2025); the Shasta inbox targets a 4-h proof submission cadence | the proposer's liveness bond is credited back in full when the batch is proved inside the window, half when outside; mainnet currently sets minBond and livenessBond to 0 | the prover earns the proving fee; two of four proofs needed (SGX Geth, SGX Reth, SP1, RISC0, at least one ZK) | 100% ZK coverage of mainnet blocks reached December 2025 (blockchain.news); preconfirmations 2 s |
| Boundless (docs.boundless.network, proof lifecycle) | one request | the requester's timeout (example 3,600 s) and a lock timeout (example 2,700 s); a reverse Dutch auction ramps the price from the minimum to the maximum over a ramp-up (example 300 s) | the locked collateral (example 5 ZKC) is slashed and used to pay another prover who fulfils the request | the prover's fee = the bid minus the market fee | not stated on the page |
| Succinct Prover Network (docs.succinct.xyz, SPN architecture and quickstart) | one request | the requester's deadline (the quickstart example: 10 minutes, 50 PROVE staked to bid, 100 PROVE maximum fee) | part or all of the winning prover's collateral slashed "according to protocol rules" | a reverse auction: the lowest bidder is assigned | "real-time", no number on the page |
| Aztec (docs.aztec.network economics; L2BEAT; forum) | an epoch of 32 blocks (38 min 24 s), a proof may cover one checkpoint (1 min 12 s) up to one epoch; maximum proof window 1 h 16 min | the epoch is declared failed only when its submission window expires | an unproven epoch is reorged out (no reward); proposals under discussion remove bonds and pay every prover that delivers on time | 400 AZTEC a slot: 70% sequencers, 30% provers (120 AZTEC), provers' share by an activity score | the public testnet proved by community provers (zkcloud blog), no page number |
| zkSync Era, Linea, Scroll (eco.com comparisons) | a batch | none on chain (the operator proves) | none | the operator | proof latency about 30 min (zkSync Era), 75 min (Linea), 90 min (Scroll), approximate |
Reading. Nobody pays an aggregator as a separate role: Aztec's 30% goes to whoever delivers the epoch proof, Taiko's fee to whoever proves the batch, the markets to the request's winner. Deadlines run from 10 minutes (Succinct's example) through 1 h (Boundless' example) to 2 h (Taiko) and 1 h 16 min (Aztec's maximum window); a miss forfeits the reward or part of a bond, and the slashed value goes to the prover who steps in (Boundless). Igneum has no bond on shards by decision (spec 7.2 item 4), so the forfeit here is the reward only.
### The decisions
| Decision | Decided | Rule and number | Evidence |
|---|---|---|---|
| `proving_v1_segment_blocks` (N) | **8** | Aggregation is a fixed cost per block, not per segment: 9.6 to 9.7 s for every chained block on a mining 5090, 7.9 s unchained (`chain-pc2-pv1c`), so N buys nothing in card time; it sets the record cadence and the forfeit. At N = 8 and 1 block/s a record every 8 s, a 1.27 MB proof gossiped every 8 s (159 KB/s per path, half of N = 4's 318 KB/s), and a missed segment forfeits 8 blocks' aggregator share (8 x 0.088 IGN at today's credit). The chain for 8 blocks cost 135.6 s cold on a mining card (66.8 s for 4), a fifth of T; pipelined per block it is 17 s after the last block. Aztec proves 32 blocks (38 min) as one; 8 blocks at 1 block/s is 8 s of chain, so the record lands well inside the minute the litepaper promises | bench-log "proving v1" chain rows; Aztec economics page |
| `proving_v1_unproven_daa` (T) | **600** DAA s (10 min) | T = p99 x 10: the measured block-to-carried-record latency of a shard record is p99 52 to 62 s (two 30-min windows), a cold chain of 8 adds 136 s and relay plus inclusion 10 to 40 s, about 240 s worst case; 600 leaves 2.5x on that and equals the 600-block record window of v0, so nothing is payable past it either way. Succinct's example deadline is the same 10 minutes; Boundless' example 1 h, Taiko 2 h, Aztec up to 1 h 16 min: Igneum's blocks are 1 s and its proofs seconds, so the shortest of the field. The forfeited aggregator share of an unproven segment STAYS IN THE POOL ESCROW (it is never paid, as an unproven shard's part today): no burn and no roll-over, the rule the pool already has, and the escrow is what later proofs are paid from | coverage rows; `chain-pc2-pv1c`; the table above |
| `proving_v1_aggregator_share_bps` | **1,000** (a tenth) | The aggregator's card time per block is 9.7 s on a mining card against 4 x 10.6 s of shard proofs at `B_p` (19% of the card time) and 2.5 s against 42.5 s with the card to itself (6%); on tonight's empty blocks it is half the card time. A tenth of every attested block's pool credit sits between the two full-block ratios, pays a role no other network pays separately (Aztec pays its 30% to whoever delivers the epoch; the markets pay the winner), and leaves the shard provers 90%, which the fast-time harness showed paid exactly (shardWei 90% of the credit). The pool's 20% emission share itself is unchanged (spec 2.5, 5.3) | `chain-pc2-pv1c`; bench-log 4 October 5090 rows; the harness |
| `proving_v1_activation_daa` (H) | **the devnet tip + 14,400 at publish** (4 h at 1 block/s), set by the 0.3.11 publisher in the same override object as `program_class_v3` | tonight's rule for consensus switches (the coordinator, 5 October 2026); the digest moves only once H is set, so the rolling update does not partition |
| Aggregator sortition | **none in v1**: the first valid record carried wins | design 5.3's VRF draw (O-7.3) with one or two aggregators on the devnet changes nothing; the segment grid and the deadline already bound the race; revisit when a second aggregator exists |
| Apple silicon default | **off** | the gate was "a shard under 60 s with the miner running": the M5 Max CPU took 41.3 and 55.4 s for EMPTY shards under tonight's load and 272 s for a 200-pgas shard on 4 October; a full shard at `S_p` was never under 60 s. Settings switches it on | bench-log "proving v1" CPU chain row; 4 October CPU rows |
| The prover profile per card and the 12 GB and 16 GB gates (the project lead: "make sure we can prove on 12gb cards"; "is there any way we can make 12gb cards mine and prove?") | **measured on PC 2, the rows below** | the SP1 6.8.1 GPU server reads `ELEMENT_THRESHOLD`, `HEIGHT_THRESHOLD`, `SHARD_SIZE` and the `SP1_WORKER_NUM_*`/`BUFFER_SIZE` knobs from the environment it inherits (`sp1-core-executor-6.8.1/src/opts.rs`, `sp1-prover-6.8.1/src/worker/config.rs`); the app passes a profile per card (`provedefault.rs`) and the host forwards it | the sweep job `memsweep-pc2-pv1` and the miner-on run |
The resume path (5 October 2026, the 0.3.11 app): `POST /api/resume` on 0.3.9 re-armed only FAULTED cards (`stop_miners("paused")` clears every slot's `restart_at`), so a healthy paused card stayed "off" at 0 MH/s until the app was relaunched: PC 2 at 21:25:11Z (the aggregation-cost job's pause and resume; `[ok] mining resumed` then `0.00 MH/s, waiting` for 20 minutes), the Mac that afternoon. Now every slot without a live worker is re-armed and its pack exported again before the start, and 90 s later `resume_check` logs `resume: <card> is not mining 90 s after resume (state ..., pid ...)` for every enabled card without a hash rate (`engine.rs`, three unit tests: the state machine, the 21:25:11Z case against the old rule, the check).
The prover-floor agent's first sweep (job `floor-sweep-1`, 22:34 to 22:38Z, PC 2's 5090, the miners stopped, this plan's per-point recipe, its patched `sp1-gpu-server` 5568108b built for sm_86, sm_89 and sm_120, every proof VERIFIED by the unpatched pv1 host): the control at upstream's sizes reproduces the curve above (empty shard 13,892 MiB and 2.2 s; the v1 shard 20,516 MiB and 4.2 s); with the core element threshold at 2^26 the v1 shard proves as four core shards in 5.3 s at **12,708 MiB** and the empty shard at 12,772 MiB; 2^25 gives 12,836 MiB at 8.5 s; 2^27 gives 15,396 MiB. The 12.7 GB left is the server's Setup (five recursion keys pre-built at a fixed 2^27 capacity plus the shrink and core keys: 9.7 GB before the first shard), which its patch v2 sizes to the need. Decided for the 12 GB profile: the split that lands under 11 GB wins (5.3 s a shard is inside the loop's own 25 to 30 s of carriage and 100x inside T); 2^27 is the second profile only if v2 leaves it under 11 GB with the miner's 1.8 GB beside it. The 12 GB row stays OPEN until the final pair (alone and beside the miner) lands and the on-order 3060 runs it.
### A self-built CUDA server (the 12 GB path), before 0.3.12 (consequences C26)
If the prover-floor agent's rebuilt `sp1-gpu-server` (the Setup sizes cut, built on PC 2 under WSL2) proves a shard under 11 GB, it becomes a shipped artefact and needs its own row of rules before 0.3.12: it is built from a pinned SP1 source tag with `CUDA_ARCHS` covering sm_86, sm_89 and sm_120 (the 12 and 16 GB tiers are Ampere and Ada, not only the 5090's Blackwell; one card family per measured row), by the packaging path that builds the Windows payload (PC 1's build job for the Linux binary, the Mac signs the manifest as it does the DMG), lands in the DMG and the WSL2 package beside the host as `wsl2/bin/sp1-gpu-server` with its sha256 in `payload-inputs.json`, is named in `evidence.md` beside the prover rows ("prover built from SP1 <tag> at <sha>"), is rebuilt and re-measured at every SP1 upgrade, and ships only after `--mode verify-segment` and `--mode verify` on proofs it made show the pinned verifying keys unchanged (the server changes allocation, not the circuit; the ids `0x2b1a81cb...` and `0x474678f3...` must still verify them). The 12 GB claim itself waits for the on-order RTX 3060 to run that server on the same fixtures and recipe as the curve; until then the public line stays at 24 GB.
The root-socket class on PC 2, the two times: 20:00:56Z (my chain job's root run; the live prover failed with Connect(PermissionDenied) until the socket was gone) and 21:25:24Z (the aggregation-cost job's root run; the prover stayed dark through the 0.3.10 restart at 21:49:41Z until `socketfix-pc2-pv1` removed the root-owned `/tmp/sp1-cuda-0.sock` at 22:01:16Z; the next shard, block 89011, was proven at 22:02:13Z and paid, and every shard since). The permanent fix in the 0.3.11 app tree: every committed playbook that runs a prove mode as root carries `pkill -f sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock` at its start and end, `tools/ci/prover-socket-check.sh` (in `ci.yml`) fails a playbook without them, and the app's prover names the cause in its log line when the host reports PermissionDenied. The app itself cannot remove a socket another user owns, so a job written outside the tree must still follow the rule.
A prover job on a shared card runs as the app's user or cleans its socket (`pkill -f sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock` at the start and the end; `tools/ci/prover-socket-check.sh`): the root-socket fault of 20:00Z, bench-log.
### The prover profiles: the tiers from the S_p curve (bench-log, "proving v1", the sweep, the miner-on pair and the curve)
The GPU server of SP1 6.8.1 sets the memory, not the shard: a floor of 13.9 GB for an empty shard, 20.4 GB for a full shard at the adopted v1 budget (30,000 pgas, 4.7 M cycles), 28.3 GB for the prototype shard (6.75 M pgas, 60 M cycles); the miner adds 1.7 GB when it shares the card; no environment knob moves the floor and the server has no options of its own; the witness is 5 to 22 KB a shard and never binds.
| Card | Alone | Beside the miner | Default (`provedefault.rs`) |
|---|---|---|---|
| 32 GB (RTX 5090) | the prototype shard, 28.3 GB, 10.8 s; the v1 shard 20.4 GB, 4.3 s | the prototype shard 30.1 GB, 33 s; the v1 shard 22.2 GB, 13.2 s | on, mine and prove, today |
| 24 GB (RTX 4090, 3090) | the v1 shard 20.4 GB; the prototype shard does NOT fit (28.3 GB) | the v1 shard 22.2 GB measured on the 5090's allocation (2.3 GB spare on a 24 GB card; approximate for the card itself) | on, mine and prove, with the line "until the devnet's fee switch its shards are the prototype size, which needs 32 GB, so this card proves from the switch on" |
| 16 GB (RTX 5080, 4080) | the shipped server: an empty shard only (13.9 GB); the patched server v3 b37defef at threshold 2^27: the v1 shard 12,915 MiB and 4.3 s, the prototype shard 13,459 MiB and 16.8 s (measured by the prover-floor agent on the 5090's allocation, job `floor-sweep-3`, 00:13 to 00:17Z 6 October; 2^27 + 2^26 gives 16,115 MiB, over the card) | the shipped server: nothing (15.7 GB for an empty shard); the patched server at 2^27 beside the miner (the 5090 mining at 95%, 338 W, same card; `floor-sweep-4`, 00:35 to 00:38Z): the v1 shard 14,786 MiB total with the miner's 3,833 MiB resident inside it, the server's own 10,953 MiB, 17.4 s a shard; on a 16 GB card that is 10.95 GB server + 1.7 GB miner = 12.7 GB plus the display, under the 15.0 GB line | off on the shipped server, with the line; on (mines and proves, 17 s a v1 shard, 4.3x the alone time) once the patched server ships (the packaging row below) |
| 12 GB (RTX 3060, 4070) | the shipped server: nothing (the floor is 13.9 GB, and the server refuses the card outright); the patched server v3 b37defef at threshold 2^26 (`SP1_GPU_ELEMENT_THRESHOLD=67108864`, the 12 GB profile): **the v1 shard 10,291 MiB and 5.7 s, an empty shard 9,971 MiB and 3.3 s**, the card's 2,089 MiB idle inside the peak and the server's own working set about 8.2 GB (6,535 MiB after Setup), so a 12 GB card proves alone with about 3 GB over it (measured by the prover-floor agent on the 5090's allocation, `floor-sweep-3`; the on-order RTX 3060 run is pending) | measured beside the miner (`floor-sweep-4`): at 2^26 the v1 shard 12,066 MiB total with the miner's 3,833 MiB inside, the server's own 8,233 MiB, 24.4 s; the empty shard 11,586 MiB, 13.0 s; 2^25 gains nothing (12,066 MiB, 43.5 s). On a 12 GB card that is 8.2 GB server + 1.7 GB miner = 9.9 GB before the display, over the 9.0 GB line the project lead set, so mine-and-prove on 12 GB is NOT claimed | off on the shipped server; "proves alone" on the patched one once it ships (the packaging row below); mine-and-prove stays off on 12 GB (9.9 GB plus the display, over the 9.0 GB line). The public gate stays "12 GB proves; 16 GB mines and proves", both on the patched server, both pending a run on the card itself. the project lead's "make sure we can prove on 12 GB cards" is answered on the 5090's allocation and OPEN on the card itself: the prover-floor agent (branch prover-floor, 5 October night) read SP1 v6.8.1's GPU server source (`sp1-gpu/crates/prover_components/src/builder.rs` lines 35 to 39): it reads the card's memory, adds 4 and panics under 24 ("Unsupported GPU memory ... must be at least 24GB"), and builds its core (ELEMENT_THRESHOLD 2^28 + 2^27 elements + 2^21), recursion (2^27), shrink (2^25) and wrap (85 M element) provers at Setup whatever the mode, which is the 13.9 GB floor; no knob reaches them, so the fix is a server rebuilt from source on PC 2 (WSL2, nvcc 12.8, CUDA_ARCHS=120) with those sizes cut, measured on the same fixtures and recipe as the curve above (D2 carries the curve) |
| under 12 GB | nothing | nothing | off, mine only |
| AMD-only and Apple machines | nothing on the GPU: no zkVM proves on an AMD GPU today (`docs/analysis/amd-proving.md`, branch amd-prove); the CPU prover is about 5 minutes a shard at a 30 GB RSS whatever the shard size (PC 1, bench-log "the SP1 CPU prover on PC 1") | | off, "mines and does not prove"; the only non-NVIDIA path with a shipped backend is RISC Zero's Metal prover behind the `ProofSystem` seam (a second guest and pinned id, a verifier for both formats, no shared aggregation): an open item, not 0.3.11 |
The three profile numbers the coordinator asked for, as measured: under 9.0 GB does not exist on this build (floor 13.9); under 15.0 GB mine-and-prove does not exist for any full shard (the v1 shard alone is 20.4); the full profile is the 32 GB card. The fleet table's "proving-only" rows therefore read 24 GB cards at the v1 budget and 32 GB cards at the prototype budget. Shards per block at the v1 budget: 1 on tonight's empty chain, 2 to 4 on blocks with transactions (`B_p` 120,000 = 4 x `S_p`); the aggregation count is one per block whatever the shard count (the chained recursion), so the aggregation-cost agent's target is per block.
The aggregation-cost agent's first rows (branch agg-cost, 5 October 2026 night, the same four live blocks on PC 2): with the miners paused an empty shard proves in 1.9 to 2.2 s and an aggregation in 1.7 to 2.2 s with the card 15.8% busy; mining, 7.4 to 7.8 s and 7.9 to 9.8 s at 93.9% busy, so the miner's kernels take the card and the prover runs 3.6x (shards) to 4.5x (aggregations) slower beside them; its batch-size curve is still open. That puts a proving-only card at about 4 s per empty block (one shard and one aggregation), 4 cards for an empty-block chain at 1 block/s, against 18 mining cards.
The re-plans of block 344 at 2.25 M and 4.5 M pgas peak at 28.3 to 28.4 GB alone (the server's buffers step up between 4.7 M and 20 M cycles and are flat to 60 M), so no shard size between the v1 budget and the prototype one changes a tier; with the miner the adopted shard proves 3.1x slower (13.2 s against 4.2 s) and the chained aggregation 9.7 s against 2.5 s: a mining 24 GB card delivers one adopted-size shard plus one aggregation in about 23 s, inside T by 25x.
## Segment-aligned proving (6 October 2026, the project lead: "find a way to solve this")
**The fault was the work order, not the rules.** The shipped loop took the newest open shard each pass (`prover::choose`), so one prover scattered one block in about 45 across the segment grid and no segment ever had all its blocks proven: pending 56, proven 0, unproven 20 at 06:28Z. No consensus parameter moves.
**The change (app branch, commit ce8f34a, app and host):**
| Part | What it does | Where |
|---|---|---|
| Whole-segment claiming | the work list (lookback 600, the record window) grouped into whole untouched segments: every block present, every shard open (past its 10-DAA exclusive window), unpaid, not in our pool | `app/igneum-app/src/segments.rs` `whole_segments` |
| The choice | candidates inside their deadline by a margin (240 DAA, or 1.5x the last segment's wall time), ranked by FNV-1a of (first block, this prover's key hash): deterministic per prover, different between provers, so several provers spread over the candidates with no coordinator; an attempted segment is not retried | `segments::candidates`, `rank`, `need_daa` |
| The statement check | for the best three: executed and pending; fresh only when the previous segment is not paid and no verified record of it waits in the pool (the chain rule would refuse a fresh record once that one pays); chained (`--prev`) when the previous is paid and its proof is in this node's pool | `prover.rs` `pick_segment` |
| The work | one export to the segment's last block, one fixture per block, one host run `--mode chain --save-shards [--prev]` that proves every shard and aggregates the segment in one process (one key setup), then every shard record signed and submitted (the shard payouts, 90% of the credit) and the segment record signed and submitted (the aggregator share, 10%) | `prover.rs` `prove_segment` |
| The fallback | when no whole segment qualifies, the per-block path as shipped | `prover::choose` |
| The host | `--mode chain` takes `--prev <aggregated.bin>` (chain_len continues; a previous proof that is not the parent block's is refused by number and parent hash) and with `--save-shards` writes per-shard records (statement, proof sha256, file, time) into the results | `proving/igneum-prove/host/src/main.rs` `run_chain` |
| The tile | "Segments: N proven whole, M paid (x IGN to the aggregator), the last in T s" and the segment path's last line; the state carries `segments_submitted`, `segments_paid`, `segment_paid_wei`, `segment_last_s` | `ui/app.js`, `state.rs` |
Tests: the grid, the grouping (missing block, paid shard, our shard in the pool, exclusive shard), the margin and the attempted set, the per-key order (deterministic, different between two keys), the margin from the last time: 6 unit tests in `segments.rs`; 120 app tests, 8 core and 9 host tests pass. The host flags were run on the Mac's CPU first (bench-log, 07:12Z to 07:17Z): per-shard records written for a chain of 2, then a chain of 1 continued from it (`base_chain_len` 2, final `chain_len` 3).
**Item 2, "own pool only", answered from the source:** a relayed proof record carries its proof bytes (`protocol/flows/src/v10/proving.rs`: `IgneumProofRecordMessage { record, proof }`, 8 MB bound, handed to the pool with `local = false`), and so does a segment record (message 75). So "this node's pool" is every record relayed to it, and any aggregator can already fold any 8 proven blocks it has received; no node or consensus change is needed for that. What does limit carriage: a block template carries only entries the node's own verifier marked verified (`template_segment_section`, `template_section`), so a node with the verifier `Off` (node 1) never carries a record and a node with `Trust` carries unverified ones; PC 2's node runs the host and verifies. With one prover the carrier is PC 2's own next block.
**What one 5090 completes (arithmetic from the 5 October rows, the measurement below replaces it):** a chain of 8 empty blocks took 135.6 s cold beside the miner; one export and one key setup per segment instead of eight; so about one segment every 150 to 200 s, 9 to 12 segments an hour out of 450 (2 to 3%), against none. The aggregator share of a segment is 8 x 0.088 IGN = 0.70 IGN plus the 8 shards' 90% share; the forfeited share of the other 97% stays in the escrow until the fleet grows (47 mining cards or 6 dedicated provers for 100% at 1 block/s).
**PC 2 measurement (job `segments-pc2-pv1c`, 07:52Z to 08:24Z, `tools/proving-v1/pc2-segments.ps1`; bench-log "the segment-aligned prover beside the miner"):**
| Figure | Value |
|---|---|
| Whole segments proven in 30 min, one 5090 beside its miner | 9, one every 210 s (export 1.5 s, cut 45 s, chain 160 s: 8 shards 63 s, 8 aggregations 80 s) |
| Shard records accepted and paid | 72 of 72, 0.91 to 2.72 IGN a shard (90% of the credit), carried 180 to 226 blocks after the block |
| Segment records accepted | 0 of 9: refused by the chain rule as shipped (below) |
| Miner's cost | 117.86 to 104.90 MH/s, 13.0 MH/s = 11.0% (the shipped prover: 5.0 MH/s, 4.0%, for 2.8x fewer shards) |
| GPU memory peak | 16.5 to 17.6 GB with the miner resident; 24 GB tier unchanged |
**The second fault, found by the measurement: the chain rule as shipped.** `check_segment_record` accepts a fresh record (chain_len = N) only when the previous segment is UNPROVEN at the carrier. A segment's own deadline is its last block's DAA plus 600, the previous segment's deadline is 8 DAA earlier, so a fresh record is valid for 8 DAA (8 s on devnet) and must be proven before and carried inside them. The refusal on the live node, 9 times: "segment 114470..114477 does not chain to segment 114462..114469 (chain_len 8), which is pending until DAA 169681". The fast-time harness passed on 5 October because its chain continued from proven segments (case 2) and its fresh case ran exactly in that window (case 3, 4 DAA at N = 4). No prover-side move escapes it: the record must be proven, submitted and carried inside the window, which the 210-s proof cannot meet.
**The fix (fork branch 0f0dda95, behind a switch, never by default):** `proving_v1_fresh_rule_daa`; from it a fresh record is valid whenever the previous segment is not proven (pending or unproven) at the carrier; after a proven one a record must still chain. Two records that do not chain each attest their own blocks against the native statement, so nothing is lost but the longer proof chain, which restarts. In the consensus digest only once set (the v1 pattern), so a 0.3.12 node on the live devnet keeps the 0.3.11 digest until the override file sets it; `igneum_getProvingStatus.v1.freshRuleDaa/freshRuleActive` and `igneum_getSegmentStatement.freshAdmissible` report it. Tests: the digest moves once set; fresh after a pending segment passes from the switch, is refused before it, never after a proven one; the harness `--fresh-rule 0` inverts case 3. This is a rule change in the execution layer, not a parameter tuning, and the code proved it unavoidable: the project lead's "no consensus parameter change unless the code proves it is unavoidable" is met by the 8-DAA window above and the nine refusals. The 0.3.12 coordinator sets the height (the same tip + 14,400 rule) in the override object with the release.
**Until the switch:** the app (272b025) holds a refused segment record and offers it again every pass until the segment's deadline, so on the shipped rule it lands only if a carrier falls inside the 8-DAA window, and from the switch it lands on the first retry; the shard records (90% of the credit) land either way, 8 per segment.
**Per tier, with this change and the switch:**
| Card | What it does | Per 30 min, empty blocks (measured on the 5090, approximate elsewhere) |
|---|---|---|
| 32 GB mining and proving (5090) | 9 whole segments, 72 shards paid, 9 aggregator shares once the switch is set | miner 11.0% down; 72 x 0.9 IGN = 65 IGN of shard payouts measured, plus 9 x 0.70 IGN aggregator share from the switch |
| 24 GB mining and proving | the same path at the 16.5 to 17.6 GB peak measured; the fee-switch shard not yet measured on a 24 GB card | approximate: the 5090's numbers |
| 16 GB | mines and proves on the patched server only (prover-floor rows); segment path untested there | pending the prover-floor agent's build |
| 12 GB prove-only | proves alone on the patched server (10.3 GB); a dedicated prover takes 4 s a block alone (agg-cost rows), so about one segment every 40 s | approximate: 45 segments per 30 min, 6 such cards for 100% |
| A rig (several cards) | one prover loop per machine today; the segment path claims one segment at a time on the aggregation card | the per-card loop is the next item |
## v1 live on devnet (6 October 2026, C47)
v1 active at 154,800 (crossed at DAA 154,814, 03:51:42Z); first segment record: none, because on a one-prover devnet no segment can be proven. The app's aggregator (`aggregate_once`, 0.3.11) needs a shard proof of every shard of every block of the segment in its node's pool, and PC 2 alone proves 13 shards per 10 minutes of about 600 blocks (2.2% coverage), so a run of 8 consecutive proven blocks never occurs: node 1 at 04:16Z reads segmentsInWindow pending 55, proven 0, unproven 20, paidSegments 0, and PC 2's app log (run win-1ccfe586-20261005-235130, 04:11Z to 04:16Z) reads every 42 s "aggregator: segment N..N+7: waiting for shard proofs N/0 ... N+7/0 in this node's pool", all 8 missing, each segment then past its 600-DAA deadline unproven. No fault in the node, the app or the record path; the fast-time harness passed because its shards ran at 90% coverage. Meanwhile the aggregator share (a tenth of every block's pool credit) stays in the escrow; shard payouts continue; miners and block watchers see nothing. What ends it: coverage at 8 consecutive blocks, 47 mining 5090-class cards with the shard loop as shipped in 0.3.11 (13 shards per 10 minutes a card), 18 mining cards through the chain mode, or 6 (approximate) proving-only cards, at 1 block/s on empty blocks (the fleet table above), or the segment length lowered on a small devnet (`proving_v1_segment_blocks`, a consensus param, so a digest change). No PC 2 job and no 0.3.12 item follow from this; the open item is the fleet, not the code.
## The empty `/api/state` reply (6 October 2026)
The aggregation-cost agent's jobs read the two bytes `{}` from `/api/state` on PC 2 at 22:22Z, 22:41Z and 00:18Z (0.3.10 and 0.3.11); the 21:01Z reply was full. Cause, from the app source and node 1's RPC: `ProvingState.paid_wei` is a `u128`, and serde_json's `to_value` refuses a u128 over u64::MAX (18,446,744,073,709,551,615 wei, 18.45 IGN); `state_json()` turned that refusal into `json!({})` with no log line. A paid shard is 1.23 IGN on average (node 1, `igneum_getProvingStatus`: 814.64 IGN over 663 shards at 00:3xZ), so the fifteenth paid shard after an app start empties the reply. PC 2's prover was blind to the root-owned socket from 20:00:56Z to 22:01Z (paid_wei stayed 0, hence the full reply at 21:01Z), proved from 22:02:13Z, and crossed 18.45 IGN inside its first 15 paid shards, before 22:22Z. Every app restart resets the counter, so the reply comes back for about 15 shards and goes again.
What it means: the dashboard on a proving machine shows nothing within about 12 minutes of its prover's first payout; every PC playbook that reads a card from `/api/state` fails the same way (the agent's job 5 reads settings.json instead). Mining, proving and payouts are untouched; it is the status page only.
Fix on the app branch: `paid_wei` serialises as a decimal string (the dashboard already reads it with `Number()`), `state_json` logs `[error] state_json: ...` once instead of answering `{}`, and the reply on any future serialisation error carries `error` and `version` rather than nothing; unit test `a_paid_total_over_u64_max_still_serialises_the_whole_state`. Not in 0.3.11 (that tree closed at 22c2363, master 630da6b, published); 6714a45 heads 0.3.12, the morning's first cut, app only, before PC 1's relaunch (coordinator, counter-asic-2-rollout.md 8a); until then the workaround is settings.json for the card keys.
## Aggregation cost (5 October, night)
the project lead, 5 October 2026: "fix everything else in the numbers tonight". Branch `agg-cost`; every number in `docs/bench-log.md`, "aggregation cost on the RTX 5090", with its job id. The proof statement and the pinned guests are unchanged: every existing fixture proof still verifies (`verify-segment` 0.027 s on the Mac, 0.036 to 0.041 s on PC 2).
| What | Before (5 October evening, `chain-pc2-pv1c`) | After (5 October night) |
|---|---|---|
| Chained aggregation, the card mining | 9.6 to 9.7 s a block | 9.6 to 9.8 s a block, the same (job `agg-cost-pc2-1`, phase A); the miner's presence is the whole cost: 2.1 to 2.2 s a block with the card to itself, 1.7 s unchained |
| Shard proof (empty shard), the card mining | 7.3 to 7.7 s | 7.4 to 7.8 s; 1.9 to 2.2 s with the card to itself |
| The miner's slowdown of the prover | 3 to 4x (against 4 October) | measured on the same fixtures 2 min apart: shards 3.6x, chained aggregation 4.5x, a block 4.2x |
| Where the time goes | not profiled | the host's share 0.000 s (the prove call is everything); the GPU server prints no timings; the second deferred proof (the chain rule) costs 0.4 to 0.5 s alone and 1.7 to 1.9 s under the miner; the prover alone keeps the card busy 15.8% of the time, the miner 93.9% |
| SP1 knobs (`SP1_WORKER_VERIFY_INTERMEDIATES=false`) | not tried | no gain: 7.8 s against 8.1 s over four aggregations, inside the spread; the shape knobs would change the recursion keys the pinned verifier accepts |
| Batch fold (K blocks in one aggregator call) | not estimated | estimate from the measured step costs: 0.9 s a block alone and 3.8 s mining at K = 4, 0.7 and 2.8 s at K = 8 (0.25 s per further deferred proof alone, 1.8 s mining); a tree fold gains nothing. A new pinned guest and program id either way, so not tonight |
| Two prover processes on one card | not tried | closed on SP1 6.8.1: both share one GPU server socket, run slower together (6.9 s a block against 4.1) and the second dies with the first (`early eof`) |
| The miner's kernel length (`--batch-log2` of the CUDA worker, 2^B nonces a launch; job `agg-cost-pc2-6`, the 5090 alone with the job's own miner) | not tried | 2^22 (the default) and 2^20: 18.1 and 18.0 s a block, 10.0 to 10.4 s a chained aggregation, 104 MH/s; 2^18: 15.6 s, 8.8 s, 99 MH/s (minus 4%); 2^16: 11.1 s, 6.1 to 6.2 s, 84 MH/s (minus 19%), reproduced |
| The GPU time-slice policy (`nvidia-smi compute-policy --set-timeslice`) | not tried | "Not Supported" on PC 2 (driver 13.3, Windows): closed |
| The chosen combination | the defaults | the defaults stay: batch-log2 22 and SP1's default knobs. The one knob that moves the prover (2^16) costs a fifth of the hash rate all the time for a prover that is busy a few seconds a minute on the devnet; it is the project lead's trade, not a default (below) |
Reading. The per-block aggregation is 2.1 s and a block 4.1 s on a 5090 that only proves, 9.7 and 17.5 s on one that also mines; no knob, fold or stream on tonight's SP1 changes the first pair, and only the miner's kernel length changes the second, at 1 MH/s per 0.37 s of block time. So "under 3 s a block" and "under 1.5x" are met on a card that is not mining and are not reachable on one that is. What that means per tier: a 5090 that mines and proves delivers a proven empty block every 17.5 s (6 cards for 1 block/s), the same card proving only every 4.1 s (2 cards, plus the shard work of full blocks: the fleet table above), and a batch fold of the aggregator (a new pinned guest) would bring the proving-only card to about 2.7 s a block and the mining one to about 12 s. What is being done: the app and host defaults are left as measured; the plan's open decision for the project lead is whether a card that holds a shard assignment should drop to 2^16 for the proof's minute (1.6x faster proof, 19% of its hash rate for that minute) or whether proving-only cards carry the aggregation (the clean 2.1 s), and the batch fold goes on the next pin's list. The state class found on the way (`/api/state` answering `{}` once `paid_wei` passes u64::MAX, fixed on the app branch at 6714a45) is in the bench log with the rest.