igneum/docs/benchmarks/proving-e2e.md
igneum-labs 0511d038be Evidence page, end-to-end proving standard, funding plan, security-budget stress, payment routes
docs/evidence.md and site/evidence.html: 28 public claims with one of five status labels (7 designed, 3 implemented, 18 tested by the team, 0 reproduced externally, 0 reviewed independently), version or commit, the reproducible test, the result with date and machine, and independent verification (none yet for every row). Evidence link in the homepage nav and the generated pages' nav.
docs/benchmarks/proving-e2e.md: replaces the 20-second shard gate with three fixed workloads, job-received-to-accepted-proof latency, cost per proof, the eligible card list with mining and proving reported separately, the verbatim acceptance standard and the three-unrelated-operator protocol.
docs/plans/funding.md: cost, what is funded (founder's means, the client's 1% fee once there is mining), what waits on revenue, what pauses.
docs/analysis/security-budget.md: emission through six halvings at three price inputs, miners and provers separate from burns, the USD 1M floor and the year it is crossed.
docs/design/payment-routes.md: Mermaid flowchart and table of every flow, with operator, app, team and protocol revenue labelled.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 22:35:03 +00:00

152 lines
13 KiB
Markdown

# End-to-end proving benchmark standard
Version 0.1, 3 October 2026. This standard replaces the "a 12 GB card proves one shard in about 20 s" gate (litepaper Proving, roadmap phase 2, spec 5.1). It is the phase 2 pass mark and the published measurement the prover customer brief points at. Nothing in it has been run. Status: designed.
## 1. Why the shard gate was the wrong gate
A per-shard time is met by shrinking the shard. The shard planner cuts a segment at transaction boundaries so each shard's proving cost is at most `S_p` (`docs/design/execution-layer.md` 5.1). `S_p` is the project's own number. Halve it and every shard proves twice as fast, the 20-second mark is passed, and nothing about the chain's capacity has changed: there are twice as many shards, twice the aggregation work, twice the witness traffic, and the per-shard fixed costs (program load, witness fetch, proof emission) are paid twice. The backlog, which is what a user and a customer feel, can grow while the gate reads green.
The trap is shrinking the shard. The standard below prevents it by fixing the unit of measurement outside the planner's reach: a whole published block, proven end to end, from the moment a prover sees the job to the moment a full node accepts the proof. The planner's choice of `S_p` is inside the measurement. A smaller shard that helps shows up as a shorter end-to-end time and a flat backlog; a smaller shard that only moves cost shows up as the same time and a rising backlog.
## 2. Fixed published workloads
Three workloads, published as files under `tools/bench-e2e/workloads/` with a genesis state, a transaction list and a SHA-256 of both, so every operator proves the identical bytes. The files do not exist yet; the execution engineer writes them with the phase 2 implementation. Each workload is one chain block's segment.
| Id | Workload | Content | Why this one |
|---|---|---|---|
| W1 | Transfers | 200 simple IGN transfers between 200 pre-funded accounts: 21,000 gas each, 4,200,000 gas total, 200 pgas each at the prototype table (40,000 pgas, to be re-stated at calibration) | The cheapest block to prove per transaction. Sets the floor of fixed cost per shard and per proof |
| W2 | Contract calls with keccak | 20 calls to a published contract: each call runs the `hashLoop` pattern of `tools/evm-smoke` at 200 iterations (98,900 gas estimated with the fold on the v3 simnet) plus one storage write with an event; about 2,000,000 gas, pgas from the calibrated table | Keccak is the opcode a zkVM pays most for relative to native execution. This block is the one a per-transaction gate would hide |
| W3 | Full budget | A block whose proving cost is exactly the consensus proving-cost budget `B_p`: the W2 call repeated until the planner refuses the next one, padded with W1 transfers to the gas limit | The worst block the chain can accept. If W3 cannot be sustained, the budget is wrong, not the gate |
Rules for the workloads:
1. The genesis state and the transaction bytes are frozen with the proof-system version. A new `ProofSystem` version (design 5.6) republishes all three with new hashes; results under one version are never compared with results under another.
2. `B_p` is set from this benchmark, not the other way round. The first run uses the provisional `B_p` of the phase 2 implementation; W3 is regenerated when `B_p` changes and the standard records the value used.
3. Nobody tunes the planner against the workloads. The planner's parameters (`S_p`, the shard count rule) are fixed in the release before the run and published with it.
## 3. The measured path
The clock starts when a prover's client receives the job and stops when a full node that did not produce the proof accepts the proof record. Every stage is timestamped by the prover and the verifying node, and every stage is reported, because an operator who proves fast but queues slow is not sustaining anything.
| Stage | Starts | Ends | Who stamps it |
|---|---|---|---|
| Queueing | Job received (shard assignment seen, or segment executed for an open shard) | A card starts on it | Prover client |
| Transfer | Witness request sent | Witness bytes complete and checked against the segment hash | Prover client |
| Proving | Card starts | Shard proof emitted | Prover client |
| Aggregation | First shard proof of the segment available | Segment proof emitted (recursion over the shards and the previous segment proof) | Aggregating client |
| Verification | Proof record gossiped | A non-producing full node verifies it and the native-execution veto passes (design 5.5) | The verifying node |
| End to end | Job received | Accepted | Both; the difference of the two wall clocks is reported with the clock-sync method used |
The wrapped proof (Groth16 or Plonk over bn254, design 5.3) is a fourth stage measured separately for the bridge and light-client case (O-10.7). It is not in the end-to-end figure for the chain, because the chain accepts the segment proof.
## 4. Reported figures
Per workload, per hardware configuration, from a continuous 2-hour run in which the workload is issued at the chain's block rate (one segment per second at launch) so that the provers must keep up, not catch up.
| Figure | Definition |
|---|---|
| Median | The 50th percentile of end-to-end time over the run |
| Slowest 5% | The 95th percentile |
| Slowest 1% | The 99th percentile |
| Per-stage medians | Queueing, transfer, proving, aggregation, verification, each at p50 and p99 |
| Failure rate | Jobs whose proof was never accepted, divided by jobs issued. A failure is counted whether the cause was a crash, an out-of-memory, a bad proof or a timeout |
| Retry rate | Jobs proven more than once by the same operator before acceptance, divided by jobs issued |
| Backlog | Shards issued and not yet accepted, sampled every 10 s for the 2 hours, reported as the series, its maximum, and the slope of a least-squares line through the second hour. A slope above zero is a growing backlog |
| Throughput | Segments accepted per second over the second hour, against the issue rate |
A run is reported whole or not at all. A run that is stopped early is reported as a failure with the time it stopped.
## 5. Full cost per proof
Cost is reported per accepted segment proof, in dollars, with each input stated so that a reader can substitute their own tariff. Nothing is netted against rewards; this is cost only.
| Component | How it is computed | Stated inputs |
|---|---|---|
| Electricity | Card and host power at the wall, measured with a meter over the run, times the tariff, divided by accepted proofs | The tariff. The standard reports two: $0.12 per kWh and $0.30 per kWh, chosen as a low and a high domestic rate, approximate. The operator also reports their own |
| Host amortisation | Purchase price of the card and the host it needs, divided by an assumed life of 3 years at 24 hours a day, times the run's hours, divided by accepted proofs | Purchase price (a receipt or a listing on the run date) and the 3-year life |
| Bandwidth | Witness bytes in plus proof bytes out, times a per-gigabyte price, divided by accepted proofs | The per-gigabyte price. The standard reports $0.01 per GB and $0.10 per GB, approximate; a metered residential line reports its own |
| Failed work | Electricity and amortisation spent on jobs that were not accepted (failures, retries, work superseded by a faster prover), spread over the accepted proofs | Counted from the same logs; nothing extra |
| Total | The sum of the four, at both tariffs and both bandwidth prices, as a range | |
The cost of a 20% pool share or an external job fee does not appear here. Those are revenue and belong in `docs/analysis/security-budget.md` and `docs/design/payment-routes.md`.
## 6. Eligible consumer cards
Mining compatibility and proving compatibility are reported separately, because they are different workloads: mining is the random-access lottery hash (bench-log), proving is the zkVM. A card can be in one column and not the other.
| Card | Memory | Mining: lottery hash | Proving: this standard |
|---|---|---|---|
| NVIDIA RTX 5090 | 32 GB | Measured: 229 Mhash/s at 1 GiB, bit-exact (bench-log, 3 October 2026) | Not run |
| NVIDIA RTX 4090 | 24 GB | Not run | Not run |
| NVIDIA RTX 4070 | 12 GB | Not run | Not run; the design's reference card class ("a 12 GB card") |
| NVIDIA RTX 3060 12 GB | 12 GB | Not run | Not run; the phase 2 gate names a "3060-class card" (ledger P1) |
| NVIDIA RTX 3080 | 10 GB | Not run | Not run; below the 12 GB design floor, listed to measure the floor |
| AMD RX 7900 XTX | 24 GB | Not run (discrete AMD never run, O-1.15) | Not run |
| AMD RX 6700 XT | 12 GB | Not run | Not run |
| AMD gfx1036 (Ryzen 7 9800X3D integrated) | shared | Measured: 4.38 Mhash/s on 1 compute unit, bit-exact (bench-log) | Not eligible: shared system memory |
| Apple M5 Max (40 GPU cores) | 64 GB unified | Measured: 45.2 Mhash/s at 1 GiB, bit-exact (bench-log) | Not run |
| Apple M-series, 16 GB unified | 16 GB unified | Not run | Not run |
| Intel Arc A770 | 16 GB | Not run | Not run |
A card enters the eligible list for proving when three unrelated operators (section 8) have sustained W1, W2 and W3 on it with a flat backlog. Until then the list is a list of candidates, and the litepaper's "12 GB or more proves full shards" is a design target.
## 7. The acceptance standard
Verbatim, the pass mark for phase 2:
> At a declared workload and hardware configuration, independent operators can sustain the advertised throughput without a growing proof backlog.
Applied to this standard:
| Term | Meaning here |
|---|---|
| Declared workload | W1, W2 or W3 by hash, under a named proof-system version |
| Declared hardware configuration | Card model, card count per host, host CPU and RAM, driver and toolchain versions, as a published fingerprint |
| Advertised throughput | The issue rate of the run: one segment per second at launch. The project advertises no higher rate than the one sustained in the published runs |
| Independent operators | Three unrelated operators as defined in section 8, each sustaining it on their own run |
| Sustain | The full 2-hour run with a failure rate under 1% and a slowest-1% end-to-end time under the proof lag the litepaper states (60 s at launch) |
| Without a growing proof backlog | The second-hour backlog slope is at or below zero and the maximum backlog is under 60 segments |
A pass is per workload and per configuration. "Phase 2 passed" means W1, W2 and W3 all passed on at least one consumer configuration from section 6 with at most one card per host. A pass on a multi-card host or a datacentre card is reported but does not pass the phase, because the chain's claim is consumer hardware.
## 8. Reproduction protocol
### 8.1 What counts as unrelated
Three operators are unrelated when every row holds for every pair:
| Test | Requirement |
|---|---|
| Person | Different natural or legal persons; none is the project, an agent of it, or paid by it for the run (a published fixed reproduction reward, equal for everyone and announced before the run, is allowed and disclosed) |
| Hardware | Bought separately; no shared host, card, rack or power meter |
| Network | Different autonomous systems, verified by the IP in the published logs; not the same residential ISP account |
| Location | Different physical sites |
| Software | The same published release by hash; nobody receives a private build |
| Money | No payment, loan or equipment between them or from the project, beyond the disclosed reproduction reward |
An operator declares each row in their report and signs it with the vote key that mined on the devnet under the same fingerprint, so a report is tied to a key with a history.
### 8.2 Steps
1. The project publishes: the release (binary hashes, source tag), the three workload files with hashes, the run script `tools/bench-e2e/run.sh`, the report format, the proof-system version, `B_p` and `S_p`, and the date range.
2. Each operator runs the 2-hour run for each workload on their configuration, with a wall-power meter and NTP-disciplined clocks, and keeps raw logs.
3. Each operator publishes the report (section 4 figures, section 5 costs, the hardware fingerprint, the declarations of 8.1) and the raw logs, signed.
4. The project publishes all reports unedited, pass or fail, next to its own run, on the bench page, and links each from `docs/evidence.md` row 15 and row 16. A failing run is as public as a passing one.
5. The eligible list (section 6) and the phase gate (section 7) are updated from the published reports only.
### 8.3 What the project may not do
- Change `S_p`, `B_p`, the planner or the workloads between the announcement and the end of the date range.
- Pick which reports to publish.
- Count its own run as one of the three.
- Advertise a throughput higher than the lowest of the three passing runs.
## 9. Open items
| Item | What closes it |
|---|---|
| The workload files and `tools/bench-e2e/` | Written with the phase 2 SP1 implementation (execution engineer) |
| The provisional `B_p` and `S_p` | The project's own first run, published before the external runs |
| The wrapped-proof stage on consumer hardware | O-10.7, measured in the same campaign |
| The reproduction reward and who pays it | `docs/plans/funding.md`, challenge reward row |
| Whether a 10 GB card can prove W3 at all | The RTX 3080 row of section 6 |