igneum/docs/benchmarks/proving-e2e.md
igneum-labs 3f558ea67a Evidence page, end-to-end proving standard, funding plan, security-budget stress, payment routes
docs/evidence.md and site/evidence.html: 28 public claims with one of five status labels (7 designed, 3 implemented, 18 tested by the team, 0 reproduced externally, 0 reviewed independently), version or commit, the reproducible test, the result with date and machine, and independent verification (none yet for every row). Evidence link in the homepage nav and the generated pages' nav.
docs/benchmarks/proving-e2e.md: replaces the 20-second shard gate with three fixed workloads, job-received-to-accepted-proof latency, cost per proof, the eligible card list with mining and proving reported separately, the verbatim acceptance standard and the three-unrelated-operator protocol.
docs/plans/funding.md: cost, what is funded (founder's means, the client's 1% fee once there is mining), what waits on revenue, what pauses.
docs/analysis/security-budget.md: emission through six halvings at three price inputs, miners and provers separate from burns, the USD 1M floor and the year it is crossed.
docs/design/payment-routes.md: Mermaid flowchart and table of every flow, with operator, app, team and protocol revenue labelled.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 22:35:03 +00:00

13 KiB

End-to-end proving benchmark standard

Version 0.1, 3 October 2026. This standard replaces the "a 12 GB card proves one shard in about 20 s" gate (litepaper Proving, roadmap phase 2, spec 5.1). It is the phase 2 pass mark and the published measurement the prover customer brief points at. Nothing in it has been run. Status: designed.

1. Why the shard gate was the wrong gate

A per-shard time is met by shrinking the shard. The shard planner cuts a segment at transaction boundaries so each shard's proving cost is at most S_p (docs/design/execution-layer.md 5.1). S_p is the project's own number. Halve it and every shard proves twice as fast, the 20-second mark is passed, and nothing about the chain's capacity has changed: there are twice as many shards, twice the aggregation work, twice the witness traffic, and the per-shard fixed costs (program load, witness fetch, proof emission) are paid twice. The backlog, which is what a user and a customer feel, can grow while the gate reads green.

The trap is shrinking the shard. The standard below prevents it by fixing the unit of measurement outside the planner's reach: a whole published block, proven end to end, from the moment a prover sees the job to the moment a full node accepts the proof. The planner's choice of S_p is inside the measurement. A smaller shard that helps shows up as a shorter end-to-end time and a flat backlog; a smaller shard that only moves cost shows up as the same time and a rising backlog.

2. Fixed published workloads

Three workloads, published as files under tools/bench-e2e/workloads/ with a genesis state, a transaction list and a SHA-256 of both, so every operator proves the identical bytes. The files do not exist yet; the execution engineer writes them with the phase 2 implementation. Each workload is one chain block's segment.

Id Workload Content Why this one
W1 Transfers 200 simple IGN transfers between 200 pre-funded accounts: 21,000 gas each, 4,200,000 gas total, 200 pgas each at the prototype table (40,000 pgas, to be re-stated at calibration) The cheapest block to prove per transaction. Sets the floor of fixed cost per shard and per proof
W2 Contract calls with keccak 20 calls to a published contract: each call runs the hashLoop pattern of tools/evm-smoke at 200 iterations (98,900 gas estimated with the fold on the v3 simnet) plus one storage write with an event; about 2,000,000 gas, pgas from the calibrated table Keccak is the opcode a zkVM pays most for relative to native execution. This block is the one a per-transaction gate would hide
W3 Full budget A block whose proving cost is exactly the consensus proving-cost budget B_p: the W2 call repeated until the planner refuses the next one, padded with W1 transfers to the gas limit The worst block the chain can accept. If W3 cannot be sustained, the budget is wrong, not the gate

Rules for the workloads:

  1. The genesis state and the transaction bytes are frozen with the proof-system version. A new ProofSystem version (design 5.6) republishes all three with new hashes; results under one version are never compared with results under another.
  2. B_p is set from this benchmark, not the other way round. The first run uses the provisional B_p of the phase 2 implementation; W3 is regenerated when B_p changes and the standard records the value used.
  3. Nobody tunes the planner against the workloads. The planner's parameters (S_p, the shard count rule) are fixed in the release before the run and published with it.

3. The measured path

The clock starts when a prover's client receives the job and stops when a full node that did not produce the proof accepts the proof record. Every stage is timestamped by the prover and the verifying node, and every stage is reported, because an operator who proves fast but queues slow is not sustaining anything.

Stage Starts Ends Who stamps it
Queueing Job received (shard assignment seen, or segment executed for an open shard) A card starts on it Prover client
Transfer Witness request sent Witness bytes complete and checked against the segment hash Prover client
Proving Card starts Shard proof emitted Prover client
Aggregation First shard proof of the segment available Segment proof emitted (recursion over the shards and the previous segment proof) Aggregating client
Verification Proof record gossiped A non-producing full node verifies it and the native-execution veto passes (design 5.5) The verifying node
End to end Job received Accepted Both; the difference of the two wall clocks is reported with the clock-sync method used

The wrapped proof (Groth16 or Plonk over bn254, design 5.3) is a fourth stage measured separately for the bridge and light-client case (O-10.7). It is not in the end-to-end figure for the chain, because the chain accepts the segment proof.

4. Reported figures

Per workload, per hardware configuration, from a continuous 2-hour run in which the workload is issued at the chain's block rate (one segment per second at launch) so that the provers must keep up, not catch up.

Figure Definition
Median The 50th percentile of end-to-end time over the run
Slowest 5% The 95th percentile
Slowest 1% The 99th percentile
Per-stage medians Queueing, transfer, proving, aggregation, verification, each at p50 and p99
Failure rate Jobs whose proof was never accepted, divided by jobs issued. A failure is counted whether the cause was a crash, an out-of-memory, a bad proof or a timeout
Retry rate Jobs proven more than once by the same operator before acceptance, divided by jobs issued
Backlog Shards issued and not yet accepted, sampled every 10 s for the 2 hours, reported as the series, its maximum, and the slope of a least-squares line through the second hour. A slope above zero is a growing backlog
Throughput Segments accepted per second over the second hour, against the issue rate

A run is reported whole or not at all. A run that is stopped early is reported as a failure with the time it stopped.

5. Full cost per proof

Cost is reported per accepted segment proof, in dollars, with each input stated so that a reader can substitute their own tariff. Nothing is netted against rewards; this is cost only.

Component How it is computed Stated inputs
Electricity Card and host power at the wall, measured with a meter over the run, times the tariff, divided by accepted proofs The tariff. The standard reports two: $0.12 per kWh and $0.30 per kWh, chosen as a low and a high domestic rate, approximate. The operator also reports their own
Host amortisation Purchase price of the card and the host it needs, divided by an assumed life of 3 years at 24 hours a day, times the run's hours, divided by accepted proofs Purchase price (a receipt or a listing on the run date) and the 3-year life
Bandwidth Witness bytes in plus proof bytes out, times a per-gigabyte price, divided by accepted proofs The per-gigabyte price. The standard reports $0.01 per GB and $0.10 per GB, approximate; a metered residential line reports its own
Failed work Electricity and amortisation spent on jobs that were not accepted (failures, retries, work superseded by a faster prover), spread over the accepted proofs Counted from the same logs; nothing extra
Total The sum of the four, at both tariffs and both bandwidth prices, as a range

The cost of a 20% pool share or an external job fee does not appear here. Those are revenue and belong in docs/analysis/security-budget.md and docs/design/payment-routes.md.

6. Eligible consumer cards

Mining compatibility and proving compatibility are reported separately, because they are different workloads: mining is the random-access lottery hash (bench-log), proving is the zkVM. A card can be in one column and not the other.

Card Memory Mining: lottery hash Proving: this standard
NVIDIA RTX 5090 32 GB Measured: 229 Mhash/s at 1 GiB, bit-exact (bench-log, 3 October 2026) Not run
NVIDIA RTX 4090 24 GB Not run Not run
NVIDIA RTX 4070 12 GB Not run Not run; the design's reference card class ("a 12 GB card")
NVIDIA RTX 3060 12 GB 12 GB Not run Not run; the phase 2 gate names a "3060-class card" (ledger P1)
NVIDIA RTX 3080 10 GB Not run Not run; below the 12 GB design floor, listed to measure the floor
AMD RX 7900 XTX 24 GB Not run (discrete AMD never run, O-1.15) Not run
AMD RX 6700 XT 12 GB Not run Not run
AMD gfx1036 (Ryzen 7 9800X3D integrated) shared Measured: 4.38 Mhash/s on 1 compute unit, bit-exact (bench-log) Not eligible: shared system memory
Apple M5 Max (40 GPU cores) 64 GB unified Measured: 45.2 Mhash/s at 1 GiB, bit-exact (bench-log) Not run
Apple M-series, 16 GB unified 16 GB unified Not run Not run
Intel Arc A770 16 GB Not run Not run

A card enters the eligible list for proving when three unrelated operators (section 8) have sustained W1, W2 and W3 on it with a flat backlog. Until then the list is a list of candidates, and the litepaper's "12 GB or more proves full shards" is a design target.

7. The acceptance standard

Verbatim, the pass mark for phase 2:

At a declared workload and hardware configuration, independent operators can sustain the advertised throughput without a growing proof backlog.

Applied to this standard:

Term Meaning here
Declared workload W1, W2 or W3 by hash, under a named proof-system version
Declared hardware configuration Card model, card count per host, host CPU and RAM, driver and toolchain versions, as a published fingerprint
Advertised throughput The issue rate of the run: one segment per second at launch. The project advertises no higher rate than the one sustained in the published runs
Independent operators Three unrelated operators as defined in section 8, each sustaining it on their own run
Sustain The full 2-hour run with a failure rate under 1% and a slowest-1% end-to-end time under the proof lag the litepaper states (60 s at launch)
Without a growing proof backlog The second-hour backlog slope is at or below zero and the maximum backlog is under 60 segments

A pass is per workload and per configuration. "Phase 2 passed" means W1, W2 and W3 all passed on at least one consumer configuration from section 6 with at most one card per host. A pass on a multi-card host or a datacentre card is reported but does not pass the phase, because the chain's claim is consumer hardware.

8. Reproduction protocol

8.1 What counts as unrelated

Three operators are unrelated when every row holds for every pair:

Test Requirement
Person Different natural or legal persons; none is the project, an agent of it, or paid by it for the run (a published fixed reproduction reward, equal for everyone and announced before the run, is allowed and disclosed)
Hardware Bought separately; no shared host, card, rack or power meter
Network Different autonomous systems, verified by the IP in the published logs; not the same residential ISP account
Location Different physical sites
Software The same published release by hash; nobody receives a private build
Money No payment, loan or equipment between them or from the project, beyond the disclosed reproduction reward

An operator declares each row in their report and signs it with the vote key that mined on the devnet under the same fingerprint, so a report is tied to a key with a history.

8.2 Steps

  1. The project publishes: the release (binary hashes, source tag), the three workload files with hashes, the run script tools/bench-e2e/run.sh, the report format, the proof-system version, B_p and S_p, and the date range.
  2. Each operator runs the 2-hour run for each workload on their configuration, with a wall-power meter and NTP-disciplined clocks, and keeps raw logs.
  3. Each operator publishes the report (section 4 figures, section 5 costs, the hardware fingerprint, the declarations of 8.1) and the raw logs, signed.
  4. The project publishes all reports unedited, pass or fail, next to its own run, on the bench page, and links each from docs/evidence.md row 15 and row 16. A failing run is as public as a passing one.
  5. The eligible list (section 6) and the phase gate (section 7) are updated from the published reports only.

8.3 What the project may not do

  • Change S_p, B_p, the planner or the workloads between the announcement and the end of the date range.
  • Pick which reports to publish.
  • Count its own run as one of the three.
  • Advertise a throughput higher than the lowest of the three passing runs.

9. Open items

Item What closes it
The workload files and tools/bench-e2e/ Written with the phase 2 SP1 implementation (execution engineer)
The provisional B_p and S_p The project's own first run, published before the external runs
The wrapped-proof stage on consumer hardware O-10.7, measured in the same campaign
The reproduction reward and who pays it docs/plans/funding.md, challenge reward row
Whether a 10 GB card can prove W3 at all The RTX 3080 row of section 6