docs: Devnet 3 proving pipeline end to end, 8 October 2026 (the fleet lane's document, landed by the research-landing hand on the coordinator's word)

The document as it stood in the fleet lane's worktree at 16:46:43 BST (the same second as gpu-fleet 9991d7ce, which carried
tools/fleet/pipeline-collect.py and left this file untracked). Documents only; the collector stays on gpu-fleet.
This commit is contained in:
igneum-labs 2026-10-08 15:51:26 +00:00
parent 61ab6bb79c
commit 714cf36138

View file

@ -0,0 +1,121 @@
# Devnet 3 proving pipeline, end to end, 8 October 2026
The external review's order (through the coordinator, 15:4x BST): every paid shard's time in each stage, separated, with the median and
the slowest 5 and 1 percent; the failure and retry rate; the queue depth over time; realised earnings and wasted work per hardware tier,
which cards complete paid work after the 10 DAA-second exclusive window, and how often a faster claimant takes an assigned prover's reward.
## The window and its limit
The measurement window is the whole of 8 October's proving on the rented fleet, 07:00:26Z (first claim) to 14:27:15Z (last claim), read
from every prover box's own log after Devnet 3 was turned off at 15:34Z. Paid work exists only between 07:46:55Z and 10:43:10Z: 93 paid
segments, 172.88 IGN. From 11:45Z the chain stalled at the class v5 crossing and then partitioned (every solo branch carried old-object
blocks, the network restarted from the stall sink at 15:27Z and was turned off at 15:34Z), so every segment claimed after 11:45Z was
submitted into a chain that never paid it. The hold's declared workload (150 tx/s) ran 11:42Z to 12:23Z with inclusion, then without, so
no paid shard carries a hold transaction: the stage columns below are the fleet's proving of the chain's own blocks, under the pre-stall
load (the DEX and faucet lanes, the hold's earlier steps), on the 0.3.24 node (5b673577). The window asked for, two hours under the hold,
does not exist in the record; this is the honest substitute, and the instrument is in place for the next chain.
## The instrument
Collector `tools/fleet/pipeline-collect.py` (hub-1's Devnet 3 node, `igneum_getProvingStatus` every 30 s; every prover's RESULT lines
every 5 min) and the per-box logs `/root/fleet/out/prover.log` written by `tools/fleet/box-prover.py`, pulled whole after the stop. Each
column names its lines.
| column | source lines in prover.log | how the number is read |
|---|---|---|
| assignment wait | `RESULT claim <t> segment A..B (n shards, fresh) margin=M tip=T` | not separable from the logs: a segment becomes claimable when its last block settles and the box claims on its next pass (passes every 15 s). The proxy recorded is the margin at claim, the DAA left before the deadline (600 DAA window): median 467, p5 (slowest) 557 is not a wait but an early claim. Block timestamps would give the wait exactly; Devnet 3 is off, so they are not read. |
| inputs | claim stamp to the chain's start (the `RESULT seg N chain` stamp minus its `wall`) | the export of the segment's records from the node, the pair check and the cuts |
| proving | `RESULT seg N chain <t> k shard records, chain_len c, proof b bytes, shards P s, aggregation A s, wall W s, peak MiB` field P | the shard proofs on the card (SP1 floor server) |
| aggregation | the same line's A | the segment chain over the shard proofs |
| verification | chain stamp to `RESULT seg N shards <t> accepted k of n` | the node's verification of each shard record at submission; it answers inside the second, so verification and submission are one column |
| inclusion and payment | `RESULT submitted <t> ... end to end E s` to `RESULT paid <t> ... after S s` | S is the prover's own clock from the record's acceptance to the payment read on its node |
| failure, retry | `RESULT seg N ... FAILED`, `RESULT segment_refused`, `RESULT unpaid`, `RESULT paid_other`, `record accepted on retry` | counted per kind |
| queue depth | `RESULT pass n <t> no whole segment inside the margin (worklist N entries, tip T)` | N is the node's assigned-shard worklist as the prover reads it on each idle pass |
## Stage columns, seconds, every segment that reached the stage
| stage | n | median | slowest 5 % | slowest 1 % | max |
|---|---|---|---|---|---|
| inputs (claim to export and cuts done) | 2,447 | 2.4 | 7.5 | 10.1 | 14.7 |
| proving (shard proofs) | 2,453 | 26.8 | 216.1 | 364.7 | 1,104.7 |
| aggregation (segment chain) | 2,453 | 22.8 | 44.2 | 96.4 | 156.6 |
| verification and shard-record submission | 2,453 | 0.0 | 2.0 | 2.0 | 10.0 |
| segment-record submission | 1,909 | 0.0 | 1.0 | 1.0 | 1.0 |
| claim to submitted (end to end) | 1,909 | 75.1 | 230.3 | 301.6 | 497.4 |
| inclusion and payment (submitted to paid) | 93 | 171.0 | 543.0 | 31,397 | 31,470 |
| claim to paid | 91 | 336.0 | 720.0 | 1,098 | 1,240 |
| peak GPU memory during the chain, MiB | 2,453 | 11,948 | 21,174 | 25,788 | 26,210 |
The two 31,000-second payments are segments submitted before the 02:4xZ pause and paid when the chain resumed; without them the
inclusion-and-payment p99 is 902 s. The assignment wait is not in the table (see the instrument row).
## Paid segments per tier
| tier | paid segments | IGN | proving median s | aggregation median s | payment median s | payment p95 s |
|---|---|---|---|---|---|---|
| RTX 4090 | 48 | 89.33 | 117.3 | 20.1 | 177.5 | 649 |
| RTX 3090 | 34 | 59.66 | 69.8 | 75.0 | 166.0 | 543 |
| L40S | 10 | 21.86 | 67.0 | 18.7 | 125.0 | 370 |
| RTX 6000 Ada | 1 | 2.03 | 20.2 | 18.4 | 116.0 | 116 |
| RTX 3060 (12 GB) | 0 | 0 | | | | |
The 3090's aggregation median (75 s) is three times the 4090's: the segment chain is memory-bound and the 3090 pays for it. The 4090's
proving median on paid segments (117 s) is above the all-segment median (27 s) because paid segments are the long ones (the short ones
were taken by a faster claimant, below).
## Outcomes per tier and wasted work
| tier | claimed | paid | stolen | refused | submitted, never paid | claimed, never submitted | work s paid | work s wasted |
|---|---|---|---|---|---|---|---|---|
| RTX 4090 | 2,515 | 48 | 98 | 93 | 1,322 | 959 | 7,025 | 183,524 |
| RTX 3060 12 GB | 313 | 0 | 0 | 0 | 34 | 280 | 0 | 6,765 |
| L40S | 273 | 10 | 25 | 28 | 147 | 63 | 1,196 | 26,026 |
| RTX 3090 | 263 | 34 | 8 | 30 | 152 | 41 | 6,484 | 34,855 |
| RTX 6000 Ada | 45 | 1 | 6 | 0 | 26 | 12 | 57 | 3,055 |
- "submitted, never paid" (1,681 segments) is the partition: records accepted into a chain that never settled them after 11:45Z. It is
the day's largest waste and is not a pipeline fault; it is the fault of the afternoon (the fleet record).
- "claimed, never submitted" (1,355): the chain step failed or the claim was abandoned. The failures are counted below.
- The 3060 tier completed no paid segment in 313 claims: at 12 GB the chain runs out of margin (its proving median on completed chains
is above the 4090's by the card's ratio, and a 4090 claimant finishes the same segment first), so the 12 GB tier is a miner, not a
prover, on this segment size. The 3060's 34 submitted segments were all after 11:45Z (never paid for the partition's reason).
- Wasted work is the end-to-end seconds of every claimed segment that was not paid; the fleet spent 254,000 card-seconds (70 card-hours)
on segments that did not pay against 14,800 (4.1 card-hours) that did. Before the stall the ratio was about 3 to 1 (the steals and
refusals below); after it, everything was waste.
## Failures and retries
- `RESULT seg N chain FAILED`: 900, median wall 2.7 s. 629 "NotFound: No such file or directory": the sm_89 floor tarball's
`igneum-prove-host` was a dangling link until the real host was served at 10:50Z (08:00 to 10:59Z, the fleet record's floor fault);
264 "invalid string length" (the host's proof-bytes string on the 48 GB cards' larger segments); 4 OutOfMemory (12 GB cards). After
10:50Z the NotFound class ended; the string-length class remains open for the node lane.
- `RESULT segment_refused`: 153. 62 "does not chain to segment N..N" (the previous segment's record moved under the claim), 54 "unproven:
the record is carried after the segment's deadline" (the chain step finished too late), 35 "segment already paid" (a faster claimant).
- Retries: 0 lines "record accepted on retry". The prover does not retry a failed chain; it drops the export and claims afresh.
- Failure rate on claims: 900 chain failures plus 153 refusals over 3,421 claims is 30.8 percent; without the floor-link class (fixed) it
is 12.5 percent.
## Steals: a faster claimant takes an assigned prover's reward
`RESULT paid_other`: 137 segments this box had claimed were paid to another key, 4.0 percent of claims (98 on 4090s, 25 on L40S, 8 on
3090s, 6 on the 6000 Ada). The time from this box's claim to the other key's payment read: median 306 s, p95 9,757 s (the long tail is the
same pause-and-resume as the payment column). The exclusive window (10 DAA seconds) does not hold a slow claimant's segment for it: a
second box that finishes its chain first is paid. The 3060 tier was never the winner and never the victim (it never finished).
## Queue depth over time
The worklist as the provers read it on idle passes, median entries per hour (tip in the line): 02Z 596, 05Z 596, 08Z 596, 11Z 713, 14Z
594. Bounded at about 600 entries (the 600 DAA unproven window times one entry a DAA) for the whole day; it did not grow, because a
segment leaves the list at its deadline whether proved or not. Growth would show the window itself lengthening; it did not.
## What this means
- With the floor link fixed, the pipeline's own stages are fast: inputs 2 s, verification and submission under 2 s, aggregation 23 s
(75 s on a 3090); the proving stage sets the pace, 27 s median and 216 s at the slowest 5 percent, and payment lands 171 s after
submission (543 s at the slowest 5 percent) when the chain settles.
- The waste is structural, not incidental: 70 card-hours wasted against 4 paid, dominated by the partition, then by the floor fault,
then by steals and late chains (13 percent of claims). Two changes would cut the pre-stall waste: a claim that is honoured for the
window it was granted (the steal rate goes to zero) and a 12 GB tier that claims only segments it can finish (the 3060's 313 claims
earned nothing).
- The next chain (igneum-devnet-4) starts this instrument from block zero; the two-hour window under the hold's declared workload is
the first measurement to run on it, with block timestamps read so the assignment wait becomes a real column.