The sweep (main's item 1): 199 tracked text files, 783 lines. The founder's full name, first name and possessive become "the founder" (sentence starts capitalised); the lowercase operating-system user name in WSL paths and commands becomes <user>; the second owner login becomes "the second owner login"; the three earlier businesses and the two other brands become "the other business", "the earlier entity", "the earlier business" and "another brand"; the Chrome profile rule names the igneum.network profile, not the profile's label. The standing commit login igneum-labs is not a founder term here: the fresh-repository step renames it in the history (docs/plans/history-rewrite.md, tools/repo/fresh-repo.sh). The patterns never appear in plain text in the tree (a plaintext list would be the hit): tools/ci/founder-strings.b64 (perl regex, tab, a sample per row) is read by tools/ci/founder-strings-check.sh (every tracked text file, perl, known-failed first: the self-test plants each row's sample in a fixture and the hit must name the file), by tools/community/discord-hooks.mjs (the guard's founder and business rows; the test takes its fixtures from the samples) and by tools/repo/fresh-repo.sh (the business names of the rewrite rules). site/forbidden-strings.txt carries the same patterns as b64: lines, decoded case-insensitive by site/scrub.mjs and tools/ci/launch-gates-check.mjs (whose fixture now plants an encoded made-up name). The check runs in the gate's tree checks on every merge. Not in this commit, by main's word: the 105 commit messages and 40 personal-identity commits that need the history rewrite (listed, not run), and the secrets found by gitleaks over the history (reported with owners). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
66 KiB
Proving v1: segment records, the chain rule, the unproven rule; the 0.3.11 rollout
5 October 2026, from 18:55 UTC (the founder: "open the proving round asap"). Branches proving-v1 in the main repository
(worktree /Users/joshm/Projects/igneum-wt-proving-v1, from master a93199a) and in the fork
(vendor/igneum-node-pv1, from release-0.3.6 a24ab01a; to be rebased onto the 0.3.10 tip when it lands on
release-0.3.6). Status words follow docs/spec/00-overview.md 0.2. Every number here is in docs/bench-log.md
with its command. Nothing ships from this plan: it delivers branches, numbers and the rollout for 0.3.11.
The gap this closes
The litepaper says every block is proven within about a minute. Proving v0 (proving-v0.md, spec 7.7) proves
some shards: one prover (PC 2's RTX 5090) takes the newest shard assigned to it, about one shard every 30 s, so
under a tenth of blocks carry a proof; consensus does not require one; the aggregator guest (design 5.3) runs
on fixtures only. Proving v1 adds the aggregated segment record on chain (spec 7.8), the chain rule (segment N's
record verifies N-1, inside the proof by recursion), the unproven rule (a segment nobody proves in T seconds pays
nothing and may be skipped), the prover on by default on every machine that can prove, and the measurements
that say how many cards cover the chain.
The round, step by step
| Step | What | State |
|---|---|---|
| 1 | Prover on by default (app/igneum-app/src/provedefault.rs, engine.rs apply_prove_default): on at install when the machine can prove (NVIDIA card with 12 GB or more; WSL2 answering on Windows; Linux native; Apple silicon off until measured), never switching an explicit on back off; the Settings switch line and the tile line say why. Unit tests (5). The cost of proving on a mining machine: PC 2 job prover-cost-pc2-pv1 (5 min mining alone, 5 min with the prover, the GPU memory peak and the host RAM peak, the sp1-gpu-server's compiled SM targets) |
Implemented; the measurement is HELD (coordinator, 19:00Z): PC 2's RTX 5090 worker has been exiting on a pack seed mismatch since 18:35Z, so the first run's "mining alone" is 0 MH/s and void; re-run after the go |
| 2 | Segment aggregation: SegmentRecord (586 bytes, the aggregator guest's 340-byte statement inline), section IGNS before the shard section, p2p message 75 at protocol 15 (14 went to the EVM transaction relay in 0.3.10), the native block statement and the veto, the credit split (split_pool_credit), the payout at the carrier, igneum-miner sign-segment-record, the RPCs; host modes chain (consecutive fixtures), aggregate (live shard proofs from the pool, a run of blocks in one process) and verify-segment (the node's verifier, pinned aggregator key); the app's aggregator step (prover.rs aggregate_once) |
Implemented, unit-tested (consensus core 2 new tests, exec 2, params 1); the GPU measurement (N = 2, 4, 8 blocks on PC 2) is HELD with step 1; the Mac CPU run of --mode chain over 2 live blocks is the known-finished case |
| 3 | Coverage: tools/proving-v1/coverage.mjs (the proven-block share and the on-chain proof latency over a window from one node's RPC, the live page beside it) |
Implemented and run for 3 min (below); the 30-min window waits for the fleet |
| 4 | The chain rule and the unproven rule in consensus behind proving_v1_activation_daa (spec 7.8 items 2, 6, 7); unit tests; the fast-time 3-node harness tools/proving-v1/net.mjs (ports 29950+, suffix 956, trust mode) with the known-finished and known-failed cases |
Implemented; the harness run waits for the Mac build of the fork (vendor/igneum-node/target-pv1) |
| 5 | This plan: the rollout for 0.3.11 and the founder's decisions | Written below |
Numbers (every one from docs/bench-log.md, "proving v1: segment records ...", 5 October 2026 evening)
| What | Number |
|---|---|
| The prover's cost to a mining 5090 (the re-run with the fleet mining, hash rate from the miner's own STATUS lines) | 124.7 MH/s alone, 119.7 MH/s with the prover on: 5.0 MH/s, 4.0%, on empty shards at 1.4 a minute |
| GPU memory on the 5090: the prover alone (empty shards) / the miner and the prover together | max 13,816 MiB / max 15,590 MiB (the miner holds 3,396 MiB); a 24 GB 4090 has 8.4 GB of headroom, a 16 GB card 0.4 GB, a 12 GB card cannot do both on this build; the full-shard peak is the chain job's row |
| Coverage, 30-min window with the fleet mining, one prover | 2.4% of blocks, latency p50 44 s, p99 52 s |
| Host RAM | host used 25.6 GB of 63 GB; the WSL2 VM 7.9 GB working set |
| sp1-gpu-server 6.8.1 compiled targets (cuobjdump) | sm_80, sm_86, sm_89, sm_90, sm_100, sm_120 and compute_120 PTX: Ada (4090) is native, no JIT; nothing for AMD |
| Shards a minute, one 5090 through the app's loop (empty shards) | 1.6 |
Chain of 2 live blocks on the Mac CPU (--mode chain) |
shard 55.4 and 41.3 s, aggregate 52.0 s then 59.1 s with the previous proof, chain_len 2, final proof 1,272,909 bytes, verify-segment 0.032 s |
| Unit tests | consensus core 13, exec 8, app 5, all passing on the Mac |
| The harness (3 nodes, fast time, trust mode) | PASSED, 21 checks in 197 s on b177718e (N = 4) and 244 s on the final tree ece42979 (N = 8): paid 1.0 s after submit, every node agreeing; the fresh chain refused after a proven segment; the unproven segment skipped after its deadline; shards at 90% |
| Coverage, 3-min window, the degraded fleet (one card, the Mac verifier down) | 4.7% of blocks proven, on-chain latency p50 39 s |
The chain of 8 live blocks on the 5090, the card also mining (chain-pc2-pv1c) |
shard 7.3 to 7.7 s, first aggregation 7.9 s, every chained one 9.6 to 9.7 s; N = 2 in 32.6 s, N = 4 in 66.8 s, N = 8 in 135.6 s (17.0 s a block); the final proof 1,272,909 bytes whatever N, the record 586 bytes, verify-segment 0.037 to 0.040 s; GPU peak 16,751 MiB with the miner resident. Against 4 October with the miner stopped (aggregate 2.2 s): the miner slows the prover 3 to 4x |
| 5090-class cards for 100% at 1 block/s, measured rows | 47 with the loop as it is, 18 through the chain mode on mining cards, 6 (approximate) on proving-only cards, at empty blocks; 14 proving-only at one full shard a block; 45 at B_p: the table in the bench log |
The rule, in one paragraph (spec 7.8)
From the first chain block A at or above proving_v1_activation_daa, chain blocks form fixed segments of N
(proving_v1_segment_blocks). A segment record carries the aggregated proof of the segment's last block, whose
chain_len says how many consecutive blocks the recursion attests. Every node checks the record's statement
against its own native block statement (every field but the provers commitment and chain_len), the chain rule
(a proof that does not chain to the previous segment, chain_len = N, is valid only for the first segment or
after an unproven one) and the deadline (T = proving_v1_unproven_daa DAA seconds after the segment's last block;
a record carried later pays nothing). The pool credit of every attested block splits: proving_v1_aggregator_share_bps
to the aggregator, the rest to the shards as v0. A block is never invalid for lack of a proof; the mandatory rule
(spec 7.8 item 10) is Designed and off, with no switch yet.
Rollout for 0.3.11 (the digest handshake pattern of 0.3.9 and tonight's switches)
The switch moves the consensus digest only once it is set (consensus_digest: the four v1 fields enter the hash
when proving_v1_activation_daa != never), so a 0.3.11 node on the unswitched devnet keeps the 0.3.10 digest and
the rolling upgrade does not partition the network. The order, each step with its check:
- Rebase and build. Fork
proving-v1rebased onto the 0.3.10 tip onrelease-0.3.6; the six node suites and the app tests as PC 2 build jobs; the Mac node and the Windows exes by the Mac cross-build; the Linux node by PC 1; the HiveOS package republished from the same fork commit (rule: the HiveOS package carries the node of the release commit and the same override object,infra/hive, as 0.3.9'shive-sync-039ochecked it). - Pinned guests. The guest ids do not change in this round (shard
0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a, aggregator0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896, pinned 2026-10-05T16:20:38Z): the aggregator guest already carried the chain rule and only the HOST gained modes. So no provers-off drain is needed for the guests;--mode idon every machine after the update must print the same two ids, and the node now reads them at start (program_ids:IGNEUM_PROOF_PROGRAM_IDSor the verifier's--mode id) and names them in the native statement. - Hand nodes and the seed first, with the UNCHANGED override object (the digest stays): observer, node 1, the
seed on the 0.3.11 node; peers back within 20 s; the
proving v1: segment records from DAA score neverline in each log. - Manifest and apps.
publish-manifest.sh --version 0.3.11with the unchanged object;update-nowto every app; every machine on 0.3.11 with a DAA score and a hash rate (the watcher takes the commit as an argument). The app's prover default applies at the first start on 0.3.11: every NVIDIA machine with WSL2 goes on; the log lineprover default: ...on each. - The switch. When every node runs 0.3.11: publish the object with
proving_v1_activation_daa= H (24 h ahead, the rule offee-switch-devnet.md) and the three parameters; read the expected digest on a scratch node first;update-now; the hand nodes and the seed with the same object; the digest sweep; the first paid segment record (igneum_getSegmentRecords) andigneum_getProvingStatus.v1.segmentsInWindowafter H. - The mandatory rule stays off: no switch exists for it yet; it gets one when the measured share is one.
The decisions (Decided 5 October 2026, delegated: the founder, "I have no idea for most of this stuff so do a lot of research and deploy what is absolute best")
Each with its rule, its number and its evidence. The deploy is the DEVNET through 0.3.11 (not the public testnet).
What the other networks do (read 5 October 2026, 20:05 to 20:15 UTC; every figure from the page named, else labelled approximate)
| Network | Unit proven | Deadline | What a miss costs | Who is paid what | Measured latency |
|---|---|---|---|---|---|
| Taiko Alethia (L2BEAT page, protocol v2.1.0 notes) | a batch of blocks | proving window 2 h, cooldown 2 h (v2.1.0, February 2025); the Shasta inbox targets a 4-h proof submission cadence | the proposer's liveness bond is credited back in full when the batch is proved inside the window, half when outside; mainnet currently sets minBond and livenessBond to 0 | the prover earns the proving fee; two of four proofs needed (SGX Geth, SGX Reth, SP1, RISC0, at least one ZK) | 100% ZK coverage of mainnet blocks reached December 2025 (blockchain.news); preconfirmations 2 s |
| Boundless (docs.boundless.network, proof lifecycle) | one request | the requester's timeout (example 3,600 s) and a lock timeout (example 2,700 s); a reverse Dutch auction ramps the price from the minimum to the maximum over a ramp-up (example 300 s) | the locked collateral (example 5 ZKC) is slashed and used to pay another prover who fulfils the request | the prover's fee = the bid minus the market fee | not stated on the page |
| Succinct Prover Network (docs.succinct.xyz, SPN architecture and quickstart) | one request | the requester's deadline (the quickstart example: 10 minutes, 50 PROVE staked to bid, 100 PROVE maximum fee) | part or all of the winning prover's collateral slashed "according to protocol rules" | a reverse auction: the lowest bidder is assigned | "real-time", no number on the page |
| Aztec (docs.aztec.network economics; L2BEAT; forum) | an epoch of 32 blocks (38 min 24 s), a proof may cover one checkpoint (1 min 12 s) up to one epoch; maximum proof window 1 h 16 min | the epoch is declared failed only when its submission window expires | an unproven epoch is reorged out (no reward); proposals under discussion remove bonds and pay every prover that delivers on time | 400 AZTEC a slot: 70% sequencers, 30% provers (120 AZTEC), provers' share by an activity score | the public testnet proved by community provers (zkcloud blog), no page number |
| zkSync Era, Linea, Scroll (eco.com comparisons) | a batch | none on chain (the operator proves) | none | the operator | proof latency about 30 min (zkSync Era), 75 min (Linea), 90 min (Scroll), approximate |
Reading. Nobody pays an aggregator as a separate role: Aztec's 30% goes to whoever delivers the epoch proof, Taiko's fee to whoever proves the batch, the markets to the request's winner. Deadlines run from 10 minutes (Succinct's example) through 1 h (Boundless' example) to 2 h (Taiko) and 1 h 16 min (Aztec's maximum window); a miss forfeits the reward or part of a bond, and the slashed value goes to the prover who steps in (Boundless). Igneum has no bond on shards by decision (spec 7.2 item 4), so the forfeit here is the reward only.
The decisions
| Decision | Decided | Rule and number | Evidence |
|---|---|---|---|
proving_v1_segment_blocks (N) |
8 | Aggregation is a fixed cost per block, not per segment: 9.6 to 9.7 s for every chained block on a mining 5090, 7.9 s unchained (chain-pc2-pv1c), so N buys nothing in card time; it sets the record cadence and the forfeit. At N = 8 and 1 block/s a record every 8 s, a 1.27 MB proof gossiped every 8 s (159 KB/s per path, half of N = 4's 318 KB/s), and a missed segment forfeits 8 blocks' aggregator share (8 x 0.088 IGN at today's credit). The chain for 8 blocks cost 135.6 s cold on a mining card (66.8 s for 4), a fifth of T; pipelined per block it is 17 s after the last block. Aztec proves 32 blocks (38 min) as one; 8 blocks at 1 block/s is 8 s of chain, so the record lands well inside the minute the litepaper promises |
bench-log "proving v1" chain rows; Aztec economics page |
proving_v1_unproven_daa (T) |
600 DAA s (10 min) | T = p99 x 10: the measured block-to-carried-record latency of a shard record is p99 52 to 62 s (two 30-min windows), a cold chain of 8 adds 136 s and relay plus inclusion 10 to 40 s, about 240 s worst case; 600 leaves 2.5x on that and equals the 600-block record window of v0, so nothing is payable past it either way. Succinct's example deadline is the same 10 minutes; Boundless' example 1 h, Taiko 2 h, Aztec up to 1 h 16 min: Igneum's blocks are 1 s and its proofs seconds, so the shortest of the field. The forfeited aggregator share of an unproven segment STAYS IN THE POOL ESCROW (it is never paid, as an unproven shard's part today): no burn and no roll-over, the rule the pool already has, and the escrow is what later proofs are paid from | coverage rows; chain-pc2-pv1c; the table above |
proving_v1_aggregator_share_bps |
1,000 (a tenth) | The aggregator's card time per block is 9.7 s on a mining card against 4 x 10.6 s of shard proofs at B_p (19% of the card time) and 2.5 s against 42.5 s with the card to itself (6%); on tonight's empty blocks it is half the card time. A tenth of every attested block's pool credit sits between the two full-block ratios, pays a role no other network pays separately (Aztec pays its 30% to whoever delivers the epoch; the markets pay the winner), and leaves the shard provers 90%, which the fast-time harness showed paid exactly (shardWei 90% of the credit). The pool's 20% emission share itself is unchanged (spec 2.5, 5.3) |
chain-pc2-pv1c; bench-log 4 October 5090 rows; the harness |
proving_v1_activation_daa (H) |
the devnet tip + 14,400 at publish (4 h at 1 block/s), set by the 0.3.11 publisher in the same override object as program_class_v3 |
tonight's rule for consensus switches (the coordinator, 5 October 2026); the digest moves only once H is set, so the rolling update does not partition | |
| Aggregator sortition | none in v1: the first valid record carried wins | design 5.3's VRF draw (O-7.3) with one or two aggregators on the devnet changes nothing; the segment grid and the deadline already bound the race; revisit when a second aggregator exists | |
| Apple silicon default | off | the gate was "a shard under 60 s with the miner running": the M5 Max CPU took 41.3 and 55.4 s for EMPTY shards under tonight's load and 272 s for a 200-pgas shard on 4 October; a full shard at S_p was never under 60 s. Settings switches it on |
bench-log "proving v1" CPU chain row; 4 October CPU rows |
| The prover profile per card and the 12 GB and 16 GB gates (the founder: "make sure we can prove on 12gb cards"; "is there any way we can make 12gb cards mine and prove?") | measured on PC 2, the rows below | the SP1 6.8.1 GPU server reads ELEMENT_THRESHOLD, HEIGHT_THRESHOLD, SHARD_SIZE and the SP1_WORKER_NUM_*/BUFFER_SIZE knobs from the environment it inherits (sp1-core-executor-6.8.1/src/opts.rs, sp1-prover-6.8.1/src/worker/config.rs); the app passes a profile per card (provedefault.rs) and the host forwards it |
the sweep job memsweep-pc2-pv1 and the miner-on run |
The resume path (5 October 2026, the 0.3.11 app): POST /api/resume on 0.3.9 re-armed only FAULTED cards (stop_miners("paused") clears every slot's restart_at), so a healthy paused card stayed "off" at 0 MH/s until the app was relaunched: PC 2 at 21:25:11Z (the aggregation-cost job's pause and resume; [ok] mining resumed then 0.00 MH/s, waiting for 20 minutes), the Mac that afternoon. Now every slot without a live worker is re-armed and its pack exported again before the start, and 90 s later resume_check logs resume: <card> is not mining 90 s after resume (state ..., pid ...) for every enabled card without a hash rate (engine.rs, three unit tests: the state machine, the 21:25:11Z case against the old rule, the check).
The prover-floor agent's first sweep (job floor-sweep-1, 22:34 to 22:38Z, PC 2's 5090, the miners stopped, this plan's per-point recipe, its patched sp1-gpu-server 5568108b built for sm_86, sm_89 and sm_120, every proof VERIFIED by the unpatched pv1 host): the control at upstream's sizes reproduces the curve above (empty shard 13,892 MiB and 2.2 s; the v1 shard 20,516 MiB and 4.2 s); with the core element threshold at 2^26 the v1 shard proves as four core shards in 5.3 s at 12,708 MiB and the empty shard at 12,772 MiB; 2^25 gives 12,836 MiB at 8.5 s; 2^27 gives 15,396 MiB. The 12.7 GB left is the server's Setup (five recursion keys pre-built at a fixed 2^27 capacity plus the shrink and core keys: 9.7 GB before the first shard), which its patch v2 sizes to the need. Decided for the 12 GB profile: the split that lands under 11 GB wins (5.3 s a shard is inside the loop's own 25 to 30 s of carriage and 100x inside T); 2^27 is the second profile only if v2 leaves it under 11 GB with the miner's 1.8 GB beside it. The 12 GB row stays OPEN until the final pair (alone and beside the miner) lands and the on-order 3060 runs it.
A self-built CUDA server (the 12 GB path), before 0.3.12 (consequences C26)
If the prover-floor agent's rebuilt sp1-gpu-server (the Setup sizes cut, built on PC 2 under WSL2) proves a shard under 11 GB, it becomes a shipped artefact and needs its own row of rules before 0.3.12: it is built from a pinned SP1 source tag with CUDA_ARCHS covering sm_86, sm_89 and sm_120 (the 12 and 16 GB tiers are Ampere and Ada, not only the 5090's Blackwell; one card family per measured row), by the packaging path that builds the Windows payload (PC 1's build job for the Linux binary, the Mac signs the manifest as it does the DMG), lands in the DMG and the WSL2 package beside the host as wsl2/bin/sp1-gpu-server with its sha256 in payload-inputs.json, is named in evidence.md beside the prover rows ("prover built from SP1 at "), is rebuilt and re-measured at every SP1 upgrade, and ships only after --mode verify-segment and --mode verify on proofs it made show the pinned verifying keys unchanged (the server changes allocation, not the circuit; the ids 0x2b1a81cb... and 0x474678f3... must still verify them). The 12 GB claim itself waits for the on-order RTX 3060 to run that server on the same fixtures and recipe as the curve; until then the public line stays at 24 GB.
The root-socket class on PC 2, the two times: 20:00:56Z (my chain job's root run; the live prover failed with Connect(PermissionDenied) until the socket was gone) and 21:25:24Z (the aggregation-cost job's root run; the prover stayed dark through the 0.3.10 restart at 21:49:41Z until socketfix-pc2-pv1 removed the root-owned /tmp/sp1-cuda-0.sock at 22:01:16Z; the next shard, block 89011, was proven at 22:02:13Z and paid, and every shard since). The permanent fix in the 0.3.11 app tree: every committed playbook that runs a prove mode as root carries pkill -f sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock at its start and end, tools/ci/prover-socket-check.sh (in ci.yml) fails a playbook without them, and the app's prover names the cause in its log line when the host reports PermissionDenied. The app itself cannot remove a socket another user owns, so a job written outside the tree must still follow the rule.
A prover job on a shared card runs as the app's user or cleans its socket (pkill -f sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock at the start and the end; tools/ci/prover-socket-check.sh): the root-socket fault of 20:00Z, bench-log.
The prover profiles: the tiers from the S_p curve (bench-log, "proving v1", the sweep, the miner-on pair and the curve)
The GPU server of SP1 6.8.1 sets the memory, not the shard: a floor of 13.9 GB for an empty shard, 20.4 GB for a full shard at the adopted v1 budget (30,000 pgas, 4.7 M cycles), 28.3 GB for the prototype shard (6.75 M pgas, 60 M cycles); the miner adds 1.7 GB when it shares the card; no environment knob moves the floor and the server has no options of its own; the witness is 5 to 22 KB a shard and never binds.
| Card | Alone | Beside the miner | Default (provedefault.rs) |
|---|---|---|---|
| 32 GB (RTX 5090) | the prototype shard, 28.3 GB, 10.8 s; the v1 shard 20.4 GB, 4.3 s | the prototype shard 30.1 GB, 33 s; the v1 shard 22.2 GB, 13.2 s | on, mine and prove, today |
| 24 GB (RTX 4090, 3090) | the v1 shard 20.4 GB; the prototype shard does NOT fit (28.3 GB) | the v1 shard 22.2 GB measured on the 5090's allocation (2.3 GB spare on a 24 GB card; approximate for the card itself) | on, mine and prove, with the line "until the devnet's fee switch its shards are the prototype size, which needs 32 GB, so this card proves from the switch on" |
| 16 GB (RTX 5080, 4080) | the shipped server: an empty shard only (13.9 GB); the patched server v3 b37defef at threshold 2^27: the v1 shard 12,915 MiB and 4.3 s, the prototype shard 13,459 MiB and 16.8 s (measured by the prover-floor agent on the 5090's allocation, job floor-sweep-3, 00:13 to 00:17Z 6 October; 2^27 + 2^26 gives 16,115 MiB, over the card) |
the shipped server: nothing (15.7 GB for an empty shard); the patched server at 2^27 beside the miner (the 5090 mining at 95%, 338 W, same card; floor-sweep-4, 00:35 to 00:38Z): the v1 shard 14,786 MiB total with the miner's 3,833 MiB resident inside it, the server's own 10,953 MiB, 17.4 s a shard; on a 16 GB card that is 10.95 GB server + 1.7 GB miner = 12.7 GB plus the display, under the 15.0 GB line |
off on the shipped server, with the line; on (mines and proves, 17 s a v1 shard, 4.3x the alone time) once the patched server ships (the packaging row below) |
| 12 GB (RTX 3060, 4070) | the shipped server: nothing (the floor is 13.9 GB, and the server refuses the card outright); the patched server v3 b37defef at threshold 2^26 (SP1_GPU_ELEMENT_THRESHOLD=67108864, the 12 GB profile): the v1 shard 10,291 MiB and 5.7 s, an empty shard 9,971 MiB and 3.3 s, the card's 2,089 MiB idle inside the peak and the server's own working set about 8.2 GB (6,535 MiB after Setup), so a 12 GB card proves alone with about 3 GB over it (measured by the prover-floor agent on the 5090's allocation, floor-sweep-3; the on-order RTX 3060 run is pending) |
measured beside the miner (floor-sweep-4): at 2^26 the v1 shard 12,066 MiB total with the miner's 3,833 MiB inside, the server's own 8,233 MiB, 24.4 s; the empty shard 11,586 MiB, 13.0 s; 2^25 gains nothing (12,066 MiB, 43.5 s). On a 12 GB card that is 8.2 GB server + 1.7 GB miner = 9.9 GB before the display, over the 9.0 GB line the founder set, so mine-and-prove on 12 GB is NOT claimed |
off on the shipped server; "proves alone" on the patched one once it ships (the packaging row below); mine-and-prove stays off on 12 GB (9.9 GB plus the display, over the 9.0 GB line). The public gate stays "12 GB proves; 16 GB mines and proves", both on the patched server, both pending a run on the card itself. The founder's "make sure we can prove on 12 GB cards" is answered on the 5090's allocation and OPEN on the card itself: the prover-floor agent (branch prover-floor, 5 October night) read SP1 v6.8.1's GPU server source (sp1-gpu/crates/prover_components/src/builder.rs lines 35 to 39): it reads the card's memory, adds 4 and panics under 24 ("Unsupported GPU memory ... must be at least 24GB"), and builds its core (ELEMENT_THRESHOLD 2^28 + 2^27 elements + 2^21), recursion (2^27), shrink (2^25) and wrap (85 M element) provers at Setup whatever the mode, which is the 13.9 GB floor; no knob reaches them, so the fix is a server rebuilt from source on PC 2 (WSL2, nvcc 12.8, CUDA_ARCHS=120) with those sizes cut, measured on the same fixtures and recipe as the curve above (D2 carries the curve) |
| under 12 GB | nothing | nothing | off, mine only |
| AMD-only and Apple machines | nothing on the GPU: no zkVM proves on an AMD GPU today (docs/analysis/amd-proving.md, branch amd-prove); the CPU prover is about 5 minutes a shard at a 30 GB RSS whatever the shard size (PC 1, bench-log "the SP1 CPU prover on PC 1") |
off, "mines and does not prove"; the only non-NVIDIA path with a shipped backend is RISC Zero's Metal prover behind the ProofSystem seam (a second guest and pinned id, a verifier for both formats, no shared aggregation): an open item, not 0.3.11 |
The three profile numbers the coordinator asked for, as measured: under 9.0 GB does not exist on this build (floor 13.9); under 15.0 GB mine-and-prove does not exist for any full shard (the v1 shard alone is 20.4); the full profile is the 32 GB card. The fleet table's "proving-only" rows therefore read 24 GB cards at the v1 budget and 32 GB cards at the prototype budget. Shards per block at the v1 budget: 1 on tonight's empty chain, 2 to 4 on blocks with transactions (B_p 120,000 = 4 x S_p); the aggregation count is one per block whatever the shard count (the chained recursion), so the aggregation-cost agent's target is per block.
The aggregation-cost agent's first rows (branch agg-cost, 5 October 2026 night, the same four live blocks on PC 2): with the miners paused an empty shard proves in 1.9 to 2.2 s and an aggregation in 1.7 to 2.2 s with the card 15.8% busy; mining, 7.4 to 7.8 s and 7.9 to 9.8 s at 93.9% busy, so the miner's kernels take the card and the prover runs 3.6x (shards) to 4.5x (aggregations) slower beside them; its batch-size curve is still open. That puts a proving-only card at about 4 s per empty block (one shard and one aggregation), 4 cards for an empty-block chain at 1 block/s, against 18 mining cards.
The re-plans of block 344 at 2.25 M and 4.5 M pgas peak at 28.3 to 28.4 GB alone (the server's buffers step up between 4.7 M and 20 M cycles and are flat to 60 M), so no shard size between the v1 budget and the prototype one changes a tier; with the miner the adopted shard proves 3.1x slower (13.2 s against 4.2 s) and the chained aggregation 9.7 s against 2.5 s: a mining 24 GB card delivers one adopted-size shard plus one aggregation in about 23 s, inside T by 25x.
Segment-aligned proving (6 October 2026, the founder: "find a way to solve this")
The fault was the work order, not the rules. The shipped loop took the newest open shard each pass (prover::choose), so one prover scattered one block in about 45 across the segment grid and no segment ever had all its blocks proven: pending 56, proven 0, unproven 20 at 06:28Z. No consensus parameter moves.
The change (app branch, commit ce8f34a, app and host):
| Part | What it does | Where |
|---|---|---|
| Whole-segment claiming | the work list (lookback 600, the record window) grouped into whole untouched segments: every block present, every shard open (past its 10-DAA exclusive window), unpaid, not in our pool | app/igneum-app/src/segments.rs whole_segments |
| The choice | candidates inside their deadline by a margin (240 DAA, or 1.5x the last segment's wall time), ranked by FNV-1a of (first block, this prover's key hash): deterministic per prover, different between provers, so several provers spread over the candidates with no coordinator; an attempted segment is not retried | segments::candidates, rank, need_daa |
| The statement check | for the best three: executed and pending; fresh only when the previous segment is not paid and no verified record of it waits in the pool (the chain rule would refuse a fresh record once that one pays); chained (--prev) when the previous is paid and its proof is in this node's pool |
prover.rs pick_segment |
| The work | one export to the segment's last block, one fixture per block, one host run --mode chain --save-shards [--prev] that proves every shard and aggregates the segment in one process (one key setup), then every shard record signed and submitted (the shard payouts, 90% of the credit) and the segment record signed and submitted (the aggregator share, 10%) |
prover.rs prove_segment |
| The fallback | when no whole segment qualifies, the per-block path as shipped | prover::choose |
| The host | --mode chain takes --prev <aggregated.bin> (chain_len continues; a previous proof that is not the parent block's is refused by number and parent hash) and with --save-shards writes per-shard records (statement, proof sha256, file, time) into the results |
proving/igneum-prove/host/src/main.rs run_chain |
| The tile | "Segments: N proven whole, M paid (x IGN to the aggregator), the last in T s" and the segment path's last line; the state carries segments_submitted, segments_paid, segment_paid_wei, segment_last_s |
ui/app.js, state.rs |
Tests: the grid, the grouping (missing block, paid shard, our shard in the pool, exclusive shard), the margin and the attempted set, the per-key order (deterministic, different between two keys), the margin from the last time: 6 unit tests in segments.rs; 120 app tests, 8 core and 9 host tests pass. The host flags were run on the Mac's CPU first (bench-log, 07:12Z to 07:17Z): per-shard records written for a chain of 2, then a chain of 1 continued from it (base_chain_len 2, final chain_len 3).
Item 2, "own pool only", answered from the source: a relayed proof record carries its proof bytes (protocol/flows/src/v10/proving.rs: IgneumProofRecordMessage { record, proof }, 8 MB bound, handed to the pool with local = false), and so does a segment record (message 75). So "this node's pool" is every record relayed to it, and any aggregator can already fold any 8 proven blocks it has received; no node or consensus change is needed for that. What does limit carriage: a block template carries only entries the node's own verifier marked verified (template_segment_section, template_section), so a node with the verifier Off (node 1) never carries a record and a node with Trust carries unverified ones; PC 2's node runs the host and verifies. With one prover the carrier is PC 2's own next block.
What one 5090 completes (arithmetic from the 5 October rows, the measurement below replaces it): a chain of 8 empty blocks took 135.6 s cold beside the miner; one export and one key setup per segment instead of eight; so about one segment every 150 to 200 s, 9 to 12 segments an hour out of 450 (2 to 3%), against none. The aggregator share of a segment is 8 x 0.088 IGN = 0.70 IGN plus the 8 shards' 90% share; the forfeited share of the other 97% stays in the escrow until the fleet grows (47 mining cards or 6 dedicated provers for 100% at 1 block/s).
PC 2 measurement (job segments-pc2-pv1c, 07:52Z to 08:24Z, tools/proving-v1/pc2-segments.ps1; bench-log "the segment-aligned prover beside the miner"):
| Figure | Value |
|---|---|
| Whole segments proven in 30 min, one 5090 beside its miner | 9, one every 210 s (export 1.5 s, cut 45 s, chain 160 s: 8 shards 63 s, 8 aggregations 80 s) |
| Shard records accepted and paid | 72 of 72, 0.91 to 2.72 IGN a shard (90% of the credit), carried 180 to 226 blocks after the block |
| Segment records accepted | 0 of 9: refused by the chain rule as shipped (below) |
| Miner's cost | 117.86 to 104.90 MH/s, 13.0 MH/s = 11.0% (the shipped prover: 5.0 MH/s, 4.0%, for 2.8x fewer shards) |
| GPU memory peak | 16.5 to 17.6 GB with the miner resident; 24 GB tier unchanged |
The second fault, found by the measurement: the chain rule as shipped. check_segment_record accepts a fresh record (chain_len = N) only when the previous segment is UNPROVEN at the carrier. A segment's own deadline is its last block's DAA plus 600, the previous segment's deadline is 8 DAA earlier, so a fresh record is valid for 8 DAA (8 s on devnet) and must be proven before and carried inside them. The refusal on the live node, 9 times: "segment 114470..114477 does not chain to segment 114462..114469 (chain_len 8), which is pending until DAA 169681". The fast-time harness passed on 5 October because its chain continued from proven segments (case 2) and its fresh case ran exactly in that window (case 3, 4 DAA at N = 4). No prover-side move escapes it: the record must be proven, submitted and carried inside the window, which the 210-s proof cannot meet.
The fix (fork branch 0f0dda95, behind a switch, never by default): proving_v1_fresh_rule_daa; from it a fresh record is valid whenever the previous segment is not proven (pending or unproven) at the carrier; after a proven one a record must still chain. Two records that do not chain each attest their own blocks against the native statement, so nothing is lost but the longer proof chain, which restarts. In the consensus digest only once set (the v1 pattern), so a 0.3.12 node on the live devnet keeps the 0.3.11 digest until the override file sets it; igneum_getProvingStatus.v1.freshRuleDaa/freshRuleActive and igneum_getSegmentStatement.freshAdmissible report it. Tests: the digest moves once set; fresh after a pending segment passes from the switch, is refused before it, never after a proven one; the harness --fresh-rule 0 inverts case 3. This is a rule change in the execution layer, not a parameter tuning, and the code proved it unavoidable: the founder's "no consensus parameter change unless the code proves it is unavoidable" is met by the 8-DAA window above and the nine refusals. The 0.3.12 coordinator sets the height (the same tip + 14,400 rule) in the override object with the release.
Until the switch: the app (272b025) holds a refused segment record and offers it again every pass until the segment's deadline, so on the shipped rule it lands only if a carrier falls inside the 8-DAA window, and from the switch it lands on the first retry; the shard records (90% of the credit) land either way, 8 per segment.
Per tier, with this change and the switch:
| Card | What it does | Per 30 min, empty blocks (measured on the 5090, approximate elsewhere) |
|---|---|---|
| 32 GB mining and proving (5090) | 9 whole segments, 72 shards paid, 9 aggregator shares once the switch is set | miner 11.0% down; 72 x 0.9 IGN = 65 IGN of shard payouts measured, plus 9 x 0.70 IGN aggregator share from the switch |
| 24 GB mining and proving | the same path at the 16.5 to 17.6 GB peak measured; the fee-switch shard not yet measured on a 24 GB card | approximate: the 5090's numbers |
| 16 GB | mines and proves on the patched server only (prover-floor rows); segment path untested there | pending the prover-floor agent's build |
| 12 GB prove-only | proves alone on the patched server (10.3 GB); a dedicated prover takes 4 s a block alone (agg-cost rows), so about one segment every 40 s | approximate: 45 segments per 30 min, 6 such cards for 100% |
| A rig (several cards) | one prover loop per machine today; the segment path claims one segment at a time on the aggregation card | the per-card loop is the next item |
v1 live on devnet (6 October 2026, C47)
v1 active at 154,800 (crossed at DAA 154,814, 03:51:42Z); first segment record: none, because on a one-prover devnet no segment can be proven. The app's aggregator (aggregate_once, 0.3.11) needs a shard proof of every shard of every block of the segment in its node's pool, and PC 2 alone proves 13 shards per 10 minutes of about 600 blocks (2.2% coverage), so a run of 8 consecutive proven blocks never occurs: node 1 at 04:16Z reads segmentsInWindow pending 55, proven 0, unproven 20, paidSegments 0, and PC 2's app log (run win-1ccfe586-20261005-235130, 04:11Z to 04:16Z) reads every 42 s "aggregator: segment N..N+7: waiting for shard proofs N/0 ... N+7/0 in this node's pool", all 8 missing, each segment then past its 600-DAA deadline unproven. No fault in the node, the app or the record path; the fast-time harness passed because its shards ran at 90% coverage. Meanwhile the aggregator share (a tenth of every block's pool credit) stays in the escrow; shard payouts continue; miners and block watchers see nothing. What ends it: coverage at 8 consecutive blocks, 47 mining 5090-class cards with the shard loop as shipped in 0.3.11 (13 shards per 10 minutes a card), 18 mining cards through the chain mode, or 6 (approximate) proving-only cards, at 1 block/s on empty blocks (the fleet table above), or the segment length lowered on a small devnet (proving_v1_segment_blocks, a consensus param, so a digest change). No PC 2 job and no 0.3.12 item follow from this; the open item is the fleet, not the code.
The empty /api/state reply (6 October 2026)
The aggregation-cost agent's jobs read the two bytes {} from /api/state on PC 2 at 22:22Z, 22:41Z and 00:18Z (0.3.10 and 0.3.11); the 21:01Z reply was full. Cause, from the app source and node 1's RPC: ProvingState.paid_wei is a u128, and serde_json's to_value refuses a u128 over u64::MAX (18,446,744,073,709,551,615 wei, 18.45 IGN); state_json() turned that refusal into json!({}) with no log line. A paid shard is 1.23 IGN on average (node 1, igneum_getProvingStatus: 814.64 IGN over 663 shards at 00:3xZ), so the fifteenth paid shard after an app start empties the reply. PC 2's prover was blind to the root-owned socket from 20:00:56Z to 22:01Z (paid_wei stayed 0, hence the full reply at 21:01Z), proved from 22:02:13Z, and crossed 18.45 IGN inside its first 15 paid shards, before 22:22Z. Every app restart resets the counter, so the reply comes back for about 15 shards and goes again.
What it means: the dashboard on a proving machine shows nothing within about 12 minutes of its prover's first payout; every PC playbook that reads a card from /api/state fails the same way (the agent's job 5 reads settings.json instead). Mining, proving and payouts are untouched; it is the status page only.
Fix on the app branch: paid_wei serialises as a decimal string (the dashboard already reads it with Number()), state_json logs [error] state_json: ... once instead of answering {}, and the reply on any future serialisation error carries error and version rather than nothing; unit test a_paid_total_over_u64_max_still_serialises_the_whole_state. Not in 0.3.11 (that tree closed at 22c2363, master 630da6b, published); 6714a45 heads 0.3.12, the morning's first cut, app only, before PC 1's relaunch (coordinator, counter-asic-2-rollout.md 8a); until then the workaround is settings.json for the card keys.
Aggregation cost (5 October, night)
The founder, 5 October 2026: "fix everything else in the numbers tonight". Branch agg-cost; every number in docs/bench-log.md, "aggregation cost on the RTX 5090", with its job id. The proof statement and the pinned guests are unchanged: every existing fixture proof still verifies (verify-segment 0.027 s on the Mac, 0.036 to 0.041 s on PC 2).
| What | Before (5 October evening, chain-pc2-pv1c) |
After (5 October night) |
|---|---|---|
| Chained aggregation, the card mining | 9.6 to 9.7 s a block | 9.6 to 9.8 s a block, the same (job agg-cost-pc2-1, phase A); the miner's presence is the whole cost: 2.1 to 2.2 s a block with the card to itself, 1.7 s unchained |
| Shard proof (empty shard), the card mining | 7.3 to 7.7 s | 7.4 to 7.8 s; 1.9 to 2.2 s with the card to itself |
| The miner's slowdown of the prover | 3 to 4x (against 4 October) | measured on the same fixtures 2 min apart: shards 3.6x, chained aggregation 4.5x, a block 4.2x |
| Where the time goes | not profiled | the host's share 0.000 s (the prove call is everything); the GPU server prints no timings; the second deferred proof (the chain rule) costs 0.4 to 0.5 s alone and 1.7 to 1.9 s under the miner; the prover alone keeps the card busy 15.8% of the time, the miner 93.9% |
SP1 knobs (SP1_WORKER_VERIFY_INTERMEDIATES=false) |
not tried | no gain: 7.8 s against 8.1 s over four aggregations, inside the spread; the shape knobs would change the recursion keys the pinned verifier accepts |
| Batch fold (K blocks in one aggregator call) | not estimated | estimate from the measured step costs: 0.9 s a block alone and 3.8 s mining at K = 4, 0.7 and 2.8 s at K = 8 (0.25 s per further deferred proof alone, 1.8 s mining); a tree fold gains nothing. A new pinned guest and program id either way, so not tonight |
| Two prover processes on one card | not tried | closed on SP1 6.8.1: both share one GPU server socket, run slower together (6.9 s a block against 4.1) and the second dies with the first (early eof) |
The miner's kernel length (--batch-log2 of the CUDA worker, 2^B nonces a launch; job agg-cost-pc2-6, the 5090 alone with the job's own miner) |
not tried | 2^22 (the default) and 2^20: 18.1 and 18.0 s a block, 10.0 to 10.4 s a chained aggregation, 104 MH/s; 2^18: 15.6 s, 8.8 s, 99 MH/s (minus 4%); 2^16: 11.1 s, 6.1 to 6.2 s, 84 MH/s (minus 19%), reproduced |
The GPU time-slice policy (nvidia-smi compute-policy --set-timeslice) |
not tried | "Not Supported" on PC 2 (driver 13.3, Windows): closed |
| The chosen combination | the defaults | the defaults stay: batch-log2 22 and SP1's default knobs. The one knob that moves the prover (2^16) costs a fifth of the hash rate all the time for a prover that is busy a few seconds a minute on the devnet; it is the founder's trade, not a default (below) |
Reading. The per-block aggregation is 2.1 s and a block 4.1 s on a 5090 that only proves, 9.7 and 17.5 s on one that also mines; no knob, fold or stream on tonight's SP1 changes the first pair, and only the miner's kernel length changes the second, at 1 MH/s per 0.37 s of block time. So "under 3 s a block" and "under 1.5x" are met on a card that is not mining and are not reachable on one that is. What that means per tier: a 5090 that mines and proves delivers a proven empty block every 17.5 s (6 cards for 1 block/s), the same card proving only every 4.1 s (2 cards, plus the shard work of full blocks: the fleet table above), and a batch fold of the aggregator (a new pinned guest) would bring the proving-only card to about 2.7 s a block and the mining one to about 12 s. What is being done: the app and host defaults are left as measured; the plan's open decision for the founder is whether a card that holds a shard assignment should drop to 2^16 for the proof's minute (1.6x faster proof, 19% of its hash rate for that minute) or whether proving-only cards carry the aggregation (the clean 2.1 s), and the batch fold goes on the next pin's list. The state class found on the way (/api/state answering {} once paid_wei passes u64::MAX, fixed on the app branch at 6714a45) is in the bench log with the rest.
The fleet night (6 October 2026, from 11:50 UTC)
The rented fleet (branch gpu-fleet, tools/fleet/, raw logs under ~/Desktop/fleet/<instance>/): 11 cards for the
memory matrix (docs/analysis/prover-tiers-real-cards.md), 15 extras for the prover night, a p2p hub on RunPod with its
port public, the proving agent's three RISC Zero boxes. Every box: the 0.3.12 Linux node 83089544 on the ten-field
override (digest 7bd98cc4...), the 0.3.12 workers, the patched SP1 server, the cuda host, a throwaway payout address
made on the box, a per-card key label. The prover is tools/fleet/box-prover.py, a port of app/igneum-app/src/prover.rs
(whole-segment claiming, FNV spread over the fleet, --mode chain --save-shards [--prev], sign and submit per shard and
per segment, held fresh records offered again every pass).
| Row | What was found | Per tier | What is being done |
|---|---|---|---|
| 1, 13:03Z: a node that joined the devnet today never executes the chain | On every fleet node eth_blockNumber reads 0x0 and igneum_getProvingStatus tipDaa 0, v1 active=false, an hour after consensus synced (blocks = headers, synced=true, blocks flowing). At exec debug the follower says cannot find header edc4fa84... (genesis) on a proof-synced node and the queried hash does not have retention root on its chain on a node started over a copy of the observer's full datadir, archival or not: the exec state is memory only (igneum/exec/src/service.rs: _db_dir unused, IgneumDb::genesis() at start, the follower walks the virtual chain from genesis), and once the devnet's pruning point left genesis (today) no node can start that walk. The observer's own exec read tipDaa 0 at 12:05Z (before the fleet touched it) and node 1's reads tipDaa 0, both restarted for publish 2 at about 11:37Z: the hand nodes' exec layer has been dead since, and with it every work list the provers read |
Home miner (any card, any OS): a 0.3.12 install today mines but never proves, its wallet and the explorer against its own node read zero. Rig: the same. Pool user: nothing visible until the pool's own node restarts. The devnet: proving v1 produced nothing after the publish-2 restarts; the fleet night's "before the switch" window cannot exist on 0.3.12 | The coordinator's decision (13:15Z): no genesis restart; the proving agent builds the follower rebuild (walk the stored blocks where the full history is on disk, persist the exec state), tested on a copy of node 1's datadir tonight, then the hand nodes swap binaries and the PCs take a node-only 0.3.13. The fleet keeps the hub and the observer's full-history tarball (/root/fleet/share/observer-datadir.tgz on the hub, sha256 e67cc649...) for the phase-2 nodes, keeps the fifteen extras building until 14:15Z, and runs phase 1 and phase 3 meanwhile |
| 2, 12:24Z: the seed drops every new node every 30 s | P2P, route error: incoming route capacity for message type IgneumFinality has been reached (peer: 188.245.5.161:26611) then P2P Connected to outgoing peer 188.245.5.161:26611 at :06, :36, :06 on every fleet node; IBD through the headers proof restarts at each drop, so 4 of 11 nodes had 0 blocks after 35 minutes while the seed served 26 nodes at once; IBD with peer 188.245.5.161:26611 completed with error: peer connection is closed |
Every joiner with the seed as its only peer syncs in pieces; a rig the same once; pools unaffected | The hub (a RunPod 4090 with 26611 public) is every fleet node's second peer; the hub itself drew 23 fleet peers within minutes through the seed's address exchange. For the node: the IgneumFinality route's capacity against the per-checkpoint burst |
| 3, 12:38Z: the patched server hangs instead of failing when a profile does not fit | The 3080 (10 GB) and the 4060 Ti 8 GB at 2^27 (a 10.3 GB allocation): the server holds the card's limit at 0% for 568 and 904 s until killed; patch v5 (prover-floor) turns it into FLOOR abort: a device allocation failed at slop/crates/tensor/src/inner.rs:51 ... AllocError { size: 486586112 } and exit 70 in 13 s on the 3080 (the known-failed case of its gate) |
Every prover run needs a wall-clock timeout (the rig unit has one; the app's prover and this fleet's loop have one) | v5 is the server the fleet ships from here |
| 4, 15:37Z to 15:41Z: the 0.3.13 swap on the fleet | The hands switched to the 0.3.13 node (bb43e9a8) with the thirteen-field override (exec_restart_number 27276, exec_restart_hash bb45cf0d..., exec_restart_trust_daa 200000; digest b18ed271...) at 15:29Z; the fleet's boxes followed in two steps (tools/fleet/box-node-swap.sh, swap.py): the binary with the ten-field file (digest 7bd98cc4 kept, the exec layer blocked by design), then the file. Replay from the restart to an executed tip at the chain tip, per box: hub (RunPod 4090) 88 s, 3080 87 s, 4070-1 113 s, 4090-3 138 s, the 8x 4090 rig 138 s, A5000 163 s, 3090-4 188 s; a second set 15:45Z to 15:57Z (3090, 5090, 3090-1, 3090-2, 3090-3, 4090-1b) about 10 to 12 min each including a 438 MB datadir pull. Seven boxes failed the first file pass on a partial pre-pull of the full-history tarball (the swap now checks its sha256) |
A joiner on 0.3.13 with a full-history datadir executes within 1.5 to 3 minutes; a joiner without one still cannot (the fresh-join path, a snapshot from a peer, is the next item); every tier | The swap and the replay times are the fleet's measurement for the 0.3.13 release note |
| 5, 15:42:57Z: DAA 198,000, the fresh-record rule armed; 15:43:42Z: a 229-block reorg reset every executing node | The hub's first PoW accepted ... daa 198000 at 15:42:57Z. One-block selected-chain reorgs at 15:42:32, :43, :53 and 15:43:05Z (heights 135,065 to 135,088), then at 15:43:42Z selected-chain reorg: 229 chain blocks removed, unwinding to height 134884 (our tip blue score 194,395, the last removed 194,117), reorg deeper than the snapshot ring; replaying from genesis, genesis executed, then exec not synced: the executor is at chain block 0 and the bodies below this node's retention root are gone: executedTip 0, persistedTip 135,028, blocked null. The same on every fleet box that was executing and on the observer and the seed (the shipper's reading). The chain ran two-sided for about a minute after the switch |
Every node operator whose node was executing at 15:43Z (home miner, rig, pool) read a zero wallet and an empty work list until a restart through the exec-restart path; the miners of the 229 losing blocks lost those rewards; a deep reorg after a snapshot ring on a pruned node is the class: the fallback must be the exec-restart point, not genesis (the proving agent's item) | The fleet restarted its boxes through the exec-restart path (90 to 190 s each, the night loop tools/fleet/night.py re-runs it on any box whose executed tip falls to 0 for 150 s) and the provers started on the executed tip from 15:56Z |
| 6, 16:38:45Z: the first segment records after the fresh-rule switch, accepted and paid | The fleet's RTX 5090 (Vast, a restart-path 0.3.13 node, the segment host from the proving-v1 bundle, the prover loop tools/fleet/box-prover.py exporting from the exec restart block) claimed segment 137142..137149 at 16:35:30Z, cut 8 fixtures, ran --mode chain --save-shards (8 shards and 8 aggregations on a card that also mines), had 8 of 8 shard records accepted and the fresh segment record accepted at 16:38:45Z: 195.8 s claim to acceptance, aggregator share 1.5076 IGN; a 4090 (137102..137109) followed at 16:39:51Z in 211.2 s, 1.6150 IGN. The hub, a node on the hands' real state, read paidSegments 2, paidSegmentWei 3.12 IGN, pool 2 entries 2 verified, paidShards 1,482 to 1,516 at 16:42Z: carried and paid |
A 24 or 32 GB card that mines earns a segment's aggregator share about every 200 s on top of its 8 shards' 90%; six such cards cover about a quarter of the chain's segments (1.8 of 7.5 a minute); 47 mining 24 GB cards or 6 proving-only ones cover it (the fleet table's arithmetic, now with a measured 200 s) | The six restart-path boxes prove through the night; the hourly rows carry segments per hour and the chain's coverage |
| 7, 16:50Z: one exec state, and the export gap on a snapshot-recovered node | The state roots at 27,276 (0xed27bb2d...), 130,272 (0xf0a762da...) and 130,273 (0x2e22e029...) are identical on a restart-path node and on a snapshot-recovered node (export segment root and eth_getBlockByNumber agree), so the two recovery paths of 16:00Z to 16:25Z are one chain state and records from either side are valid on the other (row 6 confirms it). But on a snapshot-recovered node every igneum_exportSegments (from 0, 27,276 or 130,272) carries 41 to 42 accounts, the exporter's port state root differs from the node's at the first segment, and from 0 the 27,276 segments below the restart carry zero roots: the prover kit cannot cut against the hands' kind of node. The fleet's first report of this (16:52Z) called it two states; the roots corrected it 10 minutes later |
An operator whose node recovered through the snapshot cannot run a prover until the export carries the state (the 0.3.14 account-dump export, the proving agent's item); a node that joined through the exec-restart path proves today | The nine snapshot boxes mine without provers tonight (each prover loop had pulled a 100 MB export every 6 s for nothing); the roots above are the 0.3.14 pin's numbers |
| 8, 17:45Z to 18:45Z: finality at 78 to 93 voters (the 38-pod wave, 1,748 MH/s for USD 20.44/h) | 116 checkpoints determined, 110 locked in the hour; lock delay p50 1.30 s, p90 1.54 s, max 16.6 s (n 107); the first certificate carries p50 59 voters (max 80) and the fold lifts it to 86 of 93; lock share p50 90.1% of active voters (min 76.9%), 77.2% of all (min 67.0%); 1.65 certificate replacements per checkpoint (max 5). The detector's dry run on the window: no alert, no event, the 5090 band 108/115/128 MH/s (n 1,555). The first hour of the wave (16:50Z to 17:45Z) is void: pgrep -f igneumd-0313 in box-wave.sh matched its own launching shell, so no wave node started for 55 minutes (about USD 20 of pods); the fix is the anchored pattern and the CI check tools/ci/pgrep-self-match-check.sh |
Home miner: a lock lands about every second block at 1 block/s, 0.2 s later at 93 voters than at 56; what to watch past 100 voters is the certificate's size (the fold to 86 signatures), not the vote count. Rig: the same. Pool: a pool's one node votes once for all its members; its weight is the day's blue blocks, so a pool of 38 pods is one voter with 40% of the weight | the rows are the fleet night's finality baseline for the 100-voter question; the certificate-size series continues in the hour marks |
| 9, 18:14Z: pool-v0 under load, 10 members, 0 shares | igneum-pool (PPLNS, port 4463) with 10 wave pods as members: every member's GPU shares were refused WORKER MISMATCH because the pool's job carried no program class and the member's worker hashed class v2 while the devnet's templates are class v3; 0 shares accepted in 25 minutes, the pool's stats page alive. The pool agent rebased the pool-mode miner onto 0.3.14 (pool-v0-rebase c4c92e8: the job carries program_class and era_seed, the member re-checks shares with the template's class); the rerun with both new binaries is scheduled for tonight after the block-rate runs |
Pool user: on the old pool binary every share is wasted work; on the rebased pair a member without a node takes the day and the dataset size from the pool's seeds line. Home miner and rig: unaffected (solo mining never touched the pool) | the rerun's counts (shares accepted per member inside the first vardiff interval, mismatches 0, rejected 0) go to main |
| 10, 17:25Z to 17:38Z: the 0.3.14 canary on the live devnet | the release's Linux igneumd (4c6b129d) on five restart-path boxes for ten minutes: 0 new rejected blocks on every box, exec state roots equal to the hub's at a common height, a segment record paid on the new binary, every node on the new version: PASS on all four criteria; the first run's FAIL was the gate counting the hub (whose trait is about 3 rejects a minute) and a cut -c1-140 that truncated the digest line before the match |
Operator on the restart path: the swap is a binary change with the data dir kept, seconds of downtime, no replay. Snapshot-path operator: the same binary, the snapshot recipe unchanged | the canary form (canary-next.sh) is the fleet's release gate; the 0.3.15 run is in flight on four boxes at 19:41Z |
| 11, 18:31Z to 19:35Z: the class v4 rehearsal on igneum-devnet-400 (16 nodes, 38 pods joining, one stale 0.3.13 box) | The flip by miner signal landed at DAA 1,200 (19:00:5xZ) with the identical line on every node ("epoch 2 ... seed block f0606e20..., threshold 9500 bps, 599 of 599 blue blocks"), one program id per epoch across boxes (epoch 2 0x24304f0788ea9408, epoch 3 0xcc266b4f5dbc3447, epoch 4 0x634018bab5e5f283; the Mac's igneum-pow show agrees and the v3 id differs), 0 PoW rejected all run, the stale box refused by digest at every connect (12 by 19:46Z) and refusing the sixteen-field file at parse ("unknown field program_class_v4_activation_daa"), the floor at 2,400 crossed with v4 in force and nothing to print. Two findings outside the protocol: the plan's CPU engine (5 kH/s a box) cannot move a chain whose genesis difficulty is 2^27 (the CUDA worker fixed it: 14 to 98 MH/s a box), and the 0.3.14 package's igneum-worker-cuda refuses a generator-4 pack ("not a generator version this worker runs (2 or 3)"), so the chain stood at DAA 1,202 from 19:00:5xZ until the release tree's generator-4 Linux worker (97e036e2) went on at 19:15Z; a worker restarted on its old --pack after --exit-on-seed-change also answers every job "epoch seed mismatch" until the pack is re-exported | Home miner, any card: at a class flip the old worker stops dead; the flip is safe only when the generator-4 worker ships in the hive, Mac and Windows packages before the signal window closes, and the app re-exports the pack at every seed change. Rig: the same, times eight. Pool: the pool's node decides the class; a member on an old worker mismatches every share | the cut's rule (publish 2 gated on every worker, not every node) is with the shipper; P1 PASSED is on master |
| 12, 18:39:40Z onward: the live devnet's finality paused at 93 voters | On every node read the last lock is checkpoint 6842 at 18:39:36Z; checkpoints keep being determined and no certificate is received or built anywhere after 18:39:40Z. The fleet's first reading (the hands down 18:2xZ to 19:42Z, the star topology) was wrong: the observer rows show 13 fleet keys stopped mining the live devnet between 18:27Z and 18:30Z when the rehearsal job took their GPUs (their live nodes stayed up and synced, but a voter's weight is its blue blocks), and with seven earlier leavers that was 42.7 percent of the frozen voter table; rule v3 then holds the pause for one full window, the first lock expected about 20:40Z | Home miner: a voter that stops mining stops counting within the window, and 10 percent of the weight leaving in an hour is the most the table absorbs without a pause. Rig and pool: a pool is one voter with its members' whole weight; its restart is the biggest single removal on the network | the standing-fleet rule from it (6 October 2026, 20:00Z): never remove more than 10 percent of the live devnet's 30-day weight in any hour; lib/standing.py weight_check gates every job that stops or shares a standing miner |
| 13, 19:14Z to 20:30Z: the 0.3.15 canary, FAIL on the first binary, the retry confounded | 713ef876 on four live boxes: every block a 0.3.15 node mined or relayed carried version 1026 (the class v4 signal bit stamped from its first block, no window set) and every 0.3.14 node answered "wrong block version: got 1026 but expected 2" and disconnected it (the hub: 45 such rejects in the first 14 minutes, 468 by 20:29Z), so a 0.3.15 node that fell behind could not re-sync (p2-3090-1: connected and dropped every 30 s for 32 minutes) and every 0.3.15 miner lost every block it found; FAIL, publish 1 held. Three more findings on the way back: (a) a 0.3.14 node whose datadir holds 1026 blocks keeps relaying them and stays a disconnected peer after the rollback, so the four canary datadirs are poisoned until wiped; (b) a pruned 0.3.14 node dies ("consensus/src/processes/sync/mod.rs:87 KeyNotFound(GhostdagCompact/0/)") when a peer syncing a gap asks below its retention (the hub twice, 19:58Z and 20:01Z, restarted by the standing supervisor); (c) a standalone igneum-miner keeps a dead template subscription after its node restarts (templates frozen, fetch_errors climbing, no submits). The shipper's 7961c5f1 gates the stamp on publish 2's object and fixes the serving side of (b); its retry on the poisoned boxes was confounded by (a) and (c), the clean retry runs on two untouched boxes | Home miner: an update that stamps a new block version before the network accepts it is a silent death (the app shows hashing, nothing is paid); the fix is the version gate on the object plus a datadir that never held a bad block. Rig: the same, times eight. Pool: a pool node on the bad version drops every member's share from the network's view | the canary form stays the release gate; the fresh-join line from a wiped datadir is the last read |
| 14, 20:00Z: the standing fleet | the founder's ruling (19:50 UK): rented cards stay up and are never destroyed on a job's end. 13 live-devnet boxes converted at 19:41Z (USD 3.50/h, USD 84/day), each under box-standing.sh (node, miner and prover restarted when gone, the recovery recipe on a dead exec, a status line every 10 min), lib/standing.py on the Mac (roster, check, update, re-rent in the same shape, the 10 percent weight gate), a standing block on the fleet page; the Devnet 2 six and the L4 and AMD cards owed as providers free them. The supervisor's own two faults tonight (it matched any igneumd, so it mistook the rehearsal node for the live one and restarted a canary box's dead node on the 0.3.15 file) are fixed in db58804 and 9a294b4 |
Operator of a standing box: the node comes back within a minute of dying, the miner with it, and no job takes its GPU without the weight gate | docs/plans/gpu-fleet.md carries the rule |
| 15, 19:58Z, 20:01Z, 20:42Z: the hub's live node died three times | Twice on a peer's sync request below its retention ("consensus/src/processes/sync/mod.rs:87 KeyNotFound(GhostdagCompact/0/)", while the rolled-back p2-3090-1 synced a 32-minute gap against it; the serving-side fix is in the 0.3.15 node), once on a full disk ("header_processor/processor.rs:534 IO error: No space left on device"): the prover's segment exports under /root/fleet/out/segs (50 to 500 MB a segment, never pruned by box-prover.py) had filled the hub's 60 GB and 3 to 46 GB on every standing box since 11:50Z; two boxes stood at 100 percent. The standing supervisor restarted the node each time (19:59:11Z, 20:02:36Z, 20:44:38Z; the third after 38 GB were freed by hand) and now prunes exports older than 20 minutes and trims the node log every ten minutes (353cc5f); a disk sweep every five minutes writes disk per box to the fleet page and posts one #incidents line per box per hour at 85 percent (disk-sweep.py); the exporter-side cap is the proving lane's |
Home miner: the hub is one of three public peers a fresh node dials; a dead hub means a slower first join and nothing lost; a home node's own disk is not at risk (the app's prover does not write exports). Rig: the same. Pool: a pool node that exports segments for its provers has the same disk clock | the three deaths' lines are in the hub's node.log (preserved on the box) and the finality pulls under the scratchpad |
| 16, 19:17Z to 20:58Z: the block rate on Devnet 2, 10 blocks a second against 1 | Run A (the fork's 10 blocks/s profile, fresh genesis, 42 cards, the seed on igneum-build-1, every miner dialling the seed only): 4.87 DAG blocks/s but 1.09 blue blocks/s, 77.6 percent red, tips 250 to 660, difficulty easing all hour, the exec follower at 0.05 blocks/s. Run B (1 block/s, same boxes): 1.0 blocks/s and under 2 percent red from minute six, tips 1 to 3, difficulty settled in six minutes, the follower at 0.46 blocks/s. The network lane's read: the reds came from node throughput (61 to 345 ms of CPU per accepted block), not the star | Home miner: the payout interval follows the blue rate, which the profile did not move (1.09 against 1.19 blue/s), so a 4070 at a 10 TH/s network waits about three days for a paying block at either rate; the pool, not the block rate, is the small card's shorter wait. Rig: 4 to 5 hours at 10 TH/s either way. Everyone: a wallet or prover on a 10 blocks/s chain would read state hours behind within the first hour at tonight's follower rate | docs/analysis/block-rate-devnet2.md: 1 block/s for the testnet and the launch, 10 behind three measured gates |
| 17, 21:52Z: two miners on one GPU after a restart | The read-back after publish 1 showed miner_up=2 on p1-4090, p1-a5000 and p2-4090-3: the publish killed the miner once, the supervisor's miner loop restarted it within ten seconds, and the supervisor's main loop, reading "no miner" in the same gap, started a second loop; two igneum-miner processes then shared one GPU at half rate each. The supervisor's miner loop now kills any other miner and worker before it starts its own (409bc3d); redeployed on the 14 standing boxes at 21:55Z |
Home miner: the app owns one miner per GPU and never sees this; a hand-run box with two miner loops halves its rate silently and the only sign is the STATUS line's MH/s. Rig: times eight. Pool: a member with two miners doubles its share submissions at half the rate each, the pool sees one member with a jittery rate | one miner per GPU is now the supervisor's invariant, not the operator's care |
| 18, 21:43Z to 21:56Z: two provers on one Devnet 2 box killed each other's GPU server | Repeated by-hand relaunches of box-prover.py on dn2-1 and dn2-2 left two instances on a box: each relaunch's "kill the server, remove the socket" took the other instance's SP1 GPU server away mid-proof ("CudaClientError: early eof" on every chain step from 21:43Z) and both wrote the same prover-state.json.tmp, so one lost the file (FileNotFoundError at the os.replace). box-prover.py now holds a pid file under its out directory and a second instance exits at once, and each process writes its own tmp (c29c6b9); the Devnet 2 provers relaunched clean at 21:54Z |
Home miner: the app's prover is one process by construction. Operator of a hand-run prover box: start the prover once; a second start now refuses with the first's pid instead of taking its GPU server down | the Devnet 2 gate's "paid segments" read waits on the first record from these provers |