From 80845a48a7181b586eca294ef8d6aa11e7ef2bec Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 08:56:08 +0000 Subject: [PATCH] bench-log and analysis: route 2 core-only rows (alone and beside the miner), the aggregator's cost, the hand-off as a prover-protocol change, the per-tier consequences Co-Authored-By: Claude Fable 5.1 --- docs/analysis/prover-floor.md | 11 +++++++++++ docs/bench-log.md | 32 ++++++++++++++++++++++++++++++++ 2 files changed, 43 insertions(+) diff --git a/docs/analysis/prover-floor.md b/docs/analysis/prover-floor.md index 50df3e838..32780b571 100644 --- a/docs/analysis/prover-floor.md +++ b/docs/analysis/prover-floor.md @@ -195,3 +195,14 @@ the Setup keys (6,535 MiB in use after Setup with the idle inside, about 4.4 GB peak (about 3.8 GB): cutting further means fewer keys built at Setup or a smaller recursion program, not a shard knob. The first publish of sweep 4 was refused by the publisher's kit-path-check (the wiped-jobs-folder rule of 21:49Z); the measurement template now tests the chain job's kit before use. + +### Route 2 measured (jobs `floor-core-alone`, `floor-core-miner`, 08:47 to 08:53Z) + +Core-only at 2^26 on the v1 shard: 9,874 MiB alone (7,817 own) in 3.1 s and 12,834 MiB beside the miner (7,754 own) +in 12.1 s, against the full compressed run's 8,233 own; the core proof 14,379,043 bytes, CPU verify 0.447 s; the +aggregator's extra 2.5 s alone and 12.6 s beside a miner. The empty shard 6,377 own. The core proof's hand-off is +a prover-protocol change (the pool and `igneum_submitProofRecord` carry compressed proofs today), not one line. +The 12 GB mine-and-prove line (9.0 GB own) is not met by core-only at 2^26 (7.8 + 1.7 GB): the levers left are the +core threshold at 2^25 and 2^24 for core-only proving and patch v4's `SP1_GPU_MEM_RELEASE_THRESHOLD` (the pool +returns freed memory between shards; upstream holds the high-water mark for the process's life), both in +`tools/prover-floor/pc2-floor-core2-miner.ps1` for the next PC 2 window. The table is in the bench-log entry. diff --git a/docs/bench-log.md b/docs/bench-log.md index 25444d569..62355aad5 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1593,6 +1593,38 @@ card mines and proves at 2^27 (10.95 + 1.7 GB = 12.7 plus the display, under its on-order 12 GB card runs the same two points. Shipping it is the project's own signed build of SP1's prover at every SP1 upgrade (the proving plan's packaging row before 0.3.12). +Route 2, core-only provers (the project lead, 07:40Z: can a 12 GB card mine AND prove; jobs `floor-core-alone` 08:47:07 to +08:48:37Z with the miners stopped and `floor-core-miner` 08:50:29 to 08:53:13Z with the 5090 mining at 96%; the v3 +server b37defef, the pv1 host's `--mode shard` = the core proof, its CPU verify, then the compressed proof in one +run, the 1-s sampler split at the core RESULT's timestamp so the core-only peak is the sampler's maximum up to it; +every core and compressed proof VERIFIED by the unpatched host; the resident set inside every peak: 2,057 MiB +alone, 5,080 MiB with the 0.3.11 miner): + +| Threshold | Fixture | Core-only peak MiB alone (own) | Core s | Core proof bytes | CPU verify s | Compressed s (full peak) | Beside the miner: core peak (own) | Core s | Compressed s | +|---|---|---|---|---|---|---|---|---|---| +| 2^26 | v1 shard | 9,874 (7,817) | 3.1 | 14,379,043 | 0.447 | 5.6 (10,290) | 12,834 (7,754) | 12.1 | 24.7 | +| 2^26 | empty | 8,434 (6,377) | 1.5 | 6,935,657 | 0.212 | 3.2 (9,938) | 11,458 (6,378) | 5.9 | 12.9 | +| 2^27 | v1 shard | 12,850 (10,793) | 2.5 | 10,147,581 | 0.314 | 4.1 (12,850) | 15,874 (10,794) | 8.5 | 17.5 | +| 2^27 | empty | 9,394 (7,337) | 1.4 | 5,603,823 | 0.170 | 2.4 (9,970) | 12,482 (7,402) | 4.8 | 9.8 | + +Reading. Stopping at the core proof takes 0.4 GB off the v1 shard's own working set at 2^26 (8.2 to 7.8 GB) and +1.5 GB off the empty shard's, not more, because the server still builds the compression keys at Setup (4.4 GB own +in use before any shard, mostly the allocator pool's high-water mark from the key commits) and at 2^27 the core +phase is the peak itself (the shard's own buffers). The aggregator's extra cost per shard, compressed minus core on +the same card: 2.5 s alone and 12.6 s beside a miner at 2^26, 1.6 s at 2^27, 1.0 to 1.7 s for empty shards; the +hand-off is 5.6 to 14.4 MB a shard (4 to 11x the 1.27 MB compressed proof), verified by the aggregator's CPU in +0.17 to 0.45 s before it compresses. The proof the node verifies does not change (the aggregator posts the +compressed proof; the record formats, the guest ids and the verifier stay); the pool does: today +`igneum_submitProofRecord` carries compressed proofs and the pool verifies them with `--mode verify`, so a core +proof needs a hand-off path from the small card to a compressing aggregator before submission, a prover-protocol +change (the proving agent's sizing note), not one line. Per tier, on this measurement: 12 GB mine-and-prove is +still over the 9.0 GB line (7.8 + 1.7 = 9.5 GB before the display); 8 GB cards cannot (6.4 GB own for an EMPTY +shard); 16 GB mines and proves without route 2; 24 and 32 GB cards as aggregators compress 24 shards a minute alone +or 5 beside a miner; a rig puts its small cards on core proofs and one big card on compression. The levers left for +the 12 GB line are a smaller core threshold for CORE-ONLY proving (2^25 and 2^24, where the shard's buffers are +the peak: not yet measured core-only) and the allocator pool returning freed memory between shards (patch v4, +`SP1_GPU_MEM_RELEASE_THRESHOLD`, built but not yet run), both queued for the next PC 2 window. + Consequences (the rule of 5 October 2026), as they stood after sweep 1 (superseded by the reading above for the 12 and 16 GB tiers): the shipped SP1 GPU server refuses every card under 24 GB before allocating, so 8, 12 and 16 GB NVIDIA cards cannot prove on it whatever the shard; the v1 patch takes the shard's term out (20.5 to 12.7 GB on the v1 shard) at a 1.26x time cost (5.3 s against 4.2 s,