bench-log and analysis: route 2 core-only rows (alone and beside the miner), the aggregator's cost, the hand-off as a prover-protocol change, the per-tier consequences

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 08:56:08 +00:00
parent 8953430e0c
commit 80845a48a7
2 changed files with 43 additions and 0 deletions

View file

@ -195,3 +195,14 @@ the Setup keys (6,535 MiB in use after Setup with the idle inside, about 4.4 GB
peak (about 3.8 GB): cutting further means fewer keys built at Setup or a smaller recursion program, not a shard
knob. The first publish of sweep 4 was refused by the publisher's kit-path-check (the wiped-jobs-folder rule of
21:49Z); the measurement template now tests the chain job's kit before use.
### Route 2 measured (jobs `floor-core-alone`, `floor-core-miner`, 08:47 to 08:53Z)
Core-only at 2^26 on the v1 shard: 9,874 MiB alone (7,817 own) in 3.1 s and 12,834 MiB beside the miner (7,754 own)
in 12.1 s, against the full compressed run's 8,233 own; the core proof 14,379,043 bytes, CPU verify 0.447 s; the
aggregator's extra 2.5 s alone and 12.6 s beside a miner. The empty shard 6,377 own. The core proof's hand-off is
a prover-protocol change (the pool and `igneum_submitProofRecord` carry compressed proofs today), not one line.
The 12 GB mine-and-prove line (9.0 GB own) is not met by core-only at 2^26 (7.8 + 1.7 GB): the levers left are the
core threshold at 2^25 and 2^24 for core-only proving and patch v4's `SP1_GPU_MEM_RELEASE_THRESHOLD` (the pool
returns freed memory between shards; upstream holds the high-water mark for the process's life), both in
`tools/prover-floor/pc2-floor-core2-miner.ps1` for the next PC 2 window. The table is in the bench-log entry.

View file

@ -1593,6 +1593,38 @@ card mines and proves at 2^27 (10.95 + 1.7 GB = 12.7 plus the display, under its
on-order 12 GB card runs the same two points. Shipping it is the project's own signed build of SP1's prover at
every SP1 upgrade (the proving plan's packaging row before 0.3.12).
Route 2, core-only provers (the project lead, 07:40Z: can a 12 GB card mine AND prove; jobs `floor-core-alone` 08:47:07 to
08:48:37Z with the miners stopped and `floor-core-miner` 08:50:29 to 08:53:13Z with the 5090 mining at 96%; the v3
server b37defef, the pv1 host's `--mode shard` = the core proof, its CPU verify, then the compressed proof in one
run, the 1-s sampler split at the core RESULT's timestamp so the core-only peak is the sampler's maximum up to it;
every core and compressed proof VERIFIED by the unpatched host; the resident set inside every peak: 2,057 MiB
alone, 5,080 MiB with the 0.3.11 miner):
| Threshold | Fixture | Core-only peak MiB alone (own) | Core s | Core proof bytes | CPU verify s | Compressed s (full peak) | Beside the miner: core peak (own) | Core s | Compressed s |
|---|---|---|---|---|---|---|---|---|---|
| 2^26 | v1 shard | 9,874 (7,817) | 3.1 | 14,379,043 | 0.447 | 5.6 (10,290) | 12,834 (7,754) | 12.1 | 24.7 |
| 2^26 | empty | 8,434 (6,377) | 1.5 | 6,935,657 | 0.212 | 3.2 (9,938) | 11,458 (6,378) | 5.9 | 12.9 |
| 2^27 | v1 shard | 12,850 (10,793) | 2.5 | 10,147,581 | 0.314 | 4.1 (12,850) | 15,874 (10,794) | 8.5 | 17.5 |
| 2^27 | empty | 9,394 (7,337) | 1.4 | 5,603,823 | 0.170 | 2.4 (9,970) | 12,482 (7,402) | 4.8 | 9.8 |
Reading. Stopping at the core proof takes 0.4 GB off the v1 shard's own working set at 2^26 (8.2 to 7.8 GB) and
1.5 GB off the empty shard's, not more, because the server still builds the compression keys at Setup (4.4 GB own
in use before any shard, mostly the allocator pool's high-water mark from the key commits) and at 2^27 the core
phase is the peak itself (the shard's own buffers). The aggregator's extra cost per shard, compressed minus core on
the same card: 2.5 s alone and 12.6 s beside a miner at 2^26, 1.6 s at 2^27, 1.0 to 1.7 s for empty shards; the
hand-off is 5.6 to 14.4 MB a shard (4 to 11x the 1.27 MB compressed proof), verified by the aggregator's CPU in
0.17 to 0.45 s before it compresses. The proof the node verifies does not change (the aggregator posts the
compressed proof; the record formats, the guest ids and the verifier stay); the pool does: today
`igneum_submitProofRecord` carries compressed proofs and the pool verifies them with `--mode verify`, so a core
proof needs a hand-off path from the small card to a compressing aggregator before submission, a prover-protocol
change (the proving agent's sizing note), not one line. Per tier, on this measurement: 12 GB mine-and-prove is
still over the 9.0 GB line (7.8 + 1.7 = 9.5 GB before the display); 8 GB cards cannot (6.4 GB own for an EMPTY
shard); 16 GB mines and proves without route 2; 24 and 32 GB cards as aggregators compress 24 shards a minute alone
or 5 beside a miner; a rig puts its small cards on core proofs and one big card on compression. The levers left for
the 12 GB line are a smaller core threshold for CORE-ONLY proving (2^25 and 2^24, where the shard's buffers are
the peak: not yet measured core-only) and the allocator pool returning freed memory between shards (patch v4,
`SP1_GPU_MEM_RELEASE_THRESHOLD`, built but not yet run), both queued for the next PC 2 window.
Consequences (the rule of 5 October 2026), as they stood after sweep 1 (superseded by the reading above for the 12 and 16 GB tiers): the shipped SP1 GPU server refuses every
card under 24 GB before allocating, so 8, 12 and 16 GB NVIDIA cards cannot prove on it whatever the shard; the v1
patch takes the shard's term out (20.5 to 12.7 GB on the v1 shard) at a 1.26x time cost (5.3 s against 4.2 s,