bench-log and analysis: route 2 core-only rows (alone and beside the miner), the aggregator's cost, the hand-off as a prover-protocol change, the per-tier consequences
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
8953430e0c
commit
80845a48a7
2 changed files with 43 additions and 0 deletions
|
|
@ -195,3 +195,14 @@ the Setup keys (6,535 MiB in use after Setup with the idle inside, about 4.4 GB
|
|||
peak (about 3.8 GB): cutting further means fewer keys built at Setup or a smaller recursion program, not a shard
|
||||
knob. The first publish of sweep 4 was refused by the publisher's kit-path-check (the wiped-jobs-folder rule of
|
||||
21:49Z); the measurement template now tests the chain job's kit before use.
|
||||
|
||||
### Route 2 measured (jobs `floor-core-alone`, `floor-core-miner`, 08:47 to 08:53Z)
|
||||
|
||||
Core-only at 2^26 on the v1 shard: 9,874 MiB alone (7,817 own) in 3.1 s and 12,834 MiB beside the miner (7,754 own)
|
||||
in 12.1 s, against the full compressed run's 8,233 own; the core proof 14,379,043 bytes, CPU verify 0.447 s; the
|
||||
aggregator's extra 2.5 s alone and 12.6 s beside a miner. The empty shard 6,377 own. The core proof's hand-off is
|
||||
a prover-protocol change (the pool and `igneum_submitProofRecord` carry compressed proofs today), not one line.
|
||||
The 12 GB mine-and-prove line (9.0 GB own) is not met by core-only at 2^26 (7.8 + 1.7 GB): the levers left are the
|
||||
core threshold at 2^25 and 2^24 for core-only proving and patch v4's `SP1_GPU_MEM_RELEASE_THRESHOLD` (the pool
|
||||
returns freed memory between shards; upstream holds the high-water mark for the process's life), both in
|
||||
`tools/prover-floor/pc2-floor-core2-miner.ps1` for the next PC 2 window. The table is in the bench-log entry.
|
||||
|
|
|
|||
|
|
@ -1593,6 +1593,38 @@ card mines and proves at 2^27 (10.95 + 1.7 GB = 12.7 plus the display, under its
|
|||
on-order 12 GB card runs the same two points. Shipping it is the project's own signed build of SP1's prover at
|
||||
every SP1 upgrade (the proving plan's packaging row before 0.3.12).
|
||||
|
||||
Route 2, core-only provers (the project lead, 07:40Z: can a 12 GB card mine AND prove; jobs `floor-core-alone` 08:47:07 to
|
||||
08:48:37Z with the miners stopped and `floor-core-miner` 08:50:29 to 08:53:13Z with the 5090 mining at 96%; the v3
|
||||
server b37defef, the pv1 host's `--mode shard` = the core proof, its CPU verify, then the compressed proof in one
|
||||
run, the 1-s sampler split at the core RESULT's timestamp so the core-only peak is the sampler's maximum up to it;
|
||||
every core and compressed proof VERIFIED by the unpatched host; the resident set inside every peak: 2,057 MiB
|
||||
alone, 5,080 MiB with the 0.3.11 miner):
|
||||
|
||||
| Threshold | Fixture | Core-only peak MiB alone (own) | Core s | Core proof bytes | CPU verify s | Compressed s (full peak) | Beside the miner: core peak (own) | Core s | Compressed s |
|
||||
|---|---|---|---|---|---|---|---|---|---|
|
||||
| 2^26 | v1 shard | 9,874 (7,817) | 3.1 | 14,379,043 | 0.447 | 5.6 (10,290) | 12,834 (7,754) | 12.1 | 24.7 |
|
||||
| 2^26 | empty | 8,434 (6,377) | 1.5 | 6,935,657 | 0.212 | 3.2 (9,938) | 11,458 (6,378) | 5.9 | 12.9 |
|
||||
| 2^27 | v1 shard | 12,850 (10,793) | 2.5 | 10,147,581 | 0.314 | 4.1 (12,850) | 15,874 (10,794) | 8.5 | 17.5 |
|
||||
| 2^27 | empty | 9,394 (7,337) | 1.4 | 5,603,823 | 0.170 | 2.4 (9,970) | 12,482 (7,402) | 4.8 | 9.8 |
|
||||
|
||||
Reading. Stopping at the core proof takes 0.4 GB off the v1 shard's own working set at 2^26 (8.2 to 7.8 GB) and
|
||||
1.5 GB off the empty shard's, not more, because the server still builds the compression keys at Setup (4.4 GB own
|
||||
in use before any shard, mostly the allocator pool's high-water mark from the key commits) and at 2^27 the core
|
||||
phase is the peak itself (the shard's own buffers). The aggregator's extra cost per shard, compressed minus core on
|
||||
the same card: 2.5 s alone and 12.6 s beside a miner at 2^26, 1.6 s at 2^27, 1.0 to 1.7 s for empty shards; the
|
||||
hand-off is 5.6 to 14.4 MB a shard (4 to 11x the 1.27 MB compressed proof), verified by the aggregator's CPU in
|
||||
0.17 to 0.45 s before it compresses. The proof the node verifies does not change (the aggregator posts the
|
||||
compressed proof; the record formats, the guest ids and the verifier stay); the pool does: today
|
||||
`igneum_submitProofRecord` carries compressed proofs and the pool verifies them with `--mode verify`, so a core
|
||||
proof needs a hand-off path from the small card to a compressing aggregator before submission, a prover-protocol
|
||||
change (the proving agent's sizing note), not one line. Per tier, on this measurement: 12 GB mine-and-prove is
|
||||
still over the 9.0 GB line (7.8 + 1.7 = 9.5 GB before the display); 8 GB cards cannot (6.4 GB own for an EMPTY
|
||||
shard); 16 GB mines and proves without route 2; 24 and 32 GB cards as aggregators compress 24 shards a minute alone
|
||||
or 5 beside a miner; a rig puts its small cards on core proofs and one big card on compression. The levers left for
|
||||
the 12 GB line are a smaller core threshold for CORE-ONLY proving (2^25 and 2^24, where the shard's buffers are
|
||||
the peak: not yet measured core-only) and the allocator pool returning freed memory between shards (patch v4,
|
||||
`SP1_GPU_MEM_RELEASE_THRESHOLD`, built but not yet run), both queued for the next PC 2 window.
|
||||
|
||||
Consequences (the rule of 5 October 2026), as they stood after sweep 1 (superseded by the reading above for the 12 and 16 GB tiers): the shipped SP1 GPU server refuses every
|
||||
card under 24 GB before allocating, so 8, 12 and 16 GB NVIDIA cards cannot prove on it whatever the shard; the v1
|
||||
patch takes the shard's term out (20.5 to 12.7 GB on the v1 shard) at a 1.26x time cost (5.3 s against 4.2 s,
|
||||
|
|
|
|||
Loading…
Reference in a new issue