diff --git a/docs/analysis/prover-floor.md b/docs/analysis/prover-floor.md index bd40b57b6..e87b1c5c2 100644 --- a/docs/analysis/prover-floor.md +++ b/docs/analysis/prover-floor.md @@ -177,7 +177,7 @@ threshold 16,851 MiB (20,516 stock); the server after Setup 6,535 MiB (v1 9,703, | Card | Stock SP1 6.8.1 server | The v3 server (this branch), measured on the 5090's allocation | Profile | |---|---|---|---| | 8 GB | refused (the 24 GB panic) | not measured; the empty shard alone is 9.9 GB measured (7.9 GB the server's own), so no | mine only | -| 12 GB | refused | the v1 shard alone 10.3 GB measured (8.2 GB the server's own) 5.7 s; beside the miner 8.2 GB own + the miner's 1.7 GB, 24.4 s (sweep 4); the card's own display not measured | `SP1_GPU_ELEMENT_THRESHOLD=67108864`, PROVES ALONE; mine-and-prove NOT claimed (8.2 + 1.7 = 9.9 GB before the display, over the 9.0 GB line the proving agent holds for the tier) | +| 12 GB | refused | MEASURED ON AN RTX 4070 12 GB (PC 1, 6 October): alone 7,553 MiB in 7.7 s at 2^26 (the card held nothing else); beside its own miner (1,449 MiB) 9,034 MiB of 12,282 in 24.1 s; core-only at 2^25 7,242 MiB beside the miner | `SP1_GPU_ELEMENT_THRESHOLD=67108864`: MINES AND PROVES (3.2 GB spare), no hand-off needed | | 16 GB | refused | the v1 shard alone 12.9 GB measured (10.95 own) at 2^27, 4.3 s; beside the miner 10.95 own + 1.7, 17.4 s (sweep 4) | 2^27, mine and prove (2.9 GB spare on paper) | | 24 GB | the v1 shard alone (20.4 GB), never the prototype shard (28.3) | the v1 shard 12.9 GB at 2^27, the prototype shard 13.5 GB (16.8 s) | upstream's 24 GB threshold or 2^27 | | 32 GB | everything (28.3 GB for the prototype shard beside the miner at 30.1) | 16.9 GB at upstream's threshold | unchanged | @@ -215,3 +215,13 @@ freed allocation for the process's life). Core-only beside the miner: 2^25 6,409 12 GB card mines and proves as a core-only prover at 2^25 (6.4 + 1.7 GB before the display) under the 9.0 GB line, with the hand-off to a compressing aggregator as the prover-protocol change; the full table and the per-tier consequences are in the bench-log entry. The real card's run decides the public line. + +## Route 1 measured: the RTX 4070 12 GB in PC 1 (jobs `card12-alone`, `card12-miner`, 11:53 to 12:17Z) + +Alone (the card's idle 0 MiB): the compressed v1 shard 7,553 MiB in 7.7 s at 2^26, 10,177 MiB in 5.5 s at 2^27; +core-only 5,761 MiB at 2^25. Beside its own miner (1,449 MiB): 9,034 MiB in 24.1 s at 2^26, 11,754 MiB in 17.3 s at +2^27, core-only 7,242 MiB at 2^25; a 23.6 M-cycle shard 9,066 MiB in 116.6 s at 2^26. Every proof verified by the +unpatched verifier. So a 12 GB card mines and proves compressed shards at 2^26 with 3.2 GB spare, and route 2's +core-only hand-off is the reserve, not the requirement. The 5090's allocation pattern overstated the card by about +0.65 GB. The full tables and the per-tier consequences are in the bench-log entry; the public line moves to "12 GB +mines and proves" when the packaging row ships the server. diff --git a/docs/bench-log.md b/docs/bench-log.md index 0faeb695a..57d5f890f 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1687,6 +1687,40 @@ Core-only at 2^25 is 5.76 GB on the card. Consequence for the public line: "12 G measured on a 12 GB card, so the line moves from 24 GB once the patched server ships (the packaging row before 0.3.12); mine-and-prove is the beside-the-miner job's row (below, `card12-miner`). +Beside its own miner (job `card12-miner`, 12:02:36 to 12:17:25Z, the 0.3.12 miner running on both cards, the 4070's +working set **1,449 MiB** read on the card before and after, inside every peak; the prover off then on; every +proof VERIFIED): + +| Threshold | Fixture | Core-only peak MiB | Full compressed peak MiB | Core s | Compressed s | +|---|---|---|---|---|---| +| 2^26 | v1 shard | 8,682 | **9,034** | 11.0 | **24.1** | +| 2^26 | block-72854 (23.6 M cycles) | 8,426 | 9,066 | 55.8 | 116.6 | +| 2^27 | v1 shard | 11,722 | 11,754 | 8.4 | 17.3 | +| 2^27 | block-72854 | 11,434 | 11,690 | 37.5 | 68.7 | +| 2^25 | v1 shard | **7,242** | 9,098 | 17.0 | 40.9 | +| 2^25 | block-72854 | 7,242 | 9,226 | 107.7 | 256.4 | + +VERDICT on the project lead's question, measured on the card itself: **a 12 GB card mines and proves.** Compressed proving at +2^26 beside its own miner holds 9.03 GB of the card's 12.28 GB (3.2 GB spare) at 24.1 s a v1 shard (the 600 DAA-s +deadline by 25x), with no core-only hand-off, no aggregator and no pool change; a 23.6 M-cycle shard fits the same +way at 116.6 s. 2^27 fits with 0.5 GB spare (11.75 GB), so the profile is 2^26. Core-only at 2^25 is 7.24 GB beside +the miner (5.0 GB spare), the reserve if the shard budget grows. The miner costs the 4070 3.1x in time (24.1 against +7.7 s alone; the 5090 pays 4.3x). The 5090's allocation pattern, and the 9.0 GB line read from it, overstated the +real card by about 0.65 GB: the card's own number decides, as the rule says. + +Consequences per tier, now measured on a 12 GB card: **12 GB** (RTX 4070, 3060 12 GB) mines and proves compressed +at 2^26 on the patched server (the public line moves to "12 GB mines and proves" once the packaging row ships the +server; until then the app's default stays at 24 GB); **16 GB** mines and proves at 2^27 (the 5090's allocation: +10.95 own + the miner, 2.9 GB spare, 17.4 s; the card itself not yet run); **8 GB**: not on this server (the +core-only v1 shard alone is 5.76 GB on the 4070 at 2^25, so a compressed proof at 7.55 GB does not fit beside a +miner on 8 GB; core-only at 2^25 plus a miner is about 7.2 GB and would need the hand-off path: an 8 GB row for a +later day, not claimed); **24 and 32 GB**: unchanged, with 3.7 GB more headroom at upstream's threshold; a rig: every +card proves its own compressed shards, no aggregator needed. AMD and Apple: outside SP1 (`docs/analysis/amd-proving.md`). +What ships this: the packaging row (the project's signed build of SP1's GPU server, rebuilt and re-measured at every +SP1 upgrade: `tools/prover-floor/pc2-build-server.ps1` is the recipe, the patch is 4 files and 206 lines on tag +v6.8.1), the app's per-card profile (`provedefault.rs`: 12 GB at 2^26, 16 GB at 2^27, 24 GB at upstream's 24 GB +threshold, 32 GB at the full one) and the HOME or path switch that points the SDK at the project's server. + Consequences (the rule of 5 October 2026), as they stood after sweep 1 (superseded by the reading above for the 12 and 16 GB tiers): the shipped SP1 GPU server refuses every card under 24 GB before allocating, so 8, 12 and 16 GB NVIDIA cards cannot prove on it whatever the shard; the v1 patch takes the shard's term out (20.5 to 12.7 GB on the v1 shard) at a 1.26x time cost (5.3 s against 4.2 s,