bench-log and analysis: the RTX 4070 beside its own miner (9,034 MiB compressed in 24.1 s at 2^26): a 12 GB card mines and proves, measured on the card; the per-tier consequences and what ships it

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 12:18:33 +00:00
parent 25d36336a2
commit 67bc0ebb17
2 changed files with 45 additions and 1 deletions

View file

@ -177,7 +177,7 @@ threshold 16,851 MiB (20,516 stock); the server after Setup 6,535 MiB (v1 9,703,
| Card | Stock SP1 6.8.1 server | The v3 server (this branch), measured on the 5090's allocation | Profile |
|---|---|---|---|
| 8 GB | refused (the 24 GB panic) | not measured; the empty shard alone is 9.9 GB measured (7.9 GB the server's own), so no | mine only |
| 12 GB | refused | the v1 shard alone 10.3 GB measured (8.2 GB the server's own) 5.7 s; beside the miner 8.2 GB own + the miner's 1.7 GB, 24.4 s (sweep 4); the card's own display not measured | `SP1_GPU_ELEMENT_THRESHOLD=67108864`, PROVES ALONE; mine-and-prove NOT claimed (8.2 + 1.7 = 9.9 GB before the display, over the 9.0 GB line the proving agent holds for the tier) |
| 12 GB | refused | MEASURED ON AN RTX 4070 12 GB (PC 1, 6 October): alone 7,553 MiB in 7.7 s at 2^26 (the card held nothing else); beside its own miner (1,449 MiB) 9,034 MiB of 12,282 in 24.1 s; core-only at 2^25 7,242 MiB beside the miner | `SP1_GPU_ELEMENT_THRESHOLD=67108864`: MINES AND PROVES (3.2 GB spare), no hand-off needed |
| 16 GB | refused | the v1 shard alone 12.9 GB measured (10.95 own) at 2^27, 4.3 s; beside the miner 10.95 own + 1.7, 17.4 s (sweep 4) | 2^27, mine and prove (2.9 GB spare on paper) |
| 24 GB | the v1 shard alone (20.4 GB), never the prototype shard (28.3) | the v1 shard 12.9 GB at 2^27, the prototype shard 13.5 GB (16.8 s) | upstream's 24 GB threshold or 2^27 |
| 32 GB | everything (28.3 GB for the prototype shard beside the miner at 30.1) | 16.9 GB at upstream's threshold | unchanged |
@ -215,3 +215,13 @@ freed allocation for the process's life). Core-only beside the miner: 2^25 6,409
12 GB card mines and proves as a core-only prover at 2^25 (6.4 + 1.7 GB before the display) under the 9.0 GB line,
with the hand-off to a compressing aggregator as the prover-protocol change; the full table and the per-tier
consequences are in the bench-log entry. The real card's run decides the public line.
## Route 1 measured: the RTX 4070 12 GB in PC 1 (jobs `card12-alone`, `card12-miner`, 11:53 to 12:17Z)
Alone (the card's idle 0 MiB): the compressed v1 shard 7,553 MiB in 7.7 s at 2^26, 10,177 MiB in 5.5 s at 2^27;
core-only 5,761 MiB at 2^25. Beside its own miner (1,449 MiB): 9,034 MiB in 24.1 s at 2^26, 11,754 MiB in 17.3 s at
2^27, core-only 7,242 MiB at 2^25; a 23.6 M-cycle shard 9,066 MiB in 116.6 s at 2^26. Every proof verified by the
unpatched verifier. So a 12 GB card mines and proves compressed shards at 2^26 with 3.2 GB spare, and route 2's
core-only hand-off is the reserve, not the requirement. The 5090's allocation pattern overstated the card by about
0.65 GB. The full tables and the per-tier consequences are in the bench-log entry; the public line moves to "12 GB
mines and proves" when the packaging row ships the server.

View file

@ -1687,6 +1687,40 @@ Core-only at 2^25 is 5.76 GB on the card. Consequence for the public line: "12 G
measured on a 12 GB card, so the line moves from 24 GB once the patched server ships (the packaging row before
0.3.12); mine-and-prove is the beside-the-miner job's row (below, `card12-miner`).
Beside its own miner (job `card12-miner`, 12:02:36 to 12:17:25Z, the 0.3.12 miner running on both cards, the 4070's
working set **1,449 MiB** read on the card before and after, inside every peak; the prover off then on; every
proof VERIFIED):
| Threshold | Fixture | Core-only peak MiB | Full compressed peak MiB | Core s | Compressed s |
|---|---|---|---|---|---|
| 2^26 | v1 shard | 8,682 | **9,034** | 11.0 | **24.1** |
| 2^26 | block-72854 (23.6 M cycles) | 8,426 | 9,066 | 55.8 | 116.6 |
| 2^27 | v1 shard | 11,722 | 11,754 | 8.4 | 17.3 |
| 2^27 | block-72854 | 11,434 | 11,690 | 37.5 | 68.7 |
| 2^25 | v1 shard | **7,242** | 9,098 | 17.0 | 40.9 |
| 2^25 | block-72854 | 7,242 | 9,226 | 107.7 | 256.4 |
VERDICT on the project lead's question, measured on the card itself: **a 12 GB card mines and proves.** Compressed proving at
2^26 beside its own miner holds 9.03 GB of the card's 12.28 GB (3.2 GB spare) at 24.1 s a v1 shard (the 600 DAA-s
deadline by 25x), with no core-only hand-off, no aggregator and no pool change; a 23.6 M-cycle shard fits the same
way at 116.6 s. 2^27 fits with 0.5 GB spare (11.75 GB), so the profile is 2^26. Core-only at 2^25 is 7.24 GB beside
the miner (5.0 GB spare), the reserve if the shard budget grows. The miner costs the 4070 3.1x in time (24.1 against
7.7 s alone; the 5090 pays 4.3x). The 5090's allocation pattern, and the 9.0 GB line read from it, overstated the
real card by about 0.65 GB: the card's own number decides, as the rule says.
Consequences per tier, now measured on a 12 GB card: **12 GB** (RTX 4070, 3060 12 GB) mines and proves compressed
at 2^26 on the patched server (the public line moves to "12 GB mines and proves" once the packaging row ships the
server; until then the app's default stays at 24 GB); **16 GB** mines and proves at 2^27 (the 5090's allocation:
10.95 own + the miner, 2.9 GB spare, 17.4 s; the card itself not yet run); **8 GB**: not on this server (the
core-only v1 shard alone is 5.76 GB on the 4070 at 2^25, so a compressed proof at 7.55 GB does not fit beside a
miner on 8 GB; core-only at 2^25 plus a miner is about 7.2 GB and would need the hand-off path: an 8 GB row for a
later day, not claimed); **24 and 32 GB**: unchanged, with 3.7 GB more headroom at upstream's threshold; a rig: every
card proves its own compressed shards, no aggregator needed. AMD and Apple: outside SP1 (`docs/analysis/amd-proving.md`).
What ships this: the packaging row (the project's signed build of SP1's GPU server, rebuilt and re-measured at every
SP1 upgrade: `tools/prover-floor/pc2-build-server.ps1` is the recipe, the patch is 4 files and 206 lines on tag
v6.8.1), the app's per-card profile (`provedefault.rs`: 12 GB at 2^26, 16 GB at 2^27, 24 GB at upstream's 24 GB
threshold, 32 GB at the full one) and the HOME or path switch that points the SDK at the project's server.
Consequences (the rule of 5 October 2026), as they stood after sweep 1 (superseded by the reading above for the 12 and 16 GB tiers): the shipped SP1 GPU server refuses every
card under 24 GB before allocating, so 8, 12 and 16 GB NVIDIA cards cannot prove on it whatever the shard; the v1
patch takes the shard's term out (20.5 to 12.7 GB on the v1 shard) at a 1.26x time cost (5.3 s against 4.2 s,