Prover floor: sweep 2's diagnosis, patch v3, sweep 3 (the v1 shard at 10,291 MiB in 5.7 s, verified), the server floor 6,535 MiB, the tier line and what it does not say

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 00:18:56 +00:00
parent 7ada573197
commit eff70a2cca
2 changed files with 65 additions and 2 deletions

View file

@ -150,3 +150,37 @@ So after sweep 1 the binding term is the Setup-time keys allocated at full capac
trace buffer (keys and shards) to its padded need: `padded_trace_elements` in `jagged_tracegen/src/lib.rs`
(each phase pads to the next multiple of 2^21 rows, `generate_jagged_traces`'s "final padding"), applied in
`setup_tracegen` and `full_tracegen`, one stacking height of slack, `SP1_GPU_FLOOR_EXACT=0` restoring upstream.
## Sweep 2 (hung), the diagnosis, patch v3, sweep 3: under 11 GB
Sweep 2 (`floor-sweep-2`, 23:21Z, the v2 server fb3165d8) hung on its first point. The restore job
(`floor-restore-1`, 23:57Z) read the point's log: the v2 server had panicked in a tokio worker at
`jagged_tracegen/src/lib.rs:240` ("range end index 37,428,736 out of range for slice of length 36,700,160") and the
SDK client waited on the socket for 1,740 s. The arithmetic named the bug: 36,700,160 is a recursion key's buffer
as v2 sized it (the padded preprocessed traces, 35,651,584, plus one stacking height), 37,428,736 is that end plus
one main trace of 1,777,152 elements: `prove_shard_with_pk` (`shard_prover/src/prover.rs` 332) runs
`main_tracegen`, which appends the shard's main traces INTO the key's buffer, which upstream sized for a whole shard.
Patch v3 keeps the key preprocessed-sized at Setup and lets `main_tracegen` grow it on first use (`grow_for_main`:
a bigger dense buffer and column index, the preprocessed region copied device to device, swapped into the key under
its lock; a loud `abort` instead of a hang if a bound were ever short). Build 5 (`floor-build-5`, 120 s):
`b37defef9f5de43da06eb99a5aebdb6a0db214f49f485e00b86dd6d6c8b497b4`, 166,752,944 bytes, sm_86/89/120.
Sweep 3 (`floor-sweep-3`, 00:13:46 to 00:16:57Z, idle 2,089 MiB inside every peak, every proof VERIFIED by the
unpatched host): the v1 shard at threshold 2^26 **10,291 MiB (8,202 with the idle subtracted), 5.7 s**; at 2^27
12,915 MiB 4.3 s; the empty shard 9,939 to 9,971 MiB; one transfer 10,003 MiB 3.5 s; the 60 M-cycle prototype
shard 13,459 MiB 16.8 s at the 12 GB tier (28,295 MiB on the stock server); the 16 GB tier 16,115 MiB; upstream's
threshold 16,851 MiB (20,516 stock); the server after Setup 6,535 MiB (v1 9,703, stock 11,623); the grow line
`36,700,160 -> 91,226,112 elements` on every recursion key's first use. The full table is in the bench-log entry.
### The tier line that follows (for the public copy, once a 12 GB card has run the same fixture)
| Card | Stock SP1 6.8.1 server | The v3 server (this branch), measured on the 5090's allocation | Profile |
|---|---|---|---|
| 8 GB | refused (the 24 GB panic) | not measured; the empty shard alone is 9.9 GB measured (7.9 GB the server's own), so no | mine only |
| 12 GB | refused | the v1 shard alone at 10.3 GB measured, 8.2 GB the server's own, 5.7 s; beside the miner: pending sweep 4 | `SP1_GPU_ELEMENT_THRESHOLD=67108864`; prove when not mining until sweep 4 says otherwise |
| 16 GB | refused | the v1 shard alone at 12.9 GB measured (10.8 own) at 2^27, 4.3 s; or 10.3 GB at 2^26 beside a miner's 1.8 GB (approximate: the pair is sweep 4) | 2^27 alone, 2^26 beside the miner |
| 24 GB | the v1 shard alone (20.4 GB), never the prototype shard (28.3) | the v1 shard 12.9 GB at 2^27, the prototype shard 13.5 GB (16.8 s) | upstream's 24 GB threshold or 2^27 |
| 32 GB | everything (28.3 GB for the prototype shard beside the miner at 30.1) | 16.9 GB at upstream's threshold | unchanged |
Every row is the 5090's allocation pattern under a budget; the public line keeps "24 GB" until the on-order 12 GB
card runs `fees-v1-shards2` shard 0 through this server and its own peak and time are in the bench-log.

View file

@ -1545,9 +1545,38 @@ server through `HOME=/opt/igneum-floor/home` (the SDK spawns `$HOME/.sp1/bin/sp1
| Sweep 1 (job `floor-sweep-1`, 22:34:56 to 22:37:55Z, patched server v1, card idle 1,755 MiB) | control at upstream's sizes: empty shard (block 83616, 280,706 cycles) **13,892 MiB** 2.2 s, v1 shard (fees-v1-shards2 shard 0, 4,717,439 cycles) **20,516 MiB** 4.2 s (the known curve). 12 GB tier (2^27): empty 12,740 MiB 2.4 s, v1 15,396 MiB 4.1 s. 16 GB tier (2^27+2^26): v1 18,628 MiB 3.8 s. Threshold 2^26: empty 12,772 MiB 3.1 s, v1 **12,708 MiB** 5.3 s (4 core shards). Threshold 2^25: v1 12,836 MiB 8.5 s. Normalize cache 1: 15,428 MiB (no change). Every proof VERIFIED, 1,272,897 bytes, verify 0.037 to 0.040 s |
| The second floor (the server's own `FLOOR` lines in sweep 1) | `memory after setup: 9,703 MiB` at 2^26 (11,623 at the stock sizes) before any proof: at `Setup` the server pre-builds five recursion keys at the fixed 2^27 capacity (0.75 GB each, 90,177,536 of 134,217,728 elements used: 35.6 M preprocessed, 54.5 M main), the shrink key (2^25) and the core key at the threshold; the v1 shard at 2^26 is 4 core shards (main 44.0 M then 3 x 60.8 M elements) and 4 recursion proofs |
| Patch v2 and build 4 (job `floor-build-4`, 23:18:13 to 23:20:14Z, 120 s warm) | every trace buffer sized to its padded need (`padded_trace_elements`: each phase to the next multiple of 2^21, one stacking height of slack, `SP1_GPU_FLOOR_EXACT=0` restores upstream), patch sha 08ce0555; the binary 166,748,808 bytes, sha256 `fb3165d809d2031cc05312c79a545065b2d22ced0a2cd134fb40462262449edf`, sm_86/89/120 |
| Sweep 2 (job `floor-sweep-2`, 23:21:17Z, patched server v2) | HUNG on its first point (the 12 GB tier on the empty shard): the prover off, the card idle 2,060 MiB, then no row in 17 minutes against 11 to 18 s a point in sweep 1; at 23:40Z the coordinator gave PC 2 to the 0.3.11 update, which killed the job tree. The restore-and-diagnose job (`pc2-floor-restore.ps1`) reads the point's log first; the v2 floor is UNMEASURED until it runs |
| Sweep 2 (job `floor-sweep-2`, 23:21:17Z, patched server v2) | HUNG on its first point: no row in 17 minutes against 11 to 18 s a point in sweep 1; at 23:40Z the coordinator gave PC 2 to the 0.3.11 update, which killed the job tree. The restore job (`floor-restore-1`, 23:57Z, 23 s) read the point's log: the server had PANICKED in a tokio worker ("range end index 37,428,736 out of range for slice of length 36,700,160", `jagged_tracegen/src/lib.rs:240`) and the SDK client waited on the socket for 1,740 s: a prove path (`main_tracegen`, `prove_shard_with_pk`) appends the shard's main traces INTO the key's buffer, which v2 had sized for the preprocessed phase alone (35,651,584 + 2^20) and upstream for a whole shard; before the panic v2 read 6,567 MiB after Setup. The restore killed nothing (the restart had), unlinked the root socket, switched the prover on; the live server untouched |
| Patch v3 and build 5 (job `floor-build-5`, 00:09:55 to 00:11:57Z, 120 s warm) | `main_tracegen` grows a key's buffer to the shard's need on first use (a bigger dense buffer and column index, the preprocessed region copied device to device, swapped into the key; a loud abort instead of a hang if a bound were short), patch sha 3f9d3ab0; the binary 166,752,944 bytes, sha256 `b37defef9f5de43da06eb99a5aebdb6a0db214f49f485e00b86dd6d6c8b497b4`, sm_86/89/120 |
| Sweep 3 (job `floor-sweep-3`, 00:13:46 to 00:16:57Z, patched server v3, card idle 2,089 MiB; the table below) | **the v1 shard at threshold 2^26: 10,291 MiB peak (8,202 MiB with the idle subtracted), 5.7 s, VERIFIED**: under the 11.0 GB gate on the 5090's allocation. The server after Setup: **6,535 MiB** (v1 9,703, stock 11,623). The prototype 60 M-cycle shard at the 12 GB tier: 13,459 MiB and 16.8 s (stock 28,295 MiB, 11.4 s). Upstream's threshold on v3: 16,851 MiB (stock 20,516) |
Consequences (the rule of 5 October 2026), from sweep 1 as it stands: the shipped SP1 GPU server refuses every
Sweep 3, every row (the patched server v3 `b37defef...`, sm_86, sm_89, sm_120, the RTX 5090 alone with the miners stopped and the prover off; the idle 2,089 MiB is the display and the other processes on PC 2's card, inside every peak; the unpatched host's VERIFIED on every row):
| Config (environment to the v3 server) | Fixture | Cycles | Peak MiB (idle inside) | Peak minus idle | Prove s | Verified |
|---|---|---|---|---|---|---|
| 12 GB tier (threshold 2^27) | empty shard (block 83616) | 280,706 | 9,939 | 7,850 | 2.6 | yes |
| 12 GB tier | one transfer (block 56) | 556,369 | 10,003 | 7,914 | 3.5 | yes |
| 12 GB tier | v1 shard (fees-v1-shards2 shard 0) | 4,717,439 | 12,915 | 10,826 | 4.3 | yes |
| 12 GB tier | the full PROTOTYPE shard (block-338-shard1; 28,295 MiB on the stock server) | 60,415,376 | 13,459 | 11,370 | 16.8 | yes |
| threshold 2^26 (`SP1_GPU_ELEMENT_THRESHOLD=67108864`) | v1 shard | 4,717,439 | 10,291 | 8,202 | 5.7 | yes |
| threshold 2^26 | empty shard | 280,706 | 9,971 | 7,882 | 3.3 | yes |
| 16 GB tier (2^27+2^26) | v1 shard | 4,717,439 | 16,115 | 14,026 | 4.0 | yes |
| upstream's threshold (budget 32) | v1 shard (20,516 MiB on v1 and stock) | 4,717,439 | 16,851 | 14,762 | 4.0 | yes |
| 12 GB tier + recursion allocation 100,663,296 | v1 shard | 4,717,439 | 12,883 | 10,794 | 4.3 | yes |
Reading, 00:20Z. The task's gate (a real shard under 11.0 GB peak at under 60 s, the proof format unchanged) is met
on PC 2: the adopted v1 shard proves at 10,291 MiB measured (8.2 GB the server's own) in 5.7 s on the v3 server
at threshold 2^26, and the unpatched verifier passes every proof, so the pinned guest ids and the verifying key
stand. The 12 GB tier profile is therefore `SP1_GPU_ELEMENT_THRESHOLD=67108864` on the v3 server (5.7 s a v1
shard against 4.3 s at 2^27, which the proving agent accepts for the loop); the 16 GB tier takes 2^27 (12.9 GB
measured, 10.8 the server's own); the 24 and 32 GB tiers gain 3.7 GB at upstream's own threshold and nothing
they needed. What this does NOT say: no 12 GB card has run it (every row is the 5090's allocation pattern, the
on-order 3060 or 4070 is the measurement for the public line, which stays 24 GB until then); the mine-and-prove
case on a 12 GB card is the beside-the-miner pair (sweep 4, pending the coordinator's window: the miner adds
1.7 to 1.8 GB and 3x on the 5090, so 8.2 + 1.8 = 10.0 GB of the card's 12.3 before the display, the coordinator's
"under 9.0 GB" row is not met alone and sweep 4 says whether it is met at 2^25); and shipping it is the project's
own signed build of SP1's prover at every SP1 upgrade (the proving plan's packaging row before 0.3.12).
Consequences (the rule of 5 October 2026), as they stood after sweep 1 (superseded by the reading above for the 12 and 16 GB tiers): the shipped SP1 GPU server refuses every
card under 24 GB before allocating, so 8, 12 and 16 GB NVIDIA cards cannot prove on it whatever the shard; the v1
patch takes the shard's term out (20.5 to 12.7 GB on the v1 shard) at a 1.26x time cost (5.3 s against 4.2 s,
which the proving agent accepts for the 12 GB profile: the loop's carriage is 25 to 30 s around any proof and the