Prover floor: sweep 1 table and reading (the Setup-time keys bind at 12.7 GB)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-05 22:42:17 +00:00
parent eda49ab357
commit ffbdf715fa

View file

@ -113,3 +113,40 @@ niced jobs, 16 cores). The binary: `/opt/igneum-floor/bin/sp1-gpu-server`, **166
`5568108bf7fb9b0e525d8a08926b7046e51136ffaea53f0ca858631d0e938878`**, `--version` 6.8.1, `cuobjdump --list-elf`
sm_86, sm_89, sm_120 (the stock 251,306,680-byte server lists sm_80, 86, 89, 90, 100, 120 and compute_120 PTX).
The live `/root/.sp1/bin/sp1-gpu-server` (c2642ad1...) was never touched; the miners mined throughout.
## Sweep 1 (job `floor-sweep-1`, 22:34:56 to 22:37:55Z): the shard term gone, a second floor found
PC 2's RTX 5090 (32,607 MiB, idle 1,755 MiB with the miners stopped and the live prover off), the patched server
`5568108b...` (sm_86, sm_89, sm_120; the build above), the unpatched pv1 host `dae6b006...` as client and
verifier, one `--mode compressed --shard 0` per point, every server killed and its socket unlinked around every
point, peak = `nvidia-smi memory.used` at 1 s (the idle 1,755 MiB inside it), time = the compressed proof.
Every proof VERIFIED (1,272,897 bytes, verify 0.037 to 0.040 s), so the unpatched verifier accepts every proof
of the patched server.
| Config (environment to the patched server) | Fixture | Cycles | Peak MiB | Prove s | Verified |
|---|---|---|---|---|---|
| control: `SP1_GPU_MEMORY_BUDGET_GB=32` (upstream's sizes) | empty live shard (block 83616) | 280,706 | 13,892 | 2.2 | yes |
| control | v1 shard (fees-v1-shards2 shard 0) | 4,717,439 | 20,516 | 4.2 | yes |
| 12 GB tier: budget 12 (threshold 2^27) | empty | 280,706 | 12,740 | 2.4 | yes |
| 12 GB tier | v1 shard | 4.7 M | 15,396 | 4.1 | yes |
| 12 GB tier + `SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1` | v1 shard | 4.7 M | 15,428 | 4.0 | yes |
| 16 GB tier: budget 16 (2^27 + 2^26) | v1 shard | 4.7 M | 18,628 | 3.8 | yes |
| `SP1_GPU_ELEMENT_THRESHOLD=67108864` (2^26) | empty | 280,706 | 12,772 | 3.1 | yes |
| 2^26 | v1 shard (split into 4 core shards) | 4.7 M | **12,708** | 5.3 | yes |
| `SP1_GPU_ELEMENT_THRESHOLD=33554432` (2^25) | v1 shard | 4.7 M | 12,836 | 8.5 | yes |
Reading. The control reproduces the proving agent's curve (13.9 and 20.4 GB), so the patched server behaves as
the stock one at the stock sizes. The shard's term follows the threshold as the model says (20.5 GB at 402 M
elements, 15.4 at 134 M, 12.7 at 67 M), and then stops: 2^26 and 2^25 both sit at 12.7 to 12.8 GB for the empty
shard and the v1 shard alike. The `FLOOR memory after setup` line names the rest: **9,703 MiB in use before the
first shard** (2^26; 11,623 at the stock sizes), and the `FLOOR tracegen alloc` lines at Setup are five
allocations of 134,217,728 elements (the recursion keys, 0.75 GB each, each using 90,177,536 elements: 35.6 M
preprocessed and 54.5 M main at prove time), one of 33,554,432 (the shrink key, 0.19 GB) and one core key at the
threshold. The `NORMALIZE_PROGRAM_CACHE_SIZE` knob does not reach them (they are keys built at `Setup`, not the
program LRU). The time cost of the split: the v1 shard at 2^26 is 4 core shards and 5.3 s against 4.2 s (1.26x);
at 2^27 it is 4.1 s with no split.
So after sweep 1 the binding term is the Setup-time keys allocated at full capacity, and patch v2 sizes every
trace buffer (keys and shards) to its padded need: `padded_trace_elements` in `jagged_tracegen/src/lib.rs`
(each phase pads to the next multiple of 2^21 rows, `generate_jagged_traces`'s "final padding"), applied in
`setup_tracegen` and `full_tracegen`, one stacking height of slack, `SP1_GPU_FLOOR_EXACT=0` restoring upstream.