Merge prover-floor (8e2686d) into release-0.3.11: docs and standalone probe sources the evidence rows cite (C38); nothing a built artefact reads
This commit is contained in:
commit
0059f12c98
13 changed files with 901 additions and 0 deletions
152
docs/analysis/prover-floor.md
Normal file
152
docs/analysis/prover-floor.md
Normal file
|
|
@ -0,0 +1,152 @@
|
|||
# The prover floor: why a 12 GB card cannot prove on SP1 6.8.1's GPU server, and the patch
|
||||
|
||||
5 October 2026, 22:00 UTC on (the project lead: "execute if it will solve the issue"). Branch `prover-floor`
|
||||
(worktree `igneum-wt-prover-floor`). The measured facts this starts from: `docs/plans/proving-v1.md` and the
|
||||
bench-log entry "proving v1" (branch proving-v1): the GPU server holds 13.9 GB for an empty shard, 20.4 GB at the
|
||||
adopted v1 shard, 28.3 GB flat from 20 M to 60 M cycles, and no environment knob moved the floor. Every figure
|
||||
below is from the source at tag v6.8.1 (cloned to `vendor/sp1-6.8.1`, gitignored; the fork is the patch
|
||||
`proving/prover-floor/sp1-gpu-6.8.1-floor.patch`) or from a PC 2 run named in `docs/bench-log.md`
|
||||
("prover floor"). Sizes in GiB are computed from the source constants (4-byte field elements); sizes in MiB are
|
||||
measured by `nvidia-smi` at 1 s.
|
||||
|
||||
## Where the server is built and what it reads
|
||||
|
||||
The SDK downloads `sp1_gpu_server_v6.8.1_x86_64.tar.gz` (133,750,780 bytes) from the SP1 release and runs it from
|
||||
`$HOME/.sp1/bin/sp1-gpu-server` (`crates/cuda/src/server.rs` 19 to 30, 80 to 99). The source is in the same
|
||||
repository: `sp1-gpu/crates/server` (the binary), built by `.github/workflows/release.yml` 234 to 314 on CUDA
|
||||
12.8.1 with Go and protoc (`cargo build --release --bin sp1-gpu-server`). The binary takes no options
|
||||
(`sp1-gpu/crates/server/src/main.rs` 15 to 18: `--version` only) and reads `CUDA_VISIBLE_DEVICES` (32 to 35);
|
||||
everything else comes from the environment the host process passes it, through `SP1CoreOpts::default()`
|
||||
(`crates/core/executor/src/opts.rs` 99 to 140: `SHARD_SIZE`, `ELEMENT_THRESHOLD`, `HEIGHT_THRESHOLD`,
|
||||
`MINIMAL_TRACE_CHUNK_THRESHOLD`, `TRACE_CHUNK_SLOTS`, `FULL_SIZE_SHARDS`) and the worker counts
|
||||
(`crates/prover/src/worker/config.rs`).
|
||||
|
||||
## The memory model, term by term
|
||||
|
||||
Every device buffer is sized at construction from constants, not from the shard. The server builds the prover at
|
||||
the first `Setup` request (`sp1-gpu/crates/server/src/server.rs` 126 to 137) through
|
||||
`cuda_worker_builder_with_machine` (`sp1-gpu/crates/prover_components/src/builder.rs` 102 to 148):
|
||||
|
||||
| Term | Where | Size | On the device | Moves with the shard |
|
||||
|---|---|---|---|---|
|
||||
| The gate | `builder.rs` 35 to 39: `gpu_memory_gb = ceil(total / GiB) + 4`; `panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB")` when under 24 | a 12 GB card reads 16, a 16 GB card 20: both refused before any allocation | | no |
|
||||
| The core element threshold | `builder.rs` 41 to 48: `ELEMENT_THRESHOLD` = 2^28 + 2^27 = 402,653,184 elements (`opts.rs` 12) on a card reading over 30 (a 32 GB card reads 36); minus 2^26 + 2^25 + 2^24 = 285,212,672 on a card reading 24 to 30 (a 24 GB card). The environment's `ELEMENT_THRESHOLD` is read at `opts.rs` 129 and then OVERWRITTEN at `builder.rs` 48, which is why the sweep's `ELEMENT_THRESHOLD` rows changed nothing; `HEIGHT_THRESHOLD` survives (it is not overwritten), which is why the 2^25 + 2^20 row did | sets the next two terms | | no |
|
||||
| The core trace area, one per shard in flight | `builder.rs` 70 to 71: `num_elts = element_threshold + 2^21` (`CORE_LOG_STACKING_HEIGHT` 21, `crates/prover/src/components.rs` 16) = 404,750,336; allocated on the device at `sp1-gpu/crates/jagged_tracegen/src/lib.rs` 484 to 500 (`allocate_and_initialize_traces`: `max_trace_size` felts + `max_trace_size / 2` u32 column index + 2^14 u32) | 6 bytes an element: **2.26 GiB** for the full threshold, 1.59 GiB for the 24 GB threshold | yes, in full, whatever the shard holds | no |
|
||||
| The program's preprocessed traces (the proving key) | `sp1-gpu/crates/shard_prover/src/setup.rs` 47 to 58 and 107: the same `allocate_and_initialize_traces(max_trace_size)` at `Setup`, kept in the key cache for the connection's life (`server.rs` 139 to 143) | another **2.26 GiB**, held from `Setup` on | yes | no |
|
||||
| The pinned host trace buffers | `sp1-gpu/crates/prover_components/src/components.rs` 99 to 103: 4 `PinnedBuffer` of `max_trace_size` felts per prover (core 4 x 1.51 GiB, recursion 4 x 0.5 GiB, shrink 4 x 0.125 GiB, wrap 4 x 0.32 GiB) | 9.8 GiB of pinned host RAM, not device memory (the WSL2 working set the bench saw) | no | no |
|
||||
| The recursion trace area | `builder.rs` 15 and 95: `RECURSION_TRACE_ALLOCATION` = 2^27 elements, one per recursion tracegen (the recursion program's key at setup and its shard at prove) | 0.75 GiB each | yes | no |
|
||||
| The shrink and wrap provers | `builder.rs` 16, 19, 117 to 127: 2^25 and 85,376,340 elements, built at `Setup` for every proof mode, used only by the Groth16 and PLONK path | host pinned at build; device only when a wrap runs (never, for a compressed proof) | no | no |
|
||||
| The codewords (LDE) and the Merkle trees | `sp1-gpu/crates/basefold/src/fri.rs` 92 to 97: every stacked column of 2^21 rows encoded to 2^(21 + 1) rows (`log_blowup` 1); kept until the query phase unless `drop_ldes` (`builder.rs` 52: only on a 24 GB card with `FULL_SIZE_SHARDS`); the preprocessed codewords live in the key | 2 x the padded trace, so up to 2 x the term above | yes | yes, with the padded trace |
|
||||
| The LogUp GKR layers | `builder.rs` 51: `recompute_gkr_trace = false`, so the first layer stays materialised (`sp1-gpu/crates/logup_gkr/src/tracegen.rs` 169 to 224) | of the order of the interaction count | yes | yes |
|
||||
| The allocator | `sp1-gpu/crates/cuda/src/task.rs` 152 and 196: the device's default `cudaMallocAsync` pool with its release threshold at `u64::MAX`, so nothing freed is ever returned to the driver: `nvidia-smi` reads the high-water mark of everything live at once | | | |
|
||||
|
||||
So at zero cycles the server already holds the proving key's 2.26 GiB, the shard's 2.26 GiB (both allocated at the
|
||||
threshold, not at the shard's rows), their codewords and trees, and the recursion program's key and traces (2 x
|
||||
0.75 GiB and their codewords): the 13.9 GB floor. The shard's own content only adds to the codewords, the GKR
|
||||
layers and the working buffers, which is the 13.9 to 20.4 GB step from 0.3 M to 4.7 M cycles, and the flat 28.3 GB
|
||||
from 20 M cycles is the threshold's padded area reached. The witness (5 to 22 KB) never appears.
|
||||
|
||||
## What the patch does (`proving/prover-floor/sp1-gpu-6.8.1-floor.patch`, three files)
|
||||
|
||||
1. `builder.rs`: the panic is gone; the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks the element threshold
|
||||
from a tier table (`element_threshold_for_budget`: over 30 as read, the full 402.6 M; 24 to 30, upstream's 24 GB
|
||||
figure; 18 to 24 (a 16 GB card), 2^27 + 2^26 = 201.3 M; under 18 (a 12 GB card), 2^27 = 134.2 M);
|
||||
`SP1_GPU_ELEMENT_THRESHOLD` sets it directly and `SP1_GPU_RECURSION_TRACE_ALLOCATION` the recursion buffer.
|
||||
The chosen numbers are printed as a `FLOOR opts` line. Every other option is as upstream.
|
||||
2. `jagged_tracegen/src/lib.rs`: with `SP1_GPU_FLOOR_LOG` set, every trace allocation prints its capacity and,
|
||||
after the shard's traces are in, the elements actually used and the device memory in use.
|
||||
3. `server.rs`: a `FLOOR memory` line (device used, free, total) after `Setup` and after every proof, with the
|
||||
proof's time.
|
||||
|
||||
Nothing in the proof changes: the element threshold only moves where the executor splits shards, exactly what
|
||||
upstream's own 24 GB tier does with the same verifier and the same keys; the recursion program, the verifying key
|
||||
and the pinned guest ids are untouched. The unpatched verifier (the pv1 host's SDK) is the one that verifies every
|
||||
measured proof below.
|
||||
|
||||
## The build (PC 2, WSL2 Ubuntu-24.04, job `floor-toolchain-1` then the build job)
|
||||
|
||||
Toolchain found 22:10Z (job `floor-toolchain-1`, 4 s): nvcc 12.8 at `/usr/local/cuda-12.8`, cmake 3.28.3, gcc 13.3,
|
||||
clang 18, protoc 3.21.12, cargo 1.99.0, no Go. The release workflow installs Go for the server's `native-gnark`
|
||||
feature (the Groth16 and PLONK wrap through gnark), which a compressed proof never runs, so the build drops that
|
||||
feature from `sp1-gpu/crates/server/Cargo.toml` and nothing else. `CUDA_ARCHS=86,89,120` (consequences reviewer
|
||||
C26): the 12 GB tier is sm_86 (RTX 3060) and sm_89 (RTX 4070), the 16 GB tier sm_89 and sm_120 (RTX 5080), PC 2's
|
||||
5090 is sm_120; the stock server lists sm_80, 86, 89, 90, 100 and 120, which a shipped build repeats. The recipe:
|
||||
`tools/prover-floor/pc2-build-server.ps1` (generated by `make-build-playbook.sh` from the patch, so the two cannot
|
||||
drift): clone the tag, `git apply` the patch, `touch` the three files, `cargo build --release --bin sp1-gpu-server`
|
||||
niced with 8 jobs into `/opt/igneum-floor/target`, the binary copied to `/opt/igneum-floor/home/.sp1/bin/` (the SDK
|
||||
spawns the server it finds under `$HOME/.sp1/bin`, so `HOME=/opt/igneum-floor/home` selects it and the live
|
||||
`/root/.sp1/bin/sp1-gpu-server` stays as it is).
|
||||
|
||||
What a measurement on PC 2 can and cannot say (C26). The server's allocation pattern is deterministic in the
|
||||
budget it is given, so a run with `SP1_GPU_MEMORY_BUDGET_GB=12` on the 5090 shows the peak a 12 GB card's build
|
||||
would ask for; it does not show that a 3060 proves it in time, nor what the card's display and driver hold. The
|
||||
public line keeps "24 GB" until the on-order 12 GB card runs the same fixture. Every row names the arch list and
|
||||
the card.
|
||||
|
||||
What shipping it costs (C26). A patched server means the project signs and distributes its own build of SP1's
|
||||
prover: the WSL2 package, the DMG's prover inputs, the K1-signed inputs and `evidence.md` carry it, and every SP1
|
||||
upgrade repeats the clone, patch, build and measurement. The verifying key and the pinned guest ids do not move
|
||||
(the patch changes buffer sizes and the shard split, not the circuits), which the `verify-segment` and `--mode
|
||||
compressed` VERIFIED lines of the unpatched host show on every row below. The packaging path is a row for the
|
||||
proving plan before 0.3.12, not this branch.
|
||||
|
||||
## Step 4 contingency, read not measured: RISC Zero's CUDA prover and its memory per segment
|
||||
|
||||
If SP1 could not be brought under 11 GB, the alternative's floor is read from its operators' documentation (not
|
||||
measured here; a PC 2 run would be the measurement): Boundless' prover guide
|
||||
(https://docs.boundless.network/provers/performance-optimization) sets the segment size cap by VRAM as 8 GB:
|
||||
po2 19, 16 GB: po2 20, 20 GB: po2 21, 40 GB: po2 22, with measured peaks po2 20: 13,835 MiB, po2 21: 22,905 MiB,
|
||||
po2 22: 41,089 MiB; RISC Zero's PR 3761 adds `low_vram` and `pinned_witgen` to fit po2 22 on a 24 GB 4090. So
|
||||
RISC Zero proves a 2^19-cycle segment inside 8 GB and a 2^20 one inside 16 GB, and a shard of 4.7 M cycles is
|
||||
9 segments at po2 19 plus lift and join steps (times not on the page). Adopting it would cost a second guest (the
|
||||
chain rule in the RISC Zero zkVM), a second pinned program id, a second verifier in the node and no shared
|
||||
aggregation between the two formats: `docs/analysis/amd-proving.md` and the proving plan carry that row already.
|
||||
|
||||
### The build, as it ran (job `floor-build-3`, 22:28:24 to 22:32:29Z)
|
||||
|
||||
Three runs: `floor-build-1` (22:17Z) and `floor-build-2` (22:24Z) failed in 2 to 4 minutes on
|
||||
`crates/recursion/gnark-ffi/build.rs:70`, "Failed to build Go library: NotFound" (no `go` on PC 2; the first run's
|
||||
playbook lost its own log, a bug fixed before the second). `floor-build-3` fetched go1.27.1 (tarball sha256
|
||||
`63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445`, checked before unpacking under
|
||||
`/opt/igneum-floor/go`, on the job's PATH only) and built in **240 s** (46 crates on the warm target of run 2, 8
|
||||
niced jobs, 16 cores). The binary: `/opt/igneum-floor/bin/sp1-gpu-server`, **166,768,224 bytes, sha256
|
||||
`5568108bf7fb9b0e525d8a08926b7046e51136ffaea53f0ca858631d0e938878`**, `--version` 6.8.1, `cuobjdump --list-elf`
|
||||
sm_86, sm_89, sm_120 (the stock 251,306,680-byte server lists sm_80, 86, 89, 90, 100, 120 and compute_120 PTX).
|
||||
The live `/root/.sp1/bin/sp1-gpu-server` (c2642ad1...) was never touched; the miners mined throughout.
|
||||
|
||||
## Sweep 1 (job `floor-sweep-1`, 22:34:56 to 22:37:55Z): the shard term gone, a second floor found
|
||||
|
||||
PC 2's RTX 5090 (32,607 MiB, idle 1,755 MiB with the miners stopped and the live prover off), the patched server
|
||||
`5568108b...` (sm_86, sm_89, sm_120; the build above), the unpatched pv1 host `dae6b006...` as client and
|
||||
verifier, one `--mode compressed --shard 0` per point, every server killed and its socket unlinked around every
|
||||
point, peak = `nvidia-smi memory.used` at 1 s (the idle 1,755 MiB inside it), time = the compressed proof.
|
||||
Every proof VERIFIED (1,272,897 bytes, verify 0.037 to 0.040 s), so the unpatched verifier accepts every proof
|
||||
of the patched server.
|
||||
|
||||
| Config (environment to the patched server) | Fixture | Cycles | Peak MiB | Prove s | Verified |
|
||||
|---|---|---|---|---|---|
|
||||
| control: `SP1_GPU_MEMORY_BUDGET_GB=32` (upstream's sizes) | empty live shard (block 83616) | 280,706 | 13,892 | 2.2 | yes |
|
||||
| control | v1 shard (fees-v1-shards2 shard 0) | 4,717,439 | 20,516 | 4.2 | yes |
|
||||
| 12 GB tier: budget 12 (threshold 2^27) | empty | 280,706 | 12,740 | 2.4 | yes |
|
||||
| 12 GB tier | v1 shard | 4.7 M | 15,396 | 4.1 | yes |
|
||||
| 12 GB tier + `SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1` | v1 shard | 4.7 M | 15,428 | 4.0 | yes |
|
||||
| 16 GB tier: budget 16 (2^27 + 2^26) | v1 shard | 4.7 M | 18,628 | 3.8 | yes |
|
||||
| `SP1_GPU_ELEMENT_THRESHOLD=67108864` (2^26) | empty | 280,706 | 12,772 | 3.1 | yes |
|
||||
| 2^26 | v1 shard (split into 4 core shards) | 4.7 M | **12,708** | 5.3 | yes |
|
||||
| `SP1_GPU_ELEMENT_THRESHOLD=33554432` (2^25) | v1 shard | 4.7 M | 12,836 | 8.5 | yes |
|
||||
|
||||
Reading. The control reproduces the proving agent's curve (13.9 and 20.4 GB), so the patched server behaves as
|
||||
the stock one at the stock sizes. The shard's term follows the threshold as the model says (20.5 GB at 402 M
|
||||
elements, 15.4 at 134 M, 12.7 at 67 M), and then stops: 2^26 and 2^25 both sit at 12.7 to 12.8 GB for the empty
|
||||
shard and the v1 shard alike. The `FLOOR memory after setup` line names the rest: **9,703 MiB in use before the
|
||||
first shard** (2^26; 11,623 at the stock sizes), and the `FLOOR tracegen alloc` lines at Setup are five
|
||||
allocations of 134,217,728 elements (the recursion keys, 0.75 GB each, each using 90,177,536 elements: 35.6 M
|
||||
preprocessed and 54.5 M main at prove time), one of 33,554,432 (the shrink key, 0.19 GB) and one core key at the
|
||||
threshold. The `NORMALIZE_PROGRAM_CACHE_SIZE` knob does not reach them (they are keys built at `Setup`, not the
|
||||
program LRU). The time cost of the split: the v1 shard at 2^26 is 4 core shards and 5.3 s against 4.2 s (1.26x);
|
||||
at 2^27 it is 4.1 s with no split.
|
||||
|
||||
So after sweep 1 the binding term is the Setup-time keys allocated at full capacity, and patch v2 sizes every
|
||||
trace buffer (keys and shards) to its padded need: `padded_trace_elements` in `jagged_tracegen/src/lib.rs`
|
||||
(each phase pads to the next multiple of 2^21 rows, `generate_jagged_traces`'s "final padding"), applied in
|
||||
`setup_tracegen` and `full_tracegen`, one stacking height of slack, `SP1_GPU_FLOOR_EXACT=0` restoring upstream.
|
||||
250
proving/prover-floor/sp1-gpu-6.8.1-floor.patch
Normal file
250
proving/prover-floor/sp1-gpu-6.8.1-floor.patch
Normal file
|
|
@ -0,0 +1,250 @@
|
|||
diff --git a/sp1-gpu/crates/jagged_tracegen/src/lib.rs b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
|
||||
index 579f70a..09e73e8 100644
|
||||
--- a/sp1-gpu/crates/jagged_tracegen/src/lib.rs
|
||||
+++ b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
|
||||
@@ -481,6 +481,33 @@ async fn device_preprocessed_tracegen<A: CudaTracegenAir<Felt>>(
|
||||
named_traces
|
||||
}
|
||||
|
||||
+/// Igneum prover-floor patch: the dense elements a set of traces will occupy once `generate_jagged_traces`
|
||||
+/// has laid them out, that is the sum of their buffers padded to the next multiple of 2^log_stacking_height
|
||||
+/// (the "final padding" step below). Each phase (preprocessed, then main) is padded on its own.
|
||||
+pub fn padded_trace_elements(
|
||||
+ traces: &BTreeMap<String, Trace<TaskScope>>,
|
||||
+ log_stacking_height: u32,
|
||||
+) -> usize {
|
||||
+ let total: usize = traces
|
||||
+ .values()
|
||||
+ .map(|t| match t {
|
||||
+ Trace::Real(trace) => trace.guts().as_buffer().len(),
|
||||
+ Trace::Padding(_) => 0,
|
||||
+ })
|
||||
+ .sum();
|
||||
+ total.next_multiple_of(1 << log_stacking_height)
|
||||
+}
|
||||
+
|
||||
+/// Igneum prover-floor patch: the capacity to allocate for a trace set: the exact padded size plus one
|
||||
+/// stacking height of slack, never more than the prover's `max_trace_size`. `SP1_GPU_FLOOR_EXACT=0` restores
|
||||
+/// upstream's full-capacity allocation.
|
||||
+fn floor_capacity(max_trace_size: usize, needed: usize, log_stacking_height: u32) -> usize {
|
||||
+ if std::env::var("SP1_GPU_FLOOR_EXACT").map(|v| v == "0").unwrap_or(false) {
|
||||
+ return max_trace_size;
|
||||
+ }
|
||||
+ max_trace_size.min(needed + (1 << log_stacking_height))
|
||||
+}
|
||||
+
|
||||
async fn allocate_and_initialize_traces(
|
||||
preprocessed_traces: BTreeMap<String, Trace<TaskScope>>,
|
||||
max_trace_size: usize,
|
||||
@@ -494,6 +521,11 @@ async fn allocate_and_initialize_traces(
|
||||
|
||||
let total_gb = total_bytes as f64 / (1 << 30) as f64;
|
||||
tracing::debug!("Allocating {:?} GB of traces", total_gb);
|
||||
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
|
||||
+ eprintln!(
|
||||
+ "FLOOR tracegen alloc capacity_elements={max_trace_size} bytes={total_bytes} ({total_gb:.3} GB)"
|
||||
+ );
|
||||
+ }
|
||||
let mut dense_data: Buffer<Felt, TaskScope> =
|
||||
Buffer::with_capacity_in(max_trace_size, backend.clone());
|
||||
let mut col_index: Buffer<u32, TaskScope> =
|
||||
@@ -677,9 +709,14 @@ pub async fn setup_tracegen<A: CudaTracegenAir<Felt>>(
|
||||
let preprocessed_traces =
|
||||
device_preprocessed_tracegen(program, host_phase_tracegen, backend).await;
|
||||
|
||||
+ let capacity = floor_capacity(
|
||||
+ max_trace_size,
|
||||
+ padded_trace_elements(&preprocessed_traces, log_stacking_height),
|
||||
+ log_stacking_height,
|
||||
+ );
|
||||
let jagged_traces = allocate_and_initialize_traces(
|
||||
preprocessed_traces,
|
||||
- max_trace_size,
|
||||
+ capacity,
|
||||
log_stacking_height,
|
||||
max_log_row_count,
|
||||
backend,
|
||||
@@ -984,9 +1021,15 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
|
||||
|
||||
log_chip_stats(machine, &chip_set, &main_traces);
|
||||
|
||||
+ let capacity = floor_capacity(
|
||||
+ max_trace_size,
|
||||
+ padded_trace_elements(&preprocessed_traces, log_stacking_height)
|
||||
+ + padded_trace_elements(&main_traces, log_stacking_height),
|
||||
+ log_stacking_height,
|
||||
+ );
|
||||
let mut jagged_mle = allocate_and_initialize_traces(
|
||||
preprocessed_traces,
|
||||
- max_trace_size,
|
||||
+ capacity,
|
||||
log_stacking_height,
|
||||
max_log_row_count,
|
||||
backend,
|
||||
@@ -1002,6 +1045,18 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
|
||||
)
|
||||
.await;
|
||||
|
||||
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
|
||||
+ let dense = jagged_mle.dense();
|
||||
+ let (free, total) = sp1_gpu_cudart::cuda_memory_info().unwrap_or((0, 0));
|
||||
+ eprintln!(
|
||||
+ "FLOOR tracegen used preprocessed_elements={} main_elements={} dense_len={} capacity_elements={capacity} max_trace_size={max_trace_size} device_used_mib={}",
|
||||
+ dense.preprocessed_offset,
|
||||
+ dense.main_size(),
|
||||
+ dense.dense.len(),
|
||||
+ (total - free) >> 20
|
||||
+ );
|
||||
+ }
|
||||
+
|
||||
(public_values, jagged_mle, chip_set, permit)
|
||||
}
|
||||
|
||||
diff --git a/sp1-gpu/crates/prover_components/src/builder.rs b/sp1-gpu/crates/prover_components/src/builder.rs
|
||||
index 5dccd9d..574d4fa 100644
|
||||
--- a/sp1-gpu/crates/prover_components/src/builder.rs
|
||||
+++ b/sp1-gpu/crates/prover_components/src/builder.rs
|
||||
@@ -23,28 +23,75 @@ use crate::{
|
||||
SP1CudaProverComponents,
|
||||
};
|
||||
|
||||
+/// Igneum prover-floor patch (5 October 2026). Upstream sizes every device buffer for a 24 GB card or larger
|
||||
+/// and panics below that, whatever the shard. Here the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks a
|
||||
+/// tier, and `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly.
|
||||
+/// The proof format, the verifier and the program ids do not change: the element threshold only decides where
|
||||
+/// the executor splits shards, as upstream's own 24 GB tier already does.
|
||||
+fn env_usize(name: &str) -> Option<usize> {
|
||||
+ std::env::var(name).ok().and_then(|s| s.parse::<usize>().ok())
|
||||
+}
|
||||
+
|
||||
+fn env_f64(name: &str) -> Option<f64> {
|
||||
+ std::env::var(name).ok().and_then(|s| s.parse::<f64>().ok())
|
||||
+}
|
||||
+
|
||||
+/// The core element threshold for a memory budget in GB (upstream's own figure for the budget, +4, as it
|
||||
+/// computed it: a 32 GB card is 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16).
|
||||
+pub fn element_threshold_for_budget(gpu_memory_gb: usize, full_size_shards: bool) -> u64 {
|
||||
+ if gpu_memory_gb > 30 || (full_size_shards && gpu_memory_gb >= 24) {
|
||||
+ ELEMENT_THRESHOLD
|
||||
+ } else if gpu_memory_gb >= 24 {
|
||||
+ ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24)
|
||||
+ } else if gpu_memory_gb >= 18 {
|
||||
+ (1 << 27) + (1 << 26)
|
||||
+ } else {
|
||||
+ 1 << 27
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+/// The recursion trace allocation (elements) for a memory budget.
|
||||
+pub fn recursion_trace_allocation_for_budget(gpu_memory_gb: usize) -> usize {
|
||||
+ if gpu_memory_gb >= 24 {
|
||||
+ RECURSION_TRACE_ALLOCATION
|
||||
+ } else {
|
||||
+ RECURSION_TRACE_ALLOCATION
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+pub fn gpu_memory_gb() -> usize {
|
||||
+ let gb = 1024.0 * 1024.0 * 1024.0;
|
||||
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
|
||||
+ Some(b) => (b.ceil() as usize) + 4,
|
||||
+ None => (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4,
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+pub fn recursion_trace_allocation() -> usize {
|
||||
+ env_usize("SP1_GPU_RECURSION_TRACE_ALLOCATION")
|
||||
+ .unwrap_or_else(|| recursion_trace_allocation_for_budget(gpu_memory_gb()))
|
||||
+}
|
||||
+
|
||||
pub fn local_gpu_opts() -> SP1CoreOpts {
|
||||
let mut opts = SP1CoreOpts::default();
|
||||
|
||||
let log2_shard_size = 24;
|
||||
opts.shard_size = 1 << log2_shard_size;
|
||||
|
||||
- let gb = 1024.0 * 1024.0 * 1024.0;
|
||||
-
|
||||
- // Get the amount of memory on the GPU.
|
||||
- let gpu_memory_gb: usize = (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4;
|
||||
-
|
||||
- if gpu_memory_gb < 24 {
|
||||
- panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB");
|
||||
- }
|
||||
+ // The card's memory plus 4, as upstream computed it (a 32 GB card reads 36), or the budget given.
|
||||
+ let gpu_memory_gb = gpu_memory_gb();
|
||||
|
||||
- let shard_threshold = if !opts.full_size_shards && gpu_memory_gb <= 30 {
|
||||
- ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24)
|
||||
- } else {
|
||||
- ELEMENT_THRESHOLD
|
||||
+ let shard_threshold = match env_usize("SP1_GPU_ELEMENT_THRESHOLD") {
|
||||
+ Some(t) => t as u64,
|
||||
+ None => element_threshold_for_budget(gpu_memory_gb, opts.full_size_shards),
|
||||
};
|
||||
+ let height_threshold = opts.sharding_threshold.height_threshold;
|
||||
|
||||
- tracing::debug!("Shard threshold: {shard_threshold}");
|
||||
+ eprintln!(
|
||||
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
|
||||
+ recursion_trace_allocation(),
|
||||
+ opts.full_size_shards
|
||||
+ );
|
||||
opts.sharding_threshold.element_threshold = shard_threshold;
|
||||
|
||||
opts.global_dependencies_opt = true;
|
||||
@@ -92,7 +139,7 @@ pub async fn recursion_prover_and_verifier(
|
||||
) {
|
||||
let recursion_verifier = SP1CudaProverComponents::compress_verifier();
|
||||
(
|
||||
- new_cuda_prover(&recursion_verifier, RECURSION_TRACE_ALLOCATION, 4, false, false, scope)
|
||||
+ new_cuda_prover(&recursion_verifier, recursion_trace_allocation(), 4, false, false, scope)
|
||||
.await,
|
||||
recursion_verifier,
|
||||
)
|
||||
diff --git a/sp1-gpu/crates/server/src/server.rs b/sp1-gpu/crates/server/src/server.rs
|
||||
index 4035f1f..0d0d907 100644
|
||||
--- a/sp1-gpu/crates/server/src/server.rs
|
||||
+++ b/sp1-gpu/crates/server/src/server.rs
|
||||
@@ -157,6 +157,7 @@ impl Server {
|
||||
};
|
||||
let pk = CachedProgram { elf: Arc::new(Elf::Dynamic(elf.into())), vk: vk.clone() };
|
||||
ctx.pk_cache.insert(elf_hash, pk);
|
||||
+ floor_memory_line("after setup");
|
||||
Response::Setup { id: elf_hash, vk }
|
||||
}
|
||||
Request::Destroy { key } => {
|
||||
@@ -177,15 +178,31 @@ impl Server {
|
||||
);
|
||||
};
|
||||
let context = SP1Context::builder().proof_nonce(proof_nonce).build();
|
||||
- match prover.prove_with_mode(&cached.elf, stdin, context, mode).await {
|
||||
+ let started = std::time::Instant::now();
|
||||
+ let response = match prover.prove_with_mode(&cached.elf, stdin, context, mode).await {
|
||||
Ok(proof) => Response::Proof { proof },
|
||||
Err(e) => Response::ProverError(e.to_string()),
|
||||
- }
|
||||
+ };
|
||||
+ floor_memory_line(&format!("after prove {:?} in {:.1} s", mode, started.elapsed().as_secs_f64()));
|
||||
+ response
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+/// Igneum prover-floor patch: the device memory in use (total minus free, as the driver reports it) at the
|
||||
+/// points that bound a proof, so a run's log carries the terms of the peak without a sampler.
|
||||
+fn floor_memory_line(what: &str) {
|
||||
+ if let Ok((free, total)) = sp1_gpu_cudart::cuda_memory_info() {
|
||||
+ eprintln!(
|
||||
+ "FLOOR memory {what}: device_used_mib={} free_mib={} total_mib={}",
|
||||
+ (total - free) >> 20,
|
||||
+ free >> 20,
|
||||
+ total >> 20
|
||||
+ );
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
fn sha256(data: &[u8]) -> [u8; 32] {
|
||||
use sha2::{Digest, Sha256};
|
||||
let mut hasher = Sha256::new();
|
||||
83
tools/prover-floor/make-build-playbook.sh
Executable file
83
tools/prover-floor/make-build-playbook.sh
Executable file
|
|
@ -0,0 +1,83 @@
|
|||
#!/usr/bin/env bash
|
||||
# Writes tools/prover-floor/pc2-build-server.ps1 with proving/prover-floor/sp1-gpu-6.8.1-floor.patch embedded
|
||||
# (base64), so the playbook and the patch cannot drift. Run after every change to the patch.
|
||||
set -euo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"; ROOT="$(cd "$HERE/../.." && pwd)"
|
||||
PATCH="$ROOT/proving/prover-floor/sp1-gpu-6.8.1-floor.patch"
|
||||
SHA="$(shasum -a 256 "$PATCH" | cut -c1-64)"
|
||||
B64="$(base64 < "$PATCH" | tr -d '\n')"
|
||||
OUT="$HERE/pc2-build-server.ps1"
|
||||
cat > "$OUT" <<PS
|
||||
# Prover floor (5 October 2026): builds sp1-gpu-server 6.8.1 from source with the Igneum floor patch on PC 2 inside
|
||||
# WSL2 Ubuntu-24.04 as root, into /opt/igneum-floor (the live /root/.sp1/bin/sp1-gpu-server and /opt/igneum are
|
||||
# never touched). CPU only: the miners keep mining, the card is not used. Generated by make-build-playbook.sh;
|
||||
# patch sha256 $SHA. CUDA_ARCHS=86,89,120 (the 12 GB tier is sm_86 (3060) and sm_89 (4070), the 16 GB tier sm_89 and sm_120; PC 2 runs sm_120; a shipped build lists every target the stock server does: 80, 86, 89, 90, 100, 120).
|
||||
\$ErrorActionPreference = 'Continue'
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
"RESULT start \$(Stamp)"
|
||||
\$job = \$env:IGNEUM_JOB_DIR; if (-not \$job) { \$job = Join-Path \$env:TEMP 'igneum-floor' }; New-Item -ItemType Directory -Force -Path \$job | Out-Null
|
||||
function WslPath(\$p) { \$w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a (\$p -replace '\\\\', '/') 2>\$null); if (\$w) { (\$w -replace "\`0", '').Trim() } else { '/mnt/c' + (\$p.Substring(2) -replace '\\\\', '/') } }
|
||||
\$patchB64 = '$B64'
|
||||
[IO.File]::WriteAllBytes((Join-Path \$job 'floor.patch'), [Convert]::FromBase64String(\$patchB64))
|
||||
\$patchW = WslPath (Join-Path \$job 'floor.patch')
|
||||
\$bash = @'
|
||||
set -uo pipefail
|
||||
export PATH="\$HOME/.cargo/bin:/usr/local/go/bin:\$PATH"
|
||||
CUDA_DIR="\$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"
|
||||
if [ -z "\$CUDA_DIR" ]; then echo "RESULT build_failed no /usr/local/cuda-12.*"; exit 2; fi
|
||||
export CUDA_PATH="\$CUDA_DIR" CUDACXX="\$CUDA_DIR/bin/nvcc" PATH="\$CUDA_DIR/bin:\$PATH" LD_LIBRARY_PATH="\$CUDA_DIR/lib64:/usr/lib/wsl/lib:\${LD_LIBRARY_PATH:-}"
|
||||
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
|
||||
FLOOR=/opt/igneum-floor; SRC=\$FLOOR/sp1; mkdir -p \$FLOOR/bin \$FLOOR/home/.sp1/bin \$FLOOR/logs
|
||||
PATCH="PATCH_PATH_PLACEHOLDER"
|
||||
echo "RESULT patch_sha256 \$(sha256sum "\$PATCH" | cut -c1-64) expected $SHA"
|
||||
echo "RESULT toolchain nvcc=\$(nvcc --version | grep -o 'release [0-9.]*') cargo=\$(cargo --version) go=\$(go version 2>/dev/null || echo MISSING) protoc=\$(protoc --version 2>/dev/null || echo MISSING) cmake=\$(cmake --version 2>/dev/null | head -1 || echo MISSING) nproc=\$(nproc) disk_avail=\$(df -BG /opt | awk 'NR==2 {print \$4}')"
|
||||
if [ ! -d "\$SRC/.git" ]; then
|
||||
echo "STAGE clone \$(stamp)"
|
||||
git clone -q --depth 1 --branch v6.8.1 https://github.com/succinctlabs/sp1 "\$SRC" 2>&1 | tail -2 || { echo "RESULT build_failed clone"; exit 2; }
|
||||
fi
|
||||
cd "\$SRC"
|
||||
git checkout -q -- . && git clean -qfd sp1-gpu/crates >/dev/null 2>&1
|
||||
echo "RESULT source \$(git describe --tags --always) \$(git rev-parse HEAD)"
|
||||
git apply "\$PATCH" || { echo "RESULT build_failed patch does not apply"; exit 2; }
|
||||
touch sp1-gpu/crates/prover_components/src/builder.rs sp1-gpu/crates/jagged_tracegen/src/lib.rs sp1-gpu/crates/server/src/server.rs
|
||||
echo "RESULT patched \$(git diff --stat | tail -1)"
|
||||
# Go, as the release workflow installs it (the server's native-gnark feature compiles the gnark library with go;
|
||||
# the wrap path is never run by a compressed proof, but the feature is upstream's and stays): a pinned tarball
|
||||
# unpacked under /opt/igneum-floor, on this job's PATH only, no package installed on the PC.
|
||||
if ! command -v go >/dev/null 2>&1; then
|
||||
if [ ! -x \$FLOOR/go/bin/go ]; then
|
||||
echo "STAGE go \$(stamp) fetching go1.27.1 (70,553,950 bytes, sha256 63d339f0...)"
|
||||
curl -sSL -o \$FLOOR/go.tgz https://go.dev/dl/go1.27.1.linux-amd64.tar.gz || { echo "RESULT build_failed go download"; exit 2; }
|
||||
echo "63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445 \$FLOOR/go.tgz" | sha256sum -c - >/dev/null || { echo "RESULT build_failed go sha256 mismatch: \$(sha256sum \$FLOOR/go.tgz)"; exit 2; }
|
||||
echo "RESULT go_tarball sha256 \$(sha256sum \$FLOOR/go.tgz | cut -c1-64) matches 63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445"
|
||||
tar -xzf \$FLOOR/go.tgz -C \$FLOOR && rm -f \$FLOOR/go.tgz
|
||||
fi
|
||||
export PATH="\$FLOOR/go/bin:\$PATH" GOPATH=\$FLOOR/gopath GOCACHE=\$FLOOR/gocache GOFLAGS=-mod=mod
|
||||
fi
|
||||
echo "RESULT go \$(go version 2>&1 | head -1)"
|
||||
LOG=\$FLOOR/logs/build.log; : > "\$LOG"
|
||||
echo "RESULT dirs \$(ls -ld \$FLOOR \$FLOOR/logs 2>&1 | tr '\\n' ' ') pwd=\$(pwd)"
|
||||
export CUDA_ARCHS=86,89,120 CARGO_TARGET_DIR=\$FLOOR/target
|
||||
JOBS=8
|
||||
echo "STAGE build \$(stamp) jobs=\$JOBS"
|
||||
t0=\$(date +%s)
|
||||
nice -n 19 cargo build --release --bin sp1-gpu-server -j \$JOBS >> "\$LOG" 2>&1 &
|
||||
BP=\$!
|
||||
while kill -0 \$BP 2>/dev/null; do sleep 120; echo "STAGE building \$(stamp) \$(( (\$(date +%s) - t0) / 60 )) min: \$(grep -c '^ Compiling' "\$LOG") crates compiled, last: \$(grep '^ Compiling' "\$LOG" | tail -1 | tr -s ' ' | cut -c1-80)"; done
|
||||
wait \$BP; rc=\$?
|
||||
echo "RESULT build_exit \$rc time_s=\$(( \$(date +%s) - t0 )) crates=\$(grep -c '^ Compiling' "\$LOG")"
|
||||
if [ \$rc -ne 0 ]; then echo "RESULT build_failed"; echo "== error lines =="; grep -n -B2 -A12 -E '^(error|warning: unused manifest| process didn|caused by)' "\$LOG" | head -100; echo "== log tail =="; tail -n 60 "\$LOG"; exit 2; fi
|
||||
BIN=\$FLOOR/target/release/sp1-gpu-server
|
||||
cp "\$BIN" \$FLOOR/bin/sp1-gpu-server && cp "\$BIN" \$FLOOR/home/.sp1/bin/sp1-gpu-server && chmod +x \$FLOOR/bin/sp1-gpu-server \$FLOOR/home/.sp1/bin/sp1-gpu-server
|
||||
echo "RESULT binary bytes=\$(stat -c %s "\$BIN") sha256=\$(sha256sum "\$BIN" | cut -c1-64) version=\$(\$BIN --version 2>/dev/null)"
|
||||
echo "RESULT elf_targets \$(cuobjdump --list-elf "\$BIN" 2>/dev/null | grep -o 'sm_[0-9]*' | sort -u | tr '\n' ' ')"
|
||||
echo "RESULT live_server_untouched sha256=\$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) (expected c2642ad1c42e85d8)"
|
||||
echo "RESULT end \$(stamp)"
|
||||
'@
|
||||
\$bash = \$bash.Replace('PATCH_PATH_PLACEHOLDER', \$patchW)
|
||||
\$bashFile = Join-Path \$job 'build.sh'
|
||||
[IO.File]::WriteAllText(\$bashFile, (\$bash -replace "\`r\`n", "\`n"), (New-Object System.Text.UTF8Encoding \$false))
|
||||
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath \$bashFile) 2>&1 | ForEach-Object { (\$_ -replace "\`0", '') }
|
||||
"RESULT end \$(Stamp)"
|
||||
PS
|
||||
echo "wrote $OUT ($(wc -c < "$OUT") bytes, patch $SHA)"
|
||||
15
tools/prover-floor/make-measure-playbook.sh
Executable file
15
tools/prover-floor/make-measure-playbook.sh
Executable file
|
|
@ -0,0 +1,15 @@
|
|||
#!/usr/bin/env bash
|
||||
# Writes a measurement playbook from pc2-floor-measure.ps1 with the point list given (one `run name fixture env...`
|
||||
# per line, read from the file in $1), into $2. Fixtures: $EMPTY (block 83616, an empty live shard), $V1
|
||||
# (fees-v1-shards2 shard 0, the adopted v1 shard), $FULL (block-338-shard1, the prototype shard), $ONE (block 56).
|
||||
set -euo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
POINTS="$(tr '\n' ';' < "$1" | sed 's/;$//')"
|
||||
python3 - "$HERE/pc2-floor-measure.ps1" "$2" "$POINTS" <<'PY'
|
||||
import sys
|
||||
src, out, points = sys.argv[1], sys.argv[2], sys.argv[3]
|
||||
s = open(src).read()
|
||||
s = s.replace("'POINTS_PLACEHOLDER'", "'" + points.replace("'", "''") + "'")
|
||||
open(out, 'w').write(s)
|
||||
print("wrote", out, "points:", points.count(';') + 1)
|
||||
PY
|
||||
71
tools/prover-floor/pc2-build-server.ps1
Normal file
71
tools/prover-floor/pc2-build-server.ps1
Normal file
File diff suppressed because one or more lines are too long
65
tools/prover-floor/pc2-floor-measure.ps1
Normal file
65
tools/prover-floor/pc2-floor-measure.ps1
Normal file
|
|
@ -0,0 +1,65 @@
|
|||
# Prover floor (5 October 2026): the peak GPU memory and the time of one compressed shard proof through the PATCHED
|
||||
# sp1-gpu-server (/opt/igneum-floor/home/.sp1/bin, reached by HOME=/opt/igneum-floor/home: the SDK spawns the
|
||||
# server it finds under $HOME/.sp1/bin, sp1-cuda-6.8.1/src/server.rs) on PC 2's RTX 5090, the miners STOPPED by the
|
||||
# job (--stop-miners) and the live prover switched off for the run (its server would otherwise own the socket).
|
||||
# Every point: every server killed and its socket unlinked, a 1-s nvidia-smi sampler, one `--mode compressed
|
||||
# --shard 0` of the pv1 host (/opt/igneum-pv1, the UNPATCHED SDK and verifier: its VERIFIED is the unpatched
|
||||
# verifier's word on the patched server's proof), the peak, the time, the FLOOR lines the server prints.
|
||||
# The point list comes from the FLOOR_POINTS environment the job carries, else the default sweep below.
|
||||
$ErrorActionPreference = 'Continue'
|
||||
$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/')
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } }
|
||||
"RESULT start $(Stamp) prover off for the run: $(Prove $false)"
|
||||
Start-Sleep -Seconds 45
|
||||
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,power.draw --format=csv,noheader,nounits 2>$null) -join ' | ')"
|
||||
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor-measure' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
|
||||
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
|
||||
$jobW = WslPath $job
|
||||
$emptyW = WslPath (Join-Path $env:LOCALAPPDATA 'igneum\app\jobs\chain-pc2-pv1c\block-83616.json')
|
||||
$points = if ($env:FLOOR_POINTS) { $env:FLOOR_POINTS } else { 'POINTS_PLACEHOLDER' }
|
||||
$bash = @'
|
||||
set -uo pipefail
|
||||
export PATH="$HOME/.cargo/bin:$PATH"
|
||||
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH" && export LD_LIBRARY_PATH="$CUDA_DIR/lib64:/usr/lib/wsl/lib:${LD_LIBRARY_PATH:-}"
|
||||
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
|
||||
JOB='JOBW_PLACEHOLDER'; H=/opt/igneum-pv1/igneum-prove-host; FX="/root/igneum-prove-pv1/proving/fixtures"
|
||||
FLOORHOME=/opt/igneum-floor/home; SRV=$FLOORHOME/.sp1/bin/sp1-gpu-server
|
||||
EMPTY='EMPTY_PLACEHOLDER'; V1="$FX/fees-v1-shards2.json"; FULL="$FX/block-338-shard1.json"; ONE="$FX/block-56-transfers.json"
|
||||
[ -x "$SRV" ] || { echo "RESULT measure_failed no patched server at $SRV"; exit 2; }
|
||||
[ -x "$H" ] || { echo "RESULT measure_failed no pv1 host at $H"; exit 2; }
|
||||
echo "RESULT patched_server sha256=$(sha256sum $SRV | cut -c1-64) version=$($SRV --version 2>/dev/null) host=$(sha256sum $H | cut -c1-16)"
|
||||
echo "RESULT live_server sha256=$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) untouched"
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
|
||||
echo "RESULT idle_mib $(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | head -1)"
|
||||
run() { # name fixture env...
|
||||
local name="$1" fx="$2"; shift 2
|
||||
local tag="$name-$(basename $fx .json)"
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
|
||||
local csv="$JOB/smi-$tag.csv" log="$JOB/log-$tag.txt"
|
||||
nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv,noheader,nounits -l 1 > "$csv" 2>/dev/null &
|
||||
local SMI=$!
|
||||
local t0=$(date +%s)
|
||||
env HOME=$FLOORHOME SP1_PROVER=cuda RUST_LOG=off SP1_GPU_FLOOR_LOG=1 "$@" $H "$fx" --mode compressed --shard 0 --out "$JOB/res-$tag.json" > "$log" 2>&1
|
||||
local rc=$?
|
||||
local wall=$(( $(date +%s) - t0 ))
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; kill $SMI 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
|
||||
local peak=$(awk -F', *' '{ if ($2+0 > m) m=$2+0 } END { print m+0 }' "$csv")
|
||||
local n=$(wc -l < "$csv")
|
||||
local line=$(grep -E "^RESULT compressed shard" "$log" | tail -1 | sed -E 's/.*prove ([0-9.]+) s, proof ([0-9]+) bytes, verify ([0-9.]+) s, ([A-Z ]+);.*/prove_s=\1 bytes=\2 verify_s=\3 \4/')
|
||||
local cyc=$(grep -E "^RESULT execute shard" "$log" | tail -1 | sed -E 's/.*: ([0-9]+) cycles.*/\1/')
|
||||
local err=$(grep -iE "error|panick|out of memory|OOM|unsupported" "$log" | grep -v "^FLOOR" | head -1 | cut -c1-200)
|
||||
echo "RESULT floor cfg=$name fixture=$(basename $fx .json) peak_mib=$peak samples=$n wall_s=$wall cycles=${cyc:-na} ${line:-no_result} exit=$rc env='$*' ${err:+err=$err}"
|
||||
grep -E "^FLOOR" "$log" | sed "s/^/RESULT floorline cfg=$name fixture=$(basename $fx .json) /" | head -40
|
||||
}
|
||||
POINTS_PLACEHOLDER_BASH
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
|
||||
echo "RESULT measure_end $(stamp)"
|
||||
'@
|
||||
$bash = $bash.Replace('JOBW_PLACEHOLDER', $jobW).Replace('EMPTY_PLACEHOLDER', $emptyW).Replace('POINTS_PLACEHOLDER_BASH', ($points -replace ';', "`n"))
|
||||
$bashFile = Join-Path $job 'measure.sh'
|
||||
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
|
||||
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
|
||||
"RESULT end $(Stamp) prover back on: $(Prove $true)"
|
||||
65
tools/prover-floor/pc2-floor-sweep1.ps1
Normal file
65
tools/prover-floor/pc2-floor-sweep1.ps1
Normal file
|
|
@ -0,0 +1,65 @@
|
|||
# Prover floor (5 October 2026): the peak GPU memory and the time of one compressed shard proof through the PATCHED
|
||||
# sp1-gpu-server (/opt/igneum-floor/home/.sp1/bin, reached by HOME=/opt/igneum-floor/home: the SDK spawns the
|
||||
# server it finds under $HOME/.sp1/bin, sp1-cuda-6.8.1/src/server.rs) on PC 2's RTX 5090, the miners STOPPED by the
|
||||
# job (--stop-miners) and the live prover switched off for the run (its server would otherwise own the socket).
|
||||
# Every point: every server killed and its socket unlinked, a 1-s nvidia-smi sampler, one `--mode compressed
|
||||
# --shard 0` of the pv1 host (/opt/igneum-pv1, the UNPATCHED SDK and verifier: its VERIFIED is the unpatched
|
||||
# verifier's word on the patched server's proof), the peak, the time, the FLOOR lines the server prints.
|
||||
# The point list comes from the FLOOR_POINTS environment the job carries, else the default sweep below.
|
||||
$ErrorActionPreference = 'Continue'
|
||||
$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/')
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } }
|
||||
"RESULT start $(Stamp) prover off for the run: $(Prove $false)"
|
||||
Start-Sleep -Seconds 45
|
||||
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,power.draw --format=csv,noheader,nounits 2>$null) -join ' | ')"
|
||||
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor-measure' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
|
||||
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
|
||||
$jobW = WslPath $job
|
||||
$emptyW = WslPath (Join-Path $env:LOCALAPPDATA 'igneum\app\jobs\chain-pc2-pv1c\block-83616.json')
|
||||
$points = if ($env:FLOOR_POINTS) { $env:FLOOR_POINTS } else { 'run ctrl32 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=32;run ctrl32 "$V1" SP1_GPU_MEMORY_BUDGET_GB=32;run b12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12;run b12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12;run b16 "$V1" SP1_GPU_MEMORY_BUDGET_GB=16;run e26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864;run e26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864;run e25 "$V1" SP1_GPU_ELEMENT_THRESHOLD=33554432;run b12c1 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1' }
|
||||
$bash = @'
|
||||
set -uo pipefail
|
||||
export PATH="$HOME/.cargo/bin:$PATH"
|
||||
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH" && export LD_LIBRARY_PATH="$CUDA_DIR/lib64:/usr/lib/wsl/lib:${LD_LIBRARY_PATH:-}"
|
||||
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
|
||||
JOB='JOBW_PLACEHOLDER'; H=/opt/igneum-pv1/igneum-prove-host; FX="/root/igneum-prove-pv1/proving/fixtures"
|
||||
FLOORHOME=/opt/igneum-floor/home; SRV=$FLOORHOME/.sp1/bin/sp1-gpu-server
|
||||
EMPTY='EMPTY_PLACEHOLDER'; V1="$FX/fees-v1-shards2.json"; FULL="$FX/block-338-shard1.json"; ONE="$FX/block-56-transfers.json"
|
||||
[ -x "$SRV" ] || { echo "RESULT measure_failed no patched server at $SRV"; exit 2; }
|
||||
[ -x "$H" ] || { echo "RESULT measure_failed no pv1 host at $H"; exit 2; }
|
||||
echo "RESULT patched_server sha256=$(sha256sum $SRV | cut -c1-64) version=$($SRV --version 2>/dev/null) host=$(sha256sum $H | cut -c1-16)"
|
||||
echo "RESULT live_server sha256=$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) untouched"
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
|
||||
echo "RESULT idle_mib $(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | head -1)"
|
||||
run() { # name fixture env...
|
||||
local name="$1" fx="$2"; shift 2
|
||||
local tag="$name-$(basename $fx .json)"
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
|
||||
local csv="$JOB/smi-$tag.csv" log="$JOB/log-$tag.txt"
|
||||
nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv,noheader,nounits -l 1 > "$csv" 2>/dev/null &
|
||||
local SMI=$!
|
||||
local t0=$(date +%s)
|
||||
env HOME=$FLOORHOME SP1_PROVER=cuda RUST_LOG=off SP1_GPU_FLOOR_LOG=1 "$@" $H "$fx" --mode compressed --shard 0 --out "$JOB/res-$tag.json" > "$log" 2>&1
|
||||
local rc=$?
|
||||
local wall=$(( $(date +%s) - t0 ))
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; kill $SMI 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
|
||||
local peak=$(awk -F', *' '{ if ($2+0 > m) m=$2+0 } END { print m+0 }' "$csv")
|
||||
local n=$(wc -l < "$csv")
|
||||
local line=$(grep -E "^RESULT compressed shard" "$log" | tail -1 | sed -E 's/.*prove ([0-9.]+) s, proof ([0-9]+) bytes, verify ([0-9.]+) s, ([A-Z ]+);.*/prove_s=\1 bytes=\2 verify_s=\3 \4/')
|
||||
local cyc=$(grep -E "^RESULT execute shard" "$log" | tail -1 | sed -E 's/.*: ([0-9]+) cycles.*/\1/')
|
||||
local err=$(grep -iE "error|panick|out of memory|OOM|unsupported" "$log" | grep -v "^FLOOR" | head -1 | cut -c1-200)
|
||||
echo "RESULT floor cfg=$name fixture=$(basename $fx .json) peak_mib=$peak samples=$n wall_s=$wall cycles=${cyc:-na} ${line:-no_result} exit=$rc env='$*' ${err:+err=$err}"
|
||||
grep -E "^FLOOR" "$log" | sed "s/^/RESULT floorline cfg=$name fixture=$(basename $fx .json) /" | head -40
|
||||
}
|
||||
POINTS_PLACEHOLDER_BASH
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
|
||||
echo "RESULT measure_end $(stamp)"
|
||||
'@
|
||||
$bash = $bash.Replace('JOBW_PLACEHOLDER', $jobW).Replace('EMPTY_PLACEHOLDER', $emptyW).Replace('POINTS_PLACEHOLDER_BASH', ($points -replace ';', "`n"))
|
||||
$bashFile = Join-Path $job 'measure.sh'
|
||||
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
|
||||
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
|
||||
"RESULT end $(Stamp) prover back on: $(Prove $true)"
|
||||
65
tools/prover-floor/pc2-floor-sweep2.ps1
Normal file
65
tools/prover-floor/pc2-floor-sweep2.ps1
Normal file
|
|
@ -0,0 +1,65 @@
|
|||
# Prover floor (5 October 2026): the peak GPU memory and the time of one compressed shard proof through the PATCHED
|
||||
# sp1-gpu-server (/opt/igneum-floor/home/.sp1/bin, reached by HOME=/opt/igneum-floor/home: the SDK spawns the
|
||||
# server it finds under $HOME/.sp1/bin, sp1-cuda-6.8.1/src/server.rs) on PC 2's RTX 5090, the miners STOPPED by the
|
||||
# job (--stop-miners) and the live prover switched off for the run (its server would otherwise own the socket).
|
||||
# Every point: every server killed and its socket unlinked, a 1-s nvidia-smi sampler, one `--mode compressed
|
||||
# --shard 0` of the pv1 host (/opt/igneum-pv1, the UNPATCHED SDK and verifier: its VERIFIED is the unpatched
|
||||
# verifier's word on the patched server's proof), the peak, the time, the FLOOR lines the server prints.
|
||||
# The point list comes from the FLOOR_POINTS environment the job carries, else the default sweep below.
|
||||
$ErrorActionPreference = 'Continue'
|
||||
$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/')
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } }
|
||||
"RESULT start $(Stamp) prover off for the run: $(Prove $false)"
|
||||
Start-Sleep -Seconds 45
|
||||
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,power.draw --format=csv,noheader,nounits 2>$null) -join ' | ')"
|
||||
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor-measure' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
|
||||
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
|
||||
$jobW = WslPath $job
|
||||
$emptyW = WslPath (Join-Path $env:LOCALAPPDATA 'igneum\app\jobs\chain-pc2-pv1c\block-83616.json')
|
||||
$points = if ($env:FLOOR_POINTS) { $env:FLOOR_POINTS } else { 'run x12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12;run x12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12;run x12 "$ONE" SP1_GPU_MEMORY_BUDGET_GB=12;run x12 "$FULL" SP1_GPU_MEMORY_BUDGET_GB=12;run xe26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864;run xe26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864;run x16 "$V1" SP1_GPU_MEMORY_BUDGET_GB=16;run x32 "$V1" SP1_GPU_MEMORY_BUDGET_GB=32;run x12r "$V1" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_GPU_RECURSION_TRACE_ALLOCATION=100663296' }
|
||||
$bash = @'
|
||||
set -uo pipefail
|
||||
export PATH="$HOME/.cargo/bin:$PATH"
|
||||
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH" && export LD_LIBRARY_PATH="$CUDA_DIR/lib64:/usr/lib/wsl/lib:${LD_LIBRARY_PATH:-}"
|
||||
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
|
||||
JOB='JOBW_PLACEHOLDER'; H=/opt/igneum-pv1/igneum-prove-host; FX="/root/igneum-prove-pv1/proving/fixtures"
|
||||
FLOORHOME=/opt/igneum-floor/home; SRV=$FLOORHOME/.sp1/bin/sp1-gpu-server
|
||||
EMPTY='EMPTY_PLACEHOLDER'; V1="$FX/fees-v1-shards2.json"; FULL="$FX/block-338-shard1.json"; ONE="$FX/block-56-transfers.json"
|
||||
[ -x "$SRV" ] || { echo "RESULT measure_failed no patched server at $SRV"; exit 2; }
|
||||
[ -x "$H" ] || { echo "RESULT measure_failed no pv1 host at $H"; exit 2; }
|
||||
echo "RESULT patched_server sha256=$(sha256sum $SRV | cut -c1-64) version=$($SRV --version 2>/dev/null) host=$(sha256sum $H | cut -c1-16)"
|
||||
echo "RESULT live_server sha256=$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) untouched"
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
|
||||
echo "RESULT idle_mib $(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | head -1)"
|
||||
run() { # name fixture env...
|
||||
local name="$1" fx="$2"; shift 2
|
||||
local tag="$name-$(basename $fx .json)"
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
|
||||
local csv="$JOB/smi-$tag.csv" log="$JOB/log-$tag.txt"
|
||||
nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv,noheader,nounits -l 1 > "$csv" 2>/dev/null &
|
||||
local SMI=$!
|
||||
local t0=$(date +%s)
|
||||
env HOME=$FLOORHOME SP1_PROVER=cuda RUST_LOG=off SP1_GPU_FLOOR_LOG=1 "$@" $H "$fx" --mode compressed --shard 0 --out "$JOB/res-$tag.json" > "$log" 2>&1
|
||||
local rc=$?
|
||||
local wall=$(( $(date +%s) - t0 ))
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; kill $SMI 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
|
||||
local peak=$(awk -F', *' '{ if ($2+0 > m) m=$2+0 } END { print m+0 }' "$csv")
|
||||
local n=$(wc -l < "$csv")
|
||||
local line=$(grep -E "^RESULT compressed shard" "$log" | tail -1 | sed -E 's/.*prove ([0-9.]+) s, proof ([0-9]+) bytes, verify ([0-9.]+) s, ([A-Z ]+);.*/prove_s=\1 bytes=\2 verify_s=\3 \4/')
|
||||
local cyc=$(grep -E "^RESULT execute shard" "$log" | tail -1 | sed -E 's/.*: ([0-9]+) cycles.*/\1/')
|
||||
local err=$(grep -iE "error|panick|out of memory|OOM|unsupported" "$log" | grep -v "^FLOOR" | head -1 | cut -c1-200)
|
||||
echo "RESULT floor cfg=$name fixture=$(basename $fx .json) peak_mib=$peak samples=$n wall_s=$wall cycles=${cyc:-na} ${line:-no_result} exit=$rc env='$*' ${err:+err=$err}"
|
||||
grep -E "^FLOOR" "$log" | sed "s/^/RESULT floorline cfg=$name fixture=$(basename $fx .json) /" | head -40
|
||||
}
|
||||
POINTS_PLACEHOLDER_BASH
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
|
||||
echo "RESULT measure_end $(stamp)"
|
||||
'@
|
||||
$bash = $bash.Replace('JOBW_PLACEHOLDER', $jobW).Replace('EMPTY_PLACEHOLDER', $emptyW).Replace('POINTS_PLACEHOLDER_BASH', ($points -replace ';', "`n"))
|
||||
$bashFile = Join-Path $job 'measure.sh'
|
||||
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
|
||||
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
|
||||
"RESULT end $(Stamp) prover back on: $(Prove $true)"
|
||||
66
tools/prover-floor/pc2-floor-sweep3-miner.ps1
Normal file
66
tools/prover-floor/pc2-floor-sweep3-miner.ps1
Normal file
|
|
@ -0,0 +1,66 @@
|
|||
# Published WITHOUT --stop-miners: the beside-the-miner pair (the 5090 mining at full rate on the same card).
|
||||
# Prover floor (5 October 2026): the peak GPU memory and the time of one compressed shard proof through the PATCHED
|
||||
# sp1-gpu-server (/opt/igneum-floor/home/.sp1/bin, reached by HOME=/opt/igneum-floor/home: the SDK spawns the
|
||||
# server it finds under $HOME/.sp1/bin, sp1-cuda-6.8.1/src/server.rs) on PC 2's RTX 5090, the miners STOPPED by the
|
||||
# job (--stop-miners) and the live prover switched off for the run (its server would otherwise own the socket).
|
||||
# Every point: every server killed and its socket unlinked, a 1-s nvidia-smi sampler, one `--mode compressed
|
||||
# --shard 0` of the pv1 host (/opt/igneum-pv1, the UNPATCHED SDK and verifier: its VERIFIED is the unpatched
|
||||
# verifier's word on the patched server's proof), the peak, the time, the FLOOR lines the server prints.
|
||||
# The point list comes from the FLOOR_POINTS environment the job carries, else the default sweep below.
|
||||
$ErrorActionPreference = 'Continue'
|
||||
$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/')
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } }
|
||||
"RESULT start $(Stamp) prover off for the run: $(Prove $false)"
|
||||
Start-Sleep -Seconds 45
|
||||
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,power.draw --format=csv,noheader,nounits 2>$null) -join ' | ')"
|
||||
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor-measure' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
|
||||
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
|
||||
$jobW = WslPath $job
|
||||
$emptyW = WslPath (Join-Path $env:LOCALAPPDATA 'igneum\app\jobs\chain-pc2-pv1c\block-83616.json')
|
||||
$points = if ($env:FLOOR_POINTS) { $env:FLOOR_POINTS } else { 'run m12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12;run m12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12;run me26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864;run me26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864' }
|
||||
$bash = @'
|
||||
set -uo pipefail
|
||||
export PATH="$HOME/.cargo/bin:$PATH"
|
||||
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH" && export LD_LIBRARY_PATH="$CUDA_DIR/lib64:/usr/lib/wsl/lib:${LD_LIBRARY_PATH:-}"
|
||||
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
|
||||
JOB='JOBW_PLACEHOLDER'; H=/opt/igneum-pv1/igneum-prove-host; FX="/root/igneum-prove-pv1/proving/fixtures"
|
||||
FLOORHOME=/opt/igneum-floor/home; SRV=$FLOORHOME/.sp1/bin/sp1-gpu-server
|
||||
EMPTY='EMPTY_PLACEHOLDER'; V1="$FX/fees-v1-shards2.json"; FULL="$FX/block-338-shard1.json"; ONE="$FX/block-56-transfers.json"
|
||||
[ -x "$SRV" ] || { echo "RESULT measure_failed no patched server at $SRV"; exit 2; }
|
||||
[ -x "$H" ] || { echo "RESULT measure_failed no pv1 host at $H"; exit 2; }
|
||||
echo "RESULT patched_server sha256=$(sha256sum $SRV | cut -c1-64) version=$($SRV --version 2>/dev/null) host=$(sha256sum $H | cut -c1-16)"
|
||||
echo "RESULT live_server sha256=$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) untouched"
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
|
||||
echo "RESULT idle_mib $(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | head -1)"
|
||||
run() { # name fixture env...
|
||||
local name="$1" fx="$2"; shift 2
|
||||
local tag="$name-$(basename $fx .json)"
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
|
||||
local csv="$JOB/smi-$tag.csv" log="$JOB/log-$tag.txt"
|
||||
nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv,noheader,nounits -l 1 > "$csv" 2>/dev/null &
|
||||
local SMI=$!
|
||||
local t0=$(date +%s)
|
||||
env HOME=$FLOORHOME SP1_PROVER=cuda RUST_LOG=off SP1_GPU_FLOOR_LOG=1 "$@" $H "$fx" --mode compressed --shard 0 --out "$JOB/res-$tag.json" > "$log" 2>&1
|
||||
local rc=$?
|
||||
local wall=$(( $(date +%s) - t0 ))
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; kill $SMI 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
|
||||
local peak=$(awk -F', *' '{ if ($2+0 > m) m=$2+0 } END { print m+0 }' "$csv")
|
||||
local n=$(wc -l < "$csv")
|
||||
local line=$(grep -E "^RESULT compressed shard" "$log" | tail -1 | sed -E 's/.*prove ([0-9.]+) s, proof ([0-9]+) bytes, verify ([0-9.]+) s, ([A-Z ]+);.*/prove_s=\1 bytes=\2 verify_s=\3 \4/')
|
||||
local cyc=$(grep -E "^RESULT execute shard" "$log" | tail -1 | sed -E 's/.*: ([0-9]+) cycles.*/\1/')
|
||||
local err=$(grep -iE "error|panick|out of memory|OOM|unsupported" "$log" | grep -v "^FLOOR" | head -1 | cut -c1-200)
|
||||
echo "RESULT floor cfg=$name fixture=$(basename $fx .json) peak_mib=$peak samples=$n wall_s=$wall cycles=${cyc:-na} ${line:-no_result} exit=$rc env='$*' ${err:+err=$err}"
|
||||
grep -E "^FLOOR" "$log" | sed "s/^/RESULT floorline cfg=$name fixture=$(basename $fx .json) /" | head -40
|
||||
}
|
||||
POINTS_PLACEHOLDER_BASH
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
|
||||
echo "RESULT measure_end $(stamp)"
|
||||
'@
|
||||
$bash = $bash.Replace('JOBW_PLACEHOLDER', $jobW).Replace('EMPTY_PLACEHOLDER', $emptyW).Replace('POINTS_PLACEHOLDER_BASH', ($points -replace ';', "`n"))
|
||||
$bashFile = Join-Path $job 'measure.sh'
|
||||
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
|
||||
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
|
||||
"RESULT end $(Stamp) prover back on: $(Prove $true)"
|
||||
47
tools/prover-floor/pc2-toolchain.ps1
Normal file
47
tools/prover-floor/pc2-toolchain.ps1
Normal file
|
|
@ -0,0 +1,47 @@
|
|||
# Prover floor (5 October 2026, the project lead: "execute if it will solve the issue", the 12 GB cards): the toolchain check
|
||||
# before the sp1-gpu-server source build on PC 2. Reads versions and free space inside WSL2 Ubuntu-24.04 as root.
|
||||
# Touches nothing: no build, no GPU work (one nvidia-smi query), the live host and server untouched.
|
||||
$ErrorActionPreference = 'Continue'
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
"RESULT start $(Stamp)"
|
||||
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,driver_version --format=csv,noheader,nounits 2>$null) -join ' | ')"
|
||||
"RESULT host_ram_mb $([math]::Round((Get-CimInstance Win32_OperatingSystem).TotalVisibleMemorySize / 1024))"
|
||||
$bash = @'
|
||||
set -uo pipefail
|
||||
export PATH="$HOME/.cargo/bin:$HOME/.sp1/bin:/usr/local/go/bin:$PATH"
|
||||
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"
|
||||
[ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH"
|
||||
v() { local n="$1"; shift; if command -v "$1" >/dev/null 2>&1; then echo "RESULT tool $n $("$@" 2>&1 | head -1 | tr -s ' ' | cut -c1-120)"; else echo "RESULT tool $n MISSING"; fi; }
|
||||
echo "RESULT cuda_dirs $(ls -d /usr/local/cuda* 2>/dev/null | tr '\n' ' ')"
|
||||
v nvcc nvcc --version
|
||||
[ -n "$CUDA_DIR" ] && echo "RESULT nvcc_release $(nvcc --version 2>/dev/null | grep -o 'release [0-9.]*' | head -1)"
|
||||
v cmake cmake --version
|
||||
v gcc gcc --version
|
||||
v g++ g++ --version
|
||||
v clang clang --version
|
||||
v go go version
|
||||
v protoc protoc --version
|
||||
v cargo cargo --version
|
||||
v rustc rustc --version
|
||||
v git git --version
|
||||
v cuobjdump cuobjdump --version
|
||||
v pkg-config pkg-config --version
|
||||
echo "RESULT nproc $(nproc)"
|
||||
echo "RESULT mem $(free -g | awk '/Mem:/ {print "total_gb=" $2 " available_gb=" $7}')"
|
||||
echo "RESULT disk_root $(df -BG / | awk 'NR==2 {print "size=" $2 " used=" $3 " avail=" $4}')"
|
||||
echo "RESULT disk_opt $(df -BG /opt 2>/dev/null | awk 'NR==2 {print "avail=" $4}')"
|
||||
echo "RESULT registry $(du -sh $HOME/.cargo/registry 2>/dev/null | cut -f1) target_live $(du -sh /root/igneum-prove/proving/igneum-prove/target 2>/dev/null | cut -f1)"
|
||||
echo "RESULT floor_dir $(ls -d /opt/igneum-floor 2>/dev/null || echo absent)"
|
||||
echo "RESULT live_server $(ls -l /root/.sp1/bin/sp1-gpu-server 2>/dev/null | awk '{print $5}') sha256 $(sha256sum /root/.sp1/bin/sp1-gpu-server 2>/dev/null | cut -c1-16)"
|
||||
echo "RESULT live_server_version $($HOME/.sp1/bin/sp1-gpu-server --version 2>/dev/null)"
|
||||
echo "RESULT github $(timeout 20 git ls-remote --tags https://github.com/succinctlabs/sp1 refs/tags/v6.8.1 2>&1 | cut -c1-60)"
|
||||
echo "RESULT ld_libs $(ls /usr/lib/wsl/lib/libcuda.so* 2>/dev/null | tr '\n' ' ')"
|
||||
echo "RESULT wsl_user $(id -un) home $HOME"
|
||||
echo "RESULT end $(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
||||
'@
|
||||
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
|
||||
$bashFile = Join-Path $job 'toolchain.sh'
|
||||
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
|
||||
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
|
||||
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
|
||||
"RESULT end $(Stamp)"
|
||||
9
tools/prover-floor/points-sweep1.txt
Normal file
9
tools/prover-floor/points-sweep1.txt
Normal file
|
|
@ -0,0 +1,9 @@
|
|||
run ctrl32 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=32
|
||||
run ctrl32 "$V1" SP1_GPU_MEMORY_BUDGET_GB=32
|
||||
run b12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12
|
||||
run b12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12
|
||||
run b16 "$V1" SP1_GPU_MEMORY_BUDGET_GB=16
|
||||
run e26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864
|
||||
run e26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864
|
||||
run e25 "$V1" SP1_GPU_ELEMENT_THRESHOLD=33554432
|
||||
run b12c1 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1
|
||||
9
tools/prover-floor/points-sweep2.txt
Normal file
9
tools/prover-floor/points-sweep2.txt
Normal file
|
|
@ -0,0 +1,9 @@
|
|||
run x12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12
|
||||
run x12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12
|
||||
run x12 "$ONE" SP1_GPU_MEMORY_BUDGET_GB=12
|
||||
run x12 "$FULL" SP1_GPU_MEMORY_BUDGET_GB=12
|
||||
run xe26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864
|
||||
run xe26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864
|
||||
run x16 "$V1" SP1_GPU_MEMORY_BUDGET_GB=16
|
||||
run x32 "$V1" SP1_GPU_MEMORY_BUDGET_GB=32
|
||||
run x12r "$V1" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_GPU_RECURSION_TRACE_ALLOCATION=100663296
|
||||
4
tools/prover-floor/points-sweep3-miner.txt
Normal file
4
tools/prover-floor/points-sweep3-miner.txt
Normal file
|
|
@ -0,0 +1,4 @@
|
|||
run m12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12
|
||||
run m12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12
|
||||
run me26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864
|
||||
run me26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864
|
||||
Loading…
Reference in a new issue