Prover floor: SP1 6.8.1 GPU server memory model from source (the 24 GB panic, the threshold-sized trace buffers), the floor patch for sp1-gpu, the PC 2 toolchain, build and sweep playbooks
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
19a5e391cf
commit
5e26563cd3
10 changed files with 613 additions and 0 deletions
91
docs/analysis/prover-floor.md
Normal file
91
docs/analysis/prover-floor.md
Normal file
|
|
@ -0,0 +1,91 @@
|
|||
# The prover floor: why a 12 GB card cannot prove on SP1 6.8.1's GPU server, and the patch
|
||||
|
||||
5 October 2026, 22:00 UTC on (the project lead: "execute if it will solve the issue"). Branch `prover-floor`
|
||||
(worktree `igneum-wt-prover-floor`). The measured facts this starts from: `docs/plans/proving-v1.md` and the
|
||||
bench-log entry "proving v1" (branch proving-v1): the GPU server holds 13.9 GB for an empty shard, 20.4 GB at the
|
||||
adopted v1 shard, 28.3 GB flat from 20 M to 60 M cycles, and no environment knob moved the floor. Every figure
|
||||
below is from the source at tag v6.8.1 (cloned to `vendor/sp1-6.8.1`, gitignored; the fork is the patch
|
||||
`proving/prover-floor/sp1-gpu-6.8.1-floor.patch`) or from a PC 2 run named in `docs/bench-log.md`
|
||||
("prover floor"). Sizes in GiB are computed from the source constants (4-byte field elements); sizes in MiB are
|
||||
measured by `nvidia-smi` at 1 s.
|
||||
|
||||
## Where the server is built and what it reads
|
||||
|
||||
The SDK downloads `sp1_gpu_server_v6.8.1_x86_64.tar.gz` (133,750,780 bytes) from the SP1 release and runs it from
|
||||
`$HOME/.sp1/bin/sp1-gpu-server` (`crates/cuda/src/server.rs` 19 to 30, 80 to 99). The source is in the same
|
||||
repository: `sp1-gpu/crates/server` (the binary), built by `.github/workflows/release.yml` 234 to 314 on CUDA
|
||||
12.8.1 with Go and protoc (`cargo build --release --bin sp1-gpu-server`). The binary takes no options
|
||||
(`sp1-gpu/crates/server/src/main.rs` 15 to 18: `--version` only) and reads `CUDA_VISIBLE_DEVICES` (32 to 35);
|
||||
everything else comes from the environment the host process passes it, through `SP1CoreOpts::default()`
|
||||
(`crates/core/executor/src/opts.rs` 99 to 140: `SHARD_SIZE`, `ELEMENT_THRESHOLD`, `HEIGHT_THRESHOLD`,
|
||||
`MINIMAL_TRACE_CHUNK_THRESHOLD`, `TRACE_CHUNK_SLOTS`, `FULL_SIZE_SHARDS`) and the worker counts
|
||||
(`crates/prover/src/worker/config.rs`).
|
||||
|
||||
## The memory model, term by term
|
||||
|
||||
Every device buffer is sized at construction from constants, not from the shard. The server builds the prover at
|
||||
the first `Setup` request (`sp1-gpu/crates/server/src/server.rs` 126 to 137) through
|
||||
`cuda_worker_builder_with_machine` (`sp1-gpu/crates/prover_components/src/builder.rs` 102 to 148):
|
||||
|
||||
| Term | Where | Size | On the device | Moves with the shard |
|
||||
|---|---|---|---|---|
|
||||
| The gate | `builder.rs` 35 to 39: `gpu_memory_gb = ceil(total / GiB) + 4`; `panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB")` when under 24 | a 12 GB card reads 16, a 16 GB card 20: both refused before any allocation | | no |
|
||||
| The core element threshold | `builder.rs` 41 to 48: `ELEMENT_THRESHOLD` = 2^28 + 2^27 = 402,653,184 elements (`opts.rs` 12) on a card reading over 30 (a 32 GB card reads 36); minus 2^26 + 2^25 + 2^24 = 285,212,672 on a card reading 24 to 30 (a 24 GB card). The environment's `ELEMENT_THRESHOLD` is read at `opts.rs` 129 and then OVERWRITTEN at `builder.rs` 48, which is why the sweep's `ELEMENT_THRESHOLD` rows changed nothing; `HEIGHT_THRESHOLD` survives (it is not overwritten), which is why the 2^25 + 2^20 row did | sets the next two terms | | no |
|
||||
| The core trace area, one per shard in flight | `builder.rs` 70 to 71: `num_elts = element_threshold + 2^21` (`CORE_LOG_STACKING_HEIGHT` 21, `crates/prover/src/components.rs` 16) = 404,750,336; allocated on the device at `sp1-gpu/crates/jagged_tracegen/src/lib.rs` 484 to 500 (`allocate_and_initialize_traces`: `max_trace_size` felts + `max_trace_size / 2` u32 column index + 2^14 u32) | 6 bytes an element: **2.26 GiB** for the full threshold, 1.59 GiB for the 24 GB threshold | yes, in full, whatever the shard holds | no |
|
||||
| The program's preprocessed traces (the proving key) | `sp1-gpu/crates/shard_prover/src/setup.rs` 47 to 58 and 107: the same `allocate_and_initialize_traces(max_trace_size)` at `Setup`, kept in the key cache for the connection's life (`server.rs` 139 to 143) | another **2.26 GiB**, held from `Setup` on | yes | no |
|
||||
| The pinned host trace buffers | `sp1-gpu/crates/prover_components/src/components.rs` 99 to 103: 4 `PinnedBuffer` of `max_trace_size` felts per prover (core 4 x 1.51 GiB, recursion 4 x 0.5 GiB, shrink 4 x 0.125 GiB, wrap 4 x 0.32 GiB) | 9.8 GiB of pinned host RAM, not device memory (the WSL2 working set the bench saw) | no | no |
|
||||
| The recursion trace area | `builder.rs` 15 and 95: `RECURSION_TRACE_ALLOCATION` = 2^27 elements, one per recursion tracegen (the recursion program's key at setup and its shard at prove) | 0.75 GiB each | yes | no |
|
||||
| The shrink and wrap provers | `builder.rs` 16, 19, 117 to 127: 2^25 and 85,376,340 elements, built at `Setup` for every proof mode, used only by the Groth16 and PLONK path | host pinned at build; device only when a wrap runs (never, for a compressed proof) | no | no |
|
||||
| The codewords (LDE) and the Merkle trees | `sp1-gpu/crates/basefold/src/fri.rs` 92 to 97: every stacked column of 2^21 rows encoded to 2^(21 + 1) rows (`log_blowup` 1); kept until the query phase unless `drop_ldes` (`builder.rs` 52: only on a 24 GB card with `FULL_SIZE_SHARDS`); the preprocessed codewords live in the key | 2 x the padded trace, so up to 2 x the term above | yes | yes, with the padded trace |
|
||||
| The LogUp GKR layers | `builder.rs` 51: `recompute_gkr_trace = false`, so the first layer stays materialised (`sp1-gpu/crates/logup_gkr/src/tracegen.rs` 169 to 224) | of the order of the interaction count | yes | yes |
|
||||
| The allocator | `sp1-gpu/crates/cuda/src/task.rs` 152 and 196: the device's default `cudaMallocAsync` pool with its release threshold at `u64::MAX`, so nothing freed is ever returned to the driver: `nvidia-smi` reads the high-water mark of everything live at once | | | |
|
||||
|
||||
So at zero cycles the server already holds the proving key's 2.26 GiB, the shard's 2.26 GiB (both allocated at the
|
||||
threshold, not at the shard's rows), their codewords and trees, and the recursion program's key and traces (2 x
|
||||
0.75 GiB and their codewords): the 13.9 GB floor. The shard's own content only adds to the codewords, the GKR
|
||||
layers and the working buffers, which is the 13.9 to 20.4 GB step from 0.3 M to 4.7 M cycles, and the flat 28.3 GB
|
||||
from 20 M cycles is the threshold's padded area reached. The witness (5 to 22 KB) never appears.
|
||||
|
||||
## What the patch does (`proving/prover-floor/sp1-gpu-6.8.1-floor.patch`, three files)
|
||||
|
||||
1. `builder.rs`: the panic is gone; the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks the element threshold
|
||||
from a tier table (`element_threshold_for_budget`: over 30 as read, the full 402.6 M; 24 to 30, upstream's 24 GB
|
||||
figure; 18 to 24 (a 16 GB card), 2^27 + 2^26 = 201.3 M; under 18 (a 12 GB card), 2^27 = 134.2 M);
|
||||
`SP1_GPU_ELEMENT_THRESHOLD` sets it directly and `SP1_GPU_RECURSION_TRACE_ALLOCATION` the recursion buffer.
|
||||
The chosen numbers are printed as a `FLOOR opts` line. Every other option is as upstream.
|
||||
2. `jagged_tracegen/src/lib.rs`: with `SP1_GPU_FLOOR_LOG` set, every trace allocation prints its capacity and,
|
||||
after the shard's traces are in, the elements actually used and the device memory in use.
|
||||
3. `server.rs`: a `FLOOR memory` line (device used, free, total) after `Setup` and after every proof, with the
|
||||
proof's time.
|
||||
|
||||
Nothing in the proof changes: the element threshold only moves where the executor splits shards, exactly what
|
||||
upstream's own 24 GB tier does with the same verifier and the same keys; the recursion program, the verifying key
|
||||
and the pinned guest ids are untouched. The unpatched verifier (the pv1 host's SDK) is the one that verifies every
|
||||
measured proof below.
|
||||
|
||||
## The build (PC 2, WSL2 Ubuntu-24.04, job `floor-toolchain-1` then the build job)
|
||||
|
||||
Toolchain found 22:10Z (job `floor-toolchain-1`, 4 s): nvcc 12.8 at `/usr/local/cuda-12.8`, cmake 3.28.3, gcc 13.3,
|
||||
clang 18, protoc 3.21.12, cargo 1.99.0, no Go. The release workflow installs Go for the server's `native-gnark`
|
||||
feature (the Groth16 and PLONK wrap through gnark), which a compressed proof never runs, so the build drops that
|
||||
feature from `sp1-gpu/crates/server/Cargo.toml` and nothing else. `CUDA_ARCHS=86,89,120` (consequences reviewer
|
||||
C26): the 12 GB tier is sm_86 (RTX 3060) and sm_89 (RTX 4070), the 16 GB tier sm_89 and sm_120 (RTX 5080), PC 2's
|
||||
5090 is sm_120; the stock server lists sm_80, 86, 89, 90, 100 and 120, which a shipped build repeats. The recipe:
|
||||
`tools/prover-floor/pc2-build-server.ps1` (generated by `make-build-playbook.sh` from the patch, so the two cannot
|
||||
drift): clone the tag, `git apply` the patch, `touch` the three files, `cargo build --release --bin sp1-gpu-server`
|
||||
niced with 8 jobs into `/opt/igneum-floor/target`, the binary copied to `/opt/igneum-floor/home/.sp1/bin/` (the SDK
|
||||
spawns the server it finds under `$HOME/.sp1/bin`, so `HOME=/opt/igneum-floor/home` selects it and the live
|
||||
`/root/.sp1/bin/sp1-gpu-server` stays as it is).
|
||||
|
||||
What a measurement on PC 2 can and cannot say (C26). The server's allocation pattern is deterministic in the
|
||||
budget it is given, so a run with `SP1_GPU_MEMORY_BUDGET_GB=12` on the 5090 shows the peak a 12 GB card's build
|
||||
would ask for; it does not show that a 3060 proves it in time, nor what the card's display and driver hold. The
|
||||
public line keeps "24 GB" until the on-order 12 GB card runs the same fixture. Every row names the arch list and
|
||||
the card.
|
||||
|
||||
What shipping it costs (C26). A patched server means the project signs and distributes its own build of SP1's
|
||||
prover: the WSL2 package, the DMG's prover inputs, the K1-signed inputs and `evidence.md` carry it, and every SP1
|
||||
upgrade repeats the clone, patch, build and measurement. The verifying key and the pinned guest ids do not move
|
||||
(the patch changes buffer sizes and the shard split, not the circuits), which the `verify-segment` and `--mode
|
||||
compressed` VERIFIED lines of the unpatched host show on every row below. The packaging path is a row for the
|
||||
proving plan before 0.3.12, not this branch.
|
||||
183
proving/prover-floor/sp1-gpu-6.8.1-floor.patch
Normal file
183
proving/prover-floor/sp1-gpu-6.8.1-floor.patch
Normal file
|
|
@ -0,0 +1,183 @@
|
|||
diff --git a/sp1-gpu/crates/jagged_tracegen/src/lib.rs b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
|
||||
index 579f70a..88bfc79 100644
|
||||
--- a/sp1-gpu/crates/jagged_tracegen/src/lib.rs
|
||||
+++ b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
|
||||
@@ -494,6 +494,11 @@ async fn allocate_and_initialize_traces(
|
||||
|
||||
let total_gb = total_bytes as f64 / (1 << 30) as f64;
|
||||
tracing::debug!("Allocating {:?} GB of traces", total_gb);
|
||||
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
|
||||
+ eprintln!(
|
||||
+ "FLOOR tracegen alloc capacity_elements={max_trace_size} bytes={total_bytes} ({total_gb:.3} GB)"
|
||||
+ );
|
||||
+ }
|
||||
let mut dense_data: Buffer<Felt, TaskScope> =
|
||||
Buffer::with_capacity_in(max_trace_size, backend.clone());
|
||||
let mut col_index: Buffer<u32, TaskScope> =
|
||||
@@ -1002,6 +1007,18 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
|
||||
)
|
||||
.await;
|
||||
|
||||
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
|
||||
+ let dense = jagged_mle.dense();
|
||||
+ let (free, total) = sp1_gpu_cudart::cuda_memory_info().unwrap_or((0, 0));
|
||||
+ eprintln!(
|
||||
+ "FLOOR tracegen used preprocessed_elements={} main_elements={} dense_len={} capacity_elements={max_trace_size} device_used_mib={}",
|
||||
+ dense.preprocessed_offset,
|
||||
+ dense.main_size(),
|
||||
+ dense.dense.len(),
|
||||
+ (total - free) >> 20
|
||||
+ );
|
||||
+ }
|
||||
+
|
||||
(public_values, jagged_mle, chip_set, permit)
|
||||
}
|
||||
|
||||
diff --git a/sp1-gpu/crates/prover_components/src/builder.rs b/sp1-gpu/crates/prover_components/src/builder.rs
|
||||
index 5dccd9d..574d4fa 100644
|
||||
--- a/sp1-gpu/crates/prover_components/src/builder.rs
|
||||
+++ b/sp1-gpu/crates/prover_components/src/builder.rs
|
||||
@@ -23,28 +23,75 @@ use crate::{
|
||||
SP1CudaProverComponents,
|
||||
};
|
||||
|
||||
+/// Igneum prover-floor patch (5 October 2026). Upstream sizes every device buffer for a 24 GB card or larger
|
||||
+/// and panics below that, whatever the shard. Here the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks a
|
||||
+/// tier, and `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly.
|
||||
+/// The proof format, the verifier and the program ids do not change: the element threshold only decides where
|
||||
+/// the executor splits shards, as upstream's own 24 GB tier already does.
|
||||
+fn env_usize(name: &str) -> Option<usize> {
|
||||
+ std::env::var(name).ok().and_then(|s| s.parse::<usize>().ok())
|
||||
+}
|
||||
+
|
||||
+fn env_f64(name: &str) -> Option<f64> {
|
||||
+ std::env::var(name).ok().and_then(|s| s.parse::<f64>().ok())
|
||||
+}
|
||||
+
|
||||
+/// The core element threshold for a memory budget in GB (upstream's own figure for the budget, +4, as it
|
||||
+/// computed it: a 32 GB card is 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16).
|
||||
+pub fn element_threshold_for_budget(gpu_memory_gb: usize, full_size_shards: bool) -> u64 {
|
||||
+ if gpu_memory_gb > 30 || (full_size_shards && gpu_memory_gb >= 24) {
|
||||
+ ELEMENT_THRESHOLD
|
||||
+ } else if gpu_memory_gb >= 24 {
|
||||
+ ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24)
|
||||
+ } else if gpu_memory_gb >= 18 {
|
||||
+ (1 << 27) + (1 << 26)
|
||||
+ } else {
|
||||
+ 1 << 27
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+/// The recursion trace allocation (elements) for a memory budget.
|
||||
+pub fn recursion_trace_allocation_for_budget(gpu_memory_gb: usize) -> usize {
|
||||
+ if gpu_memory_gb >= 24 {
|
||||
+ RECURSION_TRACE_ALLOCATION
|
||||
+ } else {
|
||||
+ RECURSION_TRACE_ALLOCATION
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+pub fn gpu_memory_gb() -> usize {
|
||||
+ let gb = 1024.0 * 1024.0 * 1024.0;
|
||||
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
|
||||
+ Some(b) => (b.ceil() as usize) + 4,
|
||||
+ None => (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4,
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+pub fn recursion_trace_allocation() -> usize {
|
||||
+ env_usize("SP1_GPU_RECURSION_TRACE_ALLOCATION")
|
||||
+ .unwrap_or_else(|| recursion_trace_allocation_for_budget(gpu_memory_gb()))
|
||||
+}
|
||||
+
|
||||
pub fn local_gpu_opts() -> SP1CoreOpts {
|
||||
let mut opts = SP1CoreOpts::default();
|
||||
|
||||
let log2_shard_size = 24;
|
||||
opts.shard_size = 1 << log2_shard_size;
|
||||
|
||||
- let gb = 1024.0 * 1024.0 * 1024.0;
|
||||
-
|
||||
- // Get the amount of memory on the GPU.
|
||||
- let gpu_memory_gb: usize = (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4;
|
||||
-
|
||||
- if gpu_memory_gb < 24 {
|
||||
- panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB");
|
||||
- }
|
||||
+ // The card's memory plus 4, as upstream computed it (a 32 GB card reads 36), or the budget given.
|
||||
+ let gpu_memory_gb = gpu_memory_gb();
|
||||
|
||||
- let shard_threshold = if !opts.full_size_shards && gpu_memory_gb <= 30 {
|
||||
- ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24)
|
||||
- } else {
|
||||
- ELEMENT_THRESHOLD
|
||||
+ let shard_threshold = match env_usize("SP1_GPU_ELEMENT_THRESHOLD") {
|
||||
+ Some(t) => t as u64,
|
||||
+ None => element_threshold_for_budget(gpu_memory_gb, opts.full_size_shards),
|
||||
};
|
||||
+ let height_threshold = opts.sharding_threshold.height_threshold;
|
||||
|
||||
- tracing::debug!("Shard threshold: {shard_threshold}");
|
||||
+ eprintln!(
|
||||
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
|
||||
+ recursion_trace_allocation(),
|
||||
+ opts.full_size_shards
|
||||
+ );
|
||||
opts.sharding_threshold.element_threshold = shard_threshold;
|
||||
|
||||
opts.global_dependencies_opt = true;
|
||||
@@ -92,7 +139,7 @@ pub async fn recursion_prover_and_verifier(
|
||||
) {
|
||||
let recursion_verifier = SP1CudaProverComponents::compress_verifier();
|
||||
(
|
||||
- new_cuda_prover(&recursion_verifier, RECURSION_TRACE_ALLOCATION, 4, false, false, scope)
|
||||
+ new_cuda_prover(&recursion_verifier, recursion_trace_allocation(), 4, false, false, scope)
|
||||
.await,
|
||||
recursion_verifier,
|
||||
)
|
||||
diff --git a/sp1-gpu/crates/server/src/server.rs b/sp1-gpu/crates/server/src/server.rs
|
||||
index 4035f1f..0d0d907 100644
|
||||
--- a/sp1-gpu/crates/server/src/server.rs
|
||||
+++ b/sp1-gpu/crates/server/src/server.rs
|
||||
@@ -157,6 +157,7 @@ impl Server {
|
||||
};
|
||||
let pk = CachedProgram { elf: Arc::new(Elf::Dynamic(elf.into())), vk: vk.clone() };
|
||||
ctx.pk_cache.insert(elf_hash, pk);
|
||||
+ floor_memory_line("after setup");
|
||||
Response::Setup { id: elf_hash, vk }
|
||||
}
|
||||
Request::Destroy { key } => {
|
||||
@@ -177,15 +178,31 @@ impl Server {
|
||||
);
|
||||
};
|
||||
let context = SP1Context::builder().proof_nonce(proof_nonce).build();
|
||||
- match prover.prove_with_mode(&cached.elf, stdin, context, mode).await {
|
||||
+ let started = std::time::Instant::now();
|
||||
+ let response = match prover.prove_with_mode(&cached.elf, stdin, context, mode).await {
|
||||
Ok(proof) => Response::Proof { proof },
|
||||
Err(e) => Response::ProverError(e.to_string()),
|
||||
- }
|
||||
+ };
|
||||
+ floor_memory_line(&format!("after prove {:?} in {:.1} s", mode, started.elapsed().as_secs_f64()));
|
||||
+ response
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+/// Igneum prover-floor patch: the device memory in use (total minus free, as the driver reports it) at the
|
||||
+/// points that bound a proof, so a run's log carries the terms of the peak without a sampler.
|
||||
+fn floor_memory_line(what: &str) {
|
||||
+ if let Ok((free, total)) = sp1_gpu_cudart::cuda_memory_info() {
|
||||
+ eprintln!(
|
||||
+ "FLOOR memory {what}: device_used_mib={} free_mib={} total_mib={}",
|
||||
+ (total - free) >> 20,
|
||||
+ free >> 20,
|
||||
+ total >> 20
|
||||
+ );
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
fn sha256(data: &[u8]) -> [u8; 32] {
|
||||
use sha2::{Digest, Sha256};
|
||||
let mut hasher = Sha256::new();
|
||||
72
tools/prover-floor/make-build-playbook.sh
Executable file
72
tools/prover-floor/make-build-playbook.sh
Executable file
|
|
@ -0,0 +1,72 @@
|
|||
#!/usr/bin/env bash
|
||||
# Writes tools/prover-floor/pc2-build-server.ps1 with proving/prover-floor/sp1-gpu-6.8.1-floor.patch embedded
|
||||
# (base64), so the playbook and the patch cannot drift. Run after every change to the patch.
|
||||
set -euo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"; ROOT="$(cd "$HERE/../.." && pwd)"
|
||||
PATCH="$ROOT/proving/prover-floor/sp1-gpu-6.8.1-floor.patch"
|
||||
SHA="$(shasum -a 256 "$PATCH" | cut -c1-64)"
|
||||
B64="$(base64 < "$PATCH" | tr -d '\n')"
|
||||
OUT="$HERE/pc2-build-server.ps1"
|
||||
cat > "$OUT" <<PS
|
||||
# Prover floor (5 October 2026): builds sp1-gpu-server 6.8.1 from source with the Igneum floor patch on PC 2 inside
|
||||
# WSL2 Ubuntu-24.04 as root, into /opt/igneum-floor (the live /root/.sp1/bin/sp1-gpu-server and /opt/igneum are
|
||||
# never touched). CPU only: the miners keep mining, the card is not used. Generated by make-build-playbook.sh;
|
||||
# patch sha256 $SHA. CUDA_ARCHS=86,89,120 (the 12 GB tier is sm_86 (3060) and sm_89 (4070), the 16 GB tier sm_89 and sm_120; PC 2 runs sm_120; a shipped build lists every target the stock server does: 80, 86, 89, 90, 100, 120).
|
||||
\$ErrorActionPreference = 'Continue'
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
"RESULT start \$(Stamp)"
|
||||
\$job = \$env:IGNEUM_JOB_DIR; if (-not \$job) { \$job = Join-Path \$env:TEMP 'igneum-floor' }; New-Item -ItemType Directory -Force -Path \$job | Out-Null
|
||||
function WslPath(\$p) { \$w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a (\$p -replace '\\\\', '/') 2>\$null); if (\$w) { (\$w -replace "\`0", '').Trim() } else { '/mnt/c' + (\$p.Substring(2) -replace '\\\\', '/') } }
|
||||
\$patchB64 = '$B64'
|
||||
[IO.File]::WriteAllBytes((Join-Path \$job 'floor.patch'), [Convert]::FromBase64String(\$patchB64))
|
||||
\$patchW = WslPath (Join-Path \$job 'floor.patch')
|
||||
\$bash = @'
|
||||
set -uo pipefail
|
||||
export PATH="\$HOME/.cargo/bin:/usr/local/go/bin:\$PATH"
|
||||
CUDA_DIR="\$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"
|
||||
if [ -z "\$CUDA_DIR" ]; then echo "RESULT build_failed no /usr/local/cuda-12.*"; exit 2; fi
|
||||
export CUDA_PATH="\$CUDA_DIR" CUDACXX="\$CUDA_DIR/bin/nvcc" PATH="\$CUDA_DIR/bin:\$PATH" LD_LIBRARY_PATH="\$CUDA_DIR/lib64:/usr/lib/wsl/lib:\${LD_LIBRARY_PATH:-}"
|
||||
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
|
||||
FLOOR=/opt/igneum-floor; SRC=\$FLOOR/sp1; mkdir -p \$FLOOR/bin \$FLOOR/home/.sp1/bin \$FLOOR/logs
|
||||
PATCH="PATCH_PATH_PLACEHOLDER"
|
||||
echo "RESULT patch_sha256 \$(sha256sum "\$PATCH" | cut -c1-64) expected $SHA"
|
||||
echo "RESULT toolchain nvcc=\$(nvcc --version | grep -o 'release [0-9.]*') cargo=\$(cargo --version) go=\$(go version 2>/dev/null || echo MISSING) protoc=\$(protoc --version 2>/dev/null || echo MISSING) cmake=\$(cmake --version 2>/dev/null | head -1 || echo MISSING) nproc=\$(nproc) disk_avail=\$(df -BG /opt | awk 'NR==2 {print \$4}')"
|
||||
if [ ! -d "\$SRC/.git" ]; then
|
||||
echo "STAGE clone \$(stamp)"
|
||||
git clone -q --depth 1 --branch v6.8.1 https://github.com/succinctlabs/sp1 "\$SRC" 2>&1 | tail -2 || { echo "RESULT build_failed clone"; exit 2; }
|
||||
fi
|
||||
cd "\$SRC"
|
||||
git checkout -q -- . && git clean -qfd sp1-gpu/crates >/dev/null 2>&1
|
||||
echo "RESULT source \$(git describe --tags --always) \$(git rev-parse HEAD)"
|
||||
git apply "\$PATCH" || { echo "RESULT build_failed patch does not apply"; exit 2; }
|
||||
touch sp1-gpu/crates/prover_components/src/builder.rs sp1-gpu/crates/jagged_tracegen/src/lib.rs sp1-gpu/crates/server/src/server.rs
|
||||
echo "RESULT patched \$(git diff --stat | tail -1)"
|
||||
if ! command -v go >/dev/null 2>&1; then
|
||||
sed -i 's/sp1-prover = { workspace = true, features = \["native-gnark"\] }/sp1-prover = { workspace = true }/' sp1-gpu/crates/server/Cargo.toml
|
||||
echo "RESULT native_gnark dropped: no go toolchain on PC 2 (the gnark wrap is not used by compressed proofs)"
|
||||
fi
|
||||
export CUDA_ARCHS=86,89,120 CARGO_TARGET_DIR=\$FLOOR/target
|
||||
JOBS=8
|
||||
echo "STAGE build \$(stamp) jobs=\$JOBS"
|
||||
t0=\$(date +%s)
|
||||
nice -n 19 cargo build --release --bin sp1-gpu-server -j \$JOBS > \$FLOOR/logs/build-\$(date -u +%Y%m%dT%H%M%SZ).log 2>&1 &
|
||||
BP=\$!
|
||||
LOG=\$(ls -t \$FLOOR/logs/build-*.log | head -1)
|
||||
while kill -0 \$BP 2>/dev/null; do sleep 120; echo "STAGE building \$(stamp) \$(( (\$(date +%s) - t0) / 60 )) min: \$(grep -c '^ Compiling' "\$LOG") crates compiled, last: \$(grep '^ Compiling' "\$LOG" | tail -1 | tr -s ' ' | cut -c1-80)"; done
|
||||
wait \$BP; rc=\$?
|
||||
echo "RESULT build_exit \$rc time_s=\$(( \$(date +%s) - t0 )) crates=\$(grep -c '^ Compiling' "\$LOG")"
|
||||
if [ \$rc -ne 0 ]; then echo "RESULT build_failed"; grep -n -B2 -A12 '^error' "\$LOG" | head -80; exit 2; fi
|
||||
BIN=\$FLOOR/target/release/sp1-gpu-server
|
||||
cp "\$BIN" \$FLOOR/bin/sp1-gpu-server && cp "\$BIN" \$FLOOR/home/.sp1/bin/sp1-gpu-server && chmod +x \$FLOOR/bin/sp1-gpu-server \$FLOOR/home/.sp1/bin/sp1-gpu-server
|
||||
echo "RESULT binary bytes=\$(stat -c %s "\$BIN") sha256=\$(sha256sum "\$BIN" | cut -c1-64) version=\$(\$BIN --version 2>/dev/null)"
|
||||
echo "RESULT elf_targets \$(cuobjdump --list-elf "\$BIN" 2>/dev/null | grep -o 'sm_[0-9]*' | sort -u | tr '\n' ' ')"
|
||||
echo "RESULT live_server_untouched sha256=\$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) (expected c2642ad1c42e85d8)"
|
||||
echo "RESULT end \$(stamp)"
|
||||
'@
|
||||
\$bash = \$bash.Replace('PATCH_PATH_PLACEHOLDER', \$patchW)
|
||||
\$bashFile = Join-Path \$job 'build.sh'
|
||||
[IO.File]::WriteAllText(\$bashFile, (\$bash -replace "\`r\`n", "\`n"), (New-Object System.Text.UTF8Encoding \$false))
|
||||
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath \$bashFile) 2>&1 | ForEach-Object { (\$_ -replace "\`0", '') }
|
||||
"RESULT end \$(Stamp)"
|
||||
PS
|
||||
echo "wrote $OUT ($(wc -c < "$OUT") bytes, patch $SHA)"
|
||||
15
tools/prover-floor/make-measure-playbook.sh
Executable file
15
tools/prover-floor/make-measure-playbook.sh
Executable file
|
|
@ -0,0 +1,15 @@
|
|||
#!/usr/bin/env bash
|
||||
# Writes a measurement playbook from pc2-floor-measure.ps1 with the point list given (one `run name fixture env...`
|
||||
# per line, read from the file in $1), into $2. Fixtures: $EMPTY (block 83616, an empty live shard), $V1
|
||||
# (fees-v1-shards2 shard 0, the adopted v1 shard), $FULL (block-338-shard1, the prototype shard), $ONE (block 56).
|
||||
set -euo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
POINTS="$(tr '\n' ';' < "$1" | sed 's/;$//')"
|
||||
python3 - "$HERE/pc2-floor-measure.ps1" "$2" "$POINTS" <<'PY'
|
||||
import sys
|
||||
src, out, points = sys.argv[1], sys.argv[2], sys.argv[3]
|
||||
s = open(src).read()
|
||||
s = s.replace("'POINTS_PLACEHOLDER'", "'" + points.replace("'", "''") + "'")
|
||||
open(out, 'w').write(s)
|
||||
print("wrote", out, "points:", points.count(';') + 1)
|
||||
PY
|
||||
60
tools/prover-floor/pc2-build-server.ps1
Normal file
60
tools/prover-floor/pc2-build-server.ps1
Normal file
File diff suppressed because one or more lines are too long
65
tools/prover-floor/pc2-floor-measure.ps1
Normal file
65
tools/prover-floor/pc2-floor-measure.ps1
Normal file
|
|
@ -0,0 +1,65 @@
|
|||
# Prover floor (5 October 2026): the peak GPU memory and the time of one compressed shard proof through the PATCHED
|
||||
# sp1-gpu-server (/opt/igneum-floor/home/.sp1/bin, reached by HOME=/opt/igneum-floor/home: the SDK spawns the
|
||||
# server it finds under $HOME/.sp1/bin, sp1-cuda-6.8.1/src/server.rs) on PC 2's RTX 5090, the miners STOPPED by the
|
||||
# job (--stop-miners) and the live prover switched off for the run (its server would otherwise own the socket).
|
||||
# Every point: every server killed and its socket unlinked, a 1-s nvidia-smi sampler, one `--mode compressed
|
||||
# --shard 0` of the pv1 host (/opt/igneum-pv1, the UNPATCHED SDK and verifier: its VERIFIED is the unpatched
|
||||
# verifier's word on the patched server's proof), the peak, the time, the FLOOR lines the server prints.
|
||||
# The point list comes from the FLOOR_POINTS environment the job carries, else the default sweep below.
|
||||
$ErrorActionPreference = 'Continue'
|
||||
$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/')
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } }
|
||||
"RESULT start $(Stamp) prover off for the run: $(Prove $false)"
|
||||
Start-Sleep -Seconds 45
|
||||
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,power.draw --format=csv,noheader,nounits 2>$null) -join ' | ')"
|
||||
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor-measure' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
|
||||
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
|
||||
$jobW = WslPath $job
|
||||
$emptyW = WslPath (Join-Path $env:LOCALAPPDATA 'igneum\app\jobs\chain-pc2-pv1c\block-83616.json')
|
||||
$points = if ($env:FLOOR_POINTS) { $env:FLOOR_POINTS } else { 'POINTS_PLACEHOLDER' }
|
||||
$bash = @'
|
||||
set -uo pipefail
|
||||
export PATH="$HOME/.cargo/bin:$PATH"
|
||||
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH" && export LD_LIBRARY_PATH="$CUDA_DIR/lib64:/usr/lib/wsl/lib:${LD_LIBRARY_PATH:-}"
|
||||
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
|
||||
JOB='JOBW_PLACEHOLDER'; H=/opt/igneum-pv1/igneum-prove-host; FX="/root/igneum-prove-pv1/proving/fixtures"
|
||||
FLOORHOME=/opt/igneum-floor/home; SRV=$FLOORHOME/.sp1/bin/sp1-gpu-server
|
||||
EMPTY='EMPTY_PLACEHOLDER'; V1="$FX/fees-v1-shards2.json"; FULL="$FX/block-338-shard1.json"; ONE="$FX/block-56-transfers.json"
|
||||
[ -x "$SRV" ] || { echo "RESULT measure_failed no patched server at $SRV"; exit 2; }
|
||||
[ -x "$H" ] || { echo "RESULT measure_failed no pv1 host at $H"; exit 2; }
|
||||
echo "RESULT patched_server sha256=$(sha256sum $SRV | cut -c1-64) version=$($SRV --version 2>/dev/null) host=$(sha256sum $H | cut -c1-16)"
|
||||
echo "RESULT live_server sha256=$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) untouched"
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
|
||||
echo "RESULT idle_mib $(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | head -1)"
|
||||
run() { # name fixture env...
|
||||
local name="$1" fx="$2"; shift 2
|
||||
local tag="$name-$(basename $fx .json)"
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
|
||||
local csv="$JOB/smi-$tag.csv" log="$JOB/log-$tag.txt"
|
||||
nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv,noheader,nounits -l 1 > "$csv" 2>/dev/null &
|
||||
local SMI=$!
|
||||
local t0=$(date +%s)
|
||||
env HOME=$FLOORHOME SP1_PROVER=cuda RUST_LOG=off SP1_GPU_FLOOR_LOG=1 "$@" $H "$fx" --mode compressed --shard 0 --out "$JOB/res-$tag.json" > "$log" 2>&1
|
||||
local rc=$?
|
||||
local wall=$(( $(date +%s) - t0 ))
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; kill $SMI 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
|
||||
local peak=$(awk -F', *' '{ if ($2+0 > m) m=$2+0 } END { print m+0 }' "$csv")
|
||||
local n=$(wc -l < "$csv")
|
||||
local line=$(grep -E "^RESULT compressed shard" "$log" | tail -1 | sed -E 's/.*prove ([0-9.]+) s, proof ([0-9]+) bytes, verify ([0-9.]+) s, ([A-Z ]+);.*/prove_s=\1 bytes=\2 verify_s=\3 \4/')
|
||||
local cyc=$(grep -E "^RESULT execute shard" "$log" | tail -1 | sed -E 's/.*: ([0-9]+) cycles.*/\1/')
|
||||
local err=$(grep -iE "error|panick|out of memory|OOM|unsupported" "$log" | grep -v "^FLOOR" | head -1 | cut -c1-200)
|
||||
echo "RESULT floor cfg=$name fixture=$(basename $fx .json) peak_mib=$peak samples=$n wall_s=$wall cycles=${cyc:-na} ${line:-no_result} exit=$rc env='$*' ${err:+err=$err}"
|
||||
grep -E "^FLOOR" "$log" | sed "s/^/RESULT floorline cfg=$name fixture=$(basename $fx .json) /" | head -40
|
||||
}
|
||||
POINTS_PLACEHOLDER_BASH
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
|
||||
echo "RESULT measure_end $(stamp)"
|
||||
'@
|
||||
$bash = $bash.Replace('JOBW_PLACEHOLDER', $jobW).Replace('EMPTY_PLACEHOLDER', $emptyW).Replace('POINTS_PLACEHOLDER_BASH', ($points -replace ';', "`n"))
|
||||
$bashFile = Join-Path $job 'measure.sh'
|
||||
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
|
||||
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
|
||||
"RESULT end $(Stamp) prover back on: $(Prove $true)"
|
||||
65
tools/prover-floor/pc2-floor-sweep1.ps1
Normal file
65
tools/prover-floor/pc2-floor-sweep1.ps1
Normal file
|
|
@ -0,0 +1,65 @@
|
|||
# Prover floor (5 October 2026): the peak GPU memory and the time of one compressed shard proof through the PATCHED
|
||||
# sp1-gpu-server (/opt/igneum-floor/home/.sp1/bin, reached by HOME=/opt/igneum-floor/home: the SDK spawns the
|
||||
# server it finds under $HOME/.sp1/bin, sp1-cuda-6.8.1/src/server.rs) on PC 2's RTX 5090, the miners STOPPED by the
|
||||
# job (--stop-miners) and the live prover switched off for the run (its server would otherwise own the socket).
|
||||
# Every point: every server killed and its socket unlinked, a 1-s nvidia-smi sampler, one `--mode compressed
|
||||
# --shard 0` of the pv1 host (/opt/igneum-pv1, the UNPATCHED SDK and verifier: its VERIFIED is the unpatched
|
||||
# verifier's word on the patched server's proof), the peak, the time, the FLOOR lines the server prints.
|
||||
# The point list comes from the FLOOR_POINTS environment the job carries, else the default sweep below.
|
||||
$ErrorActionPreference = 'Continue'
|
||||
$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
|
||||
$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/')
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } }
|
||||
"RESULT start $(Stamp) prover off for the run: $(Prove $false)"
|
||||
Start-Sleep -Seconds 45
|
||||
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,power.draw --format=csv,noheader,nounits 2>$null) -join ' | ')"
|
||||
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor-measure' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
|
||||
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
|
||||
$jobW = WslPath $job
|
||||
$emptyW = WslPath (Join-Path $env:LOCALAPPDATA 'igneum\app\jobs\chain-pc2-pv1c\block-83616.json')
|
||||
$points = if ($env:FLOOR_POINTS) { $env:FLOOR_POINTS } else { 'run ctrl32 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=32;run ctrl32 "$V1" SP1_GPU_MEMORY_BUDGET_GB=32;run b12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12;run b12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12;run b16 "$V1" SP1_GPU_MEMORY_BUDGET_GB=16;run e26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864;run e26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864;run e25 "$V1" SP1_GPU_ELEMENT_THRESHOLD=33554432;run b12c1 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1' }
|
||||
$bash = @'
|
||||
set -uo pipefail
|
||||
export PATH="$HOME/.cargo/bin:$PATH"
|
||||
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH" && export LD_LIBRARY_PATH="$CUDA_DIR/lib64:/usr/lib/wsl/lib:${LD_LIBRARY_PATH:-}"
|
||||
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
|
||||
JOB='JOBW_PLACEHOLDER'; H=/opt/igneum-pv1/igneum-prove-host; FX="/root/igneum-prove-pv1/proving/fixtures"
|
||||
FLOORHOME=/opt/igneum-floor/home; SRV=$FLOORHOME/.sp1/bin/sp1-gpu-server
|
||||
EMPTY='EMPTY_PLACEHOLDER'; V1="$FX/fees-v1-shards2.json"; FULL="$FX/block-338-shard1.json"; ONE="$FX/block-56-transfers.json"
|
||||
[ -x "$SRV" ] || { echo "RESULT measure_failed no patched server at $SRV"; exit 2; }
|
||||
[ -x "$H" ] || { echo "RESULT measure_failed no pv1 host at $H"; exit 2; }
|
||||
echo "RESULT patched_server sha256=$(sha256sum $SRV | cut -c1-64) version=$($SRV --version 2>/dev/null) host=$(sha256sum $H | cut -c1-16)"
|
||||
echo "RESULT live_server sha256=$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) untouched"
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
|
||||
echo "RESULT idle_mib $(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | head -1)"
|
||||
run() { # name fixture env...
|
||||
local name="$1" fx="$2"; shift 2
|
||||
local tag="$name-$(basename $fx .json)"
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
|
||||
local csv="$JOB/smi-$tag.csv" log="$JOB/log-$tag.txt"
|
||||
nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv,noheader,nounits -l 1 > "$csv" 2>/dev/null &
|
||||
local SMI=$!
|
||||
local t0=$(date +%s)
|
||||
env HOME=$FLOORHOME SP1_PROVER=cuda RUST_LOG=off SP1_GPU_FLOOR_LOG=1 "$@" $H "$fx" --mode compressed --shard 0 --out "$JOB/res-$tag.json" > "$log" 2>&1
|
||||
local rc=$?
|
||||
local wall=$(( $(date +%s) - t0 ))
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; kill $SMI 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
|
||||
local peak=$(awk -F', *' '{ if ($2+0 > m) m=$2+0 } END { print m+0 }' "$csv")
|
||||
local n=$(wc -l < "$csv")
|
||||
local line=$(grep -E "^RESULT compressed shard" "$log" | tail -1 | sed -E 's/.*prove ([0-9.]+) s, proof ([0-9]+) bytes, verify ([0-9.]+) s, ([A-Z ]+);.*/prove_s=\1 bytes=\2 verify_s=\3 \4/')
|
||||
local cyc=$(grep -E "^RESULT execute shard" "$log" | tail -1 | sed -E 's/.*: ([0-9]+) cycles.*/\1/')
|
||||
local err=$(grep -iE "error|panick|out of memory|OOM|unsupported" "$log" | grep -v "^FLOOR" | head -1 | cut -c1-200)
|
||||
echo "RESULT floor cfg=$name fixture=$(basename $fx .json) peak_mib=$peak samples=$n wall_s=$wall cycles=${cyc:-na} ${line:-no_result} exit=$rc env='$*' ${err:+err=$err}"
|
||||
grep -E "^FLOOR" "$log" | sed "s/^/RESULT floorline cfg=$name fixture=$(basename $fx .json) /" | head -40
|
||||
}
|
||||
POINTS_PLACEHOLDER_BASH
|
||||
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
|
||||
echo "RESULT measure_end $(stamp)"
|
||||
'@
|
||||
$bash = $bash.Replace('JOBW_PLACEHOLDER', $jobW).Replace('EMPTY_PLACEHOLDER', $emptyW).Replace('POINTS_PLACEHOLDER_BASH', ($points -replace ';', "`n"))
|
||||
$bashFile = Join-Path $job 'measure.sh'
|
||||
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
|
||||
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
|
||||
"RESULT end $(Stamp) prover back on: $(Prove $true)"
|
||||
47
tools/prover-floor/pc2-toolchain.ps1
Normal file
47
tools/prover-floor/pc2-toolchain.ps1
Normal file
|
|
@ -0,0 +1,47 @@
|
|||
# Prover floor (5 October 2026, the project lead: "execute if it will solve the issue", the 12 GB cards): the toolchain check
|
||||
# before the sp1-gpu-server source build on PC 2. Reads versions and free space inside WSL2 Ubuntu-24.04 as root.
|
||||
# Touches nothing: no build, no GPU work (one nvidia-smi query), the live host and server untouched.
|
||||
$ErrorActionPreference = 'Continue'
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
"RESULT start $(Stamp)"
|
||||
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,driver_version --format=csv,noheader,nounits 2>$null) -join ' | ')"
|
||||
"RESULT host_ram_mb $([math]::Round((Get-CimInstance Win32_OperatingSystem).TotalVisibleMemorySize / 1024))"
|
||||
$bash = @'
|
||||
set -uo pipefail
|
||||
export PATH="$HOME/.cargo/bin:$HOME/.sp1/bin:/usr/local/go/bin:$PATH"
|
||||
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"
|
||||
[ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH"
|
||||
v() { local n="$1"; shift; if command -v "$1" >/dev/null 2>&1; then echo "RESULT tool $n $("$@" 2>&1 | head -1 | tr -s ' ' | cut -c1-120)"; else echo "RESULT tool $n MISSING"; fi; }
|
||||
echo "RESULT cuda_dirs $(ls -d /usr/local/cuda* 2>/dev/null | tr '\n' ' ')"
|
||||
v nvcc nvcc --version
|
||||
[ -n "$CUDA_DIR" ] && echo "RESULT nvcc_release $(nvcc --version 2>/dev/null | grep -o 'release [0-9.]*' | head -1)"
|
||||
v cmake cmake --version
|
||||
v gcc gcc --version
|
||||
v g++ g++ --version
|
||||
v clang clang --version
|
||||
v go go version
|
||||
v protoc protoc --version
|
||||
v cargo cargo --version
|
||||
v rustc rustc --version
|
||||
v git git --version
|
||||
v cuobjdump cuobjdump --version
|
||||
v pkg-config pkg-config --version
|
||||
echo "RESULT nproc $(nproc)"
|
||||
echo "RESULT mem $(free -g | awk '/Mem:/ {print "total_gb=" $2 " available_gb=" $7}')"
|
||||
echo "RESULT disk_root $(df -BG / | awk 'NR==2 {print "size=" $2 " used=" $3 " avail=" $4}')"
|
||||
echo "RESULT disk_opt $(df -BG /opt 2>/dev/null | awk 'NR==2 {print "avail=" $4}')"
|
||||
echo "RESULT registry $(du -sh $HOME/.cargo/registry 2>/dev/null | cut -f1) target_live $(du -sh /root/igneum-prove/proving/igneum-prove/target 2>/dev/null | cut -f1)"
|
||||
echo "RESULT floor_dir $(ls -d /opt/igneum-floor 2>/dev/null || echo absent)"
|
||||
echo "RESULT live_server $(ls -l /root/.sp1/bin/sp1-gpu-server 2>/dev/null | awk '{print $5}') sha256 $(sha256sum /root/.sp1/bin/sp1-gpu-server 2>/dev/null | cut -c1-16)"
|
||||
echo "RESULT live_server_version $($HOME/.sp1/bin/sp1-gpu-server --version 2>/dev/null)"
|
||||
echo "RESULT github $(timeout 20 git ls-remote --tags https://github.com/succinctlabs/sp1 refs/tags/v6.8.1 2>&1 | cut -c1-60)"
|
||||
echo "RESULT ld_libs $(ls /usr/lib/wsl/lib/libcuda.so* 2>/dev/null | tr '\n' ' ')"
|
||||
echo "RESULT wsl_user $(id -un) home $HOME"
|
||||
echo "RESULT end $(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
||||
'@
|
||||
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
|
||||
$bashFile = Join-Path $job 'toolchain.sh'
|
||||
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
|
||||
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
|
||||
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
|
||||
"RESULT end $(Stamp)"
|
||||
9
tools/prover-floor/points-sweep1.txt
Normal file
9
tools/prover-floor/points-sweep1.txt
Normal file
|
|
@ -0,0 +1,9 @@
|
|||
run ctrl32 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=32
|
||||
run ctrl32 "$V1" SP1_GPU_MEMORY_BUDGET_GB=32
|
||||
run b12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12
|
||||
run b12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12
|
||||
run b16 "$V1" SP1_GPU_MEMORY_BUDGET_GB=16
|
||||
run e26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864
|
||||
run e26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864
|
||||
run e25 "$V1" SP1_GPU_ELEMENT_THRESHOLD=33554432
|
||||
run b12c1 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1
|
||||
6
tools/prover-floor/points-sweep2.txt
Normal file
6
tools/prover-floor/points-sweep2.txt
Normal file
|
|
@ -0,0 +1,6 @@
|
|||
run r26b12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_GPU_RECURSION_TRACE_ALLOCATION=67108864
|
||||
run r26b12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_GPU_RECURSION_TRACE_ALLOCATION=67108864
|
||||
run r25b12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_GPU_RECURSION_TRACE_ALLOCATION=33554432
|
||||
run r26e26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864 SP1_GPU_RECURSION_TRACE_ALLOCATION=67108864
|
||||
run r26e26 "$ONE" SP1_GPU_ELEMENT_THRESHOLD=67108864 SP1_GPU_RECURSION_TRACE_ALLOCATION=67108864
|
||||
run r26e26 "$FULL" SP1_GPU_ELEMENT_THRESHOLD=67108864 SP1_GPU_RECURSION_TRACE_ALLOCATION=67108864
|
||||
Loading…
Reference in a new issue