Merge prover-floor (8f14a5b) into release-0.3.11: docs and standalone probe sources the evidence rows cite (C38); nothing a built artefact reads

This commit is contained in:
igneum-labs 2026-10-05 23:03:46 +00:00
commit 2f3338382f
13 changed files with 901 additions and 0 deletions

View file

@ -0,0 +1,152 @@
# The prover floor: why a 12 GB card cannot prove on SP1 6.8.1's GPU server, and the patch
5 October 2026, 22:00 UTC on (the project lead: "execute if it will solve the issue"). Branch `prover-floor`
(worktree `igneum-wt-prover-floor`). The measured facts this starts from: `docs/plans/proving-v1.md` and the
bench-log entry "proving v1" (branch proving-v1): the GPU server holds 13.9 GB for an empty shard, 20.4 GB at the
adopted v1 shard, 28.3 GB flat from 20 M to 60 M cycles, and no environment knob moved the floor. Every figure
below is from the source at tag v6.8.1 (cloned to `vendor/sp1-6.8.1`, gitignored; the fork is the patch
`proving/prover-floor/sp1-gpu-6.8.1-floor.patch`) or from a PC 2 run named in `docs/bench-log.md`
("prover floor"). Sizes in GiB are computed from the source constants (4-byte field elements); sizes in MiB are
measured by `nvidia-smi` at 1 s.
## Where the server is built and what it reads
The SDK downloads `sp1_gpu_server_v6.8.1_x86_64.tar.gz` (133,750,780 bytes) from the SP1 release and runs it from
`$HOME/.sp1/bin/sp1-gpu-server` (`crates/cuda/src/server.rs` 19 to 30, 80 to 99). The source is in the same
repository: `sp1-gpu/crates/server` (the binary), built by `.github/workflows/release.yml` 234 to 314 on CUDA
12.8.1 with Go and protoc (`cargo build --release --bin sp1-gpu-server`). The binary takes no options
(`sp1-gpu/crates/server/src/main.rs` 15 to 18: `--version` only) and reads `CUDA_VISIBLE_DEVICES` (32 to 35);
everything else comes from the environment the host process passes it, through `SP1CoreOpts::default()`
(`crates/core/executor/src/opts.rs` 99 to 140: `SHARD_SIZE`, `ELEMENT_THRESHOLD`, `HEIGHT_THRESHOLD`,
`MINIMAL_TRACE_CHUNK_THRESHOLD`, `TRACE_CHUNK_SLOTS`, `FULL_SIZE_SHARDS`) and the worker counts
(`crates/prover/src/worker/config.rs`).
## The memory model, term by term
Every device buffer is sized at construction from constants, not from the shard. The server builds the prover at
the first `Setup` request (`sp1-gpu/crates/server/src/server.rs` 126 to 137) through
`cuda_worker_builder_with_machine` (`sp1-gpu/crates/prover_components/src/builder.rs` 102 to 148):
| Term | Where | Size | On the device | Moves with the shard |
|---|---|---|---|---|
| The gate | `builder.rs` 35 to 39: `gpu_memory_gb = ceil(total / GiB) + 4`; `panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB")` when under 24 | a 12 GB card reads 16, a 16 GB card 20: both refused before any allocation | | no |
| The core element threshold | `builder.rs` 41 to 48: `ELEMENT_THRESHOLD` = 2^28 + 2^27 = 402,653,184 elements (`opts.rs` 12) on a card reading over 30 (a 32 GB card reads 36); minus 2^26 + 2^25 + 2^24 = 285,212,672 on a card reading 24 to 30 (a 24 GB card). The environment's `ELEMENT_THRESHOLD` is read at `opts.rs` 129 and then OVERWRITTEN at `builder.rs` 48, which is why the sweep's `ELEMENT_THRESHOLD` rows changed nothing; `HEIGHT_THRESHOLD` survives (it is not overwritten), which is why the 2^25 + 2^20 row did | sets the next two terms | | no |
| The core trace area, one per shard in flight | `builder.rs` 70 to 71: `num_elts = element_threshold + 2^21` (`CORE_LOG_STACKING_HEIGHT` 21, `crates/prover/src/components.rs` 16) = 404,750,336; allocated on the device at `sp1-gpu/crates/jagged_tracegen/src/lib.rs` 484 to 500 (`allocate_and_initialize_traces`: `max_trace_size` felts + `max_trace_size / 2` u32 column index + 2^14 u32) | 6 bytes an element: **2.26 GiB** for the full threshold, 1.59 GiB for the 24 GB threshold | yes, in full, whatever the shard holds | no |
| The program's preprocessed traces (the proving key) | `sp1-gpu/crates/shard_prover/src/setup.rs` 47 to 58 and 107: the same `allocate_and_initialize_traces(max_trace_size)` at `Setup`, kept in the key cache for the connection's life (`server.rs` 139 to 143) | another **2.26 GiB**, held from `Setup` on | yes | no |
| The pinned host trace buffers | `sp1-gpu/crates/prover_components/src/components.rs` 99 to 103: 4 `PinnedBuffer` of `max_trace_size` felts per prover (core 4 x 1.51 GiB, recursion 4 x 0.5 GiB, shrink 4 x 0.125 GiB, wrap 4 x 0.32 GiB) | 9.8 GiB of pinned host RAM, not device memory (the WSL2 working set the bench saw) | no | no |
| The recursion trace area | `builder.rs` 15 and 95: `RECURSION_TRACE_ALLOCATION` = 2^27 elements, one per recursion tracegen (the recursion program's key at setup and its shard at prove) | 0.75 GiB each | yes | no |
| The shrink and wrap provers | `builder.rs` 16, 19, 117 to 127: 2^25 and 85,376,340 elements, built at `Setup` for every proof mode, used only by the Groth16 and PLONK path | host pinned at build; device only when a wrap runs (never, for a compressed proof) | no | no |
| The codewords (LDE) and the Merkle trees | `sp1-gpu/crates/basefold/src/fri.rs` 92 to 97: every stacked column of 2^21 rows encoded to 2^(21 + 1) rows (`log_blowup` 1); kept until the query phase unless `drop_ldes` (`builder.rs` 52: only on a 24 GB card with `FULL_SIZE_SHARDS`); the preprocessed codewords live in the key | 2 x the padded trace, so up to 2 x the term above | yes | yes, with the padded trace |
| The LogUp GKR layers | `builder.rs` 51: `recompute_gkr_trace = false`, so the first layer stays materialised (`sp1-gpu/crates/logup_gkr/src/tracegen.rs` 169 to 224) | of the order of the interaction count | yes | yes |
| The allocator | `sp1-gpu/crates/cuda/src/task.rs` 152 and 196: the device's default `cudaMallocAsync` pool with its release threshold at `u64::MAX`, so nothing freed is ever returned to the driver: `nvidia-smi` reads the high-water mark of everything live at once | | | |
So at zero cycles the server already holds the proving key's 2.26 GiB, the shard's 2.26 GiB (both allocated at the
threshold, not at the shard's rows), their codewords and trees, and the recursion program's key and traces (2 x
0.75 GiB and their codewords): the 13.9 GB floor. The shard's own content only adds to the codewords, the GKR
layers and the working buffers, which is the 13.9 to 20.4 GB step from 0.3 M to 4.7 M cycles, and the flat 28.3 GB
from 20 M cycles is the threshold's padded area reached. The witness (5 to 22 KB) never appears.
## What the patch does (`proving/prover-floor/sp1-gpu-6.8.1-floor.patch`, three files)
1. `builder.rs`: the panic is gone; the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks the element threshold
from a tier table (`element_threshold_for_budget`: over 30 as read, the full 402.6 M; 24 to 30, upstream's 24 GB
figure; 18 to 24 (a 16 GB card), 2^27 + 2^26 = 201.3 M; under 18 (a 12 GB card), 2^27 = 134.2 M);
`SP1_GPU_ELEMENT_THRESHOLD` sets it directly and `SP1_GPU_RECURSION_TRACE_ALLOCATION` the recursion buffer.
The chosen numbers are printed as a `FLOOR opts` line. Every other option is as upstream.
2. `jagged_tracegen/src/lib.rs`: with `SP1_GPU_FLOOR_LOG` set, every trace allocation prints its capacity and,
after the shard's traces are in, the elements actually used and the device memory in use.
3. `server.rs`: a `FLOOR memory` line (device used, free, total) after `Setup` and after every proof, with the
proof's time.
Nothing in the proof changes: the element threshold only moves where the executor splits shards, exactly what
upstream's own 24 GB tier does with the same verifier and the same keys; the recursion program, the verifying key
and the pinned guest ids are untouched. The unpatched verifier (the pv1 host's SDK) is the one that verifies every
measured proof below.
## The build (PC 2, WSL2 Ubuntu-24.04, job `floor-toolchain-1` then the build job)
Toolchain found 22:10Z (job `floor-toolchain-1`, 4 s): nvcc 12.8 at `/usr/local/cuda-12.8`, cmake 3.28.3, gcc 13.3,
clang 18, protoc 3.21.12, cargo 1.99.0, no Go. The release workflow installs Go for the server's `native-gnark`
feature (the Groth16 and PLONK wrap through gnark), which a compressed proof never runs, so the build drops that
feature from `sp1-gpu/crates/server/Cargo.toml` and nothing else. `CUDA_ARCHS=86,89,120` (consequences reviewer
C26): the 12 GB tier is sm_86 (RTX 3060) and sm_89 (RTX 4070), the 16 GB tier sm_89 and sm_120 (RTX 5080), PC 2's
5090 is sm_120; the stock server lists sm_80, 86, 89, 90, 100 and 120, which a shipped build repeats. The recipe:
`tools/prover-floor/pc2-build-server.ps1` (generated by `make-build-playbook.sh` from the patch, so the two cannot
drift): clone the tag, `git apply` the patch, `touch` the three files, `cargo build --release --bin sp1-gpu-server`
niced with 8 jobs into `/opt/igneum-floor/target`, the binary copied to `/opt/igneum-floor/home/.sp1/bin/` (the SDK
spawns the server it finds under `$HOME/.sp1/bin`, so `HOME=/opt/igneum-floor/home` selects it and the live
`/root/.sp1/bin/sp1-gpu-server` stays as it is).
What a measurement on PC 2 can and cannot say (C26). The server's allocation pattern is deterministic in the
budget it is given, so a run with `SP1_GPU_MEMORY_BUDGET_GB=12` on the 5090 shows the peak a 12 GB card's build
would ask for; it does not show that a 3060 proves it in time, nor what the card's display and driver hold. The
public line keeps "24 GB" until the on-order 12 GB card runs the same fixture. Every row names the arch list and
the card.
What shipping it costs (C26). A patched server means the project signs and distributes its own build of SP1's
prover: the WSL2 package, the DMG's prover inputs, the K1-signed inputs and `evidence.md` carry it, and every SP1
upgrade repeats the clone, patch, build and measurement. The verifying key and the pinned guest ids do not move
(the patch changes buffer sizes and the shard split, not the circuits), which the `verify-segment` and `--mode
compressed` VERIFIED lines of the unpatched host show on every row below. The packaging path is a row for the
proving plan before 0.3.12, not this branch.
## Step 4 contingency, read not measured: RISC Zero's CUDA prover and its memory per segment
If SP1 could not be brought under 11 GB, the alternative's floor is read from its operators' documentation (not
measured here; a PC 2 run would be the measurement): Boundless' prover guide
(https://docs.boundless.network/provers/performance-optimization) sets the segment size cap by VRAM as 8 GB:
po2 19, 16 GB: po2 20, 20 GB: po2 21, 40 GB: po2 22, with measured peaks po2 20: 13,835 MiB, po2 21: 22,905 MiB,
po2 22: 41,089 MiB; RISC Zero's PR 3761 adds `low_vram` and `pinned_witgen` to fit po2 22 on a 24 GB 4090. So
RISC Zero proves a 2^19-cycle segment inside 8 GB and a 2^20 one inside 16 GB, and a shard of 4.7 M cycles is
9 segments at po2 19 plus lift and join steps (times not on the page). Adopting it would cost a second guest (the
chain rule in the RISC Zero zkVM), a second pinned program id, a second verifier in the node and no shared
aggregation between the two formats: `docs/analysis/amd-proving.md` and the proving plan carry that row already.
### The build, as it ran (job `floor-build-3`, 22:28:24 to 22:32:29Z)
Three runs: `floor-build-1` (22:17Z) and `floor-build-2` (22:24Z) failed in 2 to 4 minutes on
`crates/recursion/gnark-ffi/build.rs:70`, "Failed to build Go library: NotFound" (no `go` on PC 2; the first run's
playbook lost its own log, a bug fixed before the second). `floor-build-3` fetched go1.27.1 (tarball sha256
`63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445`, checked before unpacking under
`/opt/igneum-floor/go`, on the job's PATH only) and built in **240 s** (46 crates on the warm target of run 2, 8
niced jobs, 16 cores). The binary: `/opt/igneum-floor/bin/sp1-gpu-server`, **166,768,224 bytes, sha256
`5568108bf7fb9b0e525d8a08926b7046e51136ffaea53f0ca858631d0e938878`**, `--version` 6.8.1, `cuobjdump --list-elf`
sm_86, sm_89, sm_120 (the stock 251,306,680-byte server lists sm_80, 86, 89, 90, 100, 120 and compute_120 PTX).
The live `/root/.sp1/bin/sp1-gpu-server` (c2642ad1...) was never touched; the miners mined throughout.
## Sweep 1 (job `floor-sweep-1`, 22:34:56 to 22:37:55Z): the shard term gone, a second floor found
PC 2's RTX 5090 (32,607 MiB, idle 1,755 MiB with the miners stopped and the live prover off), the patched server
`5568108b...` (sm_86, sm_89, sm_120; the build above), the unpatched pv1 host `dae6b006...` as client and
verifier, one `--mode compressed --shard 0` per point, every server killed and its socket unlinked around every
point, peak = `nvidia-smi memory.used` at 1 s (the idle 1,755 MiB inside it), time = the compressed proof.
Every proof VERIFIED (1,272,897 bytes, verify 0.037 to 0.040 s), so the unpatched verifier accepts every proof
of the patched server.
| Config (environment to the patched server) | Fixture | Cycles | Peak MiB | Prove s | Verified |
|---|---|---|---|---|---|
| control: `SP1_GPU_MEMORY_BUDGET_GB=32` (upstream's sizes) | empty live shard (block 83616) | 280,706 | 13,892 | 2.2 | yes |
| control | v1 shard (fees-v1-shards2 shard 0) | 4,717,439 | 20,516 | 4.2 | yes |
| 12 GB tier: budget 12 (threshold 2^27) | empty | 280,706 | 12,740 | 2.4 | yes |
| 12 GB tier | v1 shard | 4.7 M | 15,396 | 4.1 | yes |
| 12 GB tier + `SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1` | v1 shard | 4.7 M | 15,428 | 4.0 | yes |
| 16 GB tier: budget 16 (2^27 + 2^26) | v1 shard | 4.7 M | 18,628 | 3.8 | yes |
| `SP1_GPU_ELEMENT_THRESHOLD=67108864` (2^26) | empty | 280,706 | 12,772 | 3.1 | yes |
| 2^26 | v1 shard (split into 4 core shards) | 4.7 M | **12,708** | 5.3 | yes |
| `SP1_GPU_ELEMENT_THRESHOLD=33554432` (2^25) | v1 shard | 4.7 M | 12,836 | 8.5 | yes |
Reading. The control reproduces the proving agent's curve (13.9 and 20.4 GB), so the patched server behaves as
the stock one at the stock sizes. The shard's term follows the threshold as the model says (20.5 GB at 402 M
elements, 15.4 at 134 M, 12.7 at 67 M), and then stops: 2^26 and 2^25 both sit at 12.7 to 12.8 GB for the empty
shard and the v1 shard alike. The `FLOOR memory after setup` line names the rest: **9,703 MiB in use before the
first shard** (2^26; 11,623 at the stock sizes), and the `FLOOR tracegen alloc` lines at Setup are five
allocations of 134,217,728 elements (the recursion keys, 0.75 GB each, each using 90,177,536 elements: 35.6 M
preprocessed and 54.5 M main at prove time), one of 33,554,432 (the shrink key, 0.19 GB) and one core key at the
threshold. The `NORMALIZE_PROGRAM_CACHE_SIZE` knob does not reach them (they are keys built at `Setup`, not the
program LRU). The time cost of the split: the v1 shard at 2^26 is 4 core shards and 5.3 s against 4.2 s (1.26x);
at 2^27 it is 4.1 s with no split.
So after sweep 1 the binding term is the Setup-time keys allocated at full capacity, and patch v2 sizes every
trace buffer (keys and shards) to its padded need: `padded_trace_elements` in `jagged_tracegen/src/lib.rs`
(each phase pads to the next multiple of 2^21 rows, `generate_jagged_traces`'s "final padding"), applied in
`setup_tracegen` and `full_tracegen`, one stacking height of slack, `SP1_GPU_FLOOR_EXACT=0` restoring upstream.

View file

@ -0,0 +1,250 @@
diff --git a/sp1-gpu/crates/jagged_tracegen/src/lib.rs b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
index 579f70a..09e73e8 100644
--- a/sp1-gpu/crates/jagged_tracegen/src/lib.rs
+++ b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
@@ -481,6 +481,33 @@ async fn device_preprocessed_tracegen<A: CudaTracegenAir<Felt>>(
named_traces
}
+/// Igneum prover-floor patch: the dense elements a set of traces will occupy once `generate_jagged_traces`
+/// has laid them out, that is the sum of their buffers padded to the next multiple of 2^log_stacking_height
+/// (the "final padding" step below). Each phase (preprocessed, then main) is padded on its own.
+pub fn padded_trace_elements(
+ traces: &BTreeMap<String, Trace<TaskScope>>,
+ log_stacking_height: u32,
+) -> usize {
+ let total: usize = traces
+ .values()
+ .map(|t| match t {
+ Trace::Real(trace) => trace.guts().as_buffer().len(),
+ Trace::Padding(_) => 0,
+ })
+ .sum();
+ total.next_multiple_of(1 << log_stacking_height)
+}
+
+/// Igneum prover-floor patch: the capacity to allocate for a trace set: the exact padded size plus one
+/// stacking height of slack, never more than the prover's `max_trace_size`. `SP1_GPU_FLOOR_EXACT=0` restores
+/// upstream's full-capacity allocation.
+fn floor_capacity(max_trace_size: usize, needed: usize, log_stacking_height: u32) -> usize {
+ if std::env::var("SP1_GPU_FLOOR_EXACT").map(|v| v == "0").unwrap_or(false) {
+ return max_trace_size;
+ }
+ max_trace_size.min(needed + (1 << log_stacking_height))
+}
+
async fn allocate_and_initialize_traces(
preprocessed_traces: BTreeMap<String, Trace<TaskScope>>,
max_trace_size: usize,
@@ -494,6 +521,11 @@ async fn allocate_and_initialize_traces(
let total_gb = total_bytes as f64 / (1 << 30) as f64;
tracing::debug!("Allocating {:?} GB of traces", total_gb);
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
+ eprintln!(
+ "FLOOR tracegen alloc capacity_elements={max_trace_size} bytes={total_bytes} ({total_gb:.3} GB)"
+ );
+ }
let mut dense_data: Buffer<Felt, TaskScope> =
Buffer::with_capacity_in(max_trace_size, backend.clone());
let mut col_index: Buffer<u32, TaskScope> =
@@ -677,9 +709,14 @@ pub async fn setup_tracegen<A: CudaTracegenAir<Felt>>(
let preprocessed_traces =
device_preprocessed_tracegen(program, host_phase_tracegen, backend).await;
+ let capacity = floor_capacity(
+ max_trace_size,
+ padded_trace_elements(&preprocessed_traces, log_stacking_height),
+ log_stacking_height,
+ );
let jagged_traces = allocate_and_initialize_traces(
preprocessed_traces,
- max_trace_size,
+ capacity,
log_stacking_height,
max_log_row_count,
backend,
@@ -984,9 +1021,15 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
log_chip_stats(machine, &chip_set, &main_traces);
+ let capacity = floor_capacity(
+ max_trace_size,
+ padded_trace_elements(&preprocessed_traces, log_stacking_height)
+ + padded_trace_elements(&main_traces, log_stacking_height),
+ log_stacking_height,
+ );
let mut jagged_mle = allocate_and_initialize_traces(
preprocessed_traces,
- max_trace_size,
+ capacity,
log_stacking_height,
max_log_row_count,
backend,
@@ -1002,6 +1045,18 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
)
.await;
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
+ let dense = jagged_mle.dense();
+ let (free, total) = sp1_gpu_cudart::cuda_memory_info().unwrap_or((0, 0));
+ eprintln!(
+ "FLOOR tracegen used preprocessed_elements={} main_elements={} dense_len={} capacity_elements={capacity} max_trace_size={max_trace_size} device_used_mib={}",
+ dense.preprocessed_offset,
+ dense.main_size(),
+ dense.dense.len(),
+ (total - free) >> 20
+ );
+ }
+
(public_values, jagged_mle, chip_set, permit)
}
diff --git a/sp1-gpu/crates/prover_components/src/builder.rs b/sp1-gpu/crates/prover_components/src/builder.rs
index 5dccd9d..574d4fa 100644
--- a/sp1-gpu/crates/prover_components/src/builder.rs
+++ b/sp1-gpu/crates/prover_components/src/builder.rs
@@ -23,28 +23,75 @@ use crate::{
SP1CudaProverComponents,
};
+/// Igneum prover-floor patch (5 October 2026). Upstream sizes every device buffer for a 24 GB card or larger
+/// and panics below that, whatever the shard. Here the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks a
+/// tier, and `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly.
+/// The proof format, the verifier and the program ids do not change: the element threshold only decides where
+/// the executor splits shards, as upstream's own 24 GB tier already does.
+fn env_usize(name: &str) -> Option<usize> {
+ std::env::var(name).ok().and_then(|s| s.parse::<usize>().ok())
+}
+
+fn env_f64(name: &str) -> Option<f64> {
+ std::env::var(name).ok().and_then(|s| s.parse::<f64>().ok())
+}
+
+/// The core element threshold for a memory budget in GB (upstream's own figure for the budget, +4, as it
+/// computed it: a 32 GB card is 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16).
+pub fn element_threshold_for_budget(gpu_memory_gb: usize, full_size_shards: bool) -> u64 {
+ if gpu_memory_gb > 30 || (full_size_shards && gpu_memory_gb >= 24) {
+ ELEMENT_THRESHOLD
+ } else if gpu_memory_gb >= 24 {
+ ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24)
+ } else if gpu_memory_gb >= 18 {
+ (1 << 27) + (1 << 26)
+ } else {
+ 1 << 27
+ }
+}
+
+/// The recursion trace allocation (elements) for a memory budget.
+pub fn recursion_trace_allocation_for_budget(gpu_memory_gb: usize) -> usize {
+ if gpu_memory_gb >= 24 {
+ RECURSION_TRACE_ALLOCATION
+ } else {
+ RECURSION_TRACE_ALLOCATION
+ }
+}
+
+pub fn gpu_memory_gb() -> usize {
+ let gb = 1024.0 * 1024.0 * 1024.0;
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
+ Some(b) => (b.ceil() as usize) + 4,
+ None => (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4,
+ }
+}
+
+pub fn recursion_trace_allocation() -> usize {
+ env_usize("SP1_GPU_RECURSION_TRACE_ALLOCATION")
+ .unwrap_or_else(|| recursion_trace_allocation_for_budget(gpu_memory_gb()))
+}
+
pub fn local_gpu_opts() -> SP1CoreOpts {
let mut opts = SP1CoreOpts::default();
let log2_shard_size = 24;
opts.shard_size = 1 << log2_shard_size;
- let gb = 1024.0 * 1024.0 * 1024.0;
-
- // Get the amount of memory on the GPU.
- let gpu_memory_gb: usize = (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4;
-
- if gpu_memory_gb < 24 {
- panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB");
- }
+ // The card's memory plus 4, as upstream computed it (a 32 GB card reads 36), or the budget given.
+ let gpu_memory_gb = gpu_memory_gb();
- let shard_threshold = if !opts.full_size_shards && gpu_memory_gb <= 30 {
- ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24)
- } else {
- ELEMENT_THRESHOLD
+ let shard_threshold = match env_usize("SP1_GPU_ELEMENT_THRESHOLD") {
+ Some(t) => t as u64,
+ None => element_threshold_for_budget(gpu_memory_gb, opts.full_size_shards),
};
+ let height_threshold = opts.sharding_threshold.height_threshold;
- tracing::debug!("Shard threshold: {shard_threshold}");
+ eprintln!(
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
+ recursion_trace_allocation(),
+ opts.full_size_shards
+ );
opts.sharding_threshold.element_threshold = shard_threshold;
opts.global_dependencies_opt = true;
@@ -92,7 +139,7 @@ pub async fn recursion_prover_and_verifier(
) {
let recursion_verifier = SP1CudaProverComponents::compress_verifier();
(
- new_cuda_prover(&recursion_verifier, RECURSION_TRACE_ALLOCATION, 4, false, false, scope)
+ new_cuda_prover(&recursion_verifier, recursion_trace_allocation(), 4, false, false, scope)
.await,
recursion_verifier,
)
diff --git a/sp1-gpu/crates/server/src/server.rs b/sp1-gpu/crates/server/src/server.rs
index 4035f1f..0d0d907 100644
--- a/sp1-gpu/crates/server/src/server.rs
+++ b/sp1-gpu/crates/server/src/server.rs
@@ -157,6 +157,7 @@ impl Server {
};
let pk = CachedProgram { elf: Arc::new(Elf::Dynamic(elf.into())), vk: vk.clone() };
ctx.pk_cache.insert(elf_hash, pk);
+ floor_memory_line("after setup");
Response::Setup { id: elf_hash, vk }
}
Request::Destroy { key } => {
@@ -177,15 +178,31 @@ impl Server {
);
};
let context = SP1Context::builder().proof_nonce(proof_nonce).build();
- match prover.prove_with_mode(&cached.elf, stdin, context, mode).await {
+ let started = std::time::Instant::now();
+ let response = match prover.prove_with_mode(&cached.elf, stdin, context, mode).await {
Ok(proof) => Response::Proof { proof },
Err(e) => Response::ProverError(e.to_string()),
- }
+ };
+ floor_memory_line(&format!("after prove {:?} in {:.1} s", mode, started.elapsed().as_secs_f64()));
+ response
}
}
}
}
+/// Igneum prover-floor patch: the device memory in use (total minus free, as the driver reports it) at the
+/// points that bound a proof, so a run's log carries the terms of the peak without a sampler.
+fn floor_memory_line(what: &str) {
+ if let Ok((free, total)) = sp1_gpu_cudart::cuda_memory_info() {
+ eprintln!(
+ "FLOOR memory {what}: device_used_mib={} free_mib={} total_mib={}",
+ (total - free) >> 20,
+ free >> 20,
+ total >> 20
+ );
+ }
+}
+
fn sha256(data: &[u8]) -> [u8; 32] {
use sha2::{Digest, Sha256};
let mut hasher = Sha256::new();

View file

@ -0,0 +1,83 @@
#!/usr/bin/env bash
# Writes tools/prover-floor/pc2-build-server.ps1 with proving/prover-floor/sp1-gpu-6.8.1-floor.patch embedded
# (base64), so the playbook and the patch cannot drift. Run after every change to the patch.
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"; ROOT="$(cd "$HERE/../.." && pwd)"
PATCH="$ROOT/proving/prover-floor/sp1-gpu-6.8.1-floor.patch"
SHA="$(shasum -a 256 "$PATCH" | cut -c1-64)"
B64="$(base64 < "$PATCH" | tr -d '\n')"
OUT="$HERE/pc2-build-server.ps1"
cat > "$OUT" <<PS
# Prover floor (5 October 2026): builds sp1-gpu-server 6.8.1 from source with the Igneum floor patch on PC 2 inside
# WSL2 Ubuntu-24.04 as root, into /opt/igneum-floor (the live /root/.sp1/bin/sp1-gpu-server and /opt/igneum are
# never touched). CPU only: the miners keep mining, the card is not used. Generated by make-build-playbook.sh;
# patch sha256 $SHA. CUDA_ARCHS=86,89,120 (the 12 GB tier is sm_86 (3060) and sm_89 (4070), the 16 GB tier sm_89 and sm_120; PC 2 runs sm_120; a shipped build lists every target the stock server does: 80, 86, 89, 90, 100, 120).
\$ErrorActionPreference = 'Continue'
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
"RESULT start \$(Stamp)"
\$job = \$env:IGNEUM_JOB_DIR; if (-not \$job) { \$job = Join-Path \$env:TEMP 'igneum-floor' }; New-Item -ItemType Directory -Force -Path \$job | Out-Null
function WslPath(\$p) { \$w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a (\$p -replace '\\\\', '/') 2>\$null); if (\$w) { (\$w -replace "\`0", '').Trim() } else { '/mnt/c' + (\$p.Substring(2) -replace '\\\\', '/') } }
\$patchB64 = '$B64'
[IO.File]::WriteAllBytes((Join-Path \$job 'floor.patch'), [Convert]::FromBase64String(\$patchB64))
\$patchW = WslPath (Join-Path \$job 'floor.patch')
\$bash = @'
set -uo pipefail
export PATH="\$HOME/.cargo/bin:/usr/local/go/bin:\$PATH"
CUDA_DIR="\$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"
if [ -z "\$CUDA_DIR" ]; then echo "RESULT build_failed no /usr/local/cuda-12.*"; exit 2; fi
export CUDA_PATH="\$CUDA_DIR" CUDACXX="\$CUDA_DIR/bin/nvcc" PATH="\$CUDA_DIR/bin:\$PATH" LD_LIBRARY_PATH="\$CUDA_DIR/lib64:/usr/lib/wsl/lib:\${LD_LIBRARY_PATH:-}"
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
FLOOR=/opt/igneum-floor; SRC=\$FLOOR/sp1; mkdir -p \$FLOOR/bin \$FLOOR/home/.sp1/bin \$FLOOR/logs
PATCH="PATCH_PATH_PLACEHOLDER"
echo "RESULT patch_sha256 \$(sha256sum "\$PATCH" | cut -c1-64) expected $SHA"
echo "RESULT toolchain nvcc=\$(nvcc --version | grep -o 'release [0-9.]*') cargo=\$(cargo --version) go=\$(go version 2>/dev/null || echo MISSING) protoc=\$(protoc --version 2>/dev/null || echo MISSING) cmake=\$(cmake --version 2>/dev/null | head -1 || echo MISSING) nproc=\$(nproc) disk_avail=\$(df -BG /opt | awk 'NR==2 {print \$4}')"
if [ ! -d "\$SRC/.git" ]; then
echo "STAGE clone \$(stamp)"
git clone -q --depth 1 --branch v6.8.1 https://github.com/succinctlabs/sp1 "\$SRC" 2>&1 | tail -2 || { echo "RESULT build_failed clone"; exit 2; }
fi
cd "\$SRC"
git checkout -q -- . && git clean -qfd sp1-gpu/crates >/dev/null 2>&1
echo "RESULT source \$(git describe --tags --always) \$(git rev-parse HEAD)"
git apply "\$PATCH" || { echo "RESULT build_failed patch does not apply"; exit 2; }
touch sp1-gpu/crates/prover_components/src/builder.rs sp1-gpu/crates/jagged_tracegen/src/lib.rs sp1-gpu/crates/server/src/server.rs
echo "RESULT patched \$(git diff --stat | tail -1)"
# Go, as the release workflow installs it (the server's native-gnark feature compiles the gnark library with go;
# the wrap path is never run by a compressed proof, but the feature is upstream's and stays): a pinned tarball
# unpacked under /opt/igneum-floor, on this job's PATH only, no package installed on the PC.
if ! command -v go >/dev/null 2>&1; then
if [ ! -x \$FLOOR/go/bin/go ]; then
echo "STAGE go \$(stamp) fetching go1.27.1 (70,553,950 bytes, sha256 63d339f0...)"
curl -sSL -o \$FLOOR/go.tgz https://go.dev/dl/go1.27.1.linux-amd64.tar.gz || { echo "RESULT build_failed go download"; exit 2; }
echo "63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445 \$FLOOR/go.tgz" | sha256sum -c - >/dev/null || { echo "RESULT build_failed go sha256 mismatch: \$(sha256sum \$FLOOR/go.tgz)"; exit 2; }
echo "RESULT go_tarball sha256 \$(sha256sum \$FLOOR/go.tgz | cut -c1-64) matches 63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445"
tar -xzf \$FLOOR/go.tgz -C \$FLOOR && rm -f \$FLOOR/go.tgz
fi
export PATH="\$FLOOR/go/bin:\$PATH" GOPATH=\$FLOOR/gopath GOCACHE=\$FLOOR/gocache GOFLAGS=-mod=mod
fi
echo "RESULT go \$(go version 2>&1 | head -1)"
LOG=\$FLOOR/logs/build.log; : > "\$LOG"
echo "RESULT dirs \$(ls -ld \$FLOOR \$FLOOR/logs 2>&1 | tr '\\n' ' ') pwd=\$(pwd)"
export CUDA_ARCHS=86,89,120 CARGO_TARGET_DIR=\$FLOOR/target
JOBS=8
echo "STAGE build \$(stamp) jobs=\$JOBS"
t0=\$(date +%s)
nice -n 19 cargo build --release --bin sp1-gpu-server -j \$JOBS >> "\$LOG" 2>&1 &
BP=\$!
while kill -0 \$BP 2>/dev/null; do sleep 120; echo "STAGE building \$(stamp) \$(( (\$(date +%s) - t0) / 60 )) min: \$(grep -c '^ Compiling' "\$LOG") crates compiled, last: \$(grep '^ Compiling' "\$LOG" | tail -1 | tr -s ' ' | cut -c1-80)"; done
wait \$BP; rc=\$?
echo "RESULT build_exit \$rc time_s=\$(( \$(date +%s) - t0 )) crates=\$(grep -c '^ Compiling' "\$LOG")"
if [ \$rc -ne 0 ]; then echo "RESULT build_failed"; echo "== error lines =="; grep -n -B2 -A12 -E '^(error|warning: unused manifest| process didn|caused by)' "\$LOG" | head -100; echo "== log tail =="; tail -n 60 "\$LOG"; exit 2; fi
BIN=\$FLOOR/target/release/sp1-gpu-server
cp "\$BIN" \$FLOOR/bin/sp1-gpu-server && cp "\$BIN" \$FLOOR/home/.sp1/bin/sp1-gpu-server && chmod +x \$FLOOR/bin/sp1-gpu-server \$FLOOR/home/.sp1/bin/sp1-gpu-server
echo "RESULT binary bytes=\$(stat -c %s "\$BIN") sha256=\$(sha256sum "\$BIN" | cut -c1-64) version=\$(\$BIN --version 2>/dev/null)"
echo "RESULT elf_targets \$(cuobjdump --list-elf "\$BIN" 2>/dev/null | grep -o 'sm_[0-9]*' | sort -u | tr '\n' ' ')"
echo "RESULT live_server_untouched sha256=\$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) (expected c2642ad1c42e85d8)"
echo "RESULT end \$(stamp)"
'@
\$bash = \$bash.Replace('PATCH_PATH_PLACEHOLDER', \$patchW)
\$bashFile = Join-Path \$job 'build.sh'
[IO.File]::WriteAllText(\$bashFile, (\$bash -replace "\`r\`n", "\`n"), (New-Object System.Text.UTF8Encoding \$false))
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath \$bashFile) 2>&1 | ForEach-Object { (\$_ -replace "\`0", '') }
"RESULT end \$(Stamp)"
PS
echo "wrote $OUT ($(wc -c < "$OUT") bytes, patch $SHA)"

View file

@ -0,0 +1,15 @@
#!/usr/bin/env bash
# Writes a measurement playbook from pc2-floor-measure.ps1 with the point list given (one `run name fixture env...`
# per line, read from the file in $1), into $2. Fixtures: $EMPTY (block 83616, an empty live shard), $V1
# (fees-v1-shards2 shard 0, the adopted v1 shard), $FULL (block-338-shard1, the prototype shard), $ONE (block 56).
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
POINTS="$(tr '\n' ';' < "$1" | sed 's/;$//')"
python3 - "$HERE/pc2-floor-measure.ps1" "$2" "$POINTS" <<'PY'
import sys
src, out, points = sys.argv[1], sys.argv[2], sys.argv[3]
s = open(src).read()
s = s.replace("'POINTS_PLACEHOLDER'", "'" + points.replace("'", "''") + "'")
open(out, 'w').write(s)
print("wrote", out, "points:", points.count(';') + 1)
PY

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1,65 @@
# Prover floor (5 October 2026): the peak GPU memory and the time of one compressed shard proof through the PATCHED
# sp1-gpu-server (/opt/igneum-floor/home/.sp1/bin, reached by HOME=/opt/igneum-floor/home: the SDK spawns the
# server it finds under $HOME/.sp1/bin, sp1-cuda-6.8.1/src/server.rs) on PC 2's RTX 5090, the miners STOPPED by the
# job (--stop-miners) and the live prover switched off for the run (its server would otherwise own the socket).
# Every point: every server killed and its socket unlinked, a 1-s nvidia-smi sampler, one `--mode compressed
# --shard 0` of the pv1 host (/opt/igneum-pv1, the UNPATCHED SDK and verifier: its VERIFIED is the unpatched
# verifier's word on the patched server's proof), the peak, the time, the FLOOR lines the server prints.
# The point list comes from the FLOOR_POINTS environment the job carries, else the default sweep below.
$ErrorActionPreference = 'Continue'
$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/')
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } }
"RESULT start $(Stamp) prover off for the run: $(Prove $false)"
Start-Sleep -Seconds 45
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,power.draw --format=csv,noheader,nounits 2>$null) -join ' | ')"
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor-measure' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
$jobW = WslPath $job
$emptyW = WslPath (Join-Path $env:LOCALAPPDATA 'igneum\app\jobs\chain-pc2-pv1c\block-83616.json')
$points = if ($env:FLOOR_POINTS) { $env:FLOOR_POINTS } else { 'POINTS_PLACEHOLDER' }
$bash = @'
set -uo pipefail
export PATH="$HOME/.cargo/bin:$PATH"
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH" && export LD_LIBRARY_PATH="$CUDA_DIR/lib64:/usr/lib/wsl/lib:${LD_LIBRARY_PATH:-}"
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
JOB='JOBW_PLACEHOLDER'; H=/opt/igneum-pv1/igneum-prove-host; FX="/root/igneum-prove-pv1/proving/fixtures"
FLOORHOME=/opt/igneum-floor/home; SRV=$FLOORHOME/.sp1/bin/sp1-gpu-server
EMPTY='EMPTY_PLACEHOLDER'; V1="$FX/fees-v1-shards2.json"; FULL="$FX/block-338-shard1.json"; ONE="$FX/block-56-transfers.json"
[ -x "$SRV" ] || { echo "RESULT measure_failed no patched server at $SRV"; exit 2; }
[ -x "$H" ] || { echo "RESULT measure_failed no pv1 host at $H"; exit 2; }
echo "RESULT patched_server sha256=$(sha256sum $SRV | cut -c1-64) version=$($SRV --version 2>/dev/null) host=$(sha256sum $H | cut -c1-16)"
echo "RESULT live_server sha256=$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) untouched"
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
echo "RESULT idle_mib $(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | head -1)"
run() { # name fixture env...
local name="$1" fx="$2"; shift 2
local tag="$name-$(basename $fx .json)"
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
local csv="$JOB/smi-$tag.csv" log="$JOB/log-$tag.txt"
nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv,noheader,nounits -l 1 > "$csv" 2>/dev/null &
local SMI=$!
local t0=$(date +%s)
env HOME=$FLOORHOME SP1_PROVER=cuda RUST_LOG=off SP1_GPU_FLOOR_LOG=1 "$@" $H "$fx" --mode compressed --shard 0 --out "$JOB/res-$tag.json" > "$log" 2>&1
local rc=$?
local wall=$(( $(date +%s) - t0 ))
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; kill $SMI 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
local peak=$(awk -F', *' '{ if ($2+0 > m) m=$2+0 } END { print m+0 }' "$csv")
local n=$(wc -l < "$csv")
local line=$(grep -E "^RESULT compressed shard" "$log" | tail -1 | sed -E 's/.*prove ([0-9.]+) s, proof ([0-9]+) bytes, verify ([0-9.]+) s, ([A-Z ]+);.*/prove_s=\1 bytes=\2 verify_s=\3 \4/')
local cyc=$(grep -E "^RESULT execute shard" "$log" | tail -1 | sed -E 's/.*: ([0-9]+) cycles.*/\1/')
local err=$(grep -iE "error|panick|out of memory|OOM|unsupported" "$log" | grep -v "^FLOOR" | head -1 | cut -c1-200)
echo "RESULT floor cfg=$name fixture=$(basename $fx .json) peak_mib=$peak samples=$n wall_s=$wall cycles=${cyc:-na} ${line:-no_result} exit=$rc env='$*' ${err:+err=$err}"
grep -E "^FLOOR" "$log" | sed "s/^/RESULT floorline cfg=$name fixture=$(basename $fx .json) /" | head -40
}
POINTS_PLACEHOLDER_BASH
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
echo "RESULT measure_end $(stamp)"
'@
$bash = $bash.Replace('JOBW_PLACEHOLDER', $jobW).Replace('EMPTY_PLACEHOLDER', $emptyW).Replace('POINTS_PLACEHOLDER_BASH', ($points -replace ';', "`n"))
$bashFile = Join-Path $job 'measure.sh'
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
"RESULT end $(Stamp) prover back on: $(Prove $true)"

View file

@ -0,0 +1,65 @@
# Prover floor (5 October 2026): the peak GPU memory and the time of one compressed shard proof through the PATCHED
# sp1-gpu-server (/opt/igneum-floor/home/.sp1/bin, reached by HOME=/opt/igneum-floor/home: the SDK spawns the
# server it finds under $HOME/.sp1/bin, sp1-cuda-6.8.1/src/server.rs) on PC 2's RTX 5090, the miners STOPPED by the
# job (--stop-miners) and the live prover switched off for the run (its server would otherwise own the socket).
# Every point: every server killed and its socket unlinked, a 1-s nvidia-smi sampler, one `--mode compressed
# --shard 0` of the pv1 host (/opt/igneum-pv1, the UNPATCHED SDK and verifier: its VERIFIED is the unpatched
# verifier's word on the patched server's proof), the peak, the time, the FLOOR lines the server prints.
# The point list comes from the FLOOR_POINTS environment the job carries, else the default sweep below.
$ErrorActionPreference = 'Continue'
$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/')
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } }
"RESULT start $(Stamp) prover off for the run: $(Prove $false)"
Start-Sleep -Seconds 45
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,power.draw --format=csv,noheader,nounits 2>$null) -join ' | ')"
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor-measure' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
$jobW = WslPath $job
$emptyW = WslPath (Join-Path $env:LOCALAPPDATA 'igneum\app\jobs\chain-pc2-pv1c\block-83616.json')
$points = if ($env:FLOOR_POINTS) { $env:FLOOR_POINTS } else { 'run ctrl32 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=32;run ctrl32 "$V1" SP1_GPU_MEMORY_BUDGET_GB=32;run b12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12;run b12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12;run b16 "$V1" SP1_GPU_MEMORY_BUDGET_GB=16;run e26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864;run e26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864;run e25 "$V1" SP1_GPU_ELEMENT_THRESHOLD=33554432;run b12c1 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1' }
$bash = @'
set -uo pipefail
export PATH="$HOME/.cargo/bin:$PATH"
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH" && export LD_LIBRARY_PATH="$CUDA_DIR/lib64:/usr/lib/wsl/lib:${LD_LIBRARY_PATH:-}"
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
JOB='JOBW_PLACEHOLDER'; H=/opt/igneum-pv1/igneum-prove-host; FX="/root/igneum-prove-pv1/proving/fixtures"
FLOORHOME=/opt/igneum-floor/home; SRV=$FLOORHOME/.sp1/bin/sp1-gpu-server
EMPTY='EMPTY_PLACEHOLDER'; V1="$FX/fees-v1-shards2.json"; FULL="$FX/block-338-shard1.json"; ONE="$FX/block-56-transfers.json"
[ -x "$SRV" ] || { echo "RESULT measure_failed no patched server at $SRV"; exit 2; }
[ -x "$H" ] || { echo "RESULT measure_failed no pv1 host at $H"; exit 2; }
echo "RESULT patched_server sha256=$(sha256sum $SRV | cut -c1-64) version=$($SRV --version 2>/dev/null) host=$(sha256sum $H | cut -c1-16)"
echo "RESULT live_server sha256=$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) untouched"
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
echo "RESULT idle_mib $(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | head -1)"
run() { # name fixture env...
local name="$1" fx="$2"; shift 2
local tag="$name-$(basename $fx .json)"
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
local csv="$JOB/smi-$tag.csv" log="$JOB/log-$tag.txt"
nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv,noheader,nounits -l 1 > "$csv" 2>/dev/null &
local SMI=$!
local t0=$(date +%s)
env HOME=$FLOORHOME SP1_PROVER=cuda RUST_LOG=off SP1_GPU_FLOOR_LOG=1 "$@" $H "$fx" --mode compressed --shard 0 --out "$JOB/res-$tag.json" > "$log" 2>&1
local rc=$?
local wall=$(( $(date +%s) - t0 ))
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; kill $SMI 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
local peak=$(awk -F', *' '{ if ($2+0 > m) m=$2+0 } END { print m+0 }' "$csv")
local n=$(wc -l < "$csv")
local line=$(grep -E "^RESULT compressed shard" "$log" | tail -1 | sed -E 's/.*prove ([0-9.]+) s, proof ([0-9]+) bytes, verify ([0-9.]+) s, ([A-Z ]+);.*/prove_s=\1 bytes=\2 verify_s=\3 \4/')
local cyc=$(grep -E "^RESULT execute shard" "$log" | tail -1 | sed -E 's/.*: ([0-9]+) cycles.*/\1/')
local err=$(grep -iE "error|panick|out of memory|OOM|unsupported" "$log" | grep -v "^FLOOR" | head -1 | cut -c1-200)
echo "RESULT floor cfg=$name fixture=$(basename $fx .json) peak_mib=$peak samples=$n wall_s=$wall cycles=${cyc:-na} ${line:-no_result} exit=$rc env='$*' ${err:+err=$err}"
grep -E "^FLOOR" "$log" | sed "s/^/RESULT floorline cfg=$name fixture=$(basename $fx .json) /" | head -40
}
POINTS_PLACEHOLDER_BASH
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
echo "RESULT measure_end $(stamp)"
'@
$bash = $bash.Replace('JOBW_PLACEHOLDER', $jobW).Replace('EMPTY_PLACEHOLDER', $emptyW).Replace('POINTS_PLACEHOLDER_BASH', ($points -replace ';', "`n"))
$bashFile = Join-Path $job 'measure.sh'
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
"RESULT end $(Stamp) prover back on: $(Prove $true)"

View file

@ -0,0 +1,65 @@
# Prover floor (5 October 2026): the peak GPU memory and the time of one compressed shard proof through the PATCHED
# sp1-gpu-server (/opt/igneum-floor/home/.sp1/bin, reached by HOME=/opt/igneum-floor/home: the SDK spawns the
# server it finds under $HOME/.sp1/bin, sp1-cuda-6.8.1/src/server.rs) on PC 2's RTX 5090, the miners STOPPED by the
# job (--stop-miners) and the live prover switched off for the run (its server would otherwise own the socket).
# Every point: every server killed and its socket unlinked, a 1-s nvidia-smi sampler, one `--mode compressed
# --shard 0` of the pv1 host (/opt/igneum-pv1, the UNPATCHED SDK and verifier: its VERIFIED is the unpatched
# verifier's word on the patched server's proof), the peak, the time, the FLOOR lines the server prints.
# The point list comes from the FLOOR_POINTS environment the job carries, else the default sweep below.
$ErrorActionPreference = 'Continue'
$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/')
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } }
"RESULT start $(Stamp) prover off for the run: $(Prove $false)"
Start-Sleep -Seconds 45
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,power.draw --format=csv,noheader,nounits 2>$null) -join ' | ')"
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor-measure' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
$jobW = WslPath $job
$emptyW = WslPath (Join-Path $env:LOCALAPPDATA 'igneum\app\jobs\chain-pc2-pv1c\block-83616.json')
$points = if ($env:FLOOR_POINTS) { $env:FLOOR_POINTS } else { 'run x12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12;run x12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12;run x12 "$ONE" SP1_GPU_MEMORY_BUDGET_GB=12;run x12 "$FULL" SP1_GPU_MEMORY_BUDGET_GB=12;run xe26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864;run xe26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864;run x16 "$V1" SP1_GPU_MEMORY_BUDGET_GB=16;run x32 "$V1" SP1_GPU_MEMORY_BUDGET_GB=32;run x12r "$V1" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_GPU_RECURSION_TRACE_ALLOCATION=100663296' }
$bash = @'
set -uo pipefail
export PATH="$HOME/.cargo/bin:$PATH"
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH" && export LD_LIBRARY_PATH="$CUDA_DIR/lib64:/usr/lib/wsl/lib:${LD_LIBRARY_PATH:-}"
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
JOB='JOBW_PLACEHOLDER'; H=/opt/igneum-pv1/igneum-prove-host; FX="/root/igneum-prove-pv1/proving/fixtures"
FLOORHOME=/opt/igneum-floor/home; SRV=$FLOORHOME/.sp1/bin/sp1-gpu-server
EMPTY='EMPTY_PLACEHOLDER'; V1="$FX/fees-v1-shards2.json"; FULL="$FX/block-338-shard1.json"; ONE="$FX/block-56-transfers.json"
[ -x "$SRV" ] || { echo "RESULT measure_failed no patched server at $SRV"; exit 2; }
[ -x "$H" ] || { echo "RESULT measure_failed no pv1 host at $H"; exit 2; }
echo "RESULT patched_server sha256=$(sha256sum $SRV | cut -c1-64) version=$($SRV --version 2>/dev/null) host=$(sha256sum $H | cut -c1-16)"
echo "RESULT live_server sha256=$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) untouched"
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
echo "RESULT idle_mib $(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | head -1)"
run() { # name fixture env...
local name="$1" fx="$2"; shift 2
local tag="$name-$(basename $fx .json)"
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
local csv="$JOB/smi-$tag.csv" log="$JOB/log-$tag.txt"
nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv,noheader,nounits -l 1 > "$csv" 2>/dev/null &
local SMI=$!
local t0=$(date +%s)
env HOME=$FLOORHOME SP1_PROVER=cuda RUST_LOG=off SP1_GPU_FLOOR_LOG=1 "$@" $H "$fx" --mode compressed --shard 0 --out "$JOB/res-$tag.json" > "$log" 2>&1
local rc=$?
local wall=$(( $(date +%s) - t0 ))
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; kill $SMI 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
local peak=$(awk -F', *' '{ if ($2+0 > m) m=$2+0 } END { print m+0 }' "$csv")
local n=$(wc -l < "$csv")
local line=$(grep -E "^RESULT compressed shard" "$log" | tail -1 | sed -E 's/.*prove ([0-9.]+) s, proof ([0-9]+) bytes, verify ([0-9.]+) s, ([A-Z ]+);.*/prove_s=\1 bytes=\2 verify_s=\3 \4/')
local cyc=$(grep -E "^RESULT execute shard" "$log" | tail -1 | sed -E 's/.*: ([0-9]+) cycles.*/\1/')
local err=$(grep -iE "error|panick|out of memory|OOM|unsupported" "$log" | grep -v "^FLOOR" | head -1 | cut -c1-200)
echo "RESULT floor cfg=$name fixture=$(basename $fx .json) peak_mib=$peak samples=$n wall_s=$wall cycles=${cyc:-na} ${line:-no_result} exit=$rc env='$*' ${err:+err=$err}"
grep -E "^FLOOR" "$log" | sed "s/^/RESULT floorline cfg=$name fixture=$(basename $fx .json) /" | head -40
}
POINTS_PLACEHOLDER_BASH
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
echo "RESULT measure_end $(stamp)"
'@
$bash = $bash.Replace('JOBW_PLACEHOLDER', $jobW).Replace('EMPTY_PLACEHOLDER', $emptyW).Replace('POINTS_PLACEHOLDER_BASH', ($points -replace ';', "`n"))
$bashFile = Join-Path $job 'measure.sh'
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
"RESULT end $(Stamp) prover back on: $(Prove $true)"

View file

@ -0,0 +1,66 @@
# Published WITHOUT --stop-miners: the beside-the-miner pair (the 5090 mining at full rate on the same card).
# Prover floor (5 October 2026): the peak GPU memory and the time of one compressed shard proof through the PATCHED
# sp1-gpu-server (/opt/igneum-floor/home/.sp1/bin, reached by HOME=/opt/igneum-floor/home: the SDK spawns the
# server it finds under $HOME/.sp1/bin, sp1-cuda-6.8.1/src/server.rs) on PC 2's RTX 5090, the miners STOPPED by the
# job (--stop-miners) and the live prover switched off for the run (its server would otherwise own the socket).
# Every point: every server killed and its socket unlinked, a 1-s nvidia-smi sampler, one `--mode compressed
# --shard 0` of the pv1 host (/opt/igneum-pv1, the UNPATCHED SDK and verifier: its VERIFIED is the unpatched
# verifier's word on the patched server's proof), the peak, the time, the FLOOR lines the server prints.
# The point list comes from the FLOOR_POINTS environment the job carries, else the default sweep below.
$ErrorActionPreference = 'Continue'
$urlFile = if ($env:IGNEUM_APP_DIR) { Join-Path $env:IGNEUM_APP_DIR 'app.url' } else { Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
if (-not (Test-Path $urlFile)) { $urlFile = Join-Path $env:LOCALAPPDATA 'igneum\app\app.url' }
$base = (Get-Content $urlFile -Raw).Trim().TrimEnd('/')
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
function Prove($on) { try { (Invoke-RestMethod -Method Post -Uri "$base/api/prove" -ContentType 'application/json' -Body (@{on=$on} | ConvertTo-Json -Compress) -TimeoutSec 10) | ConvertTo-Json -Compress } catch { "error: $_" } }
"RESULT start $(Stamp) prover off for the run: $(Prove $false)"
Start-Sleep -Seconds 45
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu,power.draw --format=csv,noheader,nounits 2>$null) -join ' | ')"
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor-measure' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
$jobW = WslPath $job
$emptyW = WslPath (Join-Path $env:LOCALAPPDATA 'igneum\app\jobs\chain-pc2-pv1c\block-83616.json')
$points = if ($env:FLOOR_POINTS) { $env:FLOOR_POINTS } else { 'run m12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12;run m12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12;run me26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864;run me26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864' }
$bash = @'
set -uo pipefail
export PATH="$HOME/.cargo/bin:$PATH"
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"; [ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH" && export LD_LIBRARY_PATH="$CUDA_DIR/lib64:/usr/lib/wsl/lib:${LD_LIBRARY_PATH:-}"
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
JOB='JOBW_PLACEHOLDER'; H=/opt/igneum-pv1/igneum-prove-host; FX="/root/igneum-prove-pv1/proving/fixtures"
FLOORHOME=/opt/igneum-floor/home; SRV=$FLOORHOME/.sp1/bin/sp1-gpu-server
EMPTY='EMPTY_PLACEHOLDER'; V1="$FX/fees-v1-shards2.json"; FULL="$FX/block-338-shard1.json"; ONE="$FX/block-56-transfers.json"
[ -x "$SRV" ] || { echo "RESULT measure_failed no patched server at $SRV"; exit 2; }
[ -x "$H" ] || { echo "RESULT measure_failed no pv1 host at $H"; exit 2; }
echo "RESULT patched_server sha256=$(sha256sum $SRV | cut -c1-64) version=$($SRV --version 2>/dev/null) host=$(sha256sum $H | cut -c1-16)"
echo "RESULT live_server sha256=$(sha256sum /root/.sp1/bin/sp1-gpu-server | cut -c1-16) untouched"
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
echo "RESULT idle_mib $(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | head -1)"
run() { # name fixture env...
local name="$1" fx="$2"; shift 2
local tag="$name-$(basename $fx .json)"
pkill -f sp1-gpu-server 2>/dev/null; sleep 2; rm -f /tmp/sp1-cuda-*.sock
local csv="$JOB/smi-$tag.csv" log="$JOB/log-$tag.txt"
nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu --format=csv,noheader,nounits -l 1 > "$csv" 2>/dev/null &
local SMI=$!
local t0=$(date +%s)
env HOME=$FLOORHOME SP1_PROVER=cuda RUST_LOG=off SP1_GPU_FLOOR_LOG=1 "$@" $H "$fx" --mode compressed --shard 0 --out "$JOB/res-$tag.json" > "$log" 2>&1
local rc=$?
local wall=$(( $(date +%s) - t0 ))
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; kill $SMI 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
local peak=$(awk -F', *' '{ if ($2+0 > m) m=$2+0 } END { print m+0 }' "$csv")
local n=$(wc -l < "$csv")
local line=$(grep -E "^RESULT compressed shard" "$log" | tail -1 | sed -E 's/.*prove ([0-9.]+) s, proof ([0-9]+) bytes, verify ([0-9.]+) s, ([A-Z ]+);.*/prove_s=\1 bytes=\2 verify_s=\3 \4/')
local cyc=$(grep -E "^RESULT execute shard" "$log" | tail -1 | sed -E 's/.*: ([0-9]+) cycles.*/\1/')
local err=$(grep -iE "error|panick|out of memory|OOM|unsupported" "$log" | grep -v "^FLOOR" | head -1 | cut -c1-200)
echo "RESULT floor cfg=$name fixture=$(basename $fx .json) peak_mib=$peak samples=$n wall_s=$wall cycles=${cyc:-na} ${line:-no_result} exit=$rc env='$*' ${err:+err=$err}"
grep -E "^FLOOR" "$log" | sed "s/^/RESULT floorline cfg=$name fixture=$(basename $fx .json) /" | head -40
}
POINTS_PLACEHOLDER_BASH
pkill -f sp1-gpu-server 2>/dev/null; sleep 1; rm -f /tmp/sp1-cuda-*.sock
echo "RESULT measure_end $(stamp)"
'@
$bash = $bash.Replace('JOBW_PLACEHOLDER', $jobW).Replace('EMPTY_PLACEHOLDER', $emptyW).Replace('POINTS_PLACEHOLDER_BASH', ($points -replace ';', "`n"))
$bashFile = Join-Path $job 'measure.sh'
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
"RESULT end $(Stamp) prover back on: $(Prove $true)"

View file

@ -0,0 +1,47 @@
# Prover floor (5 October 2026, the project lead: "execute if it will solve the issue", the 12 GB cards): the toolchain check
# before the sp1-gpu-server source build on PC 2. Reads versions and free space inside WSL2 Ubuntu-24.04 as root.
# Touches nothing: no build, no GPU work (one nvidia-smi query), the live host and server untouched.
$ErrorActionPreference = 'Continue'
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
"RESULT start $(Stamp)"
"RESULT gpus $(Stamp) $((& nvidia-smi --query-gpu=index,name,memory.used,memory.total,driver_version --format=csv,noheader,nounits 2>$null) -join ' | ')"
"RESULT host_ram_mb $([math]::Round((Get-CimInstance Win32_OperatingSystem).TotalVisibleMemorySize / 1024))"
$bash = @'
set -uo pipefail
export PATH="$HOME/.cargo/bin:$HOME/.sp1/bin:/usr/local/go/bin:$PATH"
CUDA_DIR="$(ls -d /usr/local/cuda-12.* 2>/dev/null | sort -V | tail -1 || true)"
[ -n "$CUDA_DIR" ] && export PATH="$CUDA_DIR/bin:$PATH"
v() { local n="$1"; shift; if command -v "$1" >/dev/null 2>&1; then echo "RESULT tool $n $("$@" 2>&1 | head -1 | tr -s ' ' | cut -c1-120)"; else echo "RESULT tool $n MISSING"; fi; }
echo "RESULT cuda_dirs $(ls -d /usr/local/cuda* 2>/dev/null | tr '\n' ' ')"
v nvcc nvcc --version
[ -n "$CUDA_DIR" ] && echo "RESULT nvcc_release $(nvcc --version 2>/dev/null | grep -o 'release [0-9.]*' | head -1)"
v cmake cmake --version
v gcc gcc --version
v g++ g++ --version
v clang clang --version
v go go version
v protoc protoc --version
v cargo cargo --version
v rustc rustc --version
v git git --version
v cuobjdump cuobjdump --version
v pkg-config pkg-config --version
echo "RESULT nproc $(nproc)"
echo "RESULT mem $(free -g | awk '/Mem:/ {print "total_gb=" $2 " available_gb=" $7}')"
echo "RESULT disk_root $(df -BG / | awk 'NR==2 {print "size=" $2 " used=" $3 " avail=" $4}')"
echo "RESULT disk_opt $(df -BG /opt 2>/dev/null | awk 'NR==2 {print "avail=" $4}')"
echo "RESULT registry $(du -sh $HOME/.cargo/registry 2>/dev/null | cut -f1) target_live $(du -sh /root/igneum-prove/proving/igneum-prove/target 2>/dev/null | cut -f1)"
echo "RESULT floor_dir $(ls -d /opt/igneum-floor 2>/dev/null || echo absent)"
echo "RESULT live_server $(ls -l /root/.sp1/bin/sp1-gpu-server 2>/dev/null | awk '{print $5}') sha256 $(sha256sum /root/.sp1/bin/sp1-gpu-server 2>/dev/null | cut -c1-16)"
echo "RESULT live_server_version $($HOME/.sp1/bin/sp1-gpu-server --version 2>/dev/null)"
echo "RESULT github $(timeout 20 git ls-remote --tags https://github.com/succinctlabs/sp1 refs/tags/v6.8.1 2>&1 | cut -c1-60)"
echo "RESULT ld_libs $(ls /usr/lib/wsl/lib/libcuda.so* 2>/dev/null | tr '\n' ' ')"
echo "RESULT wsl_user $(id -un) home $HOME"
echo "RESULT end $(date -u +%Y-%m-%dT%H:%M:%SZ)"
'@
$job = $env:IGNEUM_JOB_DIR; if (-not $job) { $job = Join-Path $env:TEMP 'igneum-floor' }; New-Item -ItemType Directory -Force -Path $job | Out-Null
$bashFile = Join-Path $job 'toolchain.sh'
[IO.File]::WriteAllText($bashFile, ($bash -replace "`r`n", "`n"), (New-Object System.Text.UTF8Encoding $false))
function WslPath($p) { $w = (& wsl.exe -d Ubuntu-24.04 -u root -- wslpath -a ($p -replace '\\', '/') 2>$null); if ($w) { ($w -replace "`0", '').Trim() } else { '/mnt/c' + ($p.Substring(2) -replace '\\', '/') } }
& wsl.exe -d Ubuntu-24.04 -u root -- bash (WslPath $bashFile) 2>&1 | ForEach-Object { ($_ -replace "`0", '') }
"RESULT end $(Stamp)"

View file

@ -0,0 +1,9 @@
run ctrl32 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=32
run ctrl32 "$V1" SP1_GPU_MEMORY_BUDGET_GB=32
run b12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12
run b12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12
run b16 "$V1" SP1_GPU_MEMORY_BUDGET_GB=16
run e26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864
run e26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864
run e25 "$V1" SP1_GPU_ELEMENT_THRESHOLD=33554432
run b12c1 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_WORKER_NORMALIZE_PROGRAM_CACHE_SIZE=1

View file

@ -0,0 +1,9 @@
run x12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12
run x12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12
run x12 "$ONE" SP1_GPU_MEMORY_BUDGET_GB=12
run x12 "$FULL" SP1_GPU_MEMORY_BUDGET_GB=12
run xe26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864
run xe26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864
run x16 "$V1" SP1_GPU_MEMORY_BUDGET_GB=16
run x32 "$V1" SP1_GPU_MEMORY_BUDGET_GB=32
run x12r "$V1" SP1_GPU_MEMORY_BUDGET_GB=12 SP1_GPU_RECURSION_TRACE_ALLOCATION=100663296

View file

@ -0,0 +1,4 @@
run m12 "$EMPTY" SP1_GPU_MEMORY_BUDGET_GB=12
run m12 "$V1" SP1_GPU_MEMORY_BUDGET_GB=12
run me26 "$V1" SP1_GPU_ELEMENT_THRESHOLD=67108864
run me26 "$EMPTY" SP1_GPU_ELEMENT_THRESHOLD=67108864