Merge master aa4640901 into ca3-coord-record under the master-landing lock

This commit is contained in:
igneum-labs 2026-10-08 22:25:21 +00:00
commit dbcf1f9c64
19 changed files with 1701 additions and 224 deletions

View file

@ -0,0 +1,125 @@
# The floor's memory profile (V6-07), 8 October 2026
Master review R1, residual V6-07 (`docs/plans/igneum-2.0-master/evidence/04_full_system/IGNEUM_V6_Full_System_Review.md`, pages 210 to 212): the floor patch picked its limits from the card's total VRAM, not from what was free; its small-card element threshold defaulted to 2^27 where every passing row had set 2^26 by hand; its recursion-allocation budget returned one constant in both branches. The order: a pinned memory profile per workload read from free memory, with the app lane's device coordinator, and the 3060 and 4060 rows rerun on the default, unoverridden job path. This document is the record: the code (section 1), the table (section 2), the rows as the pods wrote them (section 3), the consequences per card tier beside each number (section 4), the coordinator hook and what is next (section 5).
Branch `v607-floor-memory` off the box mirror master cef5234b5. Code commits e6abe8c2c (the patch and the host profile), 573dad0ad (the lesser of grant and free), 8c1fb754c (one floor patch), then the floors re-pinned from the rows (the commit carrying this document). Clocks below are UTC as the pods wrote them; UK time is one hour later.
## 1. What changed
| Item | Where | Before | After |
|---|---|---|---|
| The limits' source | `proving/prover-floor/sp1-gpu-6.8.1-floor.patch`, `builder.rs` `gpu_memory_gb()` (the fleet's copies `tools/fleet/floor.patch`, `floor-v5.patch` carry the same hunk; `box-setup.sh` re-pinned) | `cuda_memory_info().1` (the total), +4 as upstream | `cuda_memory_info().0` (the FREE memory as the driver reports it), +4, read ONCE per process in a `OnceLock` so the core opts and the recursion prover, built at different moments, sit on one tier; `SP1_GPU_MEMORY_BUDGET_GB` (the host's lease) still overrides |
| The small-card element threshold | same, `element_threshold_for_budget` | `1 << 27` under the 18 tier | `1 << 26`, the value the passing rows used (section 2 cites them) |
| The recursion-allocation budget | same, `recursion_trace_allocation_for_budget` | `RECURSION_TRACE_ALLOCATION` in both branches | upstream's 2^27 on the 24 GB tier and above; `RECURSION_TRACE_ALLOCATION_SMALL` = 2^26 + 2^25 = 100,663,296 under it (a recursion key or shard uses 90,177,536 elements, `docs/analysis/prover-floor.md` sweep 1; the patch's `floor_capacity` sizes the device buffer to the need, so the constant caps the buffer and shrinks the four pinned host copies per prover); `floor_tests::recursion_branches_differ` pins that the two branches differ and that the small one clears the measured use with a stacking height of slack; `floor_tests::small_tier_is_two_to_the_26` pins the tiers |
| The FLOOR opts line | same, `local_gpu_opts` | `gpu_memory_gb`, thresholds | adds `free_mib` and `total_mib` so a log names what the server read |
| The pinned profile per workload | `proving/igneum-prove/host/src/memory_profile.rs` (new), wired in `main.rs` before the SP1 client is built | nothing: the server guessed from the total; the fleet set `SP1_GPU_ELEMENT_THRESHOLD` by hand | one table (section 2) with tests pinning every value; the row is chosen from the engine's lease or, with no engine, from `nvidia-smi memory.free` on the device, and handed to the server by environment (`SP1_GPU_ELEMENT_THRESHOLD`, `SP1_GPU_RECURSION_TRACE_ALLOCATION`, `SP1_GPU_MEMORY_BUDGET_GB`); a hand override already in the environment is kept and named; under the workload's floor the host prints one line and exits 78 before any setup |
| The fleet's default path | `tools/fleet/box-prover.py` | `SP1_GPU_ELEMENT_THRESHOLD` from `THRESHOLD`, else the server's total-VRAM guess | `IGNEUM_PROVE_WORKLOAD=chain` and `IGNEUM_PROVE_DEVICE` on the host run, no threshold unless `THRESHOLD` is set by hand; a refusal (exit 78) closes the segment as cancelled/memory |
Tests: `cargo test -p igneum-prove-host memory_profile` (box 3, section 6) and the patch's `floor_tests` (run on the build pod after the server build, section 6).
## 2. The table
The host's `TIERS` (largest first; the free-memory lines are the card tiers as the floor patch reads them, free GiB rounded up plus 4 as upstream computed its tiers from the total):
| Tier | Free memory at start | Element threshold | Recursion trace allocation | Who lands here |
|---|---|---|---|---|
| full | 26,624 MiB and up | 2^28 + 2^27 = 402,653,184 | 2^27 = 134,217,728 | a 32 GB card alone |
| 24gb | 20,480 MiB and up | 285,212,672 (upstream's 24 GB figure) | 2^27 | a 24 GB card alone |
| 16gb | 14,336 MiB and up | 2^27 + 2^26 = 201,326,592 (the patch's own figure, unmeasured on this fixture) | 2^26 + 2^25 = 100,663,296 | a 16 GB card alone |
| small | under 14,336 MiB | 2^26 = 67,108,864 | 100,663,296 | a 12 GB or 8 GB card alone; any card beside a miner's resident set |
Floors per workload (the least free memory the host runs in; under it the refusal, exit 78): shard 7,700 MiB, aggregate 8,500 MiB, chain 8,500 MiB. The shard floor is the small tier's measured peak (7,525 MiB on the 3060, 7,532 on the 4060, section 3) plus headroom for the driver's own context. The aggregate and chain floors are 8,500 MiB: the aggregation is the peak of a chain run, 8,306 MiB on the RTX 3060 (section 3, D2), and the RTX 4060 (7,807 MiB free) aborted at it, so an 8 GB card is refused the chain and the aggregation before any setup and proves shards only.
The rows cited for 2^26 on small cards: `docs/analysis/prover-tiers-real-cards.md` (6 October 2026, the matrix's `alone-comp-26-v1` point, compressed, verified: RTX 3060 7.4 GB 14.4 s; RTX 3080 8.0 GB 7.1 s; RTX 4060 7.4 GB 18.4 s; RTX 4060 Ti 8 GB 7.6 GB 9.6 s; RTX 4060 Ti 16 GB 7.8 GB 11.6 s; RTX 4070 7.6 GB 12.1 s; RTX 5070 7.6 GB 4.8 s) and `docs/analysis/class-v6/coexist-rows.md` (8 October 2026, `proof_alone` at `SP1_GPU_ELEMENT_THRESHOLD=67108864`: RTX 3060 peak 7,525 MiB 13.2 s verified; RTX 4060 peak 7,532 MiB 8.2 s verified). No small-card row ever passed at 2^27 on this host path; the 10 GB tier's 2^27 reading of 7 October (8,642 MiB alone on a 3080) ran on the segment host and is not this path.
Every value above is pinned by `memory_profile::tests::table_is_pinned`, `tiers_from_free_memory`, `refuses_under_the_floor`, `recursion_branches_differ`, `workload_from_mode` and `env_for_the_server`; a change of the table is a change of the profile and needs its rows.
## 3. The rows, rerun on the default, unoverridden job path
Every row is a RESULT line as the pod wrote it (the raw run logs, host logs and 1 Hz `nvidia-smi` samples are kept under the lane's scratch `v607/rows/<label>/`; the two bench scripts are `pod-v607.sh` for pass 1 and `ab2.sh` for pass 2). Pods: Vast.ai one-shots rented and destroyed by the lane (section 6). Fixture `fees-v1-shards2.json` (sha 20a108f159c61ff9, block 351, two shards of 22,172 and 14,762 pgas, 4,717,439 cycles on shard 0), the ds55 kit's worker (d43be4625b78baf7) on the 5.5 GiB dataset as the miner. "Default, unoverridden" = `HOME=/opt/igneum-floor/home SP1_PROVER=cuda RUST_LOG=off`, exactly `tools/fleet/box-prover.py`'s environment with no `THRESHOLD`, nothing else set by hand.
### 3.1 The default path as the fleet runs it today: the served 0317 host on the V6-07 server (pass 2, 20:53Z to 20:59Z)
The server is the V6-07 build (sm_86 + sm_89, sha db37c38b, free-memory tiers, small tier 2^26, two recursion budgets, from the one floor patch 9098c3e5); the host is the served `igneum-prove-host-0317` (71bc2438), which carries no profile table, so the row shows the SERVER's own free-memory rule deciding.
| Card | Free MiB at start | FLOOR opts (what the server read and chose) | Workload | Peak MiB | Seconds | Verdict |
|---|---|---|---|---|---|---|
| RTX 3060 12 GB (driver 595.91.07, 170 W) | 11,898 | `gpu_memory_gb=16 free_mib=11775 total_mib=11911 element_threshold=67108864 recursion_trace_allocation=100663296` | shard (compressed) | 7,601 | prove 11.4 s, verify 0.045 s | PASS, VERIFIED; 77.3 W mean |
| RTX 3060 12 GB | 11,898 | the same | chain (2 shards + aggregation) | 8,307 (after setup 4,866; after shard 0 7,504; after shard 1 7,536; after the aggregation 8,306) | shard proofs 22.0 s (11.6 + 10.4), aggregation 5.2 s, end to end 27.4 s | PASS, every proof VERIFIED, chain_len 1; 94.0 W mean |
| RTX 4060 8 GB (driver 595.91.07, 115 W, no power sensor) | 7,807 | `gpu_memory_gb=12 free_mib=7694 total_mib=7807 element_threshold=67108864 recursion_trace_allocation=100663296` | shard (compressed) | 7,504 | prove 16.3 s, verify 0.129 s | PASS, VERIFIED |
| RTX 4060 8 GB | 7,807 | the same | chain (2 shards + aggregation) | 7,792 (after setup 4,999; after each shard 7,631 with 176 MiB free) | shard 0 11.7 s VERIFIED, shard 1 12.8 s VERIFIED, then the aggregation | FAIL at the aggregation: `FLOOR abort: a device allocation failed at slop/crates/tensor/src/inner.rs:51: called Result::unwrap() on an Err value: AllocError { layout: Layout { size: 509782528, align: 4 } }` (a 486 MiB allocation on 176 MiB free) |
Beside the miner on this path (the server's rule alone, no host-side refusal): NOT RUN. Pass 2's `ab2.sh` stopped the server between points by the pids of every listening unix socket instead of the `sp1-cuda` socket's owner, and that sweep killed the miner (rc 137) at the start of D7 and D8, so those two rows ran with the card empty (3060: 11.6 s, 7,697 MiB; 4060: 16.1 s, 7,632 MiB, both VERIFIED: alone rows, not beside rows). The kill-by-pattern class, in a scratch script; the lesson is recorded here and the script corrected (`ab3.sh`). The beside rows that stand are the host-side refusals of pass 1 (3.2) and the 8 October coexistence rows (`docs/analysis/class-v6/coexist-rows.md`: the server's allocation fails beside the 6.1 GiB miner on both cards at 2^26, the miner unharmed).
### 3.2 The V6-07 host's profile and refusal rows (pass 1, 20:25Z to 20:33Z, V6-07 server db37c38b, host 66662219)
| Card | Free MiB at start | Source of the reading | Row chosen (RESULT memory_profile) | Verdict |
|---|---|---|---|---|
| RTX 3060 12 GB | 11,909 | nvidia-smi device 0 (no engine) | workload shard, tier small, element_threshold 67108864, recursion_trace_allocation 100663296, `SP1_GPU_MEMORY_BUDGET_GB=11.6`; the server read it back as `gpu_memory_gb=16 ... element_threshold=67108864 ... recursion_trace_allocation=100663296` | the profile applied; the proof itself NOT RUN on this host (3.3) |
| RTX 3060 12 GB | 11,909 | the engine's lease (`IGNEUM_PROVE_MEM_BUDGET_MB` = `IGNEUM_PROVE_MEM_FREE_MB` = 11909, `IGNEUM_PROVE_DEADLINE_S=600`) | the same row, "from 11909 MiB free (lease IGNEUM_PROVE_MEM_BUDGET_MB), deadline 600 s" | the same |
| RTX 3060 12 GB | lease 6,100 | the engine's lease under the floor | `RESULT memory_profile refused: proving needs 8 GB free on the card for the shard workload (7700 MiB floor); 6100 MiB free`, exit 78 in 0 s, peak 1 MiB (no setup, the card untouched) | PASS (the refusal) |
| RTX 3060 12 GB | 5,781 beside the ds55 miner (6,129 MiB resident, 26.83 MH/s) | nvidia-smi | `refused: ... (7700 MiB floor); 5781 MiB free`, exit 78 in 0 s; the miner alive after, 26.832 MH/s over its 253 s, self-test PASS, fingerprint 23ced07a4d28b465 | PASS (the refusal; the miner unharmed) |
| RTX 4060 8 GB | 7,807 | nvidia-smi | workload shard, tier small, 2^26, 100663296, `SP1_GPU_MEMORY_BUDGET_GB=7.6`; the server read `gpu_memory_gb=12` | the profile applied; the proof NOT RUN on this host (3.3) |
| RTX 4060 8 GB | 7,807 | the engine's lease (7807 / 7807, deadline 600) | the same row from the lease | the same |
| RTX 4060 8 GB | lease 6,100 | the lease under the floor | `refused: ... (7700 MiB floor); 6100 MiB free`, exit 78, peak 2 MiB | PASS (the refusal) |
| RTX 4060 8 GB | 1,689 beside the ds55 miner (6,120 MiB resident, 18.96 MH/s) | nvidia-smi | `refused: ... (7700 MiB floor); 1689 MiB free`, exit 78 in 0 s; the miner alive after, 18.958 MH/s over 357 s, self-test PASS | PASS (the refusal; the miner unharmed) |
### 3.3 The proof on master's host: NOT RUN, with its cause
Every proving attempt by a host built from master after the 0.3.17 build ends at the shard's execute with `Error: public values are 0 bytes, expected 328` (the V6-07 host, from cef5234b5) or `expected 392` (a plain host from master's tip dc94dda95, no V6-07 change), on both cards, with or without the threshold set by hand (pass 2 rows D4, D5, D6, D8: 14 to 31 s wall, peaks 4,579 to 4,963 MiB, the setup complete and "setup matches the manifest", then 0 bytes from the executor). The served 0317 host (71bc2438) executes and proves the same shard on the same card and server (3.1). The pinned guests are the 5 October pair (elf/manifest.json pinned_at 2026-10-05T16:20:38Z; `igneum-prove-program.elf` last changed at 15bb6cdd4) while the host's guest input changed on 8 October (6dbd5d2f8 at 17:45 the D4 proving-payment split, 2ceb09be9 at 18:34 the key-succession pairs, then 67e22b9ef and 566d0900b P22, then 7604fcf5d V6-10's format word, none with a re-pin). The bisect (section 6) names the commit: 6dbd5d2f8 (8 October 17:45, "Shard guest: the D4 proving-payment split mirrored behind FixtureEnv::proving_payment_to_pool (known-failed first; the ELF and program id not yet re-pinned)"): the host's fixture gained a field the pinned 5 October guest does not read, and the commit's own message says the re-pin is pending; a host from its parent 13b729467 proves (12.2 s VERIFIED), a host from 6dbd5d2f8 reads 0 bytes, and a host from the key-succession commit 2ceb09be9 (a side branch without 6dbd5d2f8) proves (12.1 s VERIFIED). Consequence: the host-side profile is attested tonight by its unit tests (7 of 7 on box 3) and the refusal rows (3.2), not by a proof; the proof rows on the default path are the 0317 host's (3.1) with the server's rule, which is the same table on the server side. The fleet's standing provers run the 0317 host and the served floors, so nothing on the fleet is affected until a host is rebuilt from master; a host rebuilt from master today cannot prove at all (a release gate before any 2.0.x host ships: pin-guests.sh and a CI check that a guest-input change without an elf/ change fails).
### 3.4 The recursion constant isolated (pass 1, 20:18Z to 20:20Z, RTX 3060, the served sm_86 server 224200d3 of 10:56 UTC, host 0317, threshold 2^26 set by hand)
| `SP1_GPU_RECURSION_TRACE_ALLOCATION` | Peak MiB | Seconds | Verdict |
|---|---|---|---|
| 100,663,296 (the small branch, 2^26 + 2^25) | 7,619 | prove 14.1 s, verify 0.108 s | PASS, VERIFIED |
| 134,217,728 (upstream's 2^27, the 24 GB branch) | 7,619 | prove 14.1 s, verify 0.112 s | PASS, VERIFIED |
| 117,440,512 (2^26 + 2^25 + 2^24, a control) | 7,619 | prove 14.4 s, verify 0.110 s | PASS, VERIFIED |
The device peak does not move with the constant: the patch's `floor_capacity` sizes every trace buffer to its padded need, so the constant is a cap, and the small branch's saving is the four pinned host copies per prover (4 x 100.6 M x 4 bytes = 1.5 GiB of pinned RAM against 2.0 GiB), which is what a 16 GB Windows PC under WSL2 feels. The two branches are two values by test; neither is the memory lever on the card.
### 3.5 Known-failed server build (pass 1, the first tarball)
The first V6-07 server (sha 75d0b4be, from the canonical patch copy `proving/prover-floor/sp1-gpu-6.8.1-floor.patch` as it stood) proved shard 0 on the RTX 4060 and panicked at the first recursion prove of the chain run: `sp1-gpu/crates/jagged_tracegen/src/lib.rs:240:70: range end index 37428736 out of range for slice of length 36700160` (20:57Z). That copy carried no `grow_for_main` (the key buffer sized to its preprocessed traces at setup, never grown for the main traces), which the fleet's `floor-v5.patch` has carried since 6 October; the served sm_86 and sm_89 floors behave as floor-v5. Fixed by 8c1fb754c: the three copies are one file (sha 9098c3e5), `box-setup.sh` re-pinned, and the server of 3.1 is built from it.
## 4. Consequences per card tier
Every number above, per card tier, with what the lane does about it:
| Tier | What tonight's rows say | Consequence | Done or owed |
|---|---|---|---|
| 8 GB card (RTX 4060 class: 5060, 4060, 3070 8 GB, 3060 8 GB, 5060 Ti 8 GB) alone | 7,807 MiB free; shard proof 16.3 s at 7,504 MiB VERIFIED on the default path; the chain's aggregation aborts at a 486 MiB allocation on 176 MiB free | an 8 GB card proves shards and cannot aggregate: the host refuses the chain and the aggregate before setup ("proving needs 9 GB free on the card for the chain workload (8500 MiB floor)"), so the fleet's segment path (`box-prover.py`, one chain run per segment) gets nothing from an 8 GB card until shard production and aggregation are separated (review V6-08's "optional separation", the fleet lane's item); until then an 8 GB owner mines and does not earn proving income, and the site must not say otherwise | floors pinned (this commit); the shard-only route owed to the proving lane (a morning item, not tonight) |
| 8 GB card beside its miner (6,120 MiB resident) | 1,689 MiB free; the host refuses in 0 s, the miner unharmed at 18.96 MH/s | mining-only while the miner runs; proving only in a time-share with the dataset evicted (F07's modes); nothing is lost to a failed allocation | the refusal shipped on the host side; the engine's lease (reviewb-202) refuses before the spawn |
| 12 GB card (RTX 3060 class: 3060 12 GB, 4070, 5070, 3080 Ti 12 GB, 2080 Ti 11 GB approximate) alone | 11,898 MiB free; shard 11.4 s at 7,601 MiB; the whole chain 27.4 s end to end at an 8,306 MiB aggregation peak, VERIFIED | a 12 GB card runs the fleet's segment path alone with 3.5 GB spare at the peak: shard proofs and the aggregation; the 2.0 litepaper's 12 GB prove-alone sentence holds on the default path with nothing set by hand | measured; the 10 GB tier (3080) is between the two rows and untested tonight (owed: one 3080 hour, the aggregation at 8,306 MiB against 9,885 total is the open cell) |
| 12 GB card beside its miner (6,129 MiB resident) | 5,781 MiB free; refused in 0 s, the miner unharmed at 26.83 MH/s | time-share only, as the coexistence rows of 16:10Z said; the host now says so in one line instead of dying inside an allocation 34 s later | shipped |
| 16 GB card (4060 Ti 16 GB, 5060 Ti 16 GB, 5070 Ti, 5080, 4080) alone | no row tonight; the table's 16gb tier (2^27 + 2^26, 14,336 MiB free and up) is the patch's own figure | a 16 GB card alone reads the 16gb tier; beside a 6.1 GiB miner it reads 9.9 GB free and lands on the small tier, above both floors (7,700 shard, 8,500 chain), so it is the smallest card that mines and runs the whole segment path at once (about 1.4 GB spare at the aggregation peak, approximate until measured) | owed: one rented 16 GB hour on the default path, alone and beside the miner (the fleet lane's item (h) 16 GB cell carries the V6-07 tarball) |
| 24 GB and 32 GB cards (4090, 3090, A5000, 5090) | no row tonight; the 24gb and full tiers are upstream's own figures, unchanged by V6-07 except that the server reads them from FREE memory | a 24 GB card beside its miner (6.1 GiB) reads about 17.5 GB free and now lands on the 16gb tier instead of upstream's 24 GB tier, so a mining 24 GB card proves at the 16 GB threshold (smaller core shards, more of them; the 5 October rows put the time cost of a split at 1.26x on the 5090); alone it is unchanged | measured consequence owed on a rented 4090 beside the miner (one hour) before the 2.0.2 host ships |
| Rig (8x 4090) | one process per card (`IGNEUM_CUDA_DEVICE`, `FLEET_CARD`): each host reads its own card's free memory | unchanged behaviour per card; the rig's spare cards never see another card's miner | none |
| Pool user, Windows, macOS, AMD, Intel | the profile lives in the host, which runs only where the GPU prover runs (Linux and WSL2 on NVIDIA); macOS, AMD and Intel do not prove (provedefault.rs) | the Windows app's WSL2 host takes the same env from the engine; the pinned RAM saving of the small recursion branch (0.5 GiB) helps the 32 GB RAM floor on Windows by a little, not enough to lower it | none tonight |
Cost of the rows: five Vast pods and two RunPod pods, about USD 0.40 in total (section 6), inside the day's ceiling.
## 5. The device coordinator hook and what is next
The app lane's answer (the shipper, 20:4x UK, from the window lane's F07 design on `reviewb-202`, `src/device.rs`: a Coordinator with per-device leases, `admit(device, holder, mib, mode, budget)` and a Mode per card): the engine owns admission and the host never re-reads a card the engine leased. On every spawn of `igneum-prove-host` the engine passes, as environment on the command: `IGNEUM_PROVE_DEVICE` (the card's ordinal as the host enumerates it), `IGNEUM_PROVE_WORKLOAD` (shard, aggregate or chain: the one job this process is admitted for), `IGNEUM_PROVE_MEM_FREE_MB` (the free memory the engine read on that card at admission), `IGNEUM_PROVE_MEM_BUDGET_MB` (the lease's grant, the hard ceiling for the host) and `IGNEUM_PROVE_DEADLINE_S` (the admission deadline the lease carries). The host picks its row from `IGNEUM_PROVE_MEM_BUDGET_MB`, from `IGNEUM_PROVE_MEM_FREE_MB` when no grant is given, and only without either (the fleet's `box-prover.py` path and a hand run) from `nvidia-smi memory.free` on the device (`IGNEUM_PROVE_DEVICE`, else `IGNEUM_CUDA_DEVICE`, else the first of `CUDA_VISIBLE_DEVICES`, else 0). The refusal is exit 78 with the one line `RESULT memory_profile refused: proving needs N GB free on the card for the <workload> workload (<floor> MiB floor); <free> MiB free`; the engine records it on the lease and the card reads "proving needs N GB free" on the window. No file, no socket. The host side is in this branch; the engine half (the env names exactly as written, the lease table, the refusal on the lease) is the window lane's on `reviewb-202` (a414b6bdc81d348d8), named to it at 20:5x UK.
Next: (1) the engine half lands and the `lease` rows of section 3 are rerun from the engine's own spawn; (2) the aggregate floor is re-pinned from the chain rows' aggregation peak (section 3) when the 16 GB and 24 GB chain rows exist; (3) the 16gb tier's 2^27 + 2^26 is measured on a rented 16 GB card (no row tonight); (4) the fleet's served floor tarballs (`igneum-floor-sm86.tgz`, `sm89`, `sm120`) are rebuilt from the V6-07 patch (the sm_86 + sm_89 tarball of section 6 is the first) before the standing provers restart on the default path.
## 6. Shas, pods, cost
| Item | Sha / id |
|---|---|
| Branch `v607-floor-memory` | e6abe8c2c (patch + host profile + box-prover), 573dad0ad (lease lesser rule), 8c1fb754c (one floor patch), this document's commit (floors 8,500 and the rows) |
| The one floor patch (`proving/prover-floor/sp1-gpu-6.8.1-floor.patch` = `tools/fleet/floor.patch` = `tools/fleet/floor-v5.patch`) | sha256 9098c3e5979d057031f43588668201d1f7a53ab17e980c1eaf896cd5e6f38575 |
| V6-07 floor server tarball, sm_86 + sm_89 (`/srv/workers/fleet/igneum-floor-sm8689-v607.tgz` on build-1, served at https://build.igneum.network/fleet/igneum-floor-sm8689-v607.tgz) | tarball 9c02fcc6a3f045d3167c1d9cb6e22c584ea862189d11997754a57cbe1b0d355b; `bin/sp1-gpu-server` db37c38b5feabdec5bbfd0298446da89f6a5d144125e399d6ab617f64f9cb61d; ELF targets sm_86 sm_89; the patch's `floor_tests` 2 passed on the build pod (57 s) |
| The first (known-failed) server, base patch copy | 75d0b4beb5c58fc5e9ae9902ffc9b9d64d8237474fc32a16c2830051bf3d6866 (kept only as the record of 3.5; not served) |
| V6-07 host (box 3, `--features cuda`, cef5234b5 + the branch) | 66662219a363077257c452ccf4be3e184c848d577a277e (full sha in the box's builds.jsonl); its suite: `cargo test -p igneum-prove-host memory_profile` 7 passed on box 3 (21:5x UK), re-run with the 8,500 floors (section 6 amendment) |
| The served 0317 host (the fleet's) | 71bc2438856bb141 |
| A plain master host (dc94dda95, no V6-07 change), for the attribution | f2dfe3ad0db00204 |
| Pods | fb-v607 RunPod RTX 4090 secure 6sgr7dk56z6zwc 19:41 to 19:53Z USD 0.89/h (the first server); fb-v607b RunPod RTX 5090 community 2i76gud81v440j 20:16 to 20:24Z USD 0.69/h (the served server); v607-3060 Vast 54902103 19:46 to 20:33Z USD 0.058/h; v607-4060 Vast 54902102 19:46Z, the host vanished from Vast at 20:1xZ mid-run (USD 0.04); v607-4060b Vast 54906194 20:18 to 20:34Z (USD 0.02); v607-3060c Vast 54909866 20:4x to 20:57Z (USD 0.01); v607-4060c Vast 54909865 20:4x to 21:00Z (USD 0.02); v607-3060d Vast (the bisect, section 6 amendment). Every pod destroyed and read back as unlisted; no miner on any Hetzner box. |
Amendments (22:5x UK):
- The host bisect on v607-3060d (Vast 54913403, RTX 3060, 21:20 to 21:25Z, USD 0.01, destroyed and read back unlisted), the V6-07 server db37c38b, the default path, compressed shard 0: host from 13b729467 (= 6dbd5d2f8^, sha c4b77736107bdbce, sources 8a1c7abf7b7fc378) prove 12.2 s VERIFIED; host from 6dbd5d2f8 (d0d412e9dfeaccd5, sources 600cdc322ba2028b) `Error: public values are 0 bytes, expected 328` after 13 s; host from 2ceb09be9 (89a0453895c1539c, sources 3c293d2236090449, a side branch not carrying 6dbd5d2f8) prove 12.1 s VERIFIED. The fault is 6dbd5d2f8 alone: the FixtureEnv field without the guest re-pin, as its message says.
- The host suite with the 8,500 MiB floors: `cargo test -p igneum-prove-host memory_profile` 7 passed on box 3 (22:2x UK).
- Every pod of the night is destroyed: fb-v607, fb-v607b (RunPod), v607-3060, v607-4060 (vanished), v607-4060b, v607-3060c, v607-4060c, v607-3060d (Vast); about USD 0.45 in all.

View file

@ -48,10 +48,14 @@ from 20 M cycles is the threshold's padded area reached. The witness (5 to 22 KB
## What the patch does (`proving/prover-floor/sp1-gpu-6.8.1-floor.patch`, three files)
1. `builder.rs`: the panic is gone; the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks the element threshold
from a tier table (`element_threshold_for_budget`: over 30 as read, the full 402.6 M; 24 to 30, upstream's 24 GB
figure; 18 to 24 (a 16 GB card), 2^27 + 2^26 = 201.3 M; under 18 (a 12 GB card), 2^27 = 134.2 M);
`SP1_GPU_ELEMENT_THRESHOLD` sets it directly and `SP1_GPU_RECURSION_TRACE_ALLOCATION` the recursion buffer.
1. `builder.rs`: the panic is gone; the card's FREE memory at start (read once per process; or
`SP1_GPU_MEMORY_BUDGET_GB`, the host's lease), never its total, picks the element threshold from a tier table
(`element_threshold_for_budget`: over 30 as read, the full 402.6 M; 24 to 30, upstream's 24 GB figure; 18 to 24
(a 16 GB card alone), 2^27 + 2^26 = 201.3 M; under 18 (a 12 GB or 8 GB card, any card beside a miner), 2^26 =
67.1 M, the value the passing small-card rows used) and the recursion trace allocation (upstream's 2^27 on the
24 GB tier and above, 2^26 + 2^25 under it: V6-07, 8 October 2026, `docs/analysis/floor-memory-profile-2026-10-08.md`);
`SP1_GPU_ELEMENT_THRESHOLD` sets it directly and `SP1_GPU_RECURSION_TRACE_ALLOCATION` the recursion buffer. The
host's own profile table (`host/src/memory_profile.rs`) sets both by environment before the server starts.
The chosen numbers are printed as a `FLOOR opts` line. Every other option is as upstream.
2. `jagged_tracegen/src/lib.rs`: with `SP1_GPU_FLOOR_LOG` set, every trace allocation prints its capacity and,
after the shard's traces are in, the elements actually used and the device memory in use.

View file

@ -416,6 +416,16 @@ Rule: a case maps to a cell only where the cell's tests visibly answer it; cover
- Cases:
- R2-F03-R02 Same job context produces identical accepted work in node, CPU reference, CUDA, Metal, OpenCL and pool.: partial: the worker half (CPU reference, CUDA, Metal, OpenCL) on the same program packs; the node's and the pool's accepted work on the same job context are the CI steward's and the pool lane's cells; PASS only when every listed platform reads one fingerprint and the node and pool halves are green
### rows:v607-floor-memory
- Command: `the V6-07 lane's pod scripts (scratch v607/pod-v607.sh, ab.sh, ab2.sh, ab3.sh) on rented Vast RTX 3060 12 GB and RTX 4060 8 GB one-shots: the served 0317 host and the V6-07 host on the V6-07 floor server (igneum-floor-sm8689-v607.tgz), the fixture fees-v1-shards2.json, the ds55 kit's worker as the miner, nvidia-smi at 1 Hz; recorded in docs/analysis/floor-memory-profile-2026-10-08.md`
- Box class: rented pods (Vast one-shots)
- Fixtures: F0, F1
- Cases:
- GPU-05 Test dataset fit and support-horizon costs: partial: the 8 GB and 12 GB tiers' prover headroom on the 5.5 GiB dataset (free memory, the prover's peak alone, the refusal beside the miner); fragmentation, restart and next-epoch construction not run
- CAP-02 Prove on the actual mining configuration: partial: proving alone (shard and the whole segment path) on the 8 GB and 12 GB tiers on the default job path, the time-share refusal beside the miner with the miner unharmed, memory headroom and proof latency; wall energy only on the 3060 (the 4060 host has no power sensor); induced GPU task failure and wallet control not run; 16 GB and 24 GB tiers not run
- UX-02 Make pause, stop and safe tuning reliable: partial: the prover's safe refusal under memory pressure (exit 78 in one line, the miner's worker unharmed, nothing leaked after the refusal); pause, stop, power limit and the UI's crash ownership are the suite:app cell's
## Automated cases with no harness in the matrix (NOT RUN, the reason)
- GOV-02 Approve thresholds before results: the approval is recorded in the registry's approval field; the automated half (thresholds frozen before any run_status) is the gate rule landing by 21:00
@ -566,4 +576,4 @@ Rule: a case maps to a cell only where the cell's tests visibly answer it; cover
## Count
171 automated cases: 83 mapped to a cell, 145 NOT RUN with a reason.
171 automated cases: 84 mapped to a cell, 145 NOT RUN with a reason.

View file

@ -1019,30 +1019,30 @@
"manual_page": 22,
"owner_lane": "hash lane (a690540514aa453d7)",
"run_status": "NOT RUN",
"evidence_path": "docs/analysis/class-v6/rows/run-ca3-pc1-amd-cardin-20261008T1703.md",
"run_id": "hash-lane-20261008-batch1",
"updated": "2026-10-08T18:41:01.185Z",
"evidence_path": "docs/analysis/class-v6/rows/run-ca3-pc1-amd-cardin-20261008T1703.md; docs/analysis/floor-memory-profile-2026-10-08.md",
"run_id": "v607-floor-memory-20261008T2059Z",
"updated": "2026-10-08T22:22:49.946Z",
"evidence_record": {
"cell": "bench:pc1-amd",
"manifest_sha": "c245d50b9",
"coverage": {
"GPU-01": "partial: the 8 GB AMD cell, fingerprints and rates at 1, 2, 4 and 5.5 GiB; one cell of P02's twelve",
"GPU-05": "partial: the 8 GB tier's fit rows"
},
"at": "2026-10-08T18:41:01.185Z",
"method": "GPU",
"requirement_id": "GPU-05",
"decision": "NOT RUN",
"reviewer": "",
"claim_impact": "",
"method": "GPU",
"cell": "rows:v607-floor-memory",
"manifest_sha": "c411bae9a",
"run_id": "v607-floor-memory-20261008T2059Z",
"evidence": "docs/analysis/floor-memory-profile-2026-10-08.md",
"in_progress": true,
"coverage": "partial: the 8 GB and 12 GB tiers' prover headroom on the 5.5 GiB dataset (free memory, the prover's peak alone, the refusal beside the miner); fragmentation, restart and next-epoch construction not run",
"release_identity": {
"commit": "c245d50b9",
"lockfile": "",
"binary": "",
"network_object": "",
"activation": "",
"profile_hashes": ""
}
"commit": "c411bae9a (branch v607-floor-memory off cef5234b5)",
"lockfile": "proving/igneum-prove/Cargo.lock at cef5234b5",
"binary": "sp1-gpu-server db37c38b5feabdec5bbfd0298446da89f6a5d144125e399d6ab617f64f9cb61d (igneum-floor-sm8689-v607.tgz 9c02fcc6a3f045d3167c1d9cb6e22c584ea862189d11997754a57cbe1b0d355b); igneum-prove-host-0317 71bc2438856bb141; igneum-prove-host v607 66662219a3630772",
"network_object": "none (fixture proofs on rented pods, no network)",
"activation": "none",
"profile_hashes": "floor patch 9098c3e5979d057031f43588668201d1f7a53ab17e980c1eaf896cd5e6f38575"
},
"claim_impact": "the 2.0 proving statement's memory condition per tier: a 12 GB card runs the whole segment path alone; an 8 GB card proves shards only (its aggregation fails) and is refused the chain before setup; both time-share beside the 5.5 GiB miner; the 16 GB and 24 GB tiers unmeasured tonight",
"reviewer": "",
"at": "2026-10-08T22:22:49.946Z"
},
"in_progress_since": "2026-10-08T18:41:01.185Z",
"approvals": {
@ -1076,6 +1076,28 @@
"run_id": "hash-lane-20261008-batch1",
"evidence": "docs/analysis/class-v6/rows/run-ca3-pc1-amd-cardin-20261008T1703.md",
"in_progress": true
},
"rows:v607-floor-memory": {
"requirement_id": "GPU-05",
"decision": "NOT RUN",
"method": "GPU",
"cell": "rows:v607-floor-memory",
"manifest_sha": "c411bae9a",
"run_id": "v607-floor-memory-20261008T2059Z",
"evidence": "docs/analysis/floor-memory-profile-2026-10-08.md",
"in_progress": true,
"coverage": "partial: the 8 GB and 12 GB tiers' prover headroom on the 5.5 GiB dataset (free memory, the prover's peak alone, the refusal beside the miner); fragmentation, restart and next-epoch construction not run",
"release_identity": {
"commit": "c411bae9a (branch v607-floor-memory off cef5234b5)",
"lockfile": "proving/igneum-prove/Cargo.lock at cef5234b5",
"binary": "sp1-gpu-server db37c38b5feabdec5bbfd0298446da89f6a5d144125e399d6ab617f64f9cb61d (igneum-floor-sm8689-v607.tgz 9c02fcc6a3f045d3167c1d9cb6e22c584ea862189d11997754a57cbe1b0d355b); igneum-prove-host-0317 71bc2438856bb141; igneum-prove-host v607 66662219a3630772",
"network_object": "none (fixture proofs on rented pods, no network)",
"activation": "none",
"profile_hashes": "floor patch 9098c3e5979d057031f43588668201d1f7a53ab17e980c1eaf896cd5e6f38575"
},
"claim_impact": "the 2.0 proving statement's memory condition per tier: a 12 GB card runs the whole segment path alone; an 8 GB card proves shards only (its aggregation fails) and is refused the chain before setup; both time-share beside the 5.5 GiB miner; the 16 GB and 24 GB tiers unmeasured tonight",
"reviewer": "",
"at": "2026-10-08T22:22:49.946Z"
}
}
},
@ -5194,21 +5216,21 @@
"manual_page": 38,
"owner_lane": "enforced-proving lane (a6e8f84588b809d62)",
"run_status": "PASS",
"evidence_path": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"run_id": "enforced-proving-20261008-02",
"updated": "2026-10-08T21:37:02.905Z",
"evidence_path": "docs/plans/proving-enforcement/cache-history-2.json",
"run_id": "enforced-proving-20261008-03",
"updated": "2026-10-08T22:24:23.945Z",
"evidence_record": {
"requirement_id": "EVM-08",
"decision": "PASS",
"method": "native",
"cell": "harness:proving-enforcement",
"manifest_sha": "3f672661",
"run_id": "enforced-proving-20261008-02",
"evidence": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: the controlled verifier upgrade as a key succession (both pairs embedded, the window, the start refusal) and its fast-time crossing; the mixed-client crossing on the live network and the failed-distribution step are not run",
"release_identity": {
"commit": "3f672661",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5217,7 +5239,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T21:37:02.905Z"
"at": "2026-10-08T22:24:23.945Z"
},
"dependency": "the first pin into elf/prior/ (the node lane, key-succession-pin, tonight by 21:00) and the devnet-4 succession height (main, by 22:00 under the floor rule)",
"approvals": {
@ -5232,13 +5254,13 @@
"decision": "PASS",
"method": "native",
"cell": "harness:proving-enforcement",
"manifest_sha": "3f672661",
"run_id": "enforced-proving-20261008-02",
"evidence": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: the controlled verifier upgrade as a key succession (both pairs embedded, the window, the start refusal) and its fast-time crossing; the mixed-client crossing on the live network and the failed-distribution step are not run",
"release_identity": {
"commit": "3f672661",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5247,7 +5269,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T21:37:02.905Z"
"at": "2026-10-08T22:24:23.945Z"
}
}
}
@ -5298,21 +5320,21 @@
"manual_page": 39,
"owner_lane": "enforced-proving lane (a6e8f84588b809d62)",
"run_status": "NOT RUN",
"evidence_path": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-exec.log; docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json; build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-consensus.log",
"run_id": "202-a284380b-5f50afd1",
"updated": "2026-10-08T22:11:37.839Z",
"evidence_path": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-exec.log; docs/plans/proving-enforcement/cache-history-2.json; build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-consensus.log",
"run_id": "enforced-proving-20261008-03",
"updated": "2026-10-08T22:24:23.945Z",
"evidence_record": {
"requirement_id": "ZKP-01",
"decision": "NOT RUN",
"decision": "PASS",
"method": "native",
"cell": "suite:consensus",
"manifest_sha": "a284380b",
"run_id": "202-a284380b-5f50afd1",
"evidence": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-consensus.log",
"cell": "harness:proving-enforcement",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: a proof-less or wrong-statement block refused by consensus",
"coverage": "partial: the seven refusals as named tests (no proof never inserted; a wrong proof and another program id invalidate the block; the honest record paid once) and the fast-time crossing with a modified producer; the genuine-proof positive control through ordinary network paths is the testnet's live read, not run here",
"release_identity": {
"commit": "a284380b",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5321,7 +5343,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T22:11:37.839Z"
"at": "2026-10-08T22:24:23.945Z"
},
"in_progress_since": "2026-10-08T19:32:50.856Z",
"approvals": {
@ -5358,13 +5380,13 @@
"decision": "PASS",
"method": "native",
"cell": "harness:proving-enforcement",
"manifest_sha": "3f672661",
"run_id": "enforced-proving-20261008-02",
"evidence": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: the seven refusals as named tests (no proof never inserted; a wrong proof and another program id invalidate the block; the honest record paid once) and the fast-time crossing with a modified producer; the genuine-proof positive control through ordinary network paths is the testnet's live read, not run here",
"release_identity": {
"commit": "3f672661",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5373,7 +5395,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T21:37:02.905Z"
"at": "2026-10-08T22:24:23.945Z"
},
"suite:consensus": {
"requirement_id": "ZKP-01",
@ -5428,21 +5450,21 @@
"manual_page": 39,
"owner_lane": "enforced-proving lane (a6e8f84588b809d62)",
"run_status": "PASS",
"evidence_path": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"run_id": "enforced-proving-20261008-02",
"updated": "2026-10-08T21:37:02.905Z",
"evidence_path": "docs/plans/proving-enforcement/cache-history-2.json",
"run_id": "enforced-proving-20261008-03",
"updated": "2026-10-08T22:24:23.945Z",
"evidence_record": {
"requirement_id": "ZKP-02",
"decision": "PASS",
"method": "native",
"cell": "harness:proving-enforcement",
"manifest_sha": "3f672661",
"run_id": "enforced-proving-20261008-02",
"evidence": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: the pinned-id refusal on a real proof, the daemon's start refusal, key succession (the epochs, the accepted ids, a verdict counting only where its pair is accepted, the seven refusals on both sides of the height, the start refusal under a succession, the real-proof refusal of a pair the epoch does not accept) and the fast-time crossing of a scheduled succession with a real proof under each pair; the downgrade through an old node path is not run",
"release_identity": {
"commit": "3f672661",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5451,7 +5473,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T21:37:02.905Z"
"at": "2026-10-08T22:24:23.945Z"
},
"dependency": "the mixed-pair crossing needs the node lane's first pin into elf/prior/ (key-succession-pin, tonight by 21:00) and the next-pair proof (build-2:/home/build/enforced-fixtures/next-pair-block-56-shard-0-compressed.bin)",
"approvals": {
@ -5466,13 +5488,13 @@
"decision": "PASS",
"method": "native",
"cell": "harness:proving-enforcement",
"manifest_sha": "3f672661",
"run_id": "enforced-proving-20261008-02",
"evidence": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: the pinned-id refusal on a real proof, the daemon's start refusal, key succession (the epochs, the accepted ids, a verdict counting only where its pair is accepted, the seven refusals on both sides of the height, the start refusal under a succession, the real-proof refusal of a pair the epoch does not accept) and the fast-time crossing of a scheduled succession with a real proof under each pair; the downgrade through an old node path is not run",
"release_identity": {
"commit": "3f672661",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5481,7 +5503,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T21:37:02.905Z"
"at": "2026-10-08T22:24:23.945Z"
}
}
},
@ -5514,21 +5536,21 @@
"manual_page": 39,
"owner_lane": "enforced-proving lane (a6e8f84588b809d62)",
"run_status": "NOT RUN",
"evidence_path": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-exec.log; docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"run_id": "202-a284380b-5f50afd1",
"updated": "2026-10-08T22:11:37.839Z",
"evidence_path": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-exec.log; docs/plans/proving-enforcement/cache-history-2.json",
"run_id": "enforced-proving-20261008-03",
"updated": "2026-10-08T22:24:23.945Z",
"evidence_record": {
"requirement_id": "ZKP-03",
"decision": "NOT RUN",
"decision": "PASS",
"method": "native",
"cell": "suite:exec",
"manifest_sha": "a284380b",
"run_id": "202-a284380b-5f50afd1",
"evidence": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-exec.log",
"cell": "harness:proving-enforcement",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: a snapshot under another digest refused, the epoch streams bound to the digest",
"coverage": "partial: the chain binding (a record signed for another network), the replay, a wrong block, a wrong shard, a stale record outside the window; the fork and replayed-sync steps are not run",
"release_identity": {
"commit": "a284380b",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5537,7 +5559,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T22:11:37.839Z"
"at": "2026-10-08T22:24:23.945Z"
},
"in_progress_since": "2026-10-08T19:32:50.856Z",
"approvals": {
@ -5574,13 +5596,13 @@
"decision": "PASS",
"method": "native",
"cell": "harness:proving-enforcement",
"manifest_sha": "3f672661",
"run_id": "enforced-proving-20261008-02",
"evidence": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: the chain binding (a record signed for another network), the replay, a wrong block, a wrong shard, a stale record outside the window; the fork and replayed-sync steps are not run",
"release_identity": {
"commit": "3f672661",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5589,7 +5611,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T21:37:02.905Z"
"at": "2026-10-08T22:24:23.945Z"
}
}
},
@ -5622,21 +5644,21 @@
"manual_page": 40,
"owner_lane": "enforced-proving lane (a6e8f84588b809d62)",
"run_status": "PASS",
"evidence_path": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"run_id": "enforced-proving-20261008-02",
"updated": "2026-10-08T21:37:02.905Z",
"evidence_path": "docs/plans/proving-enforcement/cache-history-2.json",
"run_id": "enforced-proving-20261008-03",
"updated": "2026-10-08T22:24:23.945Z",
"evidence_record": {
"requirement_id": "ZKP-04",
"decision": "PASS",
"method": "native",
"cell": "harness:proving-enforcement",
"manifest_sha": "3f672661",
"run_id": "enforced-proving-20261008-02",
"evidence": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: an altered payout address (after signing and re-signed), a statement over altered rewards or payouts vetoed natively, the inputs commitment of P22 stage 1 (shard 0 and the aggregator carry the commitment the node recomputes) and the derivation commitment of stage 2 (the payouts derived in the guest from the carried records); the rewards' proved derivation is P22 stage 3, PENDING",
"release_identity": {
"commit": "3f672661",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5645,7 +5667,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T21:37:02.905Z"
"at": "2026-10-08T22:24:23.945Z"
},
"dependency": "the proof side of the derivation (P22 stages 1 to 3) changes both guests' public values and so both program ids, so it ships only through a key succession (docs/design/key-succession.md, key-succession-node 291ee6ae); stage 1 (the inputs commitment carried by shard 0 and the aggregator, vetoed natively) is on branch p22-stage-1 under test since 18:46 UK; the clocks (stage 1 09:00, stage 2 14:00, stage 3's design 18:00 UK on 9 October 2026) stand because the succession's first height is asked of main by 22:00 under the floor rule; without a named height the stages are built and tested on their branch and wait for the succession that carries them",
"approvals": {
@ -5660,13 +5682,13 @@
"decision": "PASS",
"method": "native",
"cell": "harness:proving-enforcement",
"manifest_sha": "3f672661",
"run_id": "enforced-proving-20261008-02",
"evidence": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: an altered payout address (after signing and re-signed), a statement over altered rewards or payouts vetoed natively, the inputs commitment of P22 stage 1 (shard 0 and the aggregator carry the commitment the node recomputes) and the derivation commitment of stage 2 (the payouts derived in the guest from the carried records); the rewards' proved derivation is P22 stage 3, PENDING",
"release_identity": {
"commit": "3f672661",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5675,7 +5697,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T21:37:02.905Z"
"at": "2026-10-08T22:24:23.945Z"
}
}
},
@ -5708,21 +5730,21 @@
"manual_page": 40,
"owner_lane": "enforced-proving lane (a6e8f84588b809d62)",
"run_status": "PASS",
"evidence_path": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"run_id": "enforced-proving-20261008-02",
"updated": "2026-10-08T21:37:02.905Z",
"evidence_path": "docs/plans/proving-enforcement/cache-history-2.json",
"run_id": "enforced-proving-20261008-03",
"updated": "2026-10-08T22:24:23.945Z",
"evidence_record": {
"requirement_id": "ZKP-05",
"decision": "PASS",
"method": "native",
"cell": "harness:proving-enforcement",
"manifest_sha": "3f672661",
"run_id": "enforced-proving-20261008-02",
"evidence": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: a duplicate of a paid record pays nothing and the honest record pays once; the race, reorder and crash-recover steps are not run",
"release_identity": {
"commit": "3f672661",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5731,7 +5753,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T21:37:02.905Z"
"at": "2026-10-08T22:24:23.945Z"
},
"approvals": {
"scope_approved": null,
@ -5745,13 +5767,13 @@
"decision": "PASS",
"method": "native",
"cell": "harness:proving-enforcement",
"manifest_sha": "3f672661",
"run_id": "enforced-proving-20261008-02",
"evidence": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: a duplicate of a paid record pays nothing and the honest record pays once; the race, reorder and crash-recover steps are not run",
"release_identity": {
"commit": "3f672661",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5760,7 +5782,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T21:37:02.905Z"
"at": "2026-10-08T22:24:23.945Z"
}
}
},
@ -5918,21 +5940,21 @@
"manual_page": 41,
"owner_lane": "enforced-proving lane (a6e8f84588b809d62)",
"run_status": "PASS",
"evidence_path": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"run_id": "enforced-proving-20261008-02",
"updated": "2026-10-08T21:37:02.905Z",
"evidence_path": "docs/plans/proving-enforcement/cache-history-2.json",
"run_id": "enforced-proving-20261008-03",
"updated": "2026-10-08T22:24:23.945Z",
"evidence_record": {
"requirement_id": "ZKP-08",
"decision": "PASS",
"method": "native",
"cell": "harness:proving-enforcement",
"manifest_sha": "3f672661",
"run_id": "enforced-proving-20261008-02",
"evidence": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: the body rule and the payment rule read one floor, a producer with its body rule off still pays nothing on honest nodes (the fast-time crossing); the replacement-prover and authority steps are not run; the proof side of the derivation is P22 stage 3, PENDING",
"release_identity": {
"commit": "3f672661",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5941,7 +5963,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T21:37:02.905Z"
"at": "2026-10-08T22:24:23.945Z"
},
"dependency": "the proof side of the derivation (P22 stages 1 to 3) changes both guests' public values and so both program ids, so it ships only through a key succession (docs/design/key-succession.md, key-succession-node 291ee6ae); stage 1 (the inputs commitment carried by shard 0 and the aggregator, vetoed natively) is on branch p22-stage-1 under test since 18:46 UK; the clocks (stage 1 09:00, stage 2 14:00, stage 3's design 18:00 UK on 9 October 2026) stand because the succession's first height is asked of main by 22:00 under the floor rule; without a named height the stages are built and tested on their branch and wait for the succession that carries them",
"approvals": {
@ -5956,13 +5978,13 @@
"decision": "PASS",
"method": "native",
"cell": "harness:proving-enforcement",
"manifest_sha": "3f672661",
"run_id": "enforced-proving-20261008-02",
"evidence": "docs/plans/proving-enforcement/succession-360-window-300-vcache-7.json",
"manifest_sha": "c591e63b",
"run_id": "enforced-proving-20261008-03",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json",
"in_progress": false,
"coverage": "partial: the body rule and the payment rule read one floor, a producer with its body rule off still pays nothing on honest nodes (the fast-time crossing); the replacement-prover and authority steps are not run; the proof side of the derivation is P22 stage 3, PENDING",
"release_identity": {
"commit": "3f672661",
"commit": "c591e63b",
"lockfile": "",
"binary": "",
"network_object": "",
@ -5971,7 +5993,7 @@
},
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T21:37:02.905Z"
"at": "2026-10-08T22:24:23.945Z"
}
}
}
@ -6094,12 +6116,12 @@
"manual_page": 42,
"owner_lane": "fleet lane (ac055d60427caab99)",
"run_status": "NOT RUN",
"evidence_path": "",
"run_id": "",
"updated": "2026-10-08T19:22:31.798Z",
"evidence_path": "docs/analysis/floor-memory-profile-2026-10-08.md",
"run_id": "v607-floor-memory-20261008T2059Z",
"updated": "2026-10-08T22:22:49.946Z",
"evidence_record": {
"reason": "proving on the mining configuration is the fleet lane's",
"at": "2026-10-08T19:22:31.798Z",
"at": "2026-10-08T22:22:49.946Z",
"method": "static",
"requirement_id": "CAP-02",
"decision": "NOT RUN",
@ -6121,27 +6143,30 @@
"claim_authorised": null
},
"evidence_records": {
"record": {
"reason": "proving on the mining configuration is the fleet lane's",
"at": "2026-10-08T19:22:31.798Z",
"method": "static",
"rows:v607-floor-memory": {
"requirement_id": "CAP-02",
"decision": "NOT RUN",
"reviewer": "",
"claim_impact": "",
"method": "GPU",
"cell": "rows:v607-floor-memory",
"manifest_sha": "c411bae9a",
"run_id": "v607-floor-memory-20261008T2059Z",
"evidence": "docs/analysis/floor-memory-profile-2026-10-08.md",
"in_progress": true,
"coverage": "partial: proving alone (shard and the whole segment path) on the 8 GB and 12 GB tiers on the default job path, the time-share refusal beside the miner with the miner unharmed, memory headroom and proof latency; wall energy only on the 3060 (the 4060 host has no power sensor); induced GPU task failure and wallet control not run; 16 GB and 24 GB tiers not run",
"release_identity": {
"commit": "",
"lockfile": "",
"binary": "",
"network_object": "",
"activation": "",
"profile_hashes": ""
"commit": "c411bae9a (branch v607-floor-memory off cef5234b5)",
"lockfile": "proving/igneum-prove/Cargo.lock at cef5234b5",
"binary": "sp1-gpu-server db37c38b5feabdec5bbfd0298446da89f6a5d144125e399d6ab617f64f9cb61d (igneum-floor-sm8689-v607.tgz 9c02fcc6a3f045d3167c1d9cb6e22c584ea862189d11997754a57cbe1b0d355b); igneum-prove-host-0317 71bc2438856bb141; igneum-prove-host v607 66662219a3630772",
"network_object": "none (fixture proofs on rented pods, no network)",
"activation": "none",
"profile_hashes": "floor patch 9098c3e5979d057031f43588668201d1f7a53ab17e980c1eaf896cd5e6f38575"
},
"run_id": "",
"evidence": "",
"in_progress": false
"claim_impact": "the 2.0 proving statement's memory condition per tier: a 12 GB card runs the whole segment path alone; an 8 GB card proves shards only (its aggregation fails) and is refused the chain before setup; both time-share beside the 5.5 GiB miner; the 16 GB and 24 GB tiers unmeasured tonight",
"reviewer": "",
"at": "2026-10-08T22:22:49.946Z"
}
}
},
"in_progress_since": "2026-10-08T22:22:49.946Z"
},
{
"id": "CAP-03",
@ -9628,30 +9653,30 @@
"manual_page": 57,
"owner_lane": "shipper (ae892a8b0f78fe31c)",
"run_status": "NOT RUN",
"evidence_path": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/miner-77b5acc5/box4-app.log",
"run_id": "202-a284380b-5f50afd1",
"updated": "2026-10-08T22:11:37.839Z",
"evidence_path": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/miner-77b5acc5/box4-app.log; docs/analysis/floor-memory-profile-2026-10-08.md",
"run_id": "v607-floor-memory-20261008T2059Z",
"updated": "2026-10-08T22:22:49.946Z",
"evidence_record": {
"requirement_id": "UX-02",
"decision": "NOT RUN",
"method": "native",
"cell": "suite:app",
"manifest_sha": "a284380b",
"run_id": "202-a284380b-5f50afd1",
"evidence": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/miner-77b5acc5/box4-app.log",
"in_progress": false,
"coverage": "partial: the engine state machine's pause, stop and knob tests; the operator study is UX-01's",
"method": "GPU",
"cell": "rows:v607-floor-memory",
"manifest_sha": "c411bae9a",
"run_id": "v607-floor-memory-20261008T2059Z",
"evidence": "docs/analysis/floor-memory-profile-2026-10-08.md",
"in_progress": true,
"coverage": "partial: the prover's safe refusal under memory pressure (exit 78 in one line, the miner's worker unharmed, nothing leaked after the refusal); pause, stop, power limit and the UI's crash ownership are the suite:app cell's",
"release_identity": {
"commit": "a284380b",
"lockfile": "",
"binary": "",
"network_object": "",
"activation": "",
"profile_hashes": ""
"commit": "c411bae9a (branch v607-floor-memory off cef5234b5)",
"lockfile": "proving/igneum-prove/Cargo.lock at cef5234b5",
"binary": "sp1-gpu-server db37c38b5feabdec5bbfd0298446da89f6a5d144125e399d6ab617f64f9cb61d (igneum-floor-sm8689-v607.tgz 9c02fcc6a3f045d3167c1d9cb6e22c584ea862189d11997754a57cbe1b0d355b); igneum-prove-host-0317 71bc2438856bb141; igneum-prove-host v607 66662219a3630772",
"network_object": "none (fixture proofs on rented pods, no network)",
"activation": "none",
"profile_hashes": "floor patch 9098c3e5979d057031f43588668201d1f7a53ab17e980c1eaf896cd5e6f38575"
},
"claim_impact": "",
"claim_impact": "the 2.0 proving statement's memory condition per tier: a 12 GB card runs the whole segment path alone; an 8 GB card proves shards only (its aggregation fails) and is refused the chain before setup; both time-share beside the 5.5 GiB miner; the 16 GB and 24 GB tiers unmeasured tonight",
"reviewer": "",
"at": "2026-10-08T22:11:37.839Z"
"at": "2026-10-08T22:22:49.946Z"
},
"in_progress_since": "2026-10-08T19:32:50.856Z",
"approvals": {
@ -9682,6 +9707,28 @@
"claim_impact": "",
"reviewer": "",
"at": "2026-10-08T22:11:37.839Z"
},
"rows:v607-floor-memory": {
"requirement_id": "UX-02",
"decision": "NOT RUN",
"method": "GPU",
"cell": "rows:v607-floor-memory",
"manifest_sha": "c411bae9a",
"run_id": "v607-floor-memory-20261008T2059Z",
"evidence": "docs/analysis/floor-memory-profile-2026-10-08.md",
"in_progress": true,
"coverage": "partial: the prover's safe refusal under memory pressure (exit 78 in one line, the miner's worker unharmed, nothing leaked after the refusal); pause, stop, power limit and the UI's crash ownership are the suite:app cell's",
"release_identity": {
"commit": "c411bae9a (branch v607-floor-memory off cef5234b5)",
"lockfile": "proving/igneum-prove/Cargo.lock at cef5234b5",
"binary": "sp1-gpu-server db37c38b5feabdec5bbfd0298446da89f6a5d144125e399d6ab617f64f9cb61d (igneum-floor-sm8689-v607.tgz 9c02fcc6a3f045d3167c1d9cb6e22c584ea862189d11997754a57cbe1b0d355b); igneum-prove-host-0317 71bc2438856bb141; igneum-prove-host v607 66662219a3630772",
"network_object": "none (fixture proofs on rented pods, no network)",
"activation": "none",
"profile_hashes": "floor patch 9098c3e5979d057031f43588668201d1f7a53ab17e980c1eaf896cd5e6f38575"
},
"claim_impact": "the 2.0 proving statement's memory condition per tier: a 12 GB card runs the whole segment path alone; an 8 GB card proves shards only (its aggregation fails) and is refused the chain before setup; both time-share beside the 5.5 GiB miner; the 16 GB and 24 GB tiers unmeasured tonight",
"reviewer": "",
"at": "2026-10-08T22:22:49.946Z"
}
}
},

View file

@ -0,0 +1,147 @@
{
"case": "cache-history",
"node": "/srv/builds/igneum-wt-cache/vendor/igneum-node/target/release/igneumd",
"proveBin": "/srv/builds/igneum-wt-cache/proving/igneum-prove/target/release",
"pair": {
"shard": "0x51cd8cba314a32b60fac393949a4571714da6d6656717d286011601b2a163fe7",
"aggregator": "0x05b395ec238f67406084f8a449d400f64c87dc62151daebf66695864727ded8a"
},
"startedAt": "2026-10-08T21:53:32.526Z",
"phases": {
"context": {
"block": 165,
"submit": {
"accepted": true,
"block": "0x4259577a6be0f3cf3cb1f1512b8df758caa2790d5cafebc791a55076d9cd1839",
"keyHash": "0x6ce6a6cfb1e7d69f52e0764cd4ad005d244097ee9773f7918c2ab532ad73d2df",
"new": true,
"number": "0xa5",
"reason": "accepted",
"shard": 0
},
"observed": {
"paid": [],
"carried": []
},
"verifies": {
"cacheAnswers": 0,
"calls": 1,
"contextRefusals": 1,
"sp1": 0
},
"refusals": [
"a carried proof record's proof does not verify: shard record block 165 shard 0 by 6ce6a6cfb1e7d69f52e0764cd4ad005d244097ee9773f7918c2ab532ad73d2df (proof f74f7f7daf2cdcb0): shard record block 165 shard 0 by 6ce6a6cfb1e7d69f52e0764cd4ad005d244097ee9773f7918c2ab532ad73d2df: the proof's public values hash to 0x58d620e493a23ebdacdec4f3c3f46e5936a18fbdcb3880341e2bb8b1555f7955, the record's statement is 0xd8191e436529a541554168fb68a9f6663485d6b743d1b49e15ec3bbf3a5c17f0, disconnecting from peer 127.0.0.1:49762."
]
},
"honest": {
"block": 94,
"submit": {
"accepted": true,
"block": "0x6e958a1b931c6eacce696422a6611151848bb65ca8d99c36510bd7fab2925c39",
"keyHash": "0x4a28017035561e93c93032f70dba67436b68506efd3ac761f608df142c8a7594",
"new": true,
"number": "0x5e",
"reason": "accepted",
"shard": 0
},
"observed": {
"paid": [
{
"carrierNumber": "0x169",
"keyHash": "0x4a28017035561e93c93032f70dba67436b68506efd3ac761f608df142c8a7594",
"payout": "0xc3c3c3c3c3c3c3c3c3c3c3c3c3c3c3c3c3c3c3c3",
"shard": 0,
"wei": "0x8cc4d14f7888000"
}
],
"carried": [
{
"signer": "0x4a28017035561e93c93032f70dba67436b68506efd3ac761f608df142c8a7594",
"carrier": "0x169",
"rejected": "",
"paidWei": "0x8cc4d14f7888000"
},
{
"signer": "0x4a28017035561e93c93032f70dba67436b68506efd3ac761f608df142c8a7594",
"carrier": "0x16a",
"rejected": "shard already paid",
"paidWei": "0x0"
}
]
},
"verifies": {
"cacheAnswers": 1,
"calls": 2,
"contextRefusals": 1,
"sp1": 1
}
},
"invalid": {
"blocks": [
465,
591
],
"submits": [
{
"accepted": true,
"block": "0xb80c043ced2e770092ab96769fbba1533df9a9a7eb61e054cbd7a98c678f6af9",
"keyHash": "0x89470283d097209eb0287ba87717938b8642c19b12bb37b5e30ed2a4328f2aaf",
"new": true,
"number": "0x1d1",
"reason": "accepted",
"shard": 0
},
{
"accepted": true,
"block": "0xd64c1a9c8d6c901ae1ed2cb2c1afccfa35113dcf1d38704477dfbf6a810d79e3",
"keyHash": "0xaa474491707c6931a3d6d2fdb84e4d91fa3f3c57b7776b9c1fb1b01a1f96678a",
"new": true,
"number": "0x24f",
"reason": "accepted",
"shard": 0
}
],
"verifies": {
"after3a": {
"cacheAnswers": 1,
"calls": 3,
"contextRefusals": 1,
"sp1": 2
},
"after3b": {
"cacheAnswers": 2,
"calls": 3,
"contextRefusals": 1,
"sp1": 2
}
},
"refusals": {
"after3a": 2,
"after3b": 3
},
"reasons": [
"a carried proof record's proof does not verify: shard record block 165 shard 0 by 6ce6a6cfb1e7d69f52e0764cd4ad005d244097ee9773f7918c2ab532ad73d2df (proof f74f7f7daf2cdcb0): shard record block 165 shard 0 by 6ce6a6cfb1e7d69f52e0764cd4ad005d244097ee9773f7918c2ab532ad73d2df: the proof's public values hash to 0x58d620e493a23ebdacdec4f3c3f46e5936a18fbdcb3880341e2bb8b1555f7955, the record's statement is 0xd8191e436529a541554168fb68a9f6663485d6b743d1b49e15ec3bbf3a5c17f0, disconnecting from peer 127.0.0.1:49762.",
"a carried proof record's proof does not verify: shard record block 465 shard 0 by 89470283d097209eb0287ba87717938b8642c19b12bb37b5e30ed2a4328f2aaf (proof 79e76573983222b5): shard record block 465 shard 0 by 89470283d097209eb0287ba87717938b8642c19b12bb37b5e30ed2a4328f2aaf: the proof bytes are not a bincode SP1 proof: invalid value: integer `4284109422`, expected variant index 0 <= i < 4, disconnecting from peer 127.0.0.1:47690.",
"a carried proof record's proof does not verify: shard record block 591 shard 0 by aa474491707c6931a3d6d2fdb84e4d91fa3f3c57b7776b9c1fb1b01a1f96678a (proof 79e76573983222b5): shard record block 465 shard 0 by 89470283d097209eb0287ba87717938b8642c19b12bb37b5e30ed2a4328f2aaf: the proof bytes are not a bincode SP1 proof: invalid value: integer `4284109422`, expected variant index 0 <= i < 4, disconnecting from peer 127.0.0.1:51924."
]
}
},
"proof": {
"block": 94,
"statement": "0x58d620e493a23ebdacdec4f3c3f46e5936a18fbdcb3880341e2bb8b1555f7955",
"native": "0x58d620e493a23ebdacdec4f3c3f46e5936a18fbdcb3880341e2bb8b1555f7955",
"bytes": 1272961,
"sha256": "0xf74f7f7daf2cdcb031add25f28a0787d278de1079aa81aa51ac3f2e744750df7",
"hostLine": "RESULT compressed shard 0: prove 61.0 s, proof 1272961 bytes, verify 0.080 s, VERIFIED; statement 0x58d620e493a23ebdacdec4f3c3f46e5936a18fbdcb3880341e2bb8b1555f7955 proof sha256 0xf74f7f7daf2cdcb031add25f28a0787d278de1079aa81aa51ac3f2e744750df7 prover 0xc3c3c3c3c3c3c3c3c3C3C3c3C3C3C3c3C3C3c3c3 at 2026-10-08T21:56:24Z"
},
"verdict": {
"paidOnce": true,
"sp1Verifies": 2,
"cacheAnswers": 2,
"contextRefused": true,
"contextCheap": true,
"invalidRefusedTwiceNoVerify": true,
"pass": true
},
"endedAt": "2026-10-08T22:05:36.331Z"
}

View file

@ -0,0 +1,254 @@
#!/usr/bin/env node
// The two-node cache-history case (review B F01, Phase 1 item (c), 8 October 2026): on a node with the verdict cache
// of verdict-cache-fix-node 3f672661, a proof's bytes refused once for CONTEXT (another statement) must still pay
// when the same bytes arrive later as their own honest record, answered from the cached facts with no second SP1
// verify; bytes refused once for being INVALID must be refused again at a later carrier, from the cache, with no
// second verify. Three local nodes at fast time as in proving-enforcement.mjs: H1 and H2 honest (the verifier in
// consensus), A the carrier (skip rule, trust verify) whose templates carry what its pool holds.
//
// The real proof is one of this chain's own blocks: after the chain has run, block n is exported from H1
// (igneum_exportSegments), cut by igneum-prove-export and proven compressed by igneum-prove-host on the box's CPU
// (about 70 to 110 s), so its statement is H1's native statement for (n, shard 0, PAYOUT). The node binaries and the
// proving binaries must carry the SAME pair (the node's elf/ overlay and the host's embedded pin: pin C tonight), and
// the node's object names that pair (proving_shard_program_id / proving_aggregator_id from the manifest).
//
// node infra/fast-time/proving-cache-history.mjs --prove-bin <dir with igneum-prove-host and igneum-prove-export>
// [--before 90] [--after 120] [--slot 0] [--out <json>]
//
// Verdict: (1) block n shard 0 paid exactly once, to PAYOUT; (2) H1 ran the SP1 verifier exactly twice over the run (the real
// proof once, the invalid bytes once): the context refusal of phase 1 is a pre-check (the statement differs) that runs no
// verifier and is never cached, re-read at every carrier; (3) the context refusal and the invalid refusal both read in H1's
// log; (4) the second use of the invalid bytes is answered from the cache with no new verify (cacheAnswers at least 1).
import { spawn, spawnSync } from 'node:child_process';
import { mkdirSync, rmSync, writeFileSync, readFileSync, openSync, existsSync, appendFileSync } from 'node:fs';
import { createHash, randomBytes } from 'node:crypto';
const ROOT = new URL('../../', import.meta.url).pathname;
const FILE = `${ROOT}infra/fast-time/override-60x.json`;
const MANIFEST = `${ROOT}proving/igneum-prove/elf/manifest.json`;
const IGNEUMD = process.env.IGNEUMD || `${ROOT}vendor/igneum-node/target/release/igneumd`;
const CPU_MINER = process.env.IGNEUM_MINER || `${ROOT}vendor/igneum-node/target/release/igneum-miner`;
const args = process.argv.slice(2);
const flag = (name, dflt) => { const i = args.indexOf(`--${name}`); return i >= 0 ? Number(args[i + 1]) : dflt; };
const sflag = (name, dflt) => { const i = args.indexOf(`--${name}`); return i >= 0 ? args[i + 1] : dflt; };
const BEFORE = flag('before', 90), AFTER = flag('after', 120), SLOT = flag('slot', 0);
const PROVE_BIN = sflag('prove-bin', `${ROOT}proving/igneum-prove/target/release`);
const GENESIS_BITS = flag('genesis-bits', 0x1f010000);
const OUT = sflag('out') || `${ROOT}docs/plans/proving-enforcement/cache-history.json`;
const BASE = 30890 + SLOT * 40, SUFFIX = 985 + SLOT;
const CHAIN = `igneum-devnet-${SUFFIX}`;
const TMP = `/tmp/igneum-fast-time-ch${SLOT}`;
const PIDS = `${TMP}/pids`;
const NEVER = '18446744073709551615';
for (const b of [IGNEUMD, CPU_MINER, `${PROVE_BIN}/igneum-prove-host`, `${PROVE_BIN}/igneum-prove-export`]) if (!existsSync(b)) { console.error(`missing ${b}`); process.exit(2); }
const t0 = Date.now();
const log = (s) => { const t = new Date().toISOString().slice(11, 23); console.log(`${t} ch${SLOT} ${s}`); };
const sleep = (ms) => new Promise(r => setTimeout(r, ms));
const hex = (b) => '0x' + Buffer.from(b).toString('hex');
const sha256 = (b) => createHash('sha256').update(b).digest('hex');
// leftovers of an earlier run on this slot: by the pid file only
rmSync(TMP, { recursive: true, force: true }); mkdirSync(TMP, { recursive: true });
const started = [];
const track = (proc) => { started.push(proc); try { appendFileSync(PIDS, `${proc.pid}\n`); } catch { } };
// the override: the 60x file with the pair named from the manifest, the verifier in consensus from genesis, no succession
let baseText = readFileSync(FILE, 'utf8');
for (const k of ['proving_payment_activation_daa']) baseText = baseText.replace(new RegExp(`\\s*"${k}":\\s*[^,}\\n]+,?`), '');
const manifest = JSON.parse(readFileSync(MANIFEST, 'utf8'));
function mergeOverrideText(text, fields) {
let out = text;
for (const k of Object.keys(fields)) out = out.replace(new RegExp(`\\s*"${k}":\\s*[^,}\\n]+,?`), '');
const extra = Object.entries(fields).map(([k, v]) => `"${k}": ${typeof v === 'string' && !/^\d+$/.test(v) && v !== 'true' && v !== 'false' ? JSON.stringify(v) : v}`).join(', ');
return out.replace(/,?\s*}\s*$/, `,\n ${extra}\n}\n`);
}
const override = `${TMP}/override.json`;
writeFileSync(override, mergeOverrideText(baseText, {
genesis_bits: GENESIS_BITS, skip_proof_of_work: false, proving_v0_activation_daa: '0', proving_consensus_verify_daa: '0',
proving_shard_program_id: manifest.shard.program_id, proving_aggregator_id: manifest.aggregator.program_id,
verifier_in_consensus: 'true', proving_key_succession_daa: NEVER, proving_key_succession_window_daa: '120',
proving_next_shard_program_id: '', proving_next_aggregator_id: '',
program_class_v3_activation_daa: NEVER, program_class_v4_activation_daa: NEVER,
}));
const DAY_MS = (() => { const m = /"pow_day_ms":\s*([0-9]+)/.exec(baseText); return m ? +m[1] : 1440000; })();
class Node {
constructor(name, i, peers = [], env = {}) {
this.name = name; this.i = i; this.grpcPort = BASE + i * 10; this.p2pPort = BASE + i * 10 + 1; this.jsonPort = BASE + i * 10 + 2; this.evmPort = BASE + i * 10 + 3;
this.peers = peers; this.env = env; this.dir = `${TMP}/${name}`; this.logFile = `${this.dir}/node.log`;
}
get grpc() { return `grpc://127.0.0.1:${this.grpcPort}`; }
async start() {
mkdirSync(this.dir, { recursive: true });
const out = openSync(this.logFile, 'a');
const a = ['--devnet', `--devnet-suffix=${SUFFIX}`, '--nodnsseed', '--disable-upnp', '--nologfiles', '--enable-unsynced-mining', '--utxoindex',
`--appdir=${this.dir}`, `--rpclisten=127.0.0.1:${this.grpcPort}`, `--rpclisten-json=127.0.0.1:${this.jsonPort}`, `--evm-rpclisten=127.0.0.1:${this.evmPort}`,
`--listen=127.0.0.1:${this.p2pPort}`, `--override-params-file=${override}`, '--loglevel=info', '--yes'];
if (this.peers.length) for (const p of this.peers) a.push(`--addpeer=127.0.0.1:${p}`); else a.push('--outpeers=0');
this.proc = spawn(IGNEUMD, a, { stdio: ['ignore', out, out], env: { ...process.env, ...this.env } });
track(this.proc);
for (let k = 0; k < 60; k++) {
await sleep(1000);
try { await this.exec('igneum_getExecStatus', []); log(`${this.name} up (pid ${this.proc.pid}) after ${k + 1} s`); return; } catch { }
if (this.proc.exitCode !== null) throw new Error(`${this.name} exited ${this.proc.exitCode}: ${readFileSync(this.logFile, 'utf8').split('\n').slice(-5).join(' | ')}`);
}
throw new Error(`${this.name} did not answer in 60 s`);
}
async exec(method, params) {
const r = await fetch(`http://127.0.0.1:${this.evmPort}`, { method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify({ jsonrpc: '2.0', id: 1, method, params }) });
const j = await r.json();
if (j.error) throw new Error(`${method}: ${j.error.message || JSON.stringify(j.error)}`);
return j.result;
}
logText() { try { return readFileSync(this.logFile, 'utf8'); } catch { return ''; } }
stop() { try { this.proc.kill('SIGINT'); } catch { } }
}
function startMiner(node, label, secs) {
const out = openSync(`${TMP}/${label}.log`, 'a');
const p = spawn(CPU_MINER, ['mine', node.grpc, '1', String(secs), label, '--engine', 'igneum-pow', '--payout-label', label, '--status-secs', '30', '--no-vote'], { stdio: ['ignore', out, out], env: { ...process.env, IGNEUM_POW_DAY_MS: String(DAY_MS) } });
track(p);
return p;
}
function signRecord(label, chain, block, number, shard, payout, statement, proofHash) {
const r = spawnSync(CPU_MINER, ['sign-record', label, chain, block, String(number), String(shard), payout, statement, proofHash], { encoding: 'utf8' });
if (r.status !== 0) throw new Error(`sign-record failed: ${r.stderr || r.stdout}`);
return JSON.parse(r.stdout.trim().split('\n').pop());
}
const H1 = new Node('H1', 0);
const H2 = new Node('H2', 1, [H1.p2pPort]);
const A = new Node('A', 2, [H1.p2pPort, H2.p2pPort], { IGNEUM_TEST_SKIP_PROOF_RULE: '1', IGNEUM_PROOF_VERIFY: 'trust' });
const nodes = [H1, H2, A];
let miners = [];
async function stopAll() {
for (const m of miners) { try { m.kill('SIGTERM'); } catch { } }
for (const n of nodes) n.stop();
await sleep(2000);
for (const p of started) { try { p.kill('SIGKILL'); } catch { } }
}
process.on('SIGINT', async () => { await stopAll(); process.exit(130); });
const PAYOUT = '0x' + 'c3'.repeat(20);
async function tipNumber(node) { const s = await node.exec('igneum_getExecStatus', []); return Number(s.executedTip ?? 0); }
async function sameTip(a, h) {
try { const x = await a.exec('igneum_getExecStatus', []); const y = await h.exec('igneum_getExecStatus', []); return x.executedTipHash && x.executedTipHash === y.executedTipHash; } catch { return false; }
}
async function rejoin(label) {
for (let k = 0; k < 120; k++) { if (await sameTip(A, H1)) { log(`${label}: A is on H1's tip (${k * 5} s)`); return true; } await sleep(5000); }
log(`${label}: A did not rejoin H1's tip in 600 s`); return false;
}
async function verifies(node) { const s = await node.exec('igneum_getProvingStatus', []); return s.verifies || { calls: 0, sp1: 0, contextRefusals: 0, cacheAnswers: 0 }; }
const refusals = (node) => (node.logText().match(/a carried proof record's proof does not verify: [^\n]*/g) || []);
async function submitVia(node, name, record, proof) {
let out;
try { out = await node.exec('igneum_submitProofRecord', [{ record, proof }]); } catch (e) { out = { accepted: false, reason: String(e.message) }; }
log(`${name}: A's pool ${out.accepted ? 'accepted' : 'refused'} (${out.reason || ''})`);
return out;
}
async function observe(n) {
try {
const r = await H1.exec('igneum_getProofRecords', ['0x' + n.toString(16)]);
return { paid: (r.paid || []).filter(Boolean), carried: (r.carried || []).map(c => ({ signer: c.keyHash, carrier: c.carrierNumber, rejected: c.rejected, paidWei: c.paidWei })) };
} catch (e) { return { error: e.message }; }
}
/// A real proof of block n of this chain: exported from H1, cut and proven by the host (CPU), returning the proof
/// bytes and the statement the host printed (which must equal H1's native statement for (n, 0, PAYOUT)).
function proveBlock(n) {
const dir = `${TMP}/prove-${n}`; mkdirSync(dir, { recursive: true });
const body = JSON.stringify({ jsonrpc: '2.0', id: 1, method: 'igneum_exportSegments', params: ['0x0', '0x' + n.toString(16)] });
const r = spawnSync('curl', ['-s', '-m', '600', '-X', 'POST', `http://127.0.0.1:${H1.evmPort}`, '-H', 'Content-Type: application/json', '--data-binary', body], { encoding: 'utf8', maxBuffer: 1 << 30 });
if (r.status !== 0) throw new Error(`export failed: ${r.stderr}`);
writeFileSync(`${dir}/export.json`, JSON.stringify(JSON.parse(r.stdout).result));
const e = spawnSync(`${PROVE_BIN}/igneum-prove-export`, [`${dir}/export.json`, String(n), `${dir}/block-${n}.json`, '--source', `fast-time cache-history ${CHAIN}`], { encoding: 'utf8' });
if (e.status !== 0) throw new Error(`igneum-prove-export failed: ${(e.stdout + e.stderr).slice(-400)}`);
log(`block ${n} cut: ${(e.stdout.trim().split('\n').pop() || '').slice(0, 160)}`);
const h = spawnSync(`${PROVE_BIN}/igneum-prove-host`, [`${dir}/block-${n}.json`, '--mode', 'compressed', '--shard', '0', '--prover', PAYOUT, '--out', `${dir}/results.json`], { encoding: 'utf8', env: { ...process.env, SP1_PROVER: 'cpu' } });
const line = (h.stdout.match(/RESULT compressed shard 0: [^\n]*/) || [''])[0];
log(`host: ${line.slice(0, 200)}${h.status !== 0 ? ` (exit ${h.status}: ${(h.stderr || '').slice(-300)})` : ''}`);
if (h.status !== 0 || !/VERIFIED/.test(line)) throw new Error(`the host did not prove block ${n}`);
const file = `${dir}/block-${n}-shard-0-compressed.bin`;
if (!existsSync(file)) throw new Error(`no proof file ${file}`);
const statement = (line.match(/statement (0x[0-9a-f]{64})/) || [])[1];
return { bytes: readFileSync(file), statement, line };
}
const result = { case: 'cache-history', node: IGNEUMD, proveBin: PROVE_BIN, pair: { shard: manifest.shard.program_id, aggregator: manifest.aggregator.program_id }, startedAt: new Date().toISOString(), phases: {} };
try {
await H1.start(); await H2.start(); await A.start();
miners = [startMiner(H1, 'h1', BEFORE + 3 * AFTER + 900), startMiner(H2, 'h2', BEFORE + 3 * AFTER + 900), startMiner(A, 'attacker', BEFORE + 3 * AFTER + 900)];
log(`phase 0: mining ${BEFORE} s`);
await sleep(BEFORE * 1000);
const v0 = await verifies(H1);
// the block to prove: outside the exclusive window, executed on every node
let tip = await tipNumber(H1);
const n = Math.max(1, tip - 15);
const plan = await H1.exec('igneum_getShardPlan', ['0x' + n.toString(16), PAYOUT]);
const native = plan.shards[0].statement;
log(`block ${n} ${plan.hash}: H1's native statement for shard 0 and ${PAYOUT}: ${native}`);
const proof = proveBlock(n);
result.proof = { block: n, statement: proof.statement, native, bytes: proof.bytes.length, sha256: '0x' + sha256(proof.bytes), hostLine: proof.line };
if (proof.statement !== native) throw new Error(`the host's statement ${proof.statement} is not H1's ${native} (host and node layouts or pairs differ)`);
// phase 1: the real bytes under a record for another block (the statement of block m): refused for context
await rejoin('before phase 1');
tip = await tipNumber(A);
const m = Math.max(1, tip - 15);
const planM = await A.exec('igneum_getShardPlan', ['0x' + m.toString(16), PAYOUT]);
const recWrong = signRecord('ch-wrong', CHAIN, planM.hash, m, 0, PAYOUT, planM.shards[0].statement, '0x' + sha256(proof.bytes)).record;
const r1 = await submitVia(A, 'phase 1 (real bytes, another block\'s statement)', recWrong, hex(proof.bytes));
await sleep(AFTER * 1000);
const v1 = await verifies(H1);
const o1 = await observe(m);
result.phases.context = { block: m, submit: r1, observed: o1, verifies: v1, refusals: refusals(H1) };
log(`phase 1: H1 verifies ${JSON.stringify(v1)}; block ${m} paid ${JSON.stringify(o1.paid)}; refusals ${refusals(H1).length}`);
// phase 2: the same bytes as their own honest record: answered from the cached facts, paid once
await rejoin('before phase 2');
const recHonest = signRecord('ch-honest', CHAIN, plan.hash, n, 0, PAYOUT, native, '0x' + sha256(proof.bytes)).record;
const r2 = await submitVia(A, 'phase 2 (the same bytes, the honest record)', recHonest, hex(proof.bytes));
await sleep(AFTER * 1000);
const v2 = await verifies(H1);
const o2 = await observe(n);
result.phases.honest = { block: n, submit: r2, observed: o2, verifies: v2 };
log(`phase 2: H1 verifies ${JSON.stringify(v2)}; block ${n} paid ${JSON.stringify(o2.paid)}; carried ${JSON.stringify(o2.carried)}`);
// phase 3: invalid bytes twice, under two carriers: verified once, refused from the cache the second time
await rejoin('before phase 3');
const bad = randomBytes(2048);
tip = await tipNumber(A);
const p = Math.max(1, tip - 15);
const planP = await A.exec('igneum_getShardPlan', ['0x' + p.toString(16), PAYOUT]);
const recBad1 = signRecord('ch-bad1', CHAIN, planP.hash, p, 0, PAYOUT, planP.shards[0].statement, '0x' + sha256(bad)).record;
const r3 = await submitVia(A, 'phase 3a (invalid bytes)', recBad1, hex(bad));
await sleep(AFTER * 1000);
const v3a = await verifies(H1);
const ref3a = refusals(H1).length;
await rejoin('before phase 3b');
tip = await tipNumber(A);
const q = Math.max(1, tip - 15);
const planQ = await A.exec('igneum_getShardPlan', ['0x' + q.toString(16), PAYOUT]);
const recBad2 = signRecord('ch-bad2', CHAIN, planQ.hash, q, 0, PAYOUT, planQ.shards[0].statement, '0x' + sha256(bad)).record;
const r4 = await submitVia(A, 'phase 3b (the same invalid bytes, a later carrier)', recBad2, hex(bad));
await sleep(AFTER * 1000);
const v3b = await verifies(H1);
const ref3b = refusals(H1).length;
result.phases.invalid = { blocks: [p, q], submits: [r3, r4], verifies: { after3a: v3a, after3b: v3b }, refusals: { after3a: ref3a, after3b: ref3b }, reasons: refusals(H1) };
log(`phase 3: H1 verifies after 3a ${JSON.stringify(v3a)}, after 3b ${JSON.stringify(v3b)}; refusals ${ref3a} then ${ref3b}`);
// the verdict
const paidOnce = (o2.paid || []).length === 1 && (o1.paid || []).length === 0;
const runs = Number(v3b.sp1) - Number(v0.sp1); // SP1 verifies: the real proof once (phase 2), the invalid bytes once (3a); the context refusal of phase 1 runs no verifier
const cache = Number(v3b.cacheAnswers) - Number(v0.cacheAnswers);
const contextCheap = Number(v1.contextRefusals) - Number(v0.contextRefusals) >= 1 && Number(v1.sp1) === Number(v0.sp1);
const contextRefused = result.phases.context.refusals.length >= 1;
const invalidRefusedTwiceNoVerify = ref3b > ref3a && Number(v3b.sp1) === Number(v3a.sp1);
const pass = paidOnce && runs === 2 && cache >= 1 && contextRefused && contextCheap && invalidRefusedTwiceNoVerify;
result.verdict = { paidOnce, sp1Verifies: runs, cacheAnswers: cache, contextRefused, contextCheap, invalidRefusedTwiceNoVerify, pass };
result.endedAt = new Date().toISOString();
writeFileSync(OUT, JSON.stringify(result, null, 2));
log(`RESULT ${pass ? 'PASS' : 'FAIL'}: paid once=${paidOnce}; SP1 verifies=${runs} (expect 2: the real proof once, the invalid bytes once); context refusal read=${contextRefused} and cheap (no verifier)=${contextCheap}; cache answers=${cache} (expect >= 1); invalid refused again without a verify=${invalidRefusedTwiceNoVerify}; ${OUT}`);
await stopAll();
process.exit(pass ? 0 : 1);
} catch (e) {
log(`ERROR: ${e.message}`);
result.error = e.message; result.endedAt = new Date().toISOString();
try { writeFileSync(OUT, JSON.stringify(result, null, 2)); } catch { }
await stopAll();
process.exit(1);
}

View file

@ -276,7 +276,7 @@ async function forge(n, phase, which = 'pn') {
async function observe(n) {
try {
const r = await H1.exec('igneum_getProofRecords', ['0x' + n.toString(16)]);
const carried = (r.carried || []).map(c => ({ key: c.keyHash, carrier: c.carrierNumber, rejected: c.rejected, paidWei: c.paidWei }));
const carried = (r.carried || []).map(c => ({ signer: c.keyHash, carrier: c.carrierNumber, rejected: c.rejected, paidWei: c.paidWei }));
// `paid` holds one entry per shard, null when unpaid: keep the paid ones only
return { paid: (r.paid || []).filter(Boolean), carried };
} catch (e) { return { error: e.message }; }

View file

@ -20,6 +20,7 @@
//! ids with no setup. `--mode verify` uses SP1's light verifier and the pinned verifying key: no prover client,
//! no key generation (the 114 s to 138 s the Mac's node spent per proof on 5 October).
mod memory_profile;
mod pinned;
mod proof_system;
@ -69,6 +70,15 @@ fn run() -> Result<()> {
let args: Vec<String> = std::env::args().collect();
let arg = |name: &str| args.iter().position(|a| a == name).and_then(|i| args.get(i + 1)).cloned();
let mode = arg("--mode").unwrap_or_else(|| "all".into());
// V6-07 (8 October 2026): on the GPU prover the memory profile for this process's workload is chosen here, once,
// from the card's FREE memory (the engine's lease, else nvidia-smi) and handed to the floor server by environment
// before the SP1 client spawns it; a budget under the workload's floor is refused with exit 78 before any setup.
if std::env::var("SP1_PROVER").map(|p| p == "cuda").unwrap_or(false) {
if let Err(refusal) = memory_profile::apply(&mode) {
println!("RESULT memory_profile refused: {refusal}");
std::process::exit(memory_profile::EXIT_REFUSED);
}
}
let pinned = pinned::Pinned::load()?;
if mode == "id" {
println!("RESULT id: {}", pinned.describe());

View file

@ -0,0 +1,402 @@
//! The pinned memory profile per workload (master review R1 residual V6-07, 8 October 2026): the GPU server's two
//! device buffers (the core element threshold and the recursion trace allocation) are chosen HERE, once, from the
//! memory that is FREE on the card when the host starts, never from the card's total. The table below is the
//! profile; `choose` picks a row for the workload this process is admitted for; `apply` hands the row to the server
//! through the floor patch's environment (`SP1_GPU_ELEMENT_THRESHOLD`, `SP1_GPU_RECURSION_TRACE_ALLOCATION`,
//! `SP1_GPU_MEMORY_BUDGET_GB`) before the SP1 client spawns it. An explicit `SP1_GPU_ELEMENT_THRESHOLD` or
//! `SP1_GPU_RECURSION_TRACE_ALLOCATION` already in the environment is a hand override and wins, named in the line.
//!
//! Where the free memory comes from, in order (the app lane's device coordinator interface, the shipper's answer of
//! 20:4x UK 8 October 2026; the engine owns admission and the host never re-reads a card the engine leased):
//! 1. `IGNEUM_PROVE_MEM_BUDGET_MB`: the lease's grant, the hard ceiling for this process (the engine's coordinator,
//! app/igneum-app `src/device.rs`, passes it with `IGNEUM_PROVE_DEVICE`, `IGNEUM_PROVE_WORKLOAD`,
//! `IGNEUM_PROVE_MEM_FREE_MB` and `IGNEUM_PROVE_DEADLINE_S` on every spawn); with `IGNEUM_PROVE_MEM_FREE_MB`
//! beside it the LESSER of the two decides (`lease_reading`);
//! 2. `IGNEUM_PROVE_MEM_FREE_MB` alone: the engine's own read at admission, when no grant is given;
//! 3. `nvidia-smi --query-gpu=memory.free` on the device (`IGNEUM_PROVE_DEVICE`, else `IGNEUM_CUDA_DEVICE`, else the
//! first of `CUDA_VISIBLE_DEVICES`, else 0): the fleet's `tools/fleet/box-prover.py` path and a hand run, where no
//! engine admits the job.
//! When none answers (no nvidia-smi on the path), no row is applied and the server's own free-memory rule decides.
//!
//! A budget under the workload's floor is refused before the server starts: one line, exit code 78 (the engine
//! records the refusal on the lease and the window reads "proving needs N GB free").
//!
//! The rows cited for the small tier's 2^26: `docs/analysis/prover-tiers-real-cards.md` (6 October 2026,
//! `alone-comp-26-v1`: RTX 3060 7.4 GB 14.4 s, RTX 3080 8.0 GB 7.1 s, RTX 4060 7.4 GB 18.4 s, RTX 4060 Ti 8 GB 7.6 GB
//! 9.6 s, RTX 4060 Ti 16 GB 7.8 GB 11.6 s, RTX 4070 7.6 GB 12.1 s, RTX 5070 7.6 GB 4.8 s, all verified) and
//! `docs/analysis/class-v6/coexist-rows.md` (8 October 2026, `proof_alone` at `SP1_GPU_ELEMENT_THRESHOLD=67108864`:
//! RTX 3060 peak 7,525 MiB 13.2 s verified, RTX 4060 peak 7,532 MiB 8.2 s verified). No small-card row ever passed
//! at the patch's earlier default of 2^27 on this host; the 10 GB tier's 2^27 rows (8,642 MiB, 7 October 2026) ran
//! alone on the segment host and are not this path.
use std::fmt;
/// Upstream's core element threshold (sp1-core-executor 6.8.1 `opts.rs` 12): 2^28 + 2^27.
pub const ELEMENT_THRESHOLD_FULL: u64 = (1 << 28) + (1 << 27);
/// Upstream's 24 GB tier (sp1-gpu `builder.rs` 44): the full threshold less 2^26 + 2^25 + 2^24.
pub const ELEMENT_THRESHOLD_24GB: u64 = ELEMENT_THRESHOLD_FULL - (1 << 26) - (1 << 25) - (1 << 24);
/// The floor patch's 16 GB tier: 2^27 + 2^26 (unmeasured on this fixture; the patch's own figure).
pub const ELEMENT_THRESHOLD_16GB: u64 = (1 << 27) + (1 << 26);
/// The small-card threshold: 2^26, the value every passing small-card row used (module note).
pub const ELEMENT_THRESHOLD_SMALL: u64 = 1 << 26;
/// Upstream's recursion trace allocation (sp1-gpu `builder.rs` 15): 2^27 elements, 0.75 GiB each.
pub const RECURSION_TRACE_ALLOCATION: usize = 1 << 27;
/// The small-card recursion trace allocation: the recursion keys and shards use 90,177,536 elements each (35.6 M
/// preprocessed, 54.5 M main; `docs/analysis/prover-floor.md`, sweep 1), so 2^26 + 2^25 = 100,663,296 holds them
/// with the stacking slack the patch's `floor_capacity` adds. The two constants differ (`recursion_branches_differ`).
pub const RECURSION_TRACE_ALLOCATION_SMALL: usize = (1 << 26) + (1 << 25);
/// What a recursion key or shard was measured to use (elements); the small allocation must clear it with slack.
pub const RECURSION_TRACE_USED: usize = 90_177_536;
/// The one job this process is admitted for.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum Workload {
/// One shard's compressed (or core) proof.
Shard,
/// The aggregator guest over shard proofs (and the previous segment's proof).
Aggregate,
/// A whole segment in one process: every shard, then the aggregation.
Chain,
}
impl Workload {
pub fn name(self) -> &'static str {
match self {
Workload::Shard => "shard",
Workload::Aggregate => "aggregate",
Workload::Chain => "chain",
}
}
/// `IGNEUM_PROVE_WORKLOAD` when the engine names it, else the host's own mode.
pub fn from_env_or_mode(mode: &str) -> Option<Workload> {
if let Ok(w) = std::env::var("IGNEUM_PROVE_WORKLOAD") {
return match w.as_str() {
"shard" => Some(Workload::Shard),
"aggregate" => Some(Workload::Aggregate),
"chain" => Some(Workload::Chain),
_ => None,
};
}
Workload::from_mode(mode)
}
/// The workload a host mode runs on the GPU (modes that never prove return None).
pub fn from_mode(mode: &str) -> Option<Workload> {
match mode {
"shard" | "compressed" | "core" => Some(Workload::Shard),
"aggregate" => Some(Workload::Aggregate),
"chain" | "block" | "all" => Some(Workload::Chain),
_ => None,
}
}
/// The least free memory (MiB) a workload runs in. Shard: the small tier's measured peak (7,504 to 7,632 MiB
/// on the 3060 and 4060 at 2^26, 8 October 2026) plus headroom for the driver's own context. Aggregate and
/// chain: the aggregation is the peak of a chain run, 8,306 MiB on the RTX 3060 (the default path, 20:54Z
/// 8 October 2026); on the RTX 4060 (7,807 MiB free) the same run proved both shards and aborted at the
/// aggregation's 486 MiB allocation (`FLOOR abort` at `slop/crates/tensor/src/inner.rs:51`, 20:56Z), so the
/// floor sits above an 8 GB card: `docs/analysis/floor-memory-profile-2026-10-08.md` section 3.
pub fn floor_mib(self) -> u64 {
match self {
Workload::Shard => 7_700,
Workload::Aggregate => 8_500,
Workload::Chain => 8_500,
}
}
}
/// One tier of the profile table: the least free memory it needs and the two buffers it sets.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct Tier {
pub name: &'static str,
pub min_free_mib: u64,
pub element_threshold: u64,
pub recursion_trace_allocation: usize,
}
/// The pinned table, largest first. The free-memory lines are the card tiers as the floor patch reads them (free
/// GiB, ceiling, +4 as upstream computed it from the total): 26 GiB free reads over 30 (a 32 GB card alone), 20 GiB
/// reads 24 (a 24 GB card alone), 14 GiB reads 18 (a 16 GB card alone), under that the small tier (12 GB and 8 GB
/// cards alone; any card beside a miner's resident set).
pub const TIERS: [Tier; 4] = [
Tier { name: "full", min_free_mib: 26 * 1024, element_threshold: ELEMENT_THRESHOLD_FULL, recursion_trace_allocation: RECURSION_TRACE_ALLOCATION },
Tier { name: "24gb", min_free_mib: 20 * 1024, element_threshold: ELEMENT_THRESHOLD_24GB, recursion_trace_allocation: RECURSION_TRACE_ALLOCATION },
Tier { name: "16gb", min_free_mib: 14 * 1024, element_threshold: ELEMENT_THRESHOLD_16GB, recursion_trace_allocation: RECURSION_TRACE_ALLOCATION_SMALL },
Tier { name: "small", min_free_mib: 0, element_threshold: ELEMENT_THRESHOLD_SMALL, recursion_trace_allocation: RECURSION_TRACE_ALLOCATION_SMALL },
];
/// The row chosen for a process.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct Profile {
pub workload: Workload,
pub tier: Tier,
pub free_mib: u64,
}
/// Why a process is refused before the server starts.
#[derive(Clone, Debug, PartialEq, Eq)]
pub struct Refusal {
pub workload: Workload,
pub free_mib: u64,
pub floor_mib: u64,
}
impl fmt::Display for Refusal {
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
write!(
f,
"proving needs {} GB free on the card for the {} workload ({} MiB floor); {} MiB free",
(self.floor_mib + 1023) / 1024,
self.workload.name(),
self.floor_mib,
self.free_mib
)
}
}
/// The exit code of a refusal (the engine reads it on the lease).
pub const EXIT_REFUSED: i32 = 78;
/// The tier for a free-memory reading (pure; the table's first row whose line the reading clears).
pub fn tier_for(free_mib: u64) -> Tier {
*TIERS.iter().find(|t| free_mib >= t.min_free_mib).expect("the small tier has no floor")
}
/// The row for a workload at a free-memory reading, or the refusal.
pub fn choose(workload: Workload, free_mib: u64) -> Result<Profile, Refusal> {
let floor_mib = workload.floor_mib();
if free_mib < floor_mib {
return Err(Refusal { workload, free_mib, floor_mib });
}
Ok(Profile { workload, tier: tier_for(free_mib), free_mib })
}
/// Where a free-memory reading came from, for the line.
#[derive(Clone, Debug, PartialEq, Eq)]
pub enum Source {
/// `IGNEUM_PROVE_MEM_BUDGET_MB`, the engine's lease.
Lease,
/// `IGNEUM_PROVE_MEM_FREE_MB`, the engine's read without a grant.
EngineRead,
/// `nvidia-smi --query-gpu=memory.free` on the named device.
NvidiaSmi(String),
}
impl fmt::Display for Source {
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
match self {
Source::Lease => write!(f, "lease IGNEUM_PROVE_MEM_BUDGET_MB"),
Source::EngineRead => write!(f, "engine IGNEUM_PROVE_MEM_FREE_MB"),
Source::NvidiaSmi(d) => write!(f, "nvidia-smi device {d}"),
}
}
}
/// The reading from the engine's two numbers: the grant is the ceiling and the engine's free reading is a fact, so
/// the row is chosen from the LESSER of the two when both are given (the window lane's engine half, 20:5x UK
/// 8 October 2026, grants a fixed proof budget beside the miner, 7,532 MiB plus 10 percent, while a 12 GB card
/// beside the 6.1 GiB miner has 6,159 MiB free: the grant alone would pick a row the card cannot hold, and the
/// refusal must come from the free figure). One number alone is used as it is.
pub fn lease_reading(budget: Option<u64>, free: Option<u64>) -> Option<(u64, Source)> {
match (budget, free) {
(Some(b), Some(f)) => Some((b.min(f), Source::Lease)),
(Some(b), None) => Some((b, Source::Lease)),
(None, Some(f)) => Some((f, Source::EngineRead)),
(None, None) => None,
}
}
fn env_u64(name: &str) -> Option<u64> {
std::env::var(name).ok().and_then(|s| s.trim().parse::<u64>().ok())
}
/// The device ordinal the host runs on, as the module note orders it.
pub fn device_ordinal() -> String {
if let Ok(d) = std::env::var("IGNEUM_PROVE_DEVICE") {
return d;
}
if let Ok(d) = std::env::var("IGNEUM_CUDA_DEVICE") {
return d;
}
if let Ok(v) = std::env::var("CUDA_VISIBLE_DEVICES") {
if let Some(first) = v.split(',').next() {
if !first.trim().is_empty() {
return first.trim().to_string();
}
}
}
"0".into()
}
/// The free memory on the card at this call, from the sources in the module note's order.
pub fn read_free_mib() -> Option<(u64, Source)> {
let budget = env_u64("IGNEUM_PROVE_MEM_BUDGET_MB");
let free = env_u64("IGNEUM_PROVE_MEM_FREE_MB");
if let Some(choice) = lease_reading(budget, free) {
return Some(choice);
}
let dev = device_ordinal();
let out = std::process::Command::new("nvidia-smi")
.args(["--query-gpu=memory.free", "--format=csv,noheader,nounits", "-i", &dev])
.output()
.ok()?;
if !out.status.success() {
return None;
}
let text = String::from_utf8_lossy(&out.stdout);
let first = text.lines().next()?.trim();
first.parse::<u64>().ok().map(|v| (v, Source::NvidiaSmi(dev)))
}
/// The environment the row sets for the floor server (what `apply` writes), as (name, value) pairs; an explicit
/// hand override already present keeps its value and is reported.
pub fn env_for(p: &Profile) -> Vec<(&'static str, String)> {
vec![
("SP1_GPU_ELEMENT_THRESHOLD", p.tier.element_threshold.to_string()),
("SP1_GPU_RECURSION_TRACE_ALLOCATION", p.tier.recursion_trace_allocation.to_string()),
("SP1_GPU_MEMORY_BUDGET_GB", format!("{:.1}", p.free_mib as f64 / 1024.0)),
]
}
/// Chooses and applies the row for this process: prints one `RESULT memory_profile` line and returns Ok(Some) with
/// the profile, Ok(None) when no reading was possible (the server's own rule decides) or the workload never proves,
/// and Err with the refusal line (the caller exits `EXIT_REFUSED`).
pub fn apply(mode: &str) -> Result<Option<Profile>, Refusal> {
let Some(workload) = Workload::from_env_or_mode(mode) else { return Ok(None) };
let Some((free_mib, source)) = read_free_mib() else {
println!("RESULT memory_profile: no free-memory reading (no lease, no nvidia-smi on device {}); the server's own free-memory rule decides for workload {}", device_ordinal(), workload.name());
return Ok(None);
};
let profile = choose(workload, free_mib)?;
let mut set = Vec::new();
let mut kept = Vec::new();
for (k, v) in env_for(&profile) {
match std::env::var(k) {
Ok(have) if k != "SP1_GPU_MEMORY_BUDGET_GB" => kept.push(format!("{k}={have} (hand override kept, the row said {v})")),
_ => {
std::env::set_var(k, &v);
set.push(format!("{k}={v}"));
}
}
}
let deadline = std::env::var("IGNEUM_PROVE_DEADLINE_S").ok().map(|d| format!(", deadline {d} s")).unwrap_or_default();
println!(
"RESULT memory_profile: workload {} tier {} from {} MiB free ({}){}: element_threshold {} recursion_trace_allocation {}; set {}{}",
workload.name(),
profile.tier.name,
free_mib,
source,
deadline,
profile.tier.element_threshold,
profile.tier.recursion_trace_allocation,
set.join(" "),
if kept.is_empty() { String::new() } else { format!("; {}", kept.join("; ")) }
);
Ok(Some(profile))
}
#[cfg(test)]
mod tests {
use super::*;
/// The table's values are pinned: a change here is a change of the profile and needs its rows.
#[test]
fn table_is_pinned() {
assert_eq!(ELEMENT_THRESHOLD_FULL, 402_653_184);
assert_eq!(ELEMENT_THRESHOLD_24GB, 285_212_672);
assert_eq!(ELEMENT_THRESHOLD_16GB, 201_326_592);
assert_eq!(ELEMENT_THRESHOLD_SMALL, 67_108_864);
assert_eq!(RECURSION_TRACE_ALLOCATION, 134_217_728);
assert_eq!(RECURSION_TRACE_ALLOCATION_SMALL, 100_663_296);
let names: Vec<&str> = TIERS.iter().map(|t| t.name).collect();
assert_eq!(names, ["full", "24gb", "16gb", "small"]);
assert_eq!(TIERS[0].min_free_mib, 26_624);
assert_eq!(TIERS[1].min_free_mib, 20_480);
assert_eq!(TIERS[2].min_free_mib, 14_336);
assert_eq!(TIERS[3].min_free_mib, 0);
assert_eq!(TIERS[3].element_threshold, ELEMENT_THRESHOLD_SMALL);
assert_eq!(TIERS[3].recursion_trace_allocation, RECURSION_TRACE_ALLOCATION_SMALL);
assert_eq!(TIERS[0].recursion_trace_allocation, RECURSION_TRACE_ALLOCATION);
assert_eq!(Workload::Shard.floor_mib(), 7_700);
assert_eq!(Workload::Aggregate.floor_mib(), 8_500);
assert_eq!(Workload::Chain.floor_mib(), 8_500);
assert_eq!(EXIT_REFUSED, 78);
}
/// The two recursion branches are two values (the review found one constant in both), and the small one
/// clears what a recursion key or shard was measured to use, with a stacking height of slack.
#[test]
fn recursion_branches_differ() {
assert_ne!(RECURSION_TRACE_ALLOCATION, RECURSION_TRACE_ALLOCATION_SMALL);
assert!(RECURSION_TRACE_ALLOCATION_SMALL < RECURSION_TRACE_ALLOCATION);
assert!(RECURSION_TRACE_ALLOCATION_SMALL >= RECURSION_TRACE_USED + (1 << 22));
assert_ne!(tier_for(24 * 1024).recursion_trace_allocation, tier_for(12 * 1024).recursion_trace_allocation);
}
/// The small tier's threshold is 2^26, the passing rows' value; the card tiers alone read their own rows.
#[test]
fn tiers_from_free_memory() {
assert_eq!(tier_for(8_186).name, "small");
assert_eq!(tier_for(8_186).element_threshold, 1 << 26);
assert_eq!(tier_for(12_100).name, "small");
assert_eq!(tier_for(12_100).element_threshold, 1 << 26);
assert_eq!(tier_for(16_100).name, "16gb");
assert_eq!(tier_for(24_200).name, "24gb");
assert_eq!(tier_for(32_300).name, "full");
// a 12 GB card beside the 5.5 GiB miner (6,129 MiB resident) reads the small tier by what is FREE
assert_eq!(tier_for(12_288 - 6_129).name, "small");
}
/// The refusal: under the floor, before the server starts; at the floor, the small row.
#[test]
fn refuses_under_the_floor() {
let r = choose(Workload::Shard, 12_288 - 6_129).unwrap_err();
assert_eq!(r.floor_mib, 7_700);
assert_eq!(r.free_mib, 6_159);
assert_eq!(r.to_string(), "proving needs 8 GB free on the card for the shard workload (7,700 MiB floor); 6,159 MiB free".replace(",", ""));
let p = choose(Workload::Shard, 7_700).unwrap();
assert_eq!(p.tier.name, "small");
// an 8 GB card (7,807 MiB free on the RTX 4060) proves shards and is refused the chain (its aggregation
// aborted at 7,631 MiB used plus a 486 MiB allocation on 8 October 2026); a 12 GB card runs it
let p = choose(Workload::Shard, 7_807).unwrap();
assert_eq!(p.tier.element_threshold, ELEMENT_THRESHOLD_SMALL);
assert_eq!(p.tier.recursion_trace_allocation, RECURSION_TRACE_ALLOCATION_SMALL);
let r = choose(Workload::Chain, 7_807).unwrap_err();
assert_eq!(r.to_string(), "proving needs 9 GB free on the card for the chain workload (8500 MiB floor); 7807 MiB free");
assert!(choose(Workload::Aggregate, 7_807).is_err());
assert_eq!(choose(Workload::Chain, 11_898).unwrap().tier.name, "small");
}
/// The engine's grant beside its free reading: the lesser decides; one alone is used as it is.
#[test]
fn lease_takes_the_lesser_of_grant_and_free() {
assert_eq!(lease_reading(Some(8_285), Some(6_159)), Some((6_159, Source::Lease)));
assert_eq!(lease_reading(Some(8_285), Some(12_287)), Some((8_285, Source::Lease)));
assert_eq!(lease_reading(Some(8_285), None), Some((8_285, Source::Lease)));
assert_eq!(lease_reading(None, Some(12_287)), Some((12_287, Source::EngineRead)));
assert_eq!(lease_reading(None, None), None);
// a 12 GB card beside the miner under a fixed grant is refused by its free figure
assert!(choose(Workload::Shard, lease_reading(Some(8_285), Some(6_159)).unwrap().0).is_err());
}
/// The host's modes map to the three workloads; modes that never prove map to none.
#[test]
fn workload_from_mode() {
assert_eq!(Workload::from_mode("compressed"), Some(Workload::Shard));
assert_eq!(Workload::from_mode("core"), Some(Workload::Shard));
assert_eq!(Workload::from_mode("aggregate"), Some(Workload::Aggregate));
assert_eq!(Workload::from_mode("chain"), Some(Workload::Chain));
assert_eq!(Workload::from_mode("all"), Some(Workload::Chain));
assert_eq!(Workload::from_mode("verify"), None);
assert_eq!(Workload::from_mode("id"), None);
assert_eq!(Workload::from_mode("native"), None);
}
/// The environment a row hands the floor server.
#[test]
fn env_for_the_server() {
let p = choose(Workload::Shard, 8_186).unwrap();
let env = env_for(&p);
assert_eq!(env[0], ("SP1_GPU_ELEMENT_THRESHOLD", "67108864".to_string()));
assert_eq!(env[1], ("SP1_GPU_RECURSION_TRACE_ALLOCATION", "100663296".to_string()));
assert_eq!(env[2], ("SP1_GPU_MEMORY_BUDGET_GB", "8.0".to_string()));
}
}

Binary file not shown.

View file

@ -1,5 +1,26 @@
diff --git a/sp1-gpu/crates/cuda/src/task.rs b/sp1-gpu/crates/cuda/src/task.rs
index a503a86..813016b 100644
--- a/sp1-gpu/crates/cuda/src/task.rs
+++ b/sp1-gpu/crates/cuda/src/task.rs
@@ -149,7 +149,15 @@ pub enum GlobalTaskPoolBuildError {
impl TaskPoolBuilder {
pub fn new() -> Self {
- Self { capacity: None, device: CudaDevice(0), mem_release_threshold: u64::MAX }
+ // Igneum prover-floor patch: upstream keeps every freed device allocation in the pool for the process's
+ // life (threshold u64::MAX), so the prover holds its high-water mark between shards on a card it shares
+ // with a miner. `SP1_GPU_MEM_RELEASE_THRESHOLD=<bytes>` sets the pool's release threshold (0 returns
+ // freed memory to the driver at once); unset, upstream's behaviour.
+ let mem_release_threshold = std::env::var("SP1_GPU_MEM_RELEASE_THRESHOLD")
+ .ok()
+ .and_then(|s| s.parse::<u64>().ok())
+ .unwrap_or(u64::MAX);
+ Self { capacity: None, device: CudaDevice(0), mem_release_threshold }
}
pub fn num_tasks(mut self, num_tasks: usize) -> Self {
diff --git a/sp1-gpu/crates/jagged_tracegen/src/lib.rs b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
index 579f70a..09e73e8 100644
index 579f70a..2264044 100644
--- a/sp1-gpu/crates/jagged_tracegen/src/lib.rs
+++ b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
@@ -481,6 +481,33 @@ async fn device_preprocessed_tracegen<A: CudaTracegenAir<Felt>>(
@ -64,7 +85,81 @@ index 579f70a..09e73e8 100644
log_stacking_height,
max_log_row_count,
backend,
@@ -984,9 +1021,15 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
@@ -906,6 +943,11 @@ pub async fn main_tracegen<GC: IopCtx<F = Felt>, A: CudaTracegenAir<Felt>>(
log_chip_stats(machine, &chip_set, &traces);
+ // Igneum prover-floor patch: the key's buffer is sized to its preprocessed traces at setup (upstream sized it
+ // for a whole shard), so grow it here to what this shard needs before the main traces are appended: a bigger
+ // dense buffer and column index, the preprocessed region copied device to device, swapped into the key.
+ grow_for_main(&mut jagged_traces.preprocessed_traces, &traces, log_stacking_height, backend);
+
copy_main_jagged_traces(
traces,
&mut jagged_traces.preprocessed_traces,
@@ -918,6 +960,61 @@ pub async fn main_tracegen<GC: IopCtx<F = Felt>, A: CudaTracegenAir<Felt>>(
(public_values, chip_set, permit)
}
+/// Igneum prover-floor patch: see `main_tracegen`. The need is the preprocessed phase as laid out (its padded
+/// end, `preprocessed_offset`) plus the main traces padded to the stacking height plus one stacking height of
+/// slack; a buffer at least that big is left alone. The process aborts, loudly, if the copy cannot be made,
+/// because a panic inside a prover task is what left sweep 2 hanging on the client's socket.
+fn grow_for_main(
+ jagged: &mut JaggedTraceMle<Felt, TaskScope>,
+ main_traces: &BTreeMap<String, Trace<TaskScope>>,
+ log_stacking_height: u32,
+ backend: &TaskScope,
+) {
+ let pre_end = jagged.dense().preprocessed_offset;
+ let needed = pre_end
+ + padded_trace_elements(main_traces, log_stacking_height)
+ + (1 << log_stacking_height);
+ let have = jagged.dense().dense.capacity();
+ if have >= needed {
+ return;
+ }
+ let mut new_dense: Buffer<Felt, TaskScope> = Buffer::with_capacity_in(needed, backend.clone());
+ let mut new_col_index: Buffer<u32, TaskScope> =
+ Buffer::with_capacity_in(needed >> 1, backend.clone());
+ unsafe {
+ new_dense.assume_init();
+ new_col_index.assume_init();
+ }
+ {
+ let JaggedMle { dense_data, col_index, .. } = &mut **jagged;
+ let src_dense: &Slice<_, _> = &dense_data.dense[..pre_end];
+ let dst_dense: &mut Slice<_, _> = &mut new_dense[..pre_end];
+ let src_col: &Slice<_, _> = &col_index[..pre_end >> 1];
+ let dst_col: &mut Slice<_, _> = &mut new_col_index[..pre_end >> 1];
+ unsafe {
+ if dst_dense.copy_from_slice(src_dense, backend).is_err()
+ || dst_col.copy_from_slice(src_col, backend).is_err()
+ {
+ eprintln!("FLOOR grow FAILED: could not copy the preprocessed region ({pre_end} elements) into the grown buffer ({needed} elements); aborting instead of hanging");
+ std::process::abort();
+ }
+ }
+ }
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
+ eprintln!(
+ "FLOOR grow key buffer {have} -> {needed} elements (preprocessed {pre_end}, {} bytes)",
+ needed * 6
+ );
+ }
+ let JaggedMle { dense_data, col_index, .. } = &mut **jagged;
+ dense_data.dense = new_dense;
+ *col_index = new_col_index;
+ unsafe {
+ dense_data.dense.set_len(pre_end);
+ col_index.set_len(pre_end >> 1);
+ }
+}
+
#[allow(clippy::too_many_arguments)]
pub async fn main_tracegen_permit<GC: IopCtx<F = Felt>, A: CudaTracegenAir<Felt>>(
machine: &Machine<Felt, A>,
@@ -984,9 +1081,15 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
log_chip_stats(machine, &chip_set, &main_traces);
@ -81,7 +176,7 @@ index 579f70a..09e73e8 100644
log_stacking_height,
max_log_row_count,
backend,
@@ -1002,6 +1045,18 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
@@ -1002,6 +1105,18 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
)
.await;
@ -104,15 +199,16 @@ diff --git a/sp1-gpu/crates/prover_components/src/builder.rs b/sp1-gpu/crates/pr
index 5dccd9d..574d4fa 100644
--- a/sp1-gpu/crates/prover_components/src/builder.rs
+++ b/sp1-gpu/crates/prover_components/src/builder.rs
@@ -23,28 +23,75 @@ use crate::{
@@ -23,28 +23,124 @@ use crate::{
SP1CudaProverComponents,
};
+/// Igneum prover-floor patch (5 October 2026). Upstream sizes every device buffer for a 24 GB card or larger
+/// and panics below that, whatever the shard. Here the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks a
+/// tier, and `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly.
+/// The proof format, the verifier and the program ids do not change: the element threshold only decides where
+/// the executor splits shards, as upstream's own 24 GB tier already does.
+/// Igneum prover-floor patch (5 October 2026; the memory rules of V6-07, 8 October 2026). Upstream sizes every
+/// device buffer for a 24 GB card or larger and panics below that, whatever the shard. Here the card's FREE memory
+/// at start (or `SP1_GPU_MEMORY_BUDGET_GB`, the host's lease) picks a tier, read once for the process, and
+/// `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly. The proof
+/// format, the verifier and the program ids do not change: the element threshold only decides where the executor
+/// splits shards, as upstream's own 24 GB tier already does.
+fn env_usize(name: &str) -> Option<usize> {
+ std::env::var(name).ok().and_then(|s| s.parse::<usize>().ok())
+}
@ -121,8 +217,17 @@ index 5dccd9d..574d4fa 100644
+ std::env::var(name).ok().and_then(|s| s.parse::<f64>().ok())
+}
+
+/// The core element threshold for a memory budget in GB (upstream's own figure for the budget, +4, as it
+/// computed it: a 32 GB card is 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16).
+/// The recursion trace allocation under the 24 GB tier (V6-07): a recursion key or shard uses 90,177,536 elements
+/// (35.6 M preprocessed and 54.5 M main; the floor sweep of 5 October 2026), so 2^26 + 2^25 = 100,663,296 holds it
+/// with the stacking slack `floor_capacity` adds; the pinned host copies (four per prover) shrink with it. Upstream's
+/// 2^27 stays on the 24 GB tier and above. Two distinct values: `floor_tests::recursion_branches_differ`.
+pub const RECURSION_TRACE_ALLOCATION_SMALL: usize = (1 << 26) + (1 << 25);
+
+/// The core element threshold for a memory budget in GB (the budget is the FREE memory, +4, as upstream computed
+/// its tiers from the total: a 32 GB card alone reads 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16, an
+/// 8 GB card 12). Under the 18 tier the threshold is 2^26, the value every passing small-card row used (the
+/// alone-comp-26-v1 rows of 6 October 2026 on the 3060, 3080, 4060, 4060 Ti, 4070 and 5070; the 8 October 2026
+/// proof_alone rows on the 3060 at 7,525 MiB and the 4060 at 7,532 MiB); no small-card row passed at 2^27 here.
+pub fn element_threshold_for_budget(gpu_memory_gb: usize, full_size_shards: bool) -> u64 {
+ if gpu_memory_gb > 30 || (full_size_shards && gpu_memory_gb >= 24) {
+ ELEMENT_THRESHOLD
@ -131,31 +236,69 @@ index 5dccd9d..574d4fa 100644
+ } else if gpu_memory_gb >= 18 {
+ (1 << 27) + (1 << 26)
+ } else {
+ 1 << 27
+ 1 << 26
+ }
+}
+
+/// The recursion trace allocation (elements) for a memory budget.
+/// The recursion trace allocation (elements) for a memory budget: upstream's on the 24 GB tier and above, the
+/// small constant under it.
+pub fn recursion_trace_allocation_for_budget(gpu_memory_gb: usize) -> usize {
+ if gpu_memory_gb >= 24 {
+ RECURSION_TRACE_ALLOCATION
+ } else {
+ RECURSION_TRACE_ALLOCATION
+ RECURSION_TRACE_ALLOCATION_SMALL
+ }
+}
+
+/// The memory budget in GB, read ONCE for the process (the core opts and the recursion prover are built at
+/// different moments, after allocations that lower the free figure; one reading keeps both on one tier):
+/// `SP1_GPU_MEMORY_BUDGET_GB` when set (the host's lease), else the device's FREE memory as the driver reports it
+/// at the first call, never the total (a card beside a miner has the miner's resident set gone), +4 as upstream
+/// computed its tiers.
+pub fn gpu_memory_gb() -> usize {
+ let gb = 1024.0 * 1024.0 * 1024.0;
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
+ Some(b) => (b.ceil() as usize) + 4,
+ None => (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4,
+ }
+ static BUDGET: std::sync::OnceLock<usize> = std::sync::OnceLock::new();
+ *BUDGET.get_or_init(|| {
+ let gb = 1024.0 * 1024.0 * 1024.0;
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
+ Some(b) => (b.ceil() as usize) + 4,
+ None => {
+ let (free, _total) = cuda_memory_info().unwrap();
+ (((free as f64) / gb).ceil() as usize) + 4
+ }
+ }
+ })
+}
+
+pub fn recursion_trace_allocation() -> usize {
+ env_usize("SP1_GPU_RECURSION_TRACE_ALLOCATION")
+ .unwrap_or_else(|| recursion_trace_allocation_for_budget(gpu_memory_gb()))
+}
+
+#[cfg(test)]
+mod floor_tests {
+ use super::*;
+
+ #[test]
+ fn small_tier_is_two_to_the_26() {
+ assert_eq!(element_threshold_for_budget(12, false), 1 << 26, "an 8 GB card reads 12");
+ assert_eq!(element_threshold_for_budget(16, false), 1 << 26, "a 12 GB card reads 16");
+ assert_eq!(element_threshold_for_budget(17, false), 1 << 26);
+ assert_eq!(element_threshold_for_budget(20, false), (1 << 27) + (1 << 26), "a 16 GB card reads 20");
+ assert_eq!(element_threshold_for_budget(28, false), ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24));
+ assert_eq!(element_threshold_for_budget(28, true), ELEMENT_THRESHOLD);
+ assert_eq!(element_threshold_for_budget(36, false), ELEMENT_THRESHOLD);
+ }
+
+ #[test]
+ fn recursion_branches_differ() {
+ assert_ne!(recursion_trace_allocation_for_budget(28), recursion_trace_allocation_for_budget(16));
+ assert_eq!(recursion_trace_allocation_for_budget(28), RECURSION_TRACE_ALLOCATION);
+ assert_eq!(recursion_trace_allocation_for_budget(16), RECURSION_TRACE_ALLOCATION_SMALL);
+ assert_eq!(RECURSION_TRACE_ALLOCATION_SMALL, 100_663_296);
+ assert!(RECURSION_TRACE_ALLOCATION_SMALL < RECURSION_TRACE_ALLOCATION);
+ assert!(RECURSION_TRACE_ALLOCATION_SMALL >= 90_177_536 + (1 << 22));
+ }
+}
+
pub fn local_gpu_opts() -> SP1CoreOpts {
let mut opts = SP1CoreOpts::default();
@ -171,7 +314,7 @@ index 5dccd9d..574d4fa 100644
- if gpu_memory_gb < 24 {
- panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB");
- }
+ // The card's memory plus 4, as upstream computed it (a 32 GB card reads 36), or the budget given.
+ // The card's FREE memory plus 4, as upstream computed its tiers from the total, or the budget given.
+ let gpu_memory_gb = gpu_memory_gb();
- let shard_threshold = if !opts.full_size_shards && gpu_memory_gb <= 30 {
@ -183,17 +326,18 @@ index 5dccd9d..574d4fa 100644
+ None => element_threshold_for_budget(gpu_memory_gb, opts.full_size_shards),
};
+ let height_threshold = opts.sharding_threshold.height_threshold;
+ let (free_mib, total_mib) = cuda_memory_info().map(|(f, t)| (f >> 20, t >> 20)).unwrap_or((0, 0));
- tracing::debug!("Shard threshold: {shard_threshold}");
+ eprintln!(
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} free_mib={free_mib} total_mib={total_mib} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
+ recursion_trace_allocation(),
+ opts.full_size_shards
+ );
opts.sharding_threshold.element_threshold = shard_threshold;
opts.global_dependencies_opt = true;
@@ -92,7 +139,7 @@ pub async fn recursion_prover_and_verifier(
@@ -92,7 +188,7 @@ pub async fn recursion_prover_and_verifier(
) {
let recursion_verifier = SP1CudaProverComponents::compress_verifier();
(
@ -202,6 +346,86 @@ index 5dccd9d..574d4fa 100644
.await,
recursion_verifier,
)
diff --git a/sp1-gpu/crates/server/src/main.rs b/sp1-gpu/crates/server/src/main.rs
index 65e94f5..498357c 100644
--- a/sp1-gpu/crates/server/src/main.rs
+++ b/sp1-gpu/crates/server/src/main.rs
@@ -17,9 +17,55 @@ struct Args {
version: bool,
}
+/// Igneum prover-floor patch (6 October 2026, the GPU fleet's finding): a panic inside a prover task (an allocation
+/// the card cannot meet, `cudaMallocAsync` failing and `Buffer::with_capacity_in` panicking in a tokio worker) left
+/// the request's future waiting for ever, the client on its socket, the card at 0% for the 15 minutes until someone
+/// killed it (the 8 GB and 10 GB cards at threshold 2^27). The server must fail the shard instead: this hook names
+/// the stage from the panic's location and exits, so the client's proof fails at once with the server's last line.
+fn stage_of(location: &str) -> &'static str {
+ let l = location.to_ascii_lowercase();
+ if l.contains("jagged_tracegen") || l.contains("/tracegen") {
+ "trace generation"
+ } else if l.contains("commit") || l.contains("basefold") || l.contains("merkle") {
+ "the commit (codewords and Merkle trees)"
+ } else if l.contains("logup_gkr") {
+ "LogUp GKR"
+ } else if l.contains("zerocheck") {
+ "the zerocheck"
+ } else if l.contains("jagged") {
+ "the jagged sumcheck"
+ } else if l.contains("prover_components") || l.contains("recursion") || l.contains("sp1-prover") || l.contains("sp1_prover") {
+ "the recursion (compression)"
+ } else if l.contains("cuda") || l.contains("slop") {
+ "a device allocation"
+ } else {
+ "the prover"
+ }
+}
+
+fn install_abort_on_panic() {
+ std::panic::set_hook(Box::new(|info| {
+ let location = info.location().map(|l| format!("{}:{}", l.file(), l.line())).unwrap_or_else(|| "unknown".into());
+ let message = info
+ .payload()
+ .downcast_ref::<&str>()
+ .map(|s| s.to_string())
+ .or_else(|| info.payload().downcast_ref::<String>().cloned())
+ .unwrap_or_default();
+ let oom = message.to_ascii_lowercase().contains("alloc") || message.contains("MemoryAllocation") || message.contains("OUT_OF_MEMORY");
+ eprintln!(
+ "FLOOR abort: {} failed at {location}: {message}{}; the server exits so the client's proof fails instead of waiting",
+ stage_of(&location),
+ if oom { " (the card's memory could not meet an allocation: lower SP1_GPU_ELEMENT_THRESHOLD one notch)" } else { "" }
+ );
+ std::process::exit(70);
+ }));
+}
+
#[tokio::main]
#[allow(clippy::print_stdout)]
async fn main() {
+ install_abort_on_panic();
tracing_subscriber::fmt::init();
let args = Args::parse();
@@ -40,3 +86,19 @@ async fn main() {
eprintln!("Error running server: {e}");
}
}
+
+#[cfg(test)]
+mod floor_tests {
+ use super::stage_of;
+
+ #[test]
+ fn the_stage_is_named_from_the_panic_location() {
+ assert_eq!(stage_of("sp1-gpu/crates/jagged_tracegen/src/lib.rs:240"), "trace generation");
+ assert_eq!(stage_of("sp1-gpu/crates/basefold/src/fri.rs:97"), "the commit (codewords and Merkle trees)");
+ assert_eq!(stage_of("sp1-gpu/crates/logup_gkr/src/tracegen.rs:72"), "LogUp GKR");
+ assert_eq!(stage_of("sp1-gpu/crates/zerocheck/src/prover.rs:1163"), "the zerocheck");
+ assert_eq!(stage_of("sp1-gpu/crates/prover_components/src/builder.rs:70"), "the recursion (compression)");
+ assert_eq!(stage_of("sp1-gpu/crates/cuda/src/stream.rs:330"), "a device allocation");
+ assert_eq!(stage_of("somewhere/else.rs:1"), "the prover");
+ }
+}
diff --git a/sp1-gpu/crates/server/src/server.rs b/sp1-gpu/crates/server/src/server.rs
index 4035f1f..0d0d907 100644
--- a/sp1-gpu/crates/server/src/server.rs

View file

@ -0,0 +1,17 @@
{
"run_id": "enforced-proving-20261008-03",
"manifest_sha": "c591e63b",
"evidence_dir": "docs/plans/proving-enforcement",
"boxes": [
"build-9"
],
"method": "native",
"cells": [
{
"cell": "harness:proving-enforcement",
"status": "PASS",
"evidence": "docs/plans/proving-enforcement/cache-history-2.json"
}
],
"note": "The two-node cache-history case (review B F01) on the fix's rule at the 2.0.2 tip ab489403 (cache-history-node c591e63b; host and node on pin C): a real proof of the chain's own block refused for context under one carrier, then paid once as its own record under a later one from the cached facts with no second SP1 verify; the invalid bytes verified once and refused again from the cache at a later carrier; a context refusal is a cheap pre-check, never cached. build-9, 22:53 to 23:05 UK."
}

View file

@ -0,0 +1,32 @@
{
"run_id": "v607-floor-memory-20261008T2059Z",
"manifest_sha": "c411bae9a",
"evidence_dir": "docs/analysis/floor-memory-profile-2026-10-08.md",
"boxes": [
"vast:54902103",
"vast:54906194",
"vast:54909866",
"vast:54909865",
"vast:54913403",
"build-3"
],
"cells": [
{
"cell": "rows:v607-floor-memory",
"status": "RUNNING",
"evidence": "docs/analysis/floor-memory-profile-2026-10-08.md",
"note": "partial coverage of GPU-05, CAP-02 and UX-02 by the V6-07 rows (the 3060 and 4060 on the default unoverridden path, the refusal rows beside the miner); RUNNING by the map's rule since no case's coverage is full; the 4060's chain row is a FAIL at the aggregation (an 8 GB card is refused the chain before setup) and the master-host proof rows are NOT RUN (6dbd5d2f8 without the guest re-pin); island: none"
}
],
"method": "GPU",
"release_identity": {
"commit": "c411bae9a (branch v607-floor-memory off cef5234b5)",
"lockfile": "proving/igneum-prove/Cargo.lock at cef5234b5",
"binary": "sp1-gpu-server db37c38b5feabdec5bbfd0298446da89f6a5d144125e399d6ab617f64f9cb61d (igneum-floor-sm8689-v607.tgz 9c02fcc6a3f045d3167c1d9cb6e22c584ea862189d11997754a57cbe1b0d355b); igneum-prove-host-0317 71bc2438856bb141; igneum-prove-host v607 66662219a3630772",
"network_object": "none (fixture proofs on rented pods, no network)",
"activation": "none",
"profile_hashes": "floor patch 9098c3e5979d057031f43588668201d1f7a53ab17e980c1eaf896cd5e6f38575"
},
"claim_impact": "the 2.0 proving statement's memory condition per tier: a 12 GB card runs the whole segment path alone; an 8 GB card proves shards only (its aggregation fails) and is refused the chain before setup; both time-share beside the 5.5 GiB miner; the 16 GB and 24 GB tiers unmeasured tonight",
"reviewer": ""
}

View file

@ -712,6 +712,25 @@
"coverage": {
"R2-F03-R02": "partial: the worker half (CPU reference, CUDA, Metal, OpenCL) on the same program packs; the node's and the pool's accepted work on the same job context are the CI steward's and the pool lane's cells; PASS only when every listed platform reads one fingerprint and the node and pool halves are green"
}
},
"rows:v607-floor-memory": {
"command": "the V6-07 lane's pod scripts (scratch v607/pod-v607.sh, ab.sh, ab2.sh, ab3.sh) on rented Vast RTX 3060 12 GB and RTX 4060 8 GB one-shots: the served 0317 host and the V6-07 host on the V6-07 floor server (igneum-floor-sm8689-v607.tgz), the fixture fees-v1-shards2.json, the ds55 kit's worker as the miner, nvidia-smi at 1 Hz; recorded in docs/analysis/floor-memory-profile-2026-10-08.md",
"box_class": "rented pods (Vast one-shots)",
"fixtures": [
"F0",
"F1"
],
"cases": [
"GPU-05",
"CAP-02",
"UX-02"
],
"coverage": {
"GPU-05": "partial: the 8 GB and 12 GB tiers' prover headroom on the 5.5 GiB dataset (free memory, the prover's peak alone, the refusal beside the miner); fragmentation, restart and next-epoch construction not run",
"CAP-02": "partial: proving alone (shard and the whole segment path) on the 8 GB and 12 GB tiers on the default job path, the time-share refusal beside the miner with the miner unharmed, memory headroom and proof latency; wall energy only on the 3060 (the 4060 host has no power sensor); induced GPU task failure and wallet control not run; 16 GB and 24 GB tiers not run",
"UX-02": "partial: the prover's safe refusal under memory pressure (exit 78 in one line, the miner's worker unharmed, nothing leaked after the refusal); pause, stop, power limit and the UI's crash ownership are the suite:app cell's"
},
"island": "none: every row proved a fixture on a rented pod; no row read a devnet"
}
},
"not_run": {

View file

@ -11,7 +11,8 @@ proving-v1, 272b025) and tools/proving-v1/pc2-segments.ps1 around the four binar
offered again every pass until the segment's deadline (the 272b025 behaviour).
3. one export (igneum_exportSegments 0..last), one fixture per block (igneum-prove-export), one host run
(igneum-prove-host --mode chain --chain ... --save-shards [--prev]) on the patched server (HOME=/opt/igneum-floor/home,
SP1_GPU_ELEMENT_THRESHOLD from THRESHOLD), the miner paused for the run when MINER=pause (prove-alone cards).
IGNEUM_PROVE_WORKLOAD=chain and no threshold by default, the host's memory profile from the card's free memory;
SP1_GPU_ELEMENT_THRESHOLD from THRESHOLD as a hand override), the miner paused for the run when MINER=pause (prove-alone cards).
4. every shard record signed (igneum-miner sign-record) and submitted (igneum_submitProofRecord); the segment record
(sign-segment-record, igneum_submitSegmentRecord) once every shard is accepted and the statement equals the node's.
5. the paid state of every submitted segment polled each pass (igneum_getSegmentRecords); a state file for the
@ -557,7 +558,10 @@ while (time.time() - t_run0) / 3600 < RUN_HOURS:
# chain
if MINER == "pause": miner_stop()
kill_server()
env = dict(os.environ, HOME=f"{FLOOR}/home", SP1_PROVER="cuda", RUST_LOG="off")
# the default job path (V6-07, 8 October 2026): no threshold override; the host picks its memory profile for the chain
# workload from the card's FREE memory (host/src/memory_profile.rs) and hands it to the floor server; THRESHOLD, when
# set, is a hand override the host keeps and names in its RESULT memory_profile line
env = dict(os.environ, HOME=f"{FLOOR}/home", SP1_PROVER="cuda", RUST_LOG="off", IGNEUM_PROVE_WORKLOAD="chain", IGNEUM_PROVE_DEVICE=str(DEV))
if THRESHOLD: env["SP1_GPU_ELEMENT_THRESHOLD"] = THRESHOLD
args = [HOST, "--mode", "chain", "--chain", ",".join(fixtures), "--prover", WALLET, "--save-shards", "--out", f"{d}/chain-results.json"]
if prev_file: args += ["--prev", prev_file]
@ -571,6 +575,10 @@ while (time.time() - t_run0) / 3600 < RUN_HOURS:
except Exception: peak = 0
seg["peak_mib"] = peak
open(f"{d}/chain.log", "w").write((rr.stdout if rr else "") + "\n" + (rr.stderr if rr else "TIMEOUT"))
if rr and rr.returncode == 78:
# the host refused the card's free memory for the chain workload before any setup (exit 78, one line)
line = [l for l in rr.stdout.split("\n") if "memory_profile refused" in l]
say(f"RESULT seg {first} chain REFUSED {stamp()} {line[-1].strip() if line else 'memory profile refused'}"); seg["failed"] = "memory"; close(first, "cancelled", "memory"); time.sleep(60); continue
if not rr or rr.returncode != 0 or not os.path.exists(f"{d}/chain-results.json"):
say(f"RESULT seg {first} chain FAILED {stamp()} rc={rr.returncode if rr else 'timeout'} wall={seg['chain_s']} s: {((rr.stderr if rr else '') or '')[-200:].strip()}"); seg["failed"] = "chain"; close(first, "cancelled", "chain" if rr else "timeout"); continue
res = json.load(open(f"{d}/chain-results.json"))

View file

@ -22,7 +22,7 @@ fail() { echo "RESULT setup_failed $1 $(stamp)"; exit 2; }
ARCHS="${ARCHS:-86,89,120}"; LABEL="${LABEL:-box}"; WALLET="${WALLET:-}"
PKG_URL=https://dl.igneum.network/dl/public/igneum-hive-0.3.12.tar.gz
PKG_SHA=7972af92e7cd9a032303eca4d95b533f53e0e68d1b9cae5bfe406a5b7c30a454
PATCH_SHA=e81cb0d03b291f9fd4bf0a109d6da2d7c897795c9ffd7f797c0ddce723eee2b1
PATCH_SHA=9098c3e5979d057031f43588668201d1f7a53ab17e980c1eaf896cd5e6f38575
SEED=188.245.5.161:26611
echo "RESULT start $(stamp) label=$LABEL archs=$ARCHS host=$(hostname) nproc=$(nproc) ram_gb=$(( $(awk '/MemTotal/{print $2}' /proc/meminfo) / 1048576 )) disk_avail=$(df -BG /root | awk 'NR==2{print $4}')"
echo "RESULT gpu $(nvidia-smi --query-gpu=name,memory.total,driver_version,pci.bus_id,power.limit,power.min_limit,power.max_limit,clocks.max.sm --format=csv,noheader 2>&1 | tr '\n' ';')"

View file

@ -199,15 +199,16 @@ diff --git a/sp1-gpu/crates/prover_components/src/builder.rs b/sp1-gpu/crates/pr
index 5dccd9d..574d4fa 100644
--- a/sp1-gpu/crates/prover_components/src/builder.rs
+++ b/sp1-gpu/crates/prover_components/src/builder.rs
@@ -23,28 +23,75 @@ use crate::{
@@ -23,28 +23,124 @@ use crate::{
SP1CudaProverComponents,
};
+/// Igneum prover-floor patch (5 October 2026). Upstream sizes every device buffer for a 24 GB card or larger
+/// and panics below that, whatever the shard. Here the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks a
+/// tier, and `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly.
+/// The proof format, the verifier and the program ids do not change: the element threshold only decides where
+/// the executor splits shards, as upstream's own 24 GB tier already does.
+/// Igneum prover-floor patch (5 October 2026; the memory rules of V6-07, 8 October 2026). Upstream sizes every
+/// device buffer for a 24 GB card or larger and panics below that, whatever the shard. Here the card's FREE memory
+/// at start (or `SP1_GPU_MEMORY_BUDGET_GB`, the host's lease) picks a tier, read once for the process, and
+/// `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly. The proof
+/// format, the verifier and the program ids do not change: the element threshold only decides where the executor
+/// splits shards, as upstream's own 24 GB tier already does.
+fn env_usize(name: &str) -> Option<usize> {
+ std::env::var(name).ok().and_then(|s| s.parse::<usize>().ok())
+}
@ -216,8 +217,17 @@ index 5dccd9d..574d4fa 100644
+ std::env::var(name).ok().and_then(|s| s.parse::<f64>().ok())
+}
+
+/// The core element threshold for a memory budget in GB (upstream's own figure for the budget, +4, as it
+/// computed it: a 32 GB card is 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16).
+/// The recursion trace allocation under the 24 GB tier (V6-07): a recursion key or shard uses 90,177,536 elements
+/// (35.6 M preprocessed and 54.5 M main; the floor sweep of 5 October 2026), so 2^26 + 2^25 = 100,663,296 holds it
+/// with the stacking slack `floor_capacity` adds; the pinned host copies (four per prover) shrink with it. Upstream's
+/// 2^27 stays on the 24 GB tier and above. Two distinct values: `floor_tests::recursion_branches_differ`.
+pub const RECURSION_TRACE_ALLOCATION_SMALL: usize = (1 << 26) + (1 << 25);
+
+/// The core element threshold for a memory budget in GB (the budget is the FREE memory, +4, as upstream computed
+/// its tiers from the total: a 32 GB card alone reads 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16, an
+/// 8 GB card 12). Under the 18 tier the threshold is 2^26, the value every passing small-card row used (the
+/// alone-comp-26-v1 rows of 6 October 2026 on the 3060, 3080, 4060, 4060 Ti, 4070 and 5070; the 8 October 2026
+/// proof_alone rows on the 3060 at 7,525 MiB and the 4060 at 7,532 MiB); no small-card row passed at 2^27 here.
+pub fn element_threshold_for_budget(gpu_memory_gb: usize, full_size_shards: bool) -> u64 {
+ if gpu_memory_gb > 30 || (full_size_shards && gpu_memory_gb >= 24) {
+ ELEMENT_THRESHOLD
@ -226,31 +236,69 @@ index 5dccd9d..574d4fa 100644
+ } else if gpu_memory_gb >= 18 {
+ (1 << 27) + (1 << 26)
+ } else {
+ 1 << 27
+ 1 << 26
+ }
+}
+
+/// The recursion trace allocation (elements) for a memory budget.
+/// The recursion trace allocation (elements) for a memory budget: upstream's on the 24 GB tier and above, the
+/// small constant under it.
+pub fn recursion_trace_allocation_for_budget(gpu_memory_gb: usize) -> usize {
+ if gpu_memory_gb >= 24 {
+ RECURSION_TRACE_ALLOCATION
+ } else {
+ RECURSION_TRACE_ALLOCATION
+ RECURSION_TRACE_ALLOCATION_SMALL
+ }
+}
+
+/// The memory budget in GB, read ONCE for the process (the core opts and the recursion prover are built at
+/// different moments, after allocations that lower the free figure; one reading keeps both on one tier):
+/// `SP1_GPU_MEMORY_BUDGET_GB` when set (the host's lease), else the device's FREE memory as the driver reports it
+/// at the first call, never the total (a card beside a miner has the miner's resident set gone), +4 as upstream
+/// computed its tiers.
+pub fn gpu_memory_gb() -> usize {
+ let gb = 1024.0 * 1024.0 * 1024.0;
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
+ Some(b) => (b.ceil() as usize) + 4,
+ None => (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4,
+ }
+ static BUDGET: std::sync::OnceLock<usize> = std::sync::OnceLock::new();
+ *BUDGET.get_or_init(|| {
+ let gb = 1024.0 * 1024.0 * 1024.0;
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
+ Some(b) => (b.ceil() as usize) + 4,
+ None => {
+ let (free, _total) = cuda_memory_info().unwrap();
+ (((free as f64) / gb).ceil() as usize) + 4
+ }
+ }
+ })
+}
+
+pub fn recursion_trace_allocation() -> usize {
+ env_usize("SP1_GPU_RECURSION_TRACE_ALLOCATION")
+ .unwrap_or_else(|| recursion_trace_allocation_for_budget(gpu_memory_gb()))
+}
+
+#[cfg(test)]
+mod floor_tests {
+ use super::*;
+
+ #[test]
+ fn small_tier_is_two_to_the_26() {
+ assert_eq!(element_threshold_for_budget(12, false), 1 << 26, "an 8 GB card reads 12");
+ assert_eq!(element_threshold_for_budget(16, false), 1 << 26, "a 12 GB card reads 16");
+ assert_eq!(element_threshold_for_budget(17, false), 1 << 26);
+ assert_eq!(element_threshold_for_budget(20, false), (1 << 27) + (1 << 26), "a 16 GB card reads 20");
+ assert_eq!(element_threshold_for_budget(28, false), ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24));
+ assert_eq!(element_threshold_for_budget(28, true), ELEMENT_THRESHOLD);
+ assert_eq!(element_threshold_for_budget(36, false), ELEMENT_THRESHOLD);
+ }
+
+ #[test]
+ fn recursion_branches_differ() {
+ assert_ne!(recursion_trace_allocation_for_budget(28), recursion_trace_allocation_for_budget(16));
+ assert_eq!(recursion_trace_allocation_for_budget(28), RECURSION_TRACE_ALLOCATION);
+ assert_eq!(recursion_trace_allocation_for_budget(16), RECURSION_TRACE_ALLOCATION_SMALL);
+ assert_eq!(RECURSION_TRACE_ALLOCATION_SMALL, 100_663_296);
+ assert!(RECURSION_TRACE_ALLOCATION_SMALL < RECURSION_TRACE_ALLOCATION);
+ assert!(RECURSION_TRACE_ALLOCATION_SMALL >= 90_177_536 + (1 << 22));
+ }
+}
+
pub fn local_gpu_opts() -> SP1CoreOpts {
let mut opts = SP1CoreOpts::default();
@ -266,7 +314,7 @@ index 5dccd9d..574d4fa 100644
- if gpu_memory_gb < 24 {
- panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB");
- }
+ // The card's memory plus 4, as upstream computed it (a 32 GB card reads 36), or the budget given.
+ // The card's FREE memory plus 4, as upstream computed its tiers from the total, or the budget given.
+ let gpu_memory_gb = gpu_memory_gb();
- let shard_threshold = if !opts.full_size_shards && gpu_memory_gb <= 30 {
@ -278,17 +326,18 @@ index 5dccd9d..574d4fa 100644
+ None => element_threshold_for_budget(gpu_memory_gb, opts.full_size_shards),
};
+ let height_threshold = opts.sharding_threshold.height_threshold;
+ let (free_mib, total_mib) = cuda_memory_info().map(|(f, t)| (f >> 20, t >> 20)).unwrap_or((0, 0));
- tracing::debug!("Shard threshold: {shard_threshold}");
+ eprintln!(
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} free_mib={free_mib} total_mib={total_mib} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
+ recursion_trace_allocation(),
+ opts.full_size_shards
+ );
opts.sharding_threshold.element_threshold = shard_threshold;
opts.global_dependencies_opt = true;
@@ -92,7 +139,7 @@ pub async fn recursion_prover_and_verifier(
@@ -92,7 +188,7 @@ pub async fn recursion_prover_and_verifier(
) {
let recursion_verifier = SP1CudaProverComponents::compress_verifier();
(

View file

@ -199,15 +199,16 @@ diff --git a/sp1-gpu/crates/prover_components/src/builder.rs b/sp1-gpu/crates/pr
index 5dccd9d..574d4fa 100644
--- a/sp1-gpu/crates/prover_components/src/builder.rs
+++ b/sp1-gpu/crates/prover_components/src/builder.rs
@@ -23,28 +23,75 @@ use crate::{
@@ -23,28 +23,124 @@ use crate::{
SP1CudaProverComponents,
};
+/// Igneum prover-floor patch (5 October 2026). Upstream sizes every device buffer for a 24 GB card or larger
+/// and panics below that, whatever the shard. Here the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks a
+/// tier, and `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly.
+/// The proof format, the verifier and the program ids do not change: the element threshold only decides where
+/// the executor splits shards, as upstream's own 24 GB tier already does.
+/// Igneum prover-floor patch (5 October 2026; the memory rules of V6-07, 8 October 2026). Upstream sizes every
+/// device buffer for a 24 GB card or larger and panics below that, whatever the shard. Here the card's FREE memory
+/// at start (or `SP1_GPU_MEMORY_BUDGET_GB`, the host's lease) picks a tier, read once for the process, and
+/// `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly. The proof
+/// format, the verifier and the program ids do not change: the element threshold only decides where the executor
+/// splits shards, as upstream's own 24 GB tier already does.
+fn env_usize(name: &str) -> Option<usize> {
+ std::env::var(name).ok().and_then(|s| s.parse::<usize>().ok())
+}
@ -216,8 +217,17 @@ index 5dccd9d..574d4fa 100644
+ std::env::var(name).ok().and_then(|s| s.parse::<f64>().ok())
+}
+
+/// The core element threshold for a memory budget in GB (upstream's own figure for the budget, +4, as it
+/// computed it: a 32 GB card is 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16).
+/// The recursion trace allocation under the 24 GB tier (V6-07): a recursion key or shard uses 90,177,536 elements
+/// (35.6 M preprocessed and 54.5 M main; the floor sweep of 5 October 2026), so 2^26 + 2^25 = 100,663,296 holds it
+/// with the stacking slack `floor_capacity` adds; the pinned host copies (four per prover) shrink with it. Upstream's
+/// 2^27 stays on the 24 GB tier and above. Two distinct values: `floor_tests::recursion_branches_differ`.
+pub const RECURSION_TRACE_ALLOCATION_SMALL: usize = (1 << 26) + (1 << 25);
+
+/// The core element threshold for a memory budget in GB (the budget is the FREE memory, +4, as upstream computed
+/// its tiers from the total: a 32 GB card alone reads 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16, an
+/// 8 GB card 12). Under the 18 tier the threshold is 2^26, the value every passing small-card row used (the
+/// alone-comp-26-v1 rows of 6 October 2026 on the 3060, 3080, 4060, 4060 Ti, 4070 and 5070; the 8 October 2026
+/// proof_alone rows on the 3060 at 7,525 MiB and the 4060 at 7,532 MiB); no small-card row passed at 2^27 here.
+pub fn element_threshold_for_budget(gpu_memory_gb: usize, full_size_shards: bool) -> u64 {
+ if gpu_memory_gb > 30 || (full_size_shards && gpu_memory_gb >= 24) {
+ ELEMENT_THRESHOLD
@ -226,31 +236,69 @@ index 5dccd9d..574d4fa 100644
+ } else if gpu_memory_gb >= 18 {
+ (1 << 27) + (1 << 26)
+ } else {
+ 1 << 27
+ 1 << 26
+ }
+}
+
+/// The recursion trace allocation (elements) for a memory budget.
+/// The recursion trace allocation (elements) for a memory budget: upstream's on the 24 GB tier and above, the
+/// small constant under it.
+pub fn recursion_trace_allocation_for_budget(gpu_memory_gb: usize) -> usize {
+ if gpu_memory_gb >= 24 {
+ RECURSION_TRACE_ALLOCATION
+ } else {
+ RECURSION_TRACE_ALLOCATION
+ RECURSION_TRACE_ALLOCATION_SMALL
+ }
+}
+
+/// The memory budget in GB, read ONCE for the process (the core opts and the recursion prover are built at
+/// different moments, after allocations that lower the free figure; one reading keeps both on one tier):
+/// `SP1_GPU_MEMORY_BUDGET_GB` when set (the host's lease), else the device's FREE memory as the driver reports it
+/// at the first call, never the total (a card beside a miner has the miner's resident set gone), +4 as upstream
+/// computed its tiers.
+pub fn gpu_memory_gb() -> usize {
+ let gb = 1024.0 * 1024.0 * 1024.0;
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
+ Some(b) => (b.ceil() as usize) + 4,
+ None => (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4,
+ }
+ static BUDGET: std::sync::OnceLock<usize> = std::sync::OnceLock::new();
+ *BUDGET.get_or_init(|| {
+ let gb = 1024.0 * 1024.0 * 1024.0;
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
+ Some(b) => (b.ceil() as usize) + 4,
+ None => {
+ let (free, _total) = cuda_memory_info().unwrap();
+ (((free as f64) / gb).ceil() as usize) + 4
+ }
+ }
+ })
+}
+
+pub fn recursion_trace_allocation() -> usize {
+ env_usize("SP1_GPU_RECURSION_TRACE_ALLOCATION")
+ .unwrap_or_else(|| recursion_trace_allocation_for_budget(gpu_memory_gb()))
+}
+
+#[cfg(test)]
+mod floor_tests {
+ use super::*;
+
+ #[test]
+ fn small_tier_is_two_to_the_26() {
+ assert_eq!(element_threshold_for_budget(12, false), 1 << 26, "an 8 GB card reads 12");
+ assert_eq!(element_threshold_for_budget(16, false), 1 << 26, "a 12 GB card reads 16");
+ assert_eq!(element_threshold_for_budget(17, false), 1 << 26);
+ assert_eq!(element_threshold_for_budget(20, false), (1 << 27) + (1 << 26), "a 16 GB card reads 20");
+ assert_eq!(element_threshold_for_budget(28, false), ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24));
+ assert_eq!(element_threshold_for_budget(28, true), ELEMENT_THRESHOLD);
+ assert_eq!(element_threshold_for_budget(36, false), ELEMENT_THRESHOLD);
+ }
+
+ #[test]
+ fn recursion_branches_differ() {
+ assert_ne!(recursion_trace_allocation_for_budget(28), recursion_trace_allocation_for_budget(16));
+ assert_eq!(recursion_trace_allocation_for_budget(28), RECURSION_TRACE_ALLOCATION);
+ assert_eq!(recursion_trace_allocation_for_budget(16), RECURSION_TRACE_ALLOCATION_SMALL);
+ assert_eq!(RECURSION_TRACE_ALLOCATION_SMALL, 100_663_296);
+ assert!(RECURSION_TRACE_ALLOCATION_SMALL < RECURSION_TRACE_ALLOCATION);
+ assert!(RECURSION_TRACE_ALLOCATION_SMALL >= 90_177_536 + (1 << 22));
+ }
+}
+
pub fn local_gpu_opts() -> SP1CoreOpts {
let mut opts = SP1CoreOpts::default();
@ -266,7 +314,7 @@ index 5dccd9d..574d4fa 100644
- if gpu_memory_gb < 24 {
- panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB");
- }
+ // The card's memory plus 4, as upstream computed it (a 32 GB card reads 36), or the budget given.
+ // The card's FREE memory plus 4, as upstream computed its tiers from the total, or the budget given.
+ let gpu_memory_gb = gpu_memory_gb();
- let shard_threshold = if !opts.full_size_shards && gpu_memory_gb <= 30 {
@ -278,17 +326,18 @@ index 5dccd9d..574d4fa 100644
+ None => element_threshold_for_budget(gpu_memory_gb, opts.full_size_shards),
};
+ let height_threshold = opts.sharding_threshold.height_threshold;
+ let (free_mib, total_mib) = cuda_memory_info().map(|(f, t)| (f >> 20, t >> 20)).unwrap_or((0, 0));
- tracing::debug!("Shard threshold: {shard_threshold}");
+ eprintln!(
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} free_mib={free_mib} total_mib={total_mib} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
+ recursion_trace_allocation(),
+ opts.full_size_shards
+ );
opts.sharding_threshold.element_threshold = shard_threshold;
opts.global_dependencies_opt = true;
@@ -92,7 +139,7 @@ pub async fn recursion_prover_and_verifier(
@@ -92,7 +188,7 @@ pub async fn recursion_prover_and_verifier(
) {
let recursion_verifier = SP1CudaProverComponents::compress_verifier();
(
@ -297,6 +346,86 @@ index 5dccd9d..574d4fa 100644
.await,
recursion_verifier,
)
diff --git a/sp1-gpu/crates/server/src/main.rs b/sp1-gpu/crates/server/src/main.rs
index 65e94f5..498357c 100644
--- a/sp1-gpu/crates/server/src/main.rs
+++ b/sp1-gpu/crates/server/src/main.rs
@@ -17,9 +17,55 @@ struct Args {
version: bool,
}
+/// Igneum prover-floor patch (6 October 2026, the GPU fleet's finding): a panic inside a prover task (an allocation
+/// the card cannot meet, `cudaMallocAsync` failing and `Buffer::with_capacity_in` panicking in a tokio worker) left
+/// the request's future waiting for ever, the client on its socket, the card at 0% for the 15 minutes until someone
+/// killed it (the 8 GB and 10 GB cards at threshold 2^27). The server must fail the shard instead: this hook names
+/// the stage from the panic's location and exits, so the client's proof fails at once with the server's last line.
+fn stage_of(location: &str) -> &'static str {
+ let l = location.to_ascii_lowercase();
+ if l.contains("jagged_tracegen") || l.contains("/tracegen") {
+ "trace generation"
+ } else if l.contains("commit") || l.contains("basefold") || l.contains("merkle") {
+ "the commit (codewords and Merkle trees)"
+ } else if l.contains("logup_gkr") {
+ "LogUp GKR"
+ } else if l.contains("zerocheck") {
+ "the zerocheck"
+ } else if l.contains("jagged") {
+ "the jagged sumcheck"
+ } else if l.contains("prover_components") || l.contains("recursion") || l.contains("sp1-prover") || l.contains("sp1_prover") {
+ "the recursion (compression)"
+ } else if l.contains("cuda") || l.contains("slop") {
+ "a device allocation"
+ } else {
+ "the prover"
+ }
+}
+
+fn install_abort_on_panic() {
+ std::panic::set_hook(Box::new(|info| {
+ let location = info.location().map(|l| format!("{}:{}", l.file(), l.line())).unwrap_or_else(|| "unknown".into());
+ let message = info
+ .payload()
+ .downcast_ref::<&str>()
+ .map(|s| s.to_string())
+ .or_else(|| info.payload().downcast_ref::<String>().cloned())
+ .unwrap_or_default();
+ let oom = message.to_ascii_lowercase().contains("alloc") || message.contains("MemoryAllocation") || message.contains("OUT_OF_MEMORY");
+ eprintln!(
+ "FLOOR abort: {} failed at {location}: {message}{}; the server exits so the client's proof fails instead of waiting",
+ stage_of(&location),
+ if oom { " (the card's memory could not meet an allocation: lower SP1_GPU_ELEMENT_THRESHOLD one notch)" } else { "" }
+ );
+ std::process::exit(70);
+ }));
+}
+
#[tokio::main]
#[allow(clippy::print_stdout)]
async fn main() {
+ install_abort_on_panic();
tracing_subscriber::fmt::init();
let args = Args::parse();
@@ -40,3 +86,19 @@ async fn main() {
eprintln!("Error running server: {e}");
}
}
+
+#[cfg(test)]
+mod floor_tests {
+ use super::stage_of;
+
+ #[test]
+ fn the_stage_is_named_from_the_panic_location() {
+ assert_eq!(stage_of("sp1-gpu/crates/jagged_tracegen/src/lib.rs:240"), "trace generation");
+ assert_eq!(stage_of("sp1-gpu/crates/basefold/src/fri.rs:97"), "the commit (codewords and Merkle trees)");
+ assert_eq!(stage_of("sp1-gpu/crates/logup_gkr/src/tracegen.rs:72"), "LogUp GKR");
+ assert_eq!(stage_of("sp1-gpu/crates/zerocheck/src/prover.rs:1163"), "the zerocheck");
+ assert_eq!(stage_of("sp1-gpu/crates/prover_components/src/builder.rs:70"), "the recursion (compression)");
+ assert_eq!(stage_of("sp1-gpu/crates/cuda/src/stream.rs:330"), "a device allocation");
+ assert_eq!(stage_of("somewhere/else.rs:1"), "the prover");
+ }
+}
diff --git a/sp1-gpu/crates/server/src/server.rs b/sp1-gpu/crates/server/src/server.rs
index 4035f1f..0d0d907 100644
--- a/sp1-gpu/crates/server/src/server.rs