Merge master f0efd5117 into cache-history under the master-landing lock
This commit is contained in:
commit
459cdbc8ba
23 changed files with 2001 additions and 828 deletions
|
|
@ -8,3 +8,14 @@ The coordinator's blocking read: the research pack hl-v6-all (4de7b836cc40a4ea)
|
|||
| the research comparison | 442a1691b3e3507f (mx8-era93a14ac6+sh256x27+state+reg64c+fold+rw) | seed af89be5d..., era 0:af89be5d..., day 20730, node1 state | ACCEPTED; min site ratio 0.99992, under 0.98: 0, under 0.995: 0; largest bucket +5.25 sigma; worst free bit 2.91 sigma; sites over 6 sigma 0 | 256 of 256, 0 exhausted at 256, r 0.129, mean attempt 0.15, max 2, (c''') refused 0 of 294; parts a' 23 b 7 a 4 reg64 3 c 1 |
|
||||
|
||||
Verdict: the engine's draw 2a1d6caab4c24564 PASSES the sub-version 3 rule with the (c''') floor (every site clear of 0.98 and 0.995 at 0.99988, the largest 256-item bucket +6.25 sigma on one site, which the rule does not bound and the record's clean full-window sites read to +5.5, the worst free index bit 3.06 sigma, no site over 6 sigma) and the attempts census (256 of 256 chain-shaped seeds accepted inside two attempts, 0 exhausted, 0 (c''') refusals: the same 294 candidates and parts as the research pack's, since the census draws its own seeds under the pair's era and state). The research comparison reads the same on every line but the bucket (+5.25). The freeze's object on the chain's own seed is sound under the rule; the research pack's rows (census-packs.md section 7) are the same class on the genesis seed.
|
||||
|
||||
## The window trace on the signing program (22:53 to 22:56 UK, build-4)
|
||||
|
||||
The uniform tool (`tools/attack/v6-census/uniform`, with `--epoch-hex` added so one chain-drawn program is traced) on the engine's pair at 2^22 nonces, every load of the 32-site interleaved schedule through the interpreter's tracing probe on the closed form, beside the foldrw control (the same pair, no window; program 768497ac7a47599a). The tool prints the draw's id under the un-stamped generator byte (c3c7ce6d9df51c48 for the signing program); the seed bytes, class, era and attempt 0 are the CLI's, which stamps the same draw 2a1d6caab4c24564 at generator 6.
|
||||
|
||||
| Program | Per-site distinct ratio (worst site; floors 0.98 and 0.995) | Top 0.1 percent item share | Flat control | Ratio to flat | Against the foldrw control |
|
||||
|---|---|---|---|---|---|
|
||||
| the signing program (2a1d6caab4c24564; 1,073,741,824 reads) | 1.00262 (site 7) | 0.001683 | 0.001447 | 1.1631x | 1.021x |
|
||||
| the foldrw control (768497ac7a47599a; 536,870,912 reads) | 1.00266 (site 13) | 0.001875 | 0.001646 | 1.1387x | |
|
||||
|
||||
The window on the signing object reads as the class traces read (1.0026 per site, 1.015 to 1.032x of the control): uniform inside every window, the item-level share the era windows' with at most 2 percent on top.
|
||||
|
|
|
|||
|
|
@ -0,0 +1,2 @@
|
|||
class era seed attempt program_id min_site_ratio min_site top0.1_share control_share ratio_to_control max_item_reads control_max reads
|
||||
mx8+sh256x27+state+reg64c+fold+rw 0:edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07 0 0 c3c7ce6d9df51c48 1.00262 7 0.001683 0.001447 1.1631 132 111 1073741824
|
||||
|
|
|
@ -0,0 +1,2 @@
|
|||
class era seed attempt program_id min_site_ratio min_site top0.1_share control_share ratio_to_control max_item_reads control_max reads
|
||||
mx8+sh256x27+state+fold+rw 0:edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07 0 0 768497ac7a47599a 1.00266 13 0.001875 0.001646 1.1387 75 72 536870912
|
||||
|
125
docs/analysis/floor-memory-profile-2026-10-08.md
Normal file
125
docs/analysis/floor-memory-profile-2026-10-08.md
Normal file
|
|
@ -0,0 +1,125 @@
|
|||
# The floor's memory profile (V6-07), 8 October 2026
|
||||
|
||||
Master review R1, residual V6-07 (`docs/plans/igneum-2.0-master/evidence/04_full_system/IGNEUM_V6_Full_System_Review.md`, pages 210 to 212): the floor patch picked its limits from the card's total VRAM, not from what was free; its small-card element threshold defaulted to 2^27 where every passing row had set 2^26 by hand; its recursion-allocation budget returned one constant in both branches. The order: a pinned memory profile per workload read from free memory, with the app lane's device coordinator, and the 3060 and 4060 rows rerun on the default, unoverridden job path. This document is the record: the code (section 1), the table (section 2), the rows as the pods wrote them (section 3), the consequences per card tier beside each number (section 4), the coordinator hook and what is next (section 5).
|
||||
|
||||
Branch `v607-floor-memory` off the box mirror master cef5234b5. Code commits e6abe8c2c (the patch and the host profile), 573dad0ad (the lesser of grant and free), 8c1fb754c (one floor patch), then the floors re-pinned from the rows (the commit carrying this document). Clocks below are UTC as the pods wrote them; UK time is one hour later.
|
||||
|
||||
## 1. What changed
|
||||
|
||||
| Item | Where | Before | After |
|
||||
|---|---|---|---|
|
||||
| The limits' source | `proving/prover-floor/sp1-gpu-6.8.1-floor.patch`, `builder.rs` `gpu_memory_gb()` (the fleet's copies `tools/fleet/floor.patch`, `floor-v5.patch` carry the same hunk; `box-setup.sh` re-pinned) | `cuda_memory_info().1` (the total), +4 as upstream | `cuda_memory_info().0` (the FREE memory as the driver reports it), +4, read ONCE per process in a `OnceLock` so the core opts and the recursion prover, built at different moments, sit on one tier; `SP1_GPU_MEMORY_BUDGET_GB` (the host's lease) still overrides |
|
||||
| The small-card element threshold | same, `element_threshold_for_budget` | `1 << 27` under the 18 tier | `1 << 26`, the value the passing rows used (section 2 cites them) |
|
||||
| The recursion-allocation budget | same, `recursion_trace_allocation_for_budget` | `RECURSION_TRACE_ALLOCATION` in both branches | upstream's 2^27 on the 24 GB tier and above; `RECURSION_TRACE_ALLOCATION_SMALL` = 2^26 + 2^25 = 100,663,296 under it (a recursion key or shard uses 90,177,536 elements, `docs/analysis/prover-floor.md` sweep 1; the patch's `floor_capacity` sizes the device buffer to the need, so the constant caps the buffer and shrinks the four pinned host copies per prover); `floor_tests::recursion_branches_differ` pins that the two branches differ and that the small one clears the measured use with a stacking height of slack; `floor_tests::small_tier_is_two_to_the_26` pins the tiers |
|
||||
| The FLOOR opts line | same, `local_gpu_opts` | `gpu_memory_gb`, thresholds | adds `free_mib` and `total_mib` so a log names what the server read |
|
||||
| The pinned profile per workload | `proving/igneum-prove/host/src/memory_profile.rs` (new), wired in `main.rs` before the SP1 client is built | nothing: the server guessed from the total; the fleet set `SP1_GPU_ELEMENT_THRESHOLD` by hand | one table (section 2) with tests pinning every value; the row is chosen from the engine's lease or, with no engine, from `nvidia-smi memory.free` on the device, and handed to the server by environment (`SP1_GPU_ELEMENT_THRESHOLD`, `SP1_GPU_RECURSION_TRACE_ALLOCATION`, `SP1_GPU_MEMORY_BUDGET_GB`); a hand override already in the environment is kept and named; under the workload's floor the host prints one line and exits 78 before any setup |
|
||||
| The fleet's default path | `tools/fleet/box-prover.py` | `SP1_GPU_ELEMENT_THRESHOLD` from `THRESHOLD`, else the server's total-VRAM guess | `IGNEUM_PROVE_WORKLOAD=chain` and `IGNEUM_PROVE_DEVICE` on the host run, no threshold unless `THRESHOLD` is set by hand; a refusal (exit 78) closes the segment as cancelled/memory |
|
||||
|
||||
Tests: `cargo test -p igneum-prove-host memory_profile` (box 3, section 6) and the patch's `floor_tests` (run on the build pod after the server build, section 6).
|
||||
|
||||
## 2. The table
|
||||
|
||||
The host's `TIERS` (largest first; the free-memory lines are the card tiers as the floor patch reads them, free GiB rounded up plus 4 as upstream computed its tiers from the total):
|
||||
|
||||
| Tier | Free memory at start | Element threshold | Recursion trace allocation | Who lands here |
|
||||
|---|---|---|---|---|
|
||||
| full | 26,624 MiB and up | 2^28 + 2^27 = 402,653,184 | 2^27 = 134,217,728 | a 32 GB card alone |
|
||||
| 24gb | 20,480 MiB and up | 285,212,672 (upstream's 24 GB figure) | 2^27 | a 24 GB card alone |
|
||||
| 16gb | 14,336 MiB and up | 2^27 + 2^26 = 201,326,592 (the patch's own figure, unmeasured on this fixture) | 2^26 + 2^25 = 100,663,296 | a 16 GB card alone |
|
||||
| small | under 14,336 MiB | 2^26 = 67,108,864 | 100,663,296 | a 12 GB or 8 GB card alone; any card beside a miner's resident set |
|
||||
|
||||
Floors per workload (the least free memory the host runs in; under it the refusal, exit 78): shard 7,700 MiB, aggregate 8,500 MiB, chain 8,500 MiB. The shard floor is the small tier's measured peak (7,525 MiB on the 3060, 7,532 on the 4060, section 3) plus headroom for the driver's own context. The aggregate and chain floors are 8,500 MiB: the aggregation is the peak of a chain run, 8,306 MiB on the RTX 3060 (section 3, D2), and the RTX 4060 (7,807 MiB free) aborted at it, so an 8 GB card is refused the chain and the aggregation before any setup and proves shards only.
|
||||
|
||||
The rows cited for 2^26 on small cards: `docs/analysis/prover-tiers-real-cards.md` (6 October 2026, the matrix's `alone-comp-26-v1` point, compressed, verified: RTX 3060 7.4 GB 14.4 s; RTX 3080 8.0 GB 7.1 s; RTX 4060 7.4 GB 18.4 s; RTX 4060 Ti 8 GB 7.6 GB 9.6 s; RTX 4060 Ti 16 GB 7.8 GB 11.6 s; RTX 4070 7.6 GB 12.1 s; RTX 5070 7.6 GB 4.8 s) and `docs/analysis/class-v6/coexist-rows.md` (8 October 2026, `proof_alone` at `SP1_GPU_ELEMENT_THRESHOLD=67108864`: RTX 3060 peak 7,525 MiB 13.2 s verified; RTX 4060 peak 7,532 MiB 8.2 s verified). No small-card row ever passed at 2^27 on this host path; the 10 GB tier's 2^27 reading of 7 October (8,642 MiB alone on a 3080) ran on the segment host and is not this path.
|
||||
|
||||
Every value above is pinned by `memory_profile::tests::table_is_pinned`, `tiers_from_free_memory`, `refuses_under_the_floor`, `recursion_branches_differ`, `workload_from_mode` and `env_for_the_server`; a change of the table is a change of the profile and needs its rows.
|
||||
|
||||
## 3. The rows, rerun on the default, unoverridden job path
|
||||
|
||||
Every row is a RESULT line as the pod wrote it (the raw run logs, host logs and 1 Hz `nvidia-smi` samples are kept under the lane's scratch `v607/rows/<label>/`; the two bench scripts are `pod-v607.sh` for pass 1 and `ab2.sh` for pass 2). Pods: Vast.ai one-shots rented and destroyed by the lane (section 6). Fixture `fees-v1-shards2.json` (sha 20a108f159c61ff9, block 351, two shards of 22,172 and 14,762 pgas, 4,717,439 cycles on shard 0), the ds55 kit's worker (d43be4625b78baf7) on the 5.5 GiB dataset as the miner. "Default, unoverridden" = `HOME=/opt/igneum-floor/home SP1_PROVER=cuda RUST_LOG=off`, exactly `tools/fleet/box-prover.py`'s environment with no `THRESHOLD`, nothing else set by hand.
|
||||
|
||||
### 3.1 The default path as the fleet runs it today: the served 0317 host on the V6-07 server (pass 2, 20:53Z to 20:59Z)
|
||||
|
||||
The server is the V6-07 build (sm_86 + sm_89, sha db37c38b, free-memory tiers, small tier 2^26, two recursion budgets, from the one floor patch 9098c3e5); the host is the served `igneum-prove-host-0317` (71bc2438), which carries no profile table, so the row shows the SERVER's own free-memory rule deciding.
|
||||
|
||||
| Card | Free MiB at start | FLOOR opts (what the server read and chose) | Workload | Peak MiB | Seconds | Verdict |
|
||||
|---|---|---|---|---|---|---|
|
||||
| RTX 3060 12 GB (driver 595.91.07, 170 W) | 11,898 | `gpu_memory_gb=16 free_mib=11775 total_mib=11911 element_threshold=67108864 recursion_trace_allocation=100663296` | shard (compressed) | 7,601 | prove 11.4 s, verify 0.045 s | PASS, VERIFIED; 77.3 W mean |
|
||||
| RTX 3060 12 GB | 11,898 | the same | chain (2 shards + aggregation) | 8,307 (after setup 4,866; after shard 0 7,504; after shard 1 7,536; after the aggregation 8,306) | shard proofs 22.0 s (11.6 + 10.4), aggregation 5.2 s, end to end 27.4 s | PASS, every proof VERIFIED, chain_len 1; 94.0 W mean |
|
||||
| RTX 4060 8 GB (driver 595.91.07, 115 W, no power sensor) | 7,807 | `gpu_memory_gb=12 free_mib=7694 total_mib=7807 element_threshold=67108864 recursion_trace_allocation=100663296` | shard (compressed) | 7,504 | prove 16.3 s, verify 0.129 s | PASS, VERIFIED |
|
||||
| RTX 4060 8 GB | 7,807 | the same | chain (2 shards + aggregation) | 7,792 (after setup 4,999; after each shard 7,631 with 176 MiB free) | shard 0 11.7 s VERIFIED, shard 1 12.8 s VERIFIED, then the aggregation | FAIL at the aggregation: `FLOOR abort: a device allocation failed at slop/crates/tensor/src/inner.rs:51: called Result::unwrap() on an Err value: AllocError { layout: Layout { size: 509782528, align: 4 } }` (a 486 MiB allocation on 176 MiB free) |
|
||||
|
||||
Beside the miner on this path (the server's rule alone, no host-side refusal): NOT RUN. Pass 2's `ab2.sh` stopped the server between points by the pids of every listening unix socket instead of the `sp1-cuda` socket's owner, and that sweep killed the miner (rc 137) at the start of D7 and D8, so those two rows ran with the card empty (3060: 11.6 s, 7,697 MiB; 4060: 16.1 s, 7,632 MiB, both VERIFIED: alone rows, not beside rows). The kill-by-pattern class, in a scratch script; the lesson is recorded here and the script corrected (`ab3.sh`). The beside rows that stand are the host-side refusals of pass 1 (3.2) and the 8 October coexistence rows (`docs/analysis/class-v6/coexist-rows.md`: the server's allocation fails beside the 6.1 GiB miner on both cards at 2^26, the miner unharmed).
|
||||
|
||||
### 3.2 The V6-07 host's profile and refusal rows (pass 1, 20:25Z to 20:33Z, V6-07 server db37c38b, host 66662219)
|
||||
|
||||
| Card | Free MiB at start | Source of the reading | Row chosen (RESULT memory_profile) | Verdict |
|
||||
|---|---|---|---|---|
|
||||
| RTX 3060 12 GB | 11,909 | nvidia-smi device 0 (no engine) | workload shard, tier small, element_threshold 67108864, recursion_trace_allocation 100663296, `SP1_GPU_MEMORY_BUDGET_GB=11.6`; the server read it back as `gpu_memory_gb=16 ... element_threshold=67108864 ... recursion_trace_allocation=100663296` | the profile applied; the proof itself NOT RUN on this host (3.3) |
|
||||
| RTX 3060 12 GB | 11,909 | the engine's lease (`IGNEUM_PROVE_MEM_BUDGET_MB` = `IGNEUM_PROVE_MEM_FREE_MB` = 11909, `IGNEUM_PROVE_DEADLINE_S=600`) | the same row, "from 11909 MiB free (lease IGNEUM_PROVE_MEM_BUDGET_MB), deadline 600 s" | the same |
|
||||
| RTX 3060 12 GB | lease 6,100 | the engine's lease under the floor | `RESULT memory_profile refused: proving needs 8 GB free on the card for the shard workload (7700 MiB floor); 6100 MiB free`, exit 78 in 0 s, peak 1 MiB (no setup, the card untouched) | PASS (the refusal) |
|
||||
| RTX 3060 12 GB | 5,781 beside the ds55 miner (6,129 MiB resident, 26.83 MH/s) | nvidia-smi | `refused: ... (7700 MiB floor); 5781 MiB free`, exit 78 in 0 s; the miner alive after, 26.832 MH/s over its 253 s, self-test PASS, fingerprint 23ced07a4d28b465 | PASS (the refusal; the miner unharmed) |
|
||||
| RTX 4060 8 GB | 7,807 | nvidia-smi | workload shard, tier small, 2^26, 100663296, `SP1_GPU_MEMORY_BUDGET_GB=7.6`; the server read `gpu_memory_gb=12` | the profile applied; the proof NOT RUN on this host (3.3) |
|
||||
| RTX 4060 8 GB | 7,807 | the engine's lease (7807 / 7807, deadline 600) | the same row from the lease | the same |
|
||||
| RTX 4060 8 GB | lease 6,100 | the lease under the floor | `refused: ... (7700 MiB floor); 6100 MiB free`, exit 78, peak 2 MiB | PASS (the refusal) |
|
||||
| RTX 4060 8 GB | 1,689 beside the ds55 miner (6,120 MiB resident, 18.96 MH/s) | nvidia-smi | `refused: ... (7700 MiB floor); 1689 MiB free`, exit 78 in 0 s; the miner alive after, 18.958 MH/s over 357 s, self-test PASS | PASS (the refusal; the miner unharmed) |
|
||||
|
||||
### 3.3 The proof on master's host: NOT RUN, with its cause
|
||||
|
||||
Every proving attempt by a host built from master after the 0.3.17 build ends at the shard's execute with `Error: public values are 0 bytes, expected 328` (the V6-07 host, from cef5234b5) or `expected 392` (a plain host from master's tip dc94dda95, no V6-07 change), on both cards, with or without the threshold set by hand (pass 2 rows D4, D5, D6, D8: 14 to 31 s wall, peaks 4,579 to 4,963 MiB, the setup complete and "setup matches the manifest", then 0 bytes from the executor). The served 0317 host (71bc2438) executes and proves the same shard on the same card and server (3.1). The pinned guests are the 5 October pair (elf/manifest.json pinned_at 2026-10-05T16:20:38Z; `igneum-prove-program.elf` last changed at 15bb6cdd4) while the host's guest input changed on 8 October (6dbd5d2f8 at 17:45 the D4 proving-payment split, 2ceb09be9 at 18:34 the key-succession pairs, then 67e22b9ef and 566d0900b P22, then 7604fcf5d V6-10's format word, none with a re-pin). The bisect (section 6) names the commit: 6dbd5d2f8 (8 October 17:45, "Shard guest: the D4 proving-payment split mirrored behind FixtureEnv::proving_payment_to_pool (known-failed first; the ELF and program id not yet re-pinned)"): the host's fixture gained a field the pinned 5 October guest does not read, and the commit's own message says the re-pin is pending; a host from its parent 13b729467 proves (12.2 s VERIFIED), a host from 6dbd5d2f8 reads 0 bytes, and a host from the key-succession commit 2ceb09be9 (a side branch without 6dbd5d2f8) proves (12.1 s VERIFIED). Consequence: the host-side profile is attested tonight by its unit tests (7 of 7 on box 3) and the refusal rows (3.2), not by a proof; the proof rows on the default path are the 0317 host's (3.1) with the server's rule, which is the same table on the server side. The fleet's standing provers run the 0317 host and the served floors, so nothing on the fleet is affected until a host is rebuilt from master; a host rebuilt from master today cannot prove at all (a release gate before any 2.0.x host ships: pin-guests.sh and a CI check that a guest-input change without an elf/ change fails).
|
||||
|
||||
### 3.4 The recursion constant isolated (pass 1, 20:18Z to 20:20Z, RTX 3060, the served sm_86 server 224200d3 of 10:56 UTC, host 0317, threshold 2^26 set by hand)
|
||||
|
||||
| `SP1_GPU_RECURSION_TRACE_ALLOCATION` | Peak MiB | Seconds | Verdict |
|
||||
|---|---|---|---|
|
||||
| 100,663,296 (the small branch, 2^26 + 2^25) | 7,619 | prove 14.1 s, verify 0.108 s | PASS, VERIFIED |
|
||||
| 134,217,728 (upstream's 2^27, the 24 GB branch) | 7,619 | prove 14.1 s, verify 0.112 s | PASS, VERIFIED |
|
||||
| 117,440,512 (2^26 + 2^25 + 2^24, a control) | 7,619 | prove 14.4 s, verify 0.110 s | PASS, VERIFIED |
|
||||
|
||||
The device peak does not move with the constant: the patch's `floor_capacity` sizes every trace buffer to its padded need, so the constant is a cap, and the small branch's saving is the four pinned host copies per prover (4 x 100.6 M x 4 bytes = 1.5 GiB of pinned RAM against 2.0 GiB), which is what a 16 GB Windows PC under WSL2 feels. The two branches are two values by test; neither is the memory lever on the card.
|
||||
|
||||
### 3.5 Known-failed server build (pass 1, the first tarball)
|
||||
|
||||
The first V6-07 server (sha 75d0b4be, from the canonical patch copy `proving/prover-floor/sp1-gpu-6.8.1-floor.patch` as it stood) proved shard 0 on the RTX 4060 and panicked at the first recursion prove of the chain run: `sp1-gpu/crates/jagged_tracegen/src/lib.rs:240:70: range end index 37428736 out of range for slice of length 36700160` (20:57Z). That copy carried no `grow_for_main` (the key buffer sized to its preprocessed traces at setup, never grown for the main traces), which the fleet's `floor-v5.patch` has carried since 6 October; the served sm_86 and sm_89 floors behave as floor-v5. Fixed by 8c1fb754c: the three copies are one file (sha 9098c3e5), `box-setup.sh` re-pinned, and the server of 3.1 is built from it.
|
||||
|
||||
## 4. Consequences per card tier
|
||||
|
||||
Every number above, per card tier, with what the lane does about it:
|
||||
|
||||
| Tier | What tonight's rows say | Consequence | Done or owed |
|
||||
|---|---|---|---|
|
||||
| 8 GB card (RTX 4060 class: 5060, 4060, 3070 8 GB, 3060 8 GB, 5060 Ti 8 GB) alone | 7,807 MiB free; shard proof 16.3 s at 7,504 MiB VERIFIED on the default path; the chain's aggregation aborts at a 486 MiB allocation on 176 MiB free | an 8 GB card proves shards and cannot aggregate: the host refuses the chain and the aggregate before setup ("proving needs 9 GB free on the card for the chain workload (8500 MiB floor)"), so the fleet's segment path (`box-prover.py`, one chain run per segment) gets nothing from an 8 GB card until shard production and aggregation are separated (review V6-08's "optional separation", the fleet lane's item); until then an 8 GB owner mines and does not earn proving income, and the site must not say otherwise | floors pinned (this commit); the shard-only route owed to the proving lane (a morning item, not tonight) |
|
||||
| 8 GB card beside its miner (6,120 MiB resident) | 1,689 MiB free; the host refuses in 0 s, the miner unharmed at 18.96 MH/s | mining-only while the miner runs; proving only in a time-share with the dataset evicted (F07's modes); nothing is lost to a failed allocation | the refusal shipped on the host side; the engine's lease (reviewb-202) refuses before the spawn |
|
||||
| 12 GB card (RTX 3060 class: 3060 12 GB, 4070, 5070, 3080 Ti 12 GB, 2080 Ti 11 GB approximate) alone | 11,898 MiB free; shard 11.4 s at 7,601 MiB; the whole chain 27.4 s end to end at an 8,306 MiB aggregation peak, VERIFIED | a 12 GB card runs the fleet's segment path alone with 3.5 GB spare at the peak: shard proofs and the aggregation; the 2.0 litepaper's 12 GB prove-alone sentence holds on the default path with nothing set by hand | measured; the 10 GB tier (3080) is between the two rows and untested tonight (owed: one 3080 hour, the aggregation at 8,306 MiB against 9,885 total is the open cell) |
|
||||
| 12 GB card beside its miner (6,129 MiB resident) | 5,781 MiB free; refused in 0 s, the miner unharmed at 26.83 MH/s | time-share only, as the coexistence rows of 16:10Z said; the host now says so in one line instead of dying inside an allocation 34 s later | shipped |
|
||||
| 16 GB card (4060 Ti 16 GB, 5060 Ti 16 GB, 5070 Ti, 5080, 4080) alone | no row tonight; the table's 16gb tier (2^27 + 2^26, 14,336 MiB free and up) is the patch's own figure | a 16 GB card alone reads the 16gb tier; beside a 6.1 GiB miner it reads 9.9 GB free and lands on the small tier, above both floors (7,700 shard, 8,500 chain), so it is the smallest card that mines and runs the whole segment path at once (about 1.4 GB spare at the aggregation peak, approximate until measured) | owed: one rented 16 GB hour on the default path, alone and beside the miner (the fleet lane's item (h) 16 GB cell carries the V6-07 tarball) |
|
||||
| 24 GB and 32 GB cards (4090, 3090, A5000, 5090) | no row tonight; the 24gb and full tiers are upstream's own figures, unchanged by V6-07 except that the server reads them from FREE memory | a 24 GB card beside its miner (6.1 GiB) reads about 17.5 GB free and now lands on the 16gb tier instead of upstream's 24 GB tier, so a mining 24 GB card proves at the 16 GB threshold (smaller core shards, more of them; the 5 October rows put the time cost of a split at 1.26x on the 5090); alone it is unchanged | measured consequence owed on a rented 4090 beside the miner (one hour) before the 2.0.2 host ships |
|
||||
| Rig (8x 4090) | one process per card (`IGNEUM_CUDA_DEVICE`, `FLEET_CARD`): each host reads its own card's free memory | unchanged behaviour per card; the rig's spare cards never see another card's miner | none |
|
||||
| Pool user, Windows, macOS, AMD, Intel | the profile lives in the host, which runs only where the GPU prover runs (Linux and WSL2 on NVIDIA); macOS, AMD and Intel do not prove (provedefault.rs) | the Windows app's WSL2 host takes the same env from the engine; the pinned RAM saving of the small recursion branch (0.5 GiB) helps the 32 GB RAM floor on Windows by a little, not enough to lower it | none tonight |
|
||||
|
||||
Cost of the rows: five Vast pods and two RunPod pods, about USD 0.40 in total (section 6), inside the day's ceiling.
|
||||
|
||||
## 5. The device coordinator hook and what is next
|
||||
|
||||
The app lane's answer (the shipper, 20:4x UK, from the window lane's F07 design on `reviewb-202`, `src/device.rs`: a Coordinator with per-device leases, `admit(device, holder, mib, mode, budget)` and a Mode per card): the engine owns admission and the host never re-reads a card the engine leased. On every spawn of `igneum-prove-host` the engine passes, as environment on the command: `IGNEUM_PROVE_DEVICE` (the card's ordinal as the host enumerates it), `IGNEUM_PROVE_WORKLOAD` (shard, aggregate or chain: the one job this process is admitted for), `IGNEUM_PROVE_MEM_FREE_MB` (the free memory the engine read on that card at admission), `IGNEUM_PROVE_MEM_BUDGET_MB` (the lease's grant, the hard ceiling for the host) and `IGNEUM_PROVE_DEADLINE_S` (the admission deadline the lease carries). The host picks its row from `IGNEUM_PROVE_MEM_BUDGET_MB`, from `IGNEUM_PROVE_MEM_FREE_MB` when no grant is given, and only without either (the fleet's `box-prover.py` path and a hand run) from `nvidia-smi memory.free` on the device (`IGNEUM_PROVE_DEVICE`, else `IGNEUM_CUDA_DEVICE`, else the first of `CUDA_VISIBLE_DEVICES`, else 0). The refusal is exit 78 with the one line `RESULT memory_profile refused: proving needs N GB free on the card for the <workload> workload (<floor> MiB floor); <free> MiB free`; the engine records it on the lease and the card reads "proving needs N GB free" on the window. No file, no socket. The host side is in this branch; the engine half (the env names exactly as written, the lease table, the refusal on the lease) is the window lane's on `reviewb-202` (a414b6bdc81d348d8), named to it at 20:5x UK.
|
||||
|
||||
Next: (1) the engine half lands and the `lease` rows of section 3 are rerun from the engine's own spawn; (2) the aggregate floor is re-pinned from the chain rows' aggregation peak (section 3) when the 16 GB and 24 GB chain rows exist; (3) the 16gb tier's 2^27 + 2^26 is measured on a rented 16 GB card (no row tonight); (4) the fleet's served floor tarballs (`igneum-floor-sm86.tgz`, `sm89`, `sm120`) are rebuilt from the V6-07 patch (the sm_86 + sm_89 tarball of section 6 is the first) before the standing provers restart on the default path.
|
||||
|
||||
## 6. Shas, pods, cost
|
||||
|
||||
| Item | Sha / id |
|
||||
|---|---|
|
||||
| Branch `v607-floor-memory` | e6abe8c2c (patch + host profile + box-prover), 573dad0ad (lease lesser rule), 8c1fb754c (one floor patch), this document's commit (floors 8,500 and the rows) |
|
||||
| The one floor patch (`proving/prover-floor/sp1-gpu-6.8.1-floor.patch` = `tools/fleet/floor.patch` = `tools/fleet/floor-v5.patch`) | sha256 9098c3e5979d057031f43588668201d1f7a53ab17e980c1eaf896cd5e6f38575 |
|
||||
| V6-07 floor server tarball, sm_86 + sm_89 (`/srv/workers/fleet/igneum-floor-sm8689-v607.tgz` on build-1, served at https://build.igneum.network/fleet/igneum-floor-sm8689-v607.tgz) | tarball 9c02fcc6a3f045d3167c1d9cb6e22c584ea862189d11997754a57cbe1b0d355b; `bin/sp1-gpu-server` db37c38b5feabdec5bbfd0298446da89f6a5d144125e399d6ab617f64f9cb61d; ELF targets sm_86 sm_89; the patch's `floor_tests` 2 passed on the build pod (57 s) |
|
||||
| The first (known-failed) server, base patch copy | 75d0b4beb5c58fc5e9ae9902ffc9b9d64d8237474fc32a16c2830051bf3d6866 (kept only as the record of 3.5; not served) |
|
||||
| V6-07 host (box 3, `--features cuda`, cef5234b5 + the branch) | 66662219a363077257c452ccf4be3e184c848d577a277e (full sha in the box's builds.jsonl); its suite: `cargo test -p igneum-prove-host memory_profile` 7 passed on box 3 (21:5x UK), re-run with the 8,500 floors (section 6 amendment) |
|
||||
| The served 0317 host (the fleet's) | 71bc2438856bb141 |
|
||||
| A plain master host (dc94dda95, no V6-07 change), for the attribution | f2dfe3ad0db00204 |
|
||||
| Pods | fb-v607 RunPod RTX 4090 secure 6sgr7dk56z6zwc 19:41 to 19:53Z USD 0.89/h (the first server); fb-v607b RunPod RTX 5090 community 2i76gud81v440j 20:16 to 20:24Z USD 0.69/h (the served server); v607-3060 Vast 54902103 19:46 to 20:33Z USD 0.058/h; v607-4060 Vast 54902102 19:46Z, the host vanished from Vast at 20:1xZ mid-run (USD 0.04); v607-4060b Vast 54906194 20:18 to 20:34Z (USD 0.02); v607-3060c Vast 54909866 20:4x to 20:57Z (USD 0.01); v607-4060c Vast 54909865 20:4x to 21:00Z (USD 0.02); v607-3060d Vast (the bisect, section 6 amendment). Every pod destroyed and read back as unlisted; no miner on any Hetzner box. |
|
||||
|
||||
Amendments (22:5x UK):
|
||||
|
||||
- The host bisect on v607-3060d (Vast 54913403, RTX 3060, 21:20 to 21:25Z, USD 0.01, destroyed and read back unlisted), the V6-07 server db37c38b, the default path, compressed shard 0: host from 13b729467 (= 6dbd5d2f8^, sha c4b77736107bdbce, sources 8a1c7abf7b7fc378) prove 12.2 s VERIFIED; host from 6dbd5d2f8 (d0d412e9dfeaccd5, sources 600cdc322ba2028b) `Error: public values are 0 bytes, expected 328` after 13 s; host from 2ceb09be9 (89a0453895c1539c, sources 3c293d2236090449, a side branch not carrying 6dbd5d2f8) prove 12.1 s VERIFIED. The fault is 6dbd5d2f8 alone: the FixtureEnv field without the guest re-pin, as its message says.
|
||||
- The host suite with the 8,500 MiB floors: `cargo test -p igneum-prove-host memory_profile` 7 passed on box 3 (22:2x UK).
|
||||
- Every pod of the night is destroyed: fb-v607, fb-v607b (RunPod), v607-3060, v607-4060 (vanished), v607-4060b, v607-3060c, v607-4060c, v607-3060d (Vast); about USD 0.45 in all.
|
||||
|
|
@ -48,10 +48,14 @@ from 20 M cycles is the threshold's padded area reached. The witness (5 to 22 KB
|
|||
|
||||
## What the patch does (`proving/prover-floor/sp1-gpu-6.8.1-floor.patch`, three files)
|
||||
|
||||
1. `builder.rs`: the panic is gone; the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks the element threshold
|
||||
from a tier table (`element_threshold_for_budget`: over 30 as read, the full 402.6 M; 24 to 30, upstream's 24 GB
|
||||
figure; 18 to 24 (a 16 GB card), 2^27 + 2^26 = 201.3 M; under 18 (a 12 GB card), 2^27 = 134.2 M);
|
||||
`SP1_GPU_ELEMENT_THRESHOLD` sets it directly and `SP1_GPU_RECURSION_TRACE_ALLOCATION` the recursion buffer.
|
||||
1. `builder.rs`: the panic is gone; the card's FREE memory at start (read once per process; or
|
||||
`SP1_GPU_MEMORY_BUDGET_GB`, the host's lease), never its total, picks the element threshold from a tier table
|
||||
(`element_threshold_for_budget`: over 30 as read, the full 402.6 M; 24 to 30, upstream's 24 GB figure; 18 to 24
|
||||
(a 16 GB card alone), 2^27 + 2^26 = 201.3 M; under 18 (a 12 GB or 8 GB card, any card beside a miner), 2^26 =
|
||||
67.1 M, the value the passing small-card rows used) and the recursion trace allocation (upstream's 2^27 on the
|
||||
24 GB tier and above, 2^26 + 2^25 under it: V6-07, 8 October 2026, `docs/analysis/floor-memory-profile-2026-10-08.md`);
|
||||
`SP1_GPU_ELEMENT_THRESHOLD` sets it directly and `SP1_GPU_RECURSION_TRACE_ALLOCATION` the recursion buffer. The
|
||||
host's own profile table (`host/src/memory_profile.rs`) sets both by environment before the server starts.
|
||||
The chosen numbers are printed as a `FLOOR opts` line. Every other option is as upstream.
|
||||
2. `jagged_tracegen/src/lib.rs`: with `SP1_GPU_FLOOR_LOG` set, every trace allocation prints its capacity and,
|
||||
after the shard's traces are in, the elements actually used and the device memory in use.
|
||||
|
|
|
|||
|
|
@ -96,6 +96,7 @@ The daily dataset of section 2 proves the miner holds the chain once a day. Proo
|
|||
| How a miner learns it | the template's `powEpoch` carries `stateBlock` (the hash of `C_w`) and `nextStateBlock` once the next window's cut has passed (one lead before the boundary, with `nextEpochSeed`); the miner fetches the stream for that block from its node (`igneum_getPowStateLeaves [block hash]`), checks nothing when the node is its own (its node executed it), and prepares the next pair; the pack carries `IGNEUM_STATE_BLOCK_HEX`, `IGNEUM_STATE_ROOT_HEX` and `leaves.bin` | the existing next-epoch prepare: with `W = epoch_blocks` the refresh and the program swap are one swap |
|
||||
| The grace at the boundary | none in validation (a block's window is its DAA score's; a block mined with the previous window's leaves after the boundary is invalid); the grace is the lead: the next window's leaves are knowable ten minutes before it, the miner prepares the pair then, and a prepared worker swaps with no pause (the 2.0 hot-swap path). An unprepared worker rebuilds at the boundary: 32 ms of hashing lost on a 4090, under 0.001 percent of the window | no hash-rate dip by construction, as the epoch swap has none today |
|
||||
| A miner whose node is behind | the node serves no template for a window whose `C_w` it has not executed (section 5's refusal), so the miner's hash stops at the boundary rather than mining invalid blocks | a node ten minutes behind the chain is already a node without a useful template |
|
||||
| Several candidate streams per epoch (fix 2, 8 October 2026) | the streams a node serves are keyed by BLOCK HASH, never one per epoch: this chain's cut block, and the cut blocks of other chains whose headers this node validated (section 8's amended row), each with the retention of a capture (two epochs). Nothing in this design assumed one stream per epoch per node: the leaves are keyed by each stream's own root `R_w` and the pairing check is per stream | a node that can only serve its own chain's stream can never weigh another chain's post-cut headers, so a cut was a partition that never healed (igneum-devnet-4, 8 October 2026) |
|
||||
|
||||
### 2a.2 What a pool can and cannot centralise, and the farm attack, with the numbers
|
||||
|
||||
|
|
@ -141,6 +142,13 @@ What a pool can centralise: templates (as today), the day key (public), the stre
|
|||
| Timestamp or DAA grinding at the cut | the cut is a DAA score, as the epoch seed's; a producer chooses at most whether its own block is `C_w` (one bit between two honest states) |
|
||||
| Two windows of leaves in RAM on the node and the worker | the current and the next: 12 KB today, at most 4 GiB at the sample cap; the worker frees the leaves after the build (1,803 MiB resident while hashing, measured) |
|
||||
|
||||
### 2a.5 The cost of candidate streams (fix 2, the foreign-seed capture; the HEAL lane, 8 October 2026)
|
||||
|
||||
| Row | Cost | Note |
|
||||
|---|---|---|
|
||||
| Candidate streams held for validation | one dataset refresh `D` per candidate stream (6 KB of leaves today), held for the retention of a capture (two epochs) | the miner hashes on its own chain's stream only; a candidate is for validating another chain's headers, which is what lets their weight count and the node switch |
|
||||
| The grinding and DoS face | at most 2 x epoch_blocks of EVM execution per capture, one capture in flight, a 60 s memory of a refusal; the capture runs only for a block this node holds as a VALIDATED header (its chain passed proof of work up to that block), never for an unvalidated one | anyone mining a fork off a pre-cut block makes the node execute from its snapshot base; without the bound, the one-in-flight rule and the PoW-first condition the executor is a free compute oracle, so the three are named here as properties of the rule, to keep |
|
||||
|
||||
## 3. A miner with a pruned node, and what a chip must hold
|
||||
|
||||
| Who | What they hold | How they get it |
|
||||
|
|
@ -228,7 +236,7 @@ Per tier, what the 4090 rows mean: every NVIDIA card from 8 to 32 GB keeps its h
|
|||
|---|---|
|
||||
| A miner served stale state (yesterday's stream, a fork's stream, a tampered stream) | every item it builds is wrong, every block it submits is rejected by every node that holds the day's state, and it learns it from the first rejection; a worker that takes the stream from its own node checks nothing (its node executed it), a worker or pool miner that takes it from elsewhere runs `DayStream::check` against the `R_d` its node reports (`igneum_getPowDayState` returns the root with the block) |
|
||||
| A pool serving wrong state to its miners | the pool loses every share it pays for; the pool's own node rejects its miners' solutions; there is no way to profit from a wrong dataset, only to waste the pool's hash |
|
||||
| A state the verifier cannot reach (a header whose `C_d` the node never executed: a fork deeper than the lead) | the engine refuses the header with the retryable error and says which block it lacks; a fork that deep is a merge-depth-scale reorg, which the follower re-executes when the chain adopts it ("a deep reorg never resets execution", 6 October rule), after which the header validates; a chain the node never adopts is a chain it never needs the state of |
|
||||
| A state the verifier cannot reach (a header whose `C_d` the node never executed: another chain's cut block, from a partition, a ban storm or a lag across the cut) | AMENDED by fix 2 of node 2.0.2 (the HEAL lane, 8 October 2026, night; igneum-devnet-4 had fragmented into 6 chains and 11 one-box islands at the epoch-3 cut because the retryable refusal never ended: the other chain's post-cut headers never validated, their weight never counted, the virtual never moved and the executor never captured the other chain's block). The refusal is now the BOUNDED EXECUTION of the foreign selected-parent chain: a miss queues the block (16, deduplicated; the provider runs inside header validation and never takes the state lock) and returns the same wait error; the follower services one block per pass: the block must be a header this node validated (PoW first), its selected-parent chain is walked down to the fork point (the first block this executor holds a record of), a scratch state starts from the nearest held state at or below it (the ring, else the newest epoch capture; the snapshot base when started from one), this chain's blocks to the fork point and the foreign chain up to the block are executed on the scratch by the follower's own chain-block function (the proof verdict by the acceptance rule, the same the live path pays by), the stream after the block is serialised, checked against its root and published under the block's hash; the IBD's 2 s retry finds it and the lighter node switches. Bounds, which are PROPERTIES and not tunables: one capture in flight, at most 2 x epoch_blocks of execution, only a validated header; a chain that cannot be executed (a body missing, a fork below every held state, the bound passed) is refused once with a line and remembered (60 s, 10 s for a body the IBD is still bringing). Measured on two real nodes split across a cut (fork docs/igneum-foreign-seed-capture.md): without it the lighter node never left its chain in 300 s; with it, it switched at +20 s and both read one executed tip and epoch seed from +60 s on. A chain the node never adopts is still a chain it executed once on a scratch and dropped |
|
||||
| The executor is behind the cut when the day starts (a slow node, a node that restarted) | it refuses to validate and to mine v5 headers until it passes `cut(d)` and says so once per epoch at warn, then at debug (the class signal's `first_time` shape); with the lead at one hour this is a node an hour behind the chain, which is already a node that serves no useful template |
|
||||
| A node without an executor (`--evm-disable`) | states at start that it cannot validate or mine class v5; after the flip it relays headers it cannot check as it relays blocks whose bodies it has not fetched, and accepts nothing it cannot validate (the same as a node without the day's cache slot today: `BuildQueueFull`) |
|
||||
| The era draw's interaction | the era draws the layout (`t(w)`, `j(w)`, the stride and the window of each load) and nothing about the leaves; the leaf XOR sits before the first mixer of item `t`, after the layout has named `t`, so every era's dataset of a day is built from one leaf array and the hash kernel of every era is unchanged; the era draws 8 and 9 stay consumed and unused as the ladder left them |
|
||||
|
|
|
|||
|
|
@ -416,6 +416,16 @@ Rule: a case maps to a cell only where the cell's tests visibly answer it; cover
|
|||
- Cases:
|
||||
- R2-F03-R02 Same job context produces identical accepted work in node, CPU reference, CUDA, Metal, OpenCL and pool.: partial: the worker half (CPU reference, CUDA, Metal, OpenCL) on the same program packs; the node's and the pool's accepted work on the same job context are the CI steward's and the pool lane's cells; PASS only when every listed platform reads one fingerprint and the node and pool halves are green
|
||||
|
||||
### rows:v607-floor-memory
|
||||
|
||||
- Command: `the V6-07 lane's pod scripts (scratch v607/pod-v607.sh, ab.sh, ab2.sh, ab3.sh) on rented Vast RTX 3060 12 GB and RTX 4060 8 GB one-shots: the served 0317 host and the V6-07 host on the V6-07 floor server (igneum-floor-sm8689-v607.tgz), the fixture fees-v1-shards2.json, the ds55 kit's worker as the miner, nvidia-smi at 1 Hz; recorded in docs/analysis/floor-memory-profile-2026-10-08.md`
|
||||
- Box class: rented pods (Vast one-shots)
|
||||
- Fixtures: F0, F1
|
||||
- Cases:
|
||||
- GPU-05 Test dataset fit and support-horizon costs: partial: the 8 GB and 12 GB tiers' prover headroom on the 5.5 GiB dataset (free memory, the prover's peak alone, the refusal beside the miner); fragmentation, restart and next-epoch construction not run
|
||||
- CAP-02 Prove on the actual mining configuration: partial: proving alone (shard and the whole segment path) on the 8 GB and 12 GB tiers on the default job path, the time-share refusal beside the miner with the miner unharmed, memory headroom and proof latency; wall energy only on the 3060 (the 4060 host has no power sensor); induced GPU task failure and wallet control not run; 16 GB and 24 GB tiers not run
|
||||
- UX-02 Make pause, stop and safe tuning reliable: partial: the prover's safe refusal under memory pressure (exit 78 in one line, the miner's worker unharmed, nothing leaked after the refusal); pause, stop, power limit and the UI's crash ownership are the suite:app cell's
|
||||
|
||||
## Automated cases with no harness in the matrix (NOT RUN, the reason)
|
||||
|
||||
- GOV-02 Approve thresholds before results: the approval is recorded in the registry's approval field; the automated half (thresholds frozen before any run_status) is the gate rule landing by 21:00
|
||||
|
|
@ -566,4 +576,4 @@ Rule: a case maps to a cell only where the cell's tests visibly answer it; cover
|
|||
|
||||
## Count
|
||||
|
||||
171 automated cases: 83 mapped to a cell, 145 NOT RUN with a reason.
|
||||
171 automated cases: 84 mapped to a cell, 145 NOT RUN with a reason.
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
|
|
@ -20,6 +20,7 @@
|
|||
//! ids with no setup. `--mode verify` uses SP1's light verifier and the pinned verifying key: no prover client,
|
||||
//! no key generation (the 114 s to 138 s the Mac's node spent per proof on 5 October).
|
||||
|
||||
mod memory_profile;
|
||||
mod pinned;
|
||||
mod proof_system;
|
||||
|
||||
|
|
@ -69,6 +70,15 @@ fn run() -> Result<()> {
|
|||
let args: Vec<String> = std::env::args().collect();
|
||||
let arg = |name: &str| args.iter().position(|a| a == name).and_then(|i| args.get(i + 1)).cloned();
|
||||
let mode = arg("--mode").unwrap_or_else(|| "all".into());
|
||||
// V6-07 (8 October 2026): on the GPU prover the memory profile for this process's workload is chosen here, once,
|
||||
// from the card's FREE memory (the engine's lease, else nvidia-smi) and handed to the floor server by environment
|
||||
// before the SP1 client spawns it; a budget under the workload's floor is refused with exit 78 before any setup.
|
||||
if std::env::var("SP1_PROVER").map(|p| p == "cuda").unwrap_or(false) {
|
||||
if let Err(refusal) = memory_profile::apply(&mode) {
|
||||
println!("RESULT memory_profile refused: {refusal}");
|
||||
std::process::exit(memory_profile::EXIT_REFUSED);
|
||||
}
|
||||
}
|
||||
let pinned = pinned::Pinned::load()?;
|
||||
if mode == "id" {
|
||||
println!("RESULT id: {}", pinned.describe());
|
||||
|
|
|
|||
402
proving/igneum-prove/host/src/memory_profile.rs
Normal file
402
proving/igneum-prove/host/src/memory_profile.rs
Normal file
|
|
@ -0,0 +1,402 @@
|
|||
//! The pinned memory profile per workload (master review R1 residual V6-07, 8 October 2026): the GPU server's two
|
||||
//! device buffers (the core element threshold and the recursion trace allocation) are chosen HERE, once, from the
|
||||
//! memory that is FREE on the card when the host starts, never from the card's total. The table below is the
|
||||
//! profile; `choose` picks a row for the workload this process is admitted for; `apply` hands the row to the server
|
||||
//! through the floor patch's environment (`SP1_GPU_ELEMENT_THRESHOLD`, `SP1_GPU_RECURSION_TRACE_ALLOCATION`,
|
||||
//! `SP1_GPU_MEMORY_BUDGET_GB`) before the SP1 client spawns it. An explicit `SP1_GPU_ELEMENT_THRESHOLD` or
|
||||
//! `SP1_GPU_RECURSION_TRACE_ALLOCATION` already in the environment is a hand override and wins, named in the line.
|
||||
//!
|
||||
//! Where the free memory comes from, in order (the app lane's device coordinator interface, the shipper's answer of
|
||||
//! 20:4x UK 8 October 2026; the engine owns admission and the host never re-reads a card the engine leased):
|
||||
//! 1. `IGNEUM_PROVE_MEM_BUDGET_MB`: the lease's grant, the hard ceiling for this process (the engine's coordinator,
|
||||
//! app/igneum-app `src/device.rs`, passes it with `IGNEUM_PROVE_DEVICE`, `IGNEUM_PROVE_WORKLOAD`,
|
||||
//! `IGNEUM_PROVE_MEM_FREE_MB` and `IGNEUM_PROVE_DEADLINE_S` on every spawn); with `IGNEUM_PROVE_MEM_FREE_MB`
|
||||
//! beside it the LESSER of the two decides (`lease_reading`);
|
||||
//! 2. `IGNEUM_PROVE_MEM_FREE_MB` alone: the engine's own read at admission, when no grant is given;
|
||||
//! 3. `nvidia-smi --query-gpu=memory.free` on the device (`IGNEUM_PROVE_DEVICE`, else `IGNEUM_CUDA_DEVICE`, else the
|
||||
//! first of `CUDA_VISIBLE_DEVICES`, else 0): the fleet's `tools/fleet/box-prover.py` path and a hand run, where no
|
||||
//! engine admits the job.
|
||||
//! When none answers (no nvidia-smi on the path), no row is applied and the server's own free-memory rule decides.
|
||||
//!
|
||||
//! A budget under the workload's floor is refused before the server starts: one line, exit code 78 (the engine
|
||||
//! records the refusal on the lease and the window reads "proving needs N GB free").
|
||||
//!
|
||||
//! The rows cited for the small tier's 2^26: `docs/analysis/prover-tiers-real-cards.md` (6 October 2026,
|
||||
//! `alone-comp-26-v1`: RTX 3060 7.4 GB 14.4 s, RTX 3080 8.0 GB 7.1 s, RTX 4060 7.4 GB 18.4 s, RTX 4060 Ti 8 GB 7.6 GB
|
||||
//! 9.6 s, RTX 4060 Ti 16 GB 7.8 GB 11.6 s, RTX 4070 7.6 GB 12.1 s, RTX 5070 7.6 GB 4.8 s, all verified) and
|
||||
//! `docs/analysis/class-v6/coexist-rows.md` (8 October 2026, `proof_alone` at `SP1_GPU_ELEMENT_THRESHOLD=67108864`:
|
||||
//! RTX 3060 peak 7,525 MiB 13.2 s verified, RTX 4060 peak 7,532 MiB 8.2 s verified). No small-card row ever passed
|
||||
//! at the patch's earlier default of 2^27 on this host; the 10 GB tier's 2^27 rows (8,642 MiB, 7 October 2026) ran
|
||||
//! alone on the segment host and are not this path.
|
||||
|
||||
use std::fmt;
|
||||
|
||||
/// Upstream's core element threshold (sp1-core-executor 6.8.1 `opts.rs` 12): 2^28 + 2^27.
|
||||
pub const ELEMENT_THRESHOLD_FULL: u64 = (1 << 28) + (1 << 27);
|
||||
/// Upstream's 24 GB tier (sp1-gpu `builder.rs` 44): the full threshold less 2^26 + 2^25 + 2^24.
|
||||
pub const ELEMENT_THRESHOLD_24GB: u64 = ELEMENT_THRESHOLD_FULL - (1 << 26) - (1 << 25) - (1 << 24);
|
||||
/// The floor patch's 16 GB tier: 2^27 + 2^26 (unmeasured on this fixture; the patch's own figure).
|
||||
pub const ELEMENT_THRESHOLD_16GB: u64 = (1 << 27) + (1 << 26);
|
||||
/// The small-card threshold: 2^26, the value every passing small-card row used (module note).
|
||||
pub const ELEMENT_THRESHOLD_SMALL: u64 = 1 << 26;
|
||||
/// Upstream's recursion trace allocation (sp1-gpu `builder.rs` 15): 2^27 elements, 0.75 GiB each.
|
||||
pub const RECURSION_TRACE_ALLOCATION: usize = 1 << 27;
|
||||
/// The small-card recursion trace allocation: the recursion keys and shards use 90,177,536 elements each (35.6 M
|
||||
/// preprocessed, 54.5 M main; `docs/analysis/prover-floor.md`, sweep 1), so 2^26 + 2^25 = 100,663,296 holds them
|
||||
/// with the stacking slack the patch's `floor_capacity` adds. The two constants differ (`recursion_branches_differ`).
|
||||
pub const RECURSION_TRACE_ALLOCATION_SMALL: usize = (1 << 26) + (1 << 25);
|
||||
/// What a recursion key or shard was measured to use (elements); the small allocation must clear it with slack.
|
||||
pub const RECURSION_TRACE_USED: usize = 90_177_536;
|
||||
|
||||
/// The one job this process is admitted for.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub enum Workload {
|
||||
/// One shard's compressed (or core) proof.
|
||||
Shard,
|
||||
/// The aggregator guest over shard proofs (and the previous segment's proof).
|
||||
Aggregate,
|
||||
/// A whole segment in one process: every shard, then the aggregation.
|
||||
Chain,
|
||||
}
|
||||
|
||||
impl Workload {
|
||||
pub fn name(self) -> &'static str {
|
||||
match self {
|
||||
Workload::Shard => "shard",
|
||||
Workload::Aggregate => "aggregate",
|
||||
Workload::Chain => "chain",
|
||||
}
|
||||
}
|
||||
/// `IGNEUM_PROVE_WORKLOAD` when the engine names it, else the host's own mode.
|
||||
pub fn from_env_or_mode(mode: &str) -> Option<Workload> {
|
||||
if let Ok(w) = std::env::var("IGNEUM_PROVE_WORKLOAD") {
|
||||
return match w.as_str() {
|
||||
"shard" => Some(Workload::Shard),
|
||||
"aggregate" => Some(Workload::Aggregate),
|
||||
"chain" => Some(Workload::Chain),
|
||||
_ => None,
|
||||
};
|
||||
}
|
||||
Workload::from_mode(mode)
|
||||
}
|
||||
/// The workload a host mode runs on the GPU (modes that never prove return None).
|
||||
pub fn from_mode(mode: &str) -> Option<Workload> {
|
||||
match mode {
|
||||
"shard" | "compressed" | "core" => Some(Workload::Shard),
|
||||
"aggregate" => Some(Workload::Aggregate),
|
||||
"chain" | "block" | "all" => Some(Workload::Chain),
|
||||
_ => None,
|
||||
}
|
||||
}
|
||||
/// The least free memory (MiB) a workload runs in. Shard: the small tier's measured peak (7,504 to 7,632 MiB
|
||||
/// on the 3060 and 4060 at 2^26, 8 October 2026) plus headroom for the driver's own context. Aggregate and
|
||||
/// chain: the aggregation is the peak of a chain run, 8,306 MiB on the RTX 3060 (the default path, 20:54Z
|
||||
/// 8 October 2026); on the RTX 4060 (7,807 MiB free) the same run proved both shards and aborted at the
|
||||
/// aggregation's 486 MiB allocation (`FLOOR abort` at `slop/crates/tensor/src/inner.rs:51`, 20:56Z), so the
|
||||
/// floor sits above an 8 GB card: `docs/analysis/floor-memory-profile-2026-10-08.md` section 3.
|
||||
pub fn floor_mib(self) -> u64 {
|
||||
match self {
|
||||
Workload::Shard => 7_700,
|
||||
Workload::Aggregate => 8_500,
|
||||
Workload::Chain => 8_500,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// One tier of the profile table: the least free memory it needs and the two buffers it sets.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub struct Tier {
|
||||
pub name: &'static str,
|
||||
pub min_free_mib: u64,
|
||||
pub element_threshold: u64,
|
||||
pub recursion_trace_allocation: usize,
|
||||
}
|
||||
|
||||
/// The pinned table, largest first. The free-memory lines are the card tiers as the floor patch reads them (free
|
||||
/// GiB, ceiling, +4 as upstream computed it from the total): 26 GiB free reads over 30 (a 32 GB card alone), 20 GiB
|
||||
/// reads 24 (a 24 GB card alone), 14 GiB reads 18 (a 16 GB card alone), under that the small tier (12 GB and 8 GB
|
||||
/// cards alone; any card beside a miner's resident set).
|
||||
pub const TIERS: [Tier; 4] = [
|
||||
Tier { name: "full", min_free_mib: 26 * 1024, element_threshold: ELEMENT_THRESHOLD_FULL, recursion_trace_allocation: RECURSION_TRACE_ALLOCATION },
|
||||
Tier { name: "24gb", min_free_mib: 20 * 1024, element_threshold: ELEMENT_THRESHOLD_24GB, recursion_trace_allocation: RECURSION_TRACE_ALLOCATION },
|
||||
Tier { name: "16gb", min_free_mib: 14 * 1024, element_threshold: ELEMENT_THRESHOLD_16GB, recursion_trace_allocation: RECURSION_TRACE_ALLOCATION_SMALL },
|
||||
Tier { name: "small", min_free_mib: 0, element_threshold: ELEMENT_THRESHOLD_SMALL, recursion_trace_allocation: RECURSION_TRACE_ALLOCATION_SMALL },
|
||||
];
|
||||
|
||||
/// The row chosen for a process.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub struct Profile {
|
||||
pub workload: Workload,
|
||||
pub tier: Tier,
|
||||
pub free_mib: u64,
|
||||
}
|
||||
|
||||
/// Why a process is refused before the server starts.
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
pub struct Refusal {
|
||||
pub workload: Workload,
|
||||
pub free_mib: u64,
|
||||
pub floor_mib: u64,
|
||||
}
|
||||
|
||||
impl fmt::Display for Refusal {
|
||||
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
|
||||
write!(
|
||||
f,
|
||||
"proving needs {} GB free on the card for the {} workload ({} MiB floor); {} MiB free",
|
||||
(self.floor_mib + 1023) / 1024,
|
||||
self.workload.name(),
|
||||
self.floor_mib,
|
||||
self.free_mib
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
/// The exit code of a refusal (the engine reads it on the lease).
|
||||
pub const EXIT_REFUSED: i32 = 78;
|
||||
|
||||
/// The tier for a free-memory reading (pure; the table's first row whose line the reading clears).
|
||||
pub fn tier_for(free_mib: u64) -> Tier {
|
||||
*TIERS.iter().find(|t| free_mib >= t.min_free_mib).expect("the small tier has no floor")
|
||||
}
|
||||
|
||||
/// The row for a workload at a free-memory reading, or the refusal.
|
||||
pub fn choose(workload: Workload, free_mib: u64) -> Result<Profile, Refusal> {
|
||||
let floor_mib = workload.floor_mib();
|
||||
if free_mib < floor_mib {
|
||||
return Err(Refusal { workload, free_mib, floor_mib });
|
||||
}
|
||||
Ok(Profile { workload, tier: tier_for(free_mib), free_mib })
|
||||
}
|
||||
|
||||
/// Where a free-memory reading came from, for the line.
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
pub enum Source {
|
||||
/// `IGNEUM_PROVE_MEM_BUDGET_MB`, the engine's lease.
|
||||
Lease,
|
||||
/// `IGNEUM_PROVE_MEM_FREE_MB`, the engine's read without a grant.
|
||||
EngineRead,
|
||||
/// `nvidia-smi --query-gpu=memory.free` on the named device.
|
||||
NvidiaSmi(String),
|
||||
}
|
||||
|
||||
impl fmt::Display for Source {
|
||||
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
|
||||
match self {
|
||||
Source::Lease => write!(f, "lease IGNEUM_PROVE_MEM_BUDGET_MB"),
|
||||
Source::EngineRead => write!(f, "engine IGNEUM_PROVE_MEM_FREE_MB"),
|
||||
Source::NvidiaSmi(d) => write!(f, "nvidia-smi device {d}"),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The reading from the engine's two numbers: the grant is the ceiling and the engine's free reading is a fact, so
|
||||
/// the row is chosen from the LESSER of the two when both are given (the window lane's engine half, 20:5x UK
|
||||
/// 8 October 2026, grants a fixed proof budget beside the miner, 7,532 MiB plus 10 percent, while a 12 GB card
|
||||
/// beside the 6.1 GiB miner has 6,159 MiB free: the grant alone would pick a row the card cannot hold, and the
|
||||
/// refusal must come from the free figure). One number alone is used as it is.
|
||||
pub fn lease_reading(budget: Option<u64>, free: Option<u64>) -> Option<(u64, Source)> {
|
||||
match (budget, free) {
|
||||
(Some(b), Some(f)) => Some((b.min(f), Source::Lease)),
|
||||
(Some(b), None) => Some((b, Source::Lease)),
|
||||
(None, Some(f)) => Some((f, Source::EngineRead)),
|
||||
(None, None) => None,
|
||||
}
|
||||
}
|
||||
|
||||
fn env_u64(name: &str) -> Option<u64> {
|
||||
std::env::var(name).ok().and_then(|s| s.trim().parse::<u64>().ok())
|
||||
}
|
||||
|
||||
/// The device ordinal the host runs on, as the module note orders it.
|
||||
pub fn device_ordinal() -> String {
|
||||
if let Ok(d) = std::env::var("IGNEUM_PROVE_DEVICE") {
|
||||
return d;
|
||||
}
|
||||
if let Ok(d) = std::env::var("IGNEUM_CUDA_DEVICE") {
|
||||
return d;
|
||||
}
|
||||
if let Ok(v) = std::env::var("CUDA_VISIBLE_DEVICES") {
|
||||
if let Some(first) = v.split(',').next() {
|
||||
if !first.trim().is_empty() {
|
||||
return first.trim().to_string();
|
||||
}
|
||||
}
|
||||
}
|
||||
"0".into()
|
||||
}
|
||||
|
||||
/// The free memory on the card at this call, from the sources in the module note's order.
|
||||
pub fn read_free_mib() -> Option<(u64, Source)> {
|
||||
let budget = env_u64("IGNEUM_PROVE_MEM_BUDGET_MB");
|
||||
let free = env_u64("IGNEUM_PROVE_MEM_FREE_MB");
|
||||
if let Some(choice) = lease_reading(budget, free) {
|
||||
return Some(choice);
|
||||
}
|
||||
let dev = device_ordinal();
|
||||
let out = std::process::Command::new("nvidia-smi")
|
||||
.args(["--query-gpu=memory.free", "--format=csv,noheader,nounits", "-i", &dev])
|
||||
.output()
|
||||
.ok()?;
|
||||
if !out.status.success() {
|
||||
return None;
|
||||
}
|
||||
let text = String::from_utf8_lossy(&out.stdout);
|
||||
let first = text.lines().next()?.trim();
|
||||
first.parse::<u64>().ok().map(|v| (v, Source::NvidiaSmi(dev)))
|
||||
}
|
||||
|
||||
/// The environment the row sets for the floor server (what `apply` writes), as (name, value) pairs; an explicit
|
||||
/// hand override already present keeps its value and is reported.
|
||||
pub fn env_for(p: &Profile) -> Vec<(&'static str, String)> {
|
||||
vec![
|
||||
("SP1_GPU_ELEMENT_THRESHOLD", p.tier.element_threshold.to_string()),
|
||||
("SP1_GPU_RECURSION_TRACE_ALLOCATION", p.tier.recursion_trace_allocation.to_string()),
|
||||
("SP1_GPU_MEMORY_BUDGET_GB", format!("{:.1}", p.free_mib as f64 / 1024.0)),
|
||||
]
|
||||
}
|
||||
|
||||
/// Chooses and applies the row for this process: prints one `RESULT memory_profile` line and returns Ok(Some) with
|
||||
/// the profile, Ok(None) when no reading was possible (the server's own rule decides) or the workload never proves,
|
||||
/// and Err with the refusal line (the caller exits `EXIT_REFUSED`).
|
||||
pub fn apply(mode: &str) -> Result<Option<Profile>, Refusal> {
|
||||
let Some(workload) = Workload::from_env_or_mode(mode) else { return Ok(None) };
|
||||
let Some((free_mib, source)) = read_free_mib() else {
|
||||
println!("RESULT memory_profile: no free-memory reading (no lease, no nvidia-smi on device {}); the server's own free-memory rule decides for workload {}", device_ordinal(), workload.name());
|
||||
return Ok(None);
|
||||
};
|
||||
let profile = choose(workload, free_mib)?;
|
||||
let mut set = Vec::new();
|
||||
let mut kept = Vec::new();
|
||||
for (k, v) in env_for(&profile) {
|
||||
match std::env::var(k) {
|
||||
Ok(have) if k != "SP1_GPU_MEMORY_BUDGET_GB" => kept.push(format!("{k}={have} (hand override kept, the row said {v})")),
|
||||
_ => {
|
||||
std::env::set_var(k, &v);
|
||||
set.push(format!("{k}={v}"));
|
||||
}
|
||||
}
|
||||
}
|
||||
let deadline = std::env::var("IGNEUM_PROVE_DEADLINE_S").ok().map(|d| format!(", deadline {d} s")).unwrap_or_default();
|
||||
println!(
|
||||
"RESULT memory_profile: workload {} tier {} from {} MiB free ({}){}: element_threshold {} recursion_trace_allocation {}; set {}{}",
|
||||
workload.name(),
|
||||
profile.tier.name,
|
||||
free_mib,
|
||||
source,
|
||||
deadline,
|
||||
profile.tier.element_threshold,
|
||||
profile.tier.recursion_trace_allocation,
|
||||
set.join(" "),
|
||||
if kept.is_empty() { String::new() } else { format!("; {}", kept.join("; ")) }
|
||||
);
|
||||
Ok(Some(profile))
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
/// The table's values are pinned: a change here is a change of the profile and needs its rows.
|
||||
#[test]
|
||||
fn table_is_pinned() {
|
||||
assert_eq!(ELEMENT_THRESHOLD_FULL, 402_653_184);
|
||||
assert_eq!(ELEMENT_THRESHOLD_24GB, 285_212_672);
|
||||
assert_eq!(ELEMENT_THRESHOLD_16GB, 201_326_592);
|
||||
assert_eq!(ELEMENT_THRESHOLD_SMALL, 67_108_864);
|
||||
assert_eq!(RECURSION_TRACE_ALLOCATION, 134_217_728);
|
||||
assert_eq!(RECURSION_TRACE_ALLOCATION_SMALL, 100_663_296);
|
||||
let names: Vec<&str> = TIERS.iter().map(|t| t.name).collect();
|
||||
assert_eq!(names, ["full", "24gb", "16gb", "small"]);
|
||||
assert_eq!(TIERS[0].min_free_mib, 26_624);
|
||||
assert_eq!(TIERS[1].min_free_mib, 20_480);
|
||||
assert_eq!(TIERS[2].min_free_mib, 14_336);
|
||||
assert_eq!(TIERS[3].min_free_mib, 0);
|
||||
assert_eq!(TIERS[3].element_threshold, ELEMENT_THRESHOLD_SMALL);
|
||||
assert_eq!(TIERS[3].recursion_trace_allocation, RECURSION_TRACE_ALLOCATION_SMALL);
|
||||
assert_eq!(TIERS[0].recursion_trace_allocation, RECURSION_TRACE_ALLOCATION);
|
||||
assert_eq!(Workload::Shard.floor_mib(), 7_700);
|
||||
assert_eq!(Workload::Aggregate.floor_mib(), 8_500);
|
||||
assert_eq!(Workload::Chain.floor_mib(), 8_500);
|
||||
assert_eq!(EXIT_REFUSED, 78);
|
||||
}
|
||||
|
||||
/// The two recursion branches are two values (the review found one constant in both), and the small one
|
||||
/// clears what a recursion key or shard was measured to use, with a stacking height of slack.
|
||||
#[test]
|
||||
fn recursion_branches_differ() {
|
||||
assert_ne!(RECURSION_TRACE_ALLOCATION, RECURSION_TRACE_ALLOCATION_SMALL);
|
||||
assert!(RECURSION_TRACE_ALLOCATION_SMALL < RECURSION_TRACE_ALLOCATION);
|
||||
assert!(RECURSION_TRACE_ALLOCATION_SMALL >= RECURSION_TRACE_USED + (1 << 22));
|
||||
assert_ne!(tier_for(24 * 1024).recursion_trace_allocation, tier_for(12 * 1024).recursion_trace_allocation);
|
||||
}
|
||||
|
||||
/// The small tier's threshold is 2^26, the passing rows' value; the card tiers alone read their own rows.
|
||||
#[test]
|
||||
fn tiers_from_free_memory() {
|
||||
assert_eq!(tier_for(8_186).name, "small");
|
||||
assert_eq!(tier_for(8_186).element_threshold, 1 << 26);
|
||||
assert_eq!(tier_for(12_100).name, "small");
|
||||
assert_eq!(tier_for(12_100).element_threshold, 1 << 26);
|
||||
assert_eq!(tier_for(16_100).name, "16gb");
|
||||
assert_eq!(tier_for(24_200).name, "24gb");
|
||||
assert_eq!(tier_for(32_300).name, "full");
|
||||
// a 12 GB card beside the 5.5 GiB miner (6,129 MiB resident) reads the small tier by what is FREE
|
||||
assert_eq!(tier_for(12_288 - 6_129).name, "small");
|
||||
}
|
||||
|
||||
/// The refusal: under the floor, before the server starts; at the floor, the small row.
|
||||
#[test]
|
||||
fn refuses_under_the_floor() {
|
||||
let r = choose(Workload::Shard, 12_288 - 6_129).unwrap_err();
|
||||
assert_eq!(r.floor_mib, 7_700);
|
||||
assert_eq!(r.free_mib, 6_159);
|
||||
assert_eq!(r.to_string(), "proving needs 8 GB free on the card for the shard workload (7,700 MiB floor); 6,159 MiB free".replace(",", ""));
|
||||
let p = choose(Workload::Shard, 7_700).unwrap();
|
||||
assert_eq!(p.tier.name, "small");
|
||||
// an 8 GB card (7,807 MiB free on the RTX 4060) proves shards and is refused the chain (its aggregation
|
||||
// aborted at 7,631 MiB used plus a 486 MiB allocation on 8 October 2026); a 12 GB card runs it
|
||||
let p = choose(Workload::Shard, 7_807).unwrap();
|
||||
assert_eq!(p.tier.element_threshold, ELEMENT_THRESHOLD_SMALL);
|
||||
assert_eq!(p.tier.recursion_trace_allocation, RECURSION_TRACE_ALLOCATION_SMALL);
|
||||
let r = choose(Workload::Chain, 7_807).unwrap_err();
|
||||
assert_eq!(r.to_string(), "proving needs 9 GB free on the card for the chain workload (8500 MiB floor); 7807 MiB free");
|
||||
assert!(choose(Workload::Aggregate, 7_807).is_err());
|
||||
assert_eq!(choose(Workload::Chain, 11_898).unwrap().tier.name, "small");
|
||||
}
|
||||
|
||||
/// The engine's grant beside its free reading: the lesser decides; one alone is used as it is.
|
||||
#[test]
|
||||
fn lease_takes_the_lesser_of_grant_and_free() {
|
||||
assert_eq!(lease_reading(Some(8_285), Some(6_159)), Some((6_159, Source::Lease)));
|
||||
assert_eq!(lease_reading(Some(8_285), Some(12_287)), Some((8_285, Source::Lease)));
|
||||
assert_eq!(lease_reading(Some(8_285), None), Some((8_285, Source::Lease)));
|
||||
assert_eq!(lease_reading(None, Some(12_287)), Some((12_287, Source::EngineRead)));
|
||||
assert_eq!(lease_reading(None, None), None);
|
||||
// a 12 GB card beside the miner under a fixed grant is refused by its free figure
|
||||
assert!(choose(Workload::Shard, lease_reading(Some(8_285), Some(6_159)).unwrap().0).is_err());
|
||||
}
|
||||
|
||||
/// The host's modes map to the three workloads; modes that never prove map to none.
|
||||
#[test]
|
||||
fn workload_from_mode() {
|
||||
assert_eq!(Workload::from_mode("compressed"), Some(Workload::Shard));
|
||||
assert_eq!(Workload::from_mode("core"), Some(Workload::Shard));
|
||||
assert_eq!(Workload::from_mode("aggregate"), Some(Workload::Aggregate));
|
||||
assert_eq!(Workload::from_mode("chain"), Some(Workload::Chain));
|
||||
assert_eq!(Workload::from_mode("all"), Some(Workload::Chain));
|
||||
assert_eq!(Workload::from_mode("verify"), None);
|
||||
assert_eq!(Workload::from_mode("id"), None);
|
||||
assert_eq!(Workload::from_mode("native"), None);
|
||||
}
|
||||
|
||||
/// The environment a row hands the floor server.
|
||||
#[test]
|
||||
fn env_for_the_server() {
|
||||
let p = choose(Workload::Shard, 8_186).unwrap();
|
||||
let env = env_for(&p);
|
||||
assert_eq!(env[0], ("SP1_GPU_ELEMENT_THRESHOLD", "67108864".to_string()));
|
||||
assert_eq!(env[1], ("SP1_GPU_RECURSION_TRACE_ALLOCATION", "100663296".to_string()));
|
||||
assert_eq!(env[2], ("SP1_GPU_MEMORY_BUDGET_GB", "8.0".to_string()));
|
||||
}
|
||||
}
|
||||
BIN
proving/igneum-prove/target-remote/release/igneum-prove-export
Executable file
BIN
proving/igneum-prove/target-remote/release/igneum-prove-export
Executable file
Binary file not shown.
BIN
proving/igneum-prove/target-remote/release/igneum-prove-host
Executable file
BIN
proving/igneum-prove/target-remote/release/igneum-prove-host
Executable file
Binary file not shown.
|
|
@ -1,5 +1,26 @@
|
|||
diff --git a/sp1-gpu/crates/cuda/src/task.rs b/sp1-gpu/crates/cuda/src/task.rs
|
||||
index a503a86..813016b 100644
|
||||
--- a/sp1-gpu/crates/cuda/src/task.rs
|
||||
+++ b/sp1-gpu/crates/cuda/src/task.rs
|
||||
@@ -149,7 +149,15 @@ pub enum GlobalTaskPoolBuildError {
|
||||
|
||||
impl TaskPoolBuilder {
|
||||
pub fn new() -> Self {
|
||||
- Self { capacity: None, device: CudaDevice(0), mem_release_threshold: u64::MAX }
|
||||
+ // Igneum prover-floor patch: upstream keeps every freed device allocation in the pool for the process's
|
||||
+ // life (threshold u64::MAX), so the prover holds its high-water mark between shards on a card it shares
|
||||
+ // with a miner. `SP1_GPU_MEM_RELEASE_THRESHOLD=<bytes>` sets the pool's release threshold (0 returns
|
||||
+ // freed memory to the driver at once); unset, upstream's behaviour.
|
||||
+ let mem_release_threshold = std::env::var("SP1_GPU_MEM_RELEASE_THRESHOLD")
|
||||
+ .ok()
|
||||
+ .and_then(|s| s.parse::<u64>().ok())
|
||||
+ .unwrap_or(u64::MAX);
|
||||
+ Self { capacity: None, device: CudaDevice(0), mem_release_threshold }
|
||||
}
|
||||
|
||||
pub fn num_tasks(mut self, num_tasks: usize) -> Self {
|
||||
diff --git a/sp1-gpu/crates/jagged_tracegen/src/lib.rs b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
|
||||
index 579f70a..09e73e8 100644
|
||||
index 579f70a..2264044 100644
|
||||
--- a/sp1-gpu/crates/jagged_tracegen/src/lib.rs
|
||||
+++ b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
|
||||
@@ -481,6 +481,33 @@ async fn device_preprocessed_tracegen<A: CudaTracegenAir<Felt>>(
|
||||
|
|
@ -64,7 +85,81 @@ index 579f70a..09e73e8 100644
|
|||
log_stacking_height,
|
||||
max_log_row_count,
|
||||
backend,
|
||||
@@ -984,9 +1021,15 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
|
||||
@@ -906,6 +943,11 @@ pub async fn main_tracegen<GC: IopCtx<F = Felt>, A: CudaTracegenAir<Felt>>(
|
||||
|
||||
log_chip_stats(machine, &chip_set, &traces);
|
||||
|
||||
+ // Igneum prover-floor patch: the key's buffer is sized to its preprocessed traces at setup (upstream sized it
|
||||
+ // for a whole shard), so grow it here to what this shard needs before the main traces are appended: a bigger
|
||||
+ // dense buffer and column index, the preprocessed region copied device to device, swapped into the key.
|
||||
+ grow_for_main(&mut jagged_traces.preprocessed_traces, &traces, log_stacking_height, backend);
|
||||
+
|
||||
copy_main_jagged_traces(
|
||||
traces,
|
||||
&mut jagged_traces.preprocessed_traces,
|
||||
@@ -918,6 +960,61 @@ pub async fn main_tracegen<GC: IopCtx<F = Felt>, A: CudaTracegenAir<Felt>>(
|
||||
(public_values, chip_set, permit)
|
||||
}
|
||||
|
||||
+/// Igneum prover-floor patch: see `main_tracegen`. The need is the preprocessed phase as laid out (its padded
|
||||
+/// end, `preprocessed_offset`) plus the main traces padded to the stacking height plus one stacking height of
|
||||
+/// slack; a buffer at least that big is left alone. The process aborts, loudly, if the copy cannot be made,
|
||||
+/// because a panic inside a prover task is what left sweep 2 hanging on the client's socket.
|
||||
+fn grow_for_main(
|
||||
+ jagged: &mut JaggedTraceMle<Felt, TaskScope>,
|
||||
+ main_traces: &BTreeMap<String, Trace<TaskScope>>,
|
||||
+ log_stacking_height: u32,
|
||||
+ backend: &TaskScope,
|
||||
+) {
|
||||
+ let pre_end = jagged.dense().preprocessed_offset;
|
||||
+ let needed = pre_end
|
||||
+ + padded_trace_elements(main_traces, log_stacking_height)
|
||||
+ + (1 << log_stacking_height);
|
||||
+ let have = jagged.dense().dense.capacity();
|
||||
+ if have >= needed {
|
||||
+ return;
|
||||
+ }
|
||||
+ let mut new_dense: Buffer<Felt, TaskScope> = Buffer::with_capacity_in(needed, backend.clone());
|
||||
+ let mut new_col_index: Buffer<u32, TaskScope> =
|
||||
+ Buffer::with_capacity_in(needed >> 1, backend.clone());
|
||||
+ unsafe {
|
||||
+ new_dense.assume_init();
|
||||
+ new_col_index.assume_init();
|
||||
+ }
|
||||
+ {
|
||||
+ let JaggedMle { dense_data, col_index, .. } = &mut **jagged;
|
||||
+ let src_dense: &Slice<_, _> = &dense_data.dense[..pre_end];
|
||||
+ let dst_dense: &mut Slice<_, _> = &mut new_dense[..pre_end];
|
||||
+ let src_col: &Slice<_, _> = &col_index[..pre_end >> 1];
|
||||
+ let dst_col: &mut Slice<_, _> = &mut new_col_index[..pre_end >> 1];
|
||||
+ unsafe {
|
||||
+ if dst_dense.copy_from_slice(src_dense, backend).is_err()
|
||||
+ || dst_col.copy_from_slice(src_col, backend).is_err()
|
||||
+ {
|
||||
+ eprintln!("FLOOR grow FAILED: could not copy the preprocessed region ({pre_end} elements) into the grown buffer ({needed} elements); aborting instead of hanging");
|
||||
+ std::process::abort();
|
||||
+ }
|
||||
+ }
|
||||
+ }
|
||||
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
|
||||
+ eprintln!(
|
||||
+ "FLOOR grow key buffer {have} -> {needed} elements (preprocessed {pre_end}, {} bytes)",
|
||||
+ needed * 6
|
||||
+ );
|
||||
+ }
|
||||
+ let JaggedMle { dense_data, col_index, .. } = &mut **jagged;
|
||||
+ dense_data.dense = new_dense;
|
||||
+ *col_index = new_col_index;
|
||||
+ unsafe {
|
||||
+ dense_data.dense.set_len(pre_end);
|
||||
+ col_index.set_len(pre_end >> 1);
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
pub async fn main_tracegen_permit<GC: IopCtx<F = Felt>, A: CudaTracegenAir<Felt>>(
|
||||
machine: &Machine<Felt, A>,
|
||||
@@ -984,9 +1081,15 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
|
||||
|
||||
log_chip_stats(machine, &chip_set, &main_traces);
|
||||
|
||||
|
|
@ -81,7 +176,7 @@ index 579f70a..09e73e8 100644
|
|||
log_stacking_height,
|
||||
max_log_row_count,
|
||||
backend,
|
||||
@@ -1002,6 +1045,18 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
|
||||
@@ -1002,6 +1105,18 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
|
||||
)
|
||||
.await;
|
||||
|
||||
|
|
@ -104,15 +199,16 @@ diff --git a/sp1-gpu/crates/prover_components/src/builder.rs b/sp1-gpu/crates/pr
|
|||
index 5dccd9d..574d4fa 100644
|
||||
--- a/sp1-gpu/crates/prover_components/src/builder.rs
|
||||
+++ b/sp1-gpu/crates/prover_components/src/builder.rs
|
||||
@@ -23,28 +23,75 @@ use crate::{
|
||||
@@ -23,28 +23,124 @@ use crate::{
|
||||
SP1CudaProverComponents,
|
||||
};
|
||||
|
||||
+/// Igneum prover-floor patch (5 October 2026). Upstream sizes every device buffer for a 24 GB card or larger
|
||||
+/// and panics below that, whatever the shard. Here the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks a
|
||||
+/// tier, and `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly.
|
||||
+/// The proof format, the verifier and the program ids do not change: the element threshold only decides where
|
||||
+/// the executor splits shards, as upstream's own 24 GB tier already does.
|
||||
+/// Igneum prover-floor patch (5 October 2026; the memory rules of V6-07, 8 October 2026). Upstream sizes every
|
||||
+/// device buffer for a 24 GB card or larger and panics below that, whatever the shard. Here the card's FREE memory
|
||||
+/// at start (or `SP1_GPU_MEMORY_BUDGET_GB`, the host's lease) picks a tier, read once for the process, and
|
||||
+/// `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly. The proof
|
||||
+/// format, the verifier and the program ids do not change: the element threshold only decides where the executor
|
||||
+/// splits shards, as upstream's own 24 GB tier already does.
|
||||
+fn env_usize(name: &str) -> Option<usize> {
|
||||
+ std::env::var(name).ok().and_then(|s| s.parse::<usize>().ok())
|
||||
+}
|
||||
|
|
@ -121,8 +217,17 @@ index 5dccd9d..574d4fa 100644
|
|||
+ std::env::var(name).ok().and_then(|s| s.parse::<f64>().ok())
|
||||
+}
|
||||
+
|
||||
+/// The core element threshold for a memory budget in GB (upstream's own figure for the budget, +4, as it
|
||||
+/// computed it: a 32 GB card is 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16).
|
||||
+/// The recursion trace allocation under the 24 GB tier (V6-07): a recursion key or shard uses 90,177,536 elements
|
||||
+/// (35.6 M preprocessed and 54.5 M main; the floor sweep of 5 October 2026), so 2^26 + 2^25 = 100,663,296 holds it
|
||||
+/// with the stacking slack `floor_capacity` adds; the pinned host copies (four per prover) shrink with it. Upstream's
|
||||
+/// 2^27 stays on the 24 GB tier and above. Two distinct values: `floor_tests::recursion_branches_differ`.
|
||||
+pub const RECURSION_TRACE_ALLOCATION_SMALL: usize = (1 << 26) + (1 << 25);
|
||||
+
|
||||
+/// The core element threshold for a memory budget in GB (the budget is the FREE memory, +4, as upstream computed
|
||||
+/// its tiers from the total: a 32 GB card alone reads 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16, an
|
||||
+/// 8 GB card 12). Under the 18 tier the threshold is 2^26, the value every passing small-card row used (the
|
||||
+/// alone-comp-26-v1 rows of 6 October 2026 on the 3060, 3080, 4060, 4060 Ti, 4070 and 5070; the 8 October 2026
|
||||
+/// proof_alone rows on the 3060 at 7,525 MiB and the 4060 at 7,532 MiB); no small-card row passed at 2^27 here.
|
||||
+pub fn element_threshold_for_budget(gpu_memory_gb: usize, full_size_shards: bool) -> u64 {
|
||||
+ if gpu_memory_gb > 30 || (full_size_shards && gpu_memory_gb >= 24) {
|
||||
+ ELEMENT_THRESHOLD
|
||||
|
|
@ -131,31 +236,69 @@ index 5dccd9d..574d4fa 100644
|
|||
+ } else if gpu_memory_gb >= 18 {
|
||||
+ (1 << 27) + (1 << 26)
|
||||
+ } else {
|
||||
+ 1 << 27
|
||||
+ 1 << 26
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+/// The recursion trace allocation (elements) for a memory budget.
|
||||
+/// The recursion trace allocation (elements) for a memory budget: upstream's on the 24 GB tier and above, the
|
||||
+/// small constant under it.
|
||||
+pub fn recursion_trace_allocation_for_budget(gpu_memory_gb: usize) -> usize {
|
||||
+ if gpu_memory_gb >= 24 {
|
||||
+ RECURSION_TRACE_ALLOCATION
|
||||
+ } else {
|
||||
+ RECURSION_TRACE_ALLOCATION
|
||||
+ RECURSION_TRACE_ALLOCATION_SMALL
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+/// The memory budget in GB, read ONCE for the process (the core opts and the recursion prover are built at
|
||||
+/// different moments, after allocations that lower the free figure; one reading keeps both on one tier):
|
||||
+/// `SP1_GPU_MEMORY_BUDGET_GB` when set (the host's lease), else the device's FREE memory as the driver reports it
|
||||
+/// at the first call, never the total (a card beside a miner has the miner's resident set gone), +4 as upstream
|
||||
+/// computed its tiers.
|
||||
+pub fn gpu_memory_gb() -> usize {
|
||||
+ let gb = 1024.0 * 1024.0 * 1024.0;
|
||||
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
|
||||
+ Some(b) => (b.ceil() as usize) + 4,
|
||||
+ None => (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4,
|
||||
+ }
|
||||
+ static BUDGET: std::sync::OnceLock<usize> = std::sync::OnceLock::new();
|
||||
+ *BUDGET.get_or_init(|| {
|
||||
+ let gb = 1024.0 * 1024.0 * 1024.0;
|
||||
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
|
||||
+ Some(b) => (b.ceil() as usize) + 4,
|
||||
+ None => {
|
||||
+ let (free, _total) = cuda_memory_info().unwrap();
|
||||
+ (((free as f64) / gb).ceil() as usize) + 4
|
||||
+ }
|
||||
+ }
|
||||
+ })
|
||||
+}
|
||||
+
|
||||
+pub fn recursion_trace_allocation() -> usize {
|
||||
+ env_usize("SP1_GPU_RECURSION_TRACE_ALLOCATION")
|
||||
+ .unwrap_or_else(|| recursion_trace_allocation_for_budget(gpu_memory_gb()))
|
||||
+}
|
||||
+
|
||||
+#[cfg(test)]
|
||||
+mod floor_tests {
|
||||
+ use super::*;
|
||||
+
|
||||
+ #[test]
|
||||
+ fn small_tier_is_two_to_the_26() {
|
||||
+ assert_eq!(element_threshold_for_budget(12, false), 1 << 26, "an 8 GB card reads 12");
|
||||
+ assert_eq!(element_threshold_for_budget(16, false), 1 << 26, "a 12 GB card reads 16");
|
||||
+ assert_eq!(element_threshold_for_budget(17, false), 1 << 26);
|
||||
+ assert_eq!(element_threshold_for_budget(20, false), (1 << 27) + (1 << 26), "a 16 GB card reads 20");
|
||||
+ assert_eq!(element_threshold_for_budget(28, false), ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24));
|
||||
+ assert_eq!(element_threshold_for_budget(28, true), ELEMENT_THRESHOLD);
|
||||
+ assert_eq!(element_threshold_for_budget(36, false), ELEMENT_THRESHOLD);
|
||||
+ }
|
||||
+
|
||||
+ #[test]
|
||||
+ fn recursion_branches_differ() {
|
||||
+ assert_ne!(recursion_trace_allocation_for_budget(28), recursion_trace_allocation_for_budget(16));
|
||||
+ assert_eq!(recursion_trace_allocation_for_budget(28), RECURSION_TRACE_ALLOCATION);
|
||||
+ assert_eq!(recursion_trace_allocation_for_budget(16), RECURSION_TRACE_ALLOCATION_SMALL);
|
||||
+ assert_eq!(RECURSION_TRACE_ALLOCATION_SMALL, 100_663_296);
|
||||
+ assert!(RECURSION_TRACE_ALLOCATION_SMALL < RECURSION_TRACE_ALLOCATION);
|
||||
+ assert!(RECURSION_TRACE_ALLOCATION_SMALL >= 90_177_536 + (1 << 22));
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
pub fn local_gpu_opts() -> SP1CoreOpts {
|
||||
let mut opts = SP1CoreOpts::default();
|
||||
|
|
@ -171,7 +314,7 @@ index 5dccd9d..574d4fa 100644
|
|||
- if gpu_memory_gb < 24 {
|
||||
- panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB");
|
||||
- }
|
||||
+ // The card's memory plus 4, as upstream computed it (a 32 GB card reads 36), or the budget given.
|
||||
+ // The card's FREE memory plus 4, as upstream computed its tiers from the total, or the budget given.
|
||||
+ let gpu_memory_gb = gpu_memory_gb();
|
||||
|
||||
- let shard_threshold = if !opts.full_size_shards && gpu_memory_gb <= 30 {
|
||||
|
|
@ -183,17 +326,18 @@ index 5dccd9d..574d4fa 100644
|
|||
+ None => element_threshold_for_budget(gpu_memory_gb, opts.full_size_shards),
|
||||
};
|
||||
+ let height_threshold = opts.sharding_threshold.height_threshold;
|
||||
+ let (free_mib, total_mib) = cuda_memory_info().map(|(f, t)| (f >> 20, t >> 20)).unwrap_or((0, 0));
|
||||
|
||||
- tracing::debug!("Shard threshold: {shard_threshold}");
|
||||
+ eprintln!(
|
||||
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
|
||||
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} free_mib={free_mib} total_mib={total_mib} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
|
||||
+ recursion_trace_allocation(),
|
||||
+ opts.full_size_shards
|
||||
+ );
|
||||
opts.sharding_threshold.element_threshold = shard_threshold;
|
||||
|
||||
opts.global_dependencies_opt = true;
|
||||
@@ -92,7 +139,7 @@ pub async fn recursion_prover_and_verifier(
|
||||
@@ -92,7 +188,7 @@ pub async fn recursion_prover_and_verifier(
|
||||
) {
|
||||
let recursion_verifier = SP1CudaProverComponents::compress_verifier();
|
||||
(
|
||||
|
|
@ -202,6 +346,86 @@ index 5dccd9d..574d4fa 100644
|
|||
.await,
|
||||
recursion_verifier,
|
||||
)
|
||||
diff --git a/sp1-gpu/crates/server/src/main.rs b/sp1-gpu/crates/server/src/main.rs
|
||||
index 65e94f5..498357c 100644
|
||||
--- a/sp1-gpu/crates/server/src/main.rs
|
||||
+++ b/sp1-gpu/crates/server/src/main.rs
|
||||
@@ -17,9 +17,55 @@ struct Args {
|
||||
version: bool,
|
||||
}
|
||||
|
||||
+/// Igneum prover-floor patch (6 October 2026, the GPU fleet's finding): a panic inside a prover task (an allocation
|
||||
+/// the card cannot meet, `cudaMallocAsync` failing and `Buffer::with_capacity_in` panicking in a tokio worker) left
|
||||
+/// the request's future waiting for ever, the client on its socket, the card at 0% for the 15 minutes until someone
|
||||
+/// killed it (the 8 GB and 10 GB cards at threshold 2^27). The server must fail the shard instead: this hook names
|
||||
+/// the stage from the panic's location and exits, so the client's proof fails at once with the server's last line.
|
||||
+fn stage_of(location: &str) -> &'static str {
|
||||
+ let l = location.to_ascii_lowercase();
|
||||
+ if l.contains("jagged_tracegen") || l.contains("/tracegen") {
|
||||
+ "trace generation"
|
||||
+ } else if l.contains("commit") || l.contains("basefold") || l.contains("merkle") {
|
||||
+ "the commit (codewords and Merkle trees)"
|
||||
+ } else if l.contains("logup_gkr") {
|
||||
+ "LogUp GKR"
|
||||
+ } else if l.contains("zerocheck") {
|
||||
+ "the zerocheck"
|
||||
+ } else if l.contains("jagged") {
|
||||
+ "the jagged sumcheck"
|
||||
+ } else if l.contains("prover_components") || l.contains("recursion") || l.contains("sp1-prover") || l.contains("sp1_prover") {
|
||||
+ "the recursion (compression)"
|
||||
+ } else if l.contains("cuda") || l.contains("slop") {
|
||||
+ "a device allocation"
|
||||
+ } else {
|
||||
+ "the prover"
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+fn install_abort_on_panic() {
|
||||
+ std::panic::set_hook(Box::new(|info| {
|
||||
+ let location = info.location().map(|l| format!("{}:{}", l.file(), l.line())).unwrap_or_else(|| "unknown".into());
|
||||
+ let message = info
|
||||
+ .payload()
|
||||
+ .downcast_ref::<&str>()
|
||||
+ .map(|s| s.to_string())
|
||||
+ .or_else(|| info.payload().downcast_ref::<String>().cloned())
|
||||
+ .unwrap_or_default();
|
||||
+ let oom = message.to_ascii_lowercase().contains("alloc") || message.contains("MemoryAllocation") || message.contains("OUT_OF_MEMORY");
|
||||
+ eprintln!(
|
||||
+ "FLOOR abort: {} failed at {location}: {message}{}; the server exits so the client's proof fails instead of waiting",
|
||||
+ stage_of(&location),
|
||||
+ if oom { " (the card's memory could not meet an allocation: lower SP1_GPU_ELEMENT_THRESHOLD one notch)" } else { "" }
|
||||
+ );
|
||||
+ std::process::exit(70);
|
||||
+ }));
|
||||
+}
|
||||
+
|
||||
#[tokio::main]
|
||||
#[allow(clippy::print_stdout)]
|
||||
async fn main() {
|
||||
+ install_abort_on_panic();
|
||||
tracing_subscriber::fmt::init();
|
||||
|
||||
let args = Args::parse();
|
||||
@@ -40,3 +86,19 @@ async fn main() {
|
||||
eprintln!("Error running server: {e}");
|
||||
}
|
||||
}
|
||||
+
|
||||
+#[cfg(test)]
|
||||
+mod floor_tests {
|
||||
+ use super::stage_of;
|
||||
+
|
||||
+ #[test]
|
||||
+ fn the_stage_is_named_from_the_panic_location() {
|
||||
+ assert_eq!(stage_of("sp1-gpu/crates/jagged_tracegen/src/lib.rs:240"), "trace generation");
|
||||
+ assert_eq!(stage_of("sp1-gpu/crates/basefold/src/fri.rs:97"), "the commit (codewords and Merkle trees)");
|
||||
+ assert_eq!(stage_of("sp1-gpu/crates/logup_gkr/src/tracegen.rs:72"), "LogUp GKR");
|
||||
+ assert_eq!(stage_of("sp1-gpu/crates/zerocheck/src/prover.rs:1163"), "the zerocheck");
|
||||
+ assert_eq!(stage_of("sp1-gpu/crates/prover_components/src/builder.rs:70"), "the recursion (compression)");
|
||||
+ assert_eq!(stage_of("sp1-gpu/crates/cuda/src/stream.rs:330"), "a device allocation");
|
||||
+ assert_eq!(stage_of("somewhere/else.rs:1"), "the prover");
|
||||
+ }
|
||||
+}
|
||||
diff --git a/sp1-gpu/crates/server/src/server.rs b/sp1-gpu/crates/server/src/server.rs
|
||||
index 4035f1f..0d0d907 100644
|
||||
--- a/sp1-gpu/crates/server/src/server.rs
|
||||
|
|
|
|||
69
tools/ci/batches/202-a284380b-5f50afd1.json
Normal file
69
tools/ci/batches/202-a284380b-5f50afd1.json
Normal file
|
|
@ -0,0 +1,69 @@
|
|||
{
|
||||
"run_id": "202-a284380b-5f50afd1",
|
||||
"manifest_sha": "a284380b",
|
||||
"cut_tip": "5f50afd1 (miner; pow and app cells on 77b5acc5, no crate move since); node a284380b on the key-succession pairing with the class v5 freeze (fingerprint cbc5bd0a)",
|
||||
"evidence_dir": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1",
|
||||
"boxes": [
|
||||
"build-4"
|
||||
],
|
||||
"method": "native",
|
||||
"note": "2.0.2 pair gate on build-4, 22:03 UTC: release-manifest-check green on 5f50afd1 (node block release-2.0.2-node a284380b, the 2.0.1 pairing); 77b5acc5..5f50afd1 moves pool/ (2 files) and packaging only Every suite cell maps its cases partially (the map's coverage words), so the green suites read NOT RUN in progress with their logs as evidence, never PASS by inference (the standard's rule; the 2.0.1 batch 201-7cfa422a-aa0e0f45 reads the same way).",
|
||||
"cells": [
|
||||
{
|
||||
"cell": "suite:pow",
|
||||
"status": "NOT RUN",
|
||||
"evidence": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/miner-77b5acc5/box4-pow.log",
|
||||
"note": "suite GREEN (rc 0, every test passed); the cell covers its cases partially, so the case reads NOT RUN in progress with this evidence until its own acceptance is exercised; pow on release-2.0.2 77b5acc5 (box4, 226s); 77b5acc5..5f50afd1 moves no crate, so the cell covers the tip by content",
|
||||
"in_progress": true
|
||||
},
|
||||
{
|
||||
"cell": "suite:app",
|
||||
"status": "NOT RUN",
|
||||
"evidence": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/miner-77b5acc5/box4-app.log",
|
||||
"note": "suite GREEN (rc 0, every test passed); the cell covers its cases partially, so the case reads NOT RUN in progress with this evidence until its own acceptance is exercised; app on release-2.0.2 77b5acc5 (box4, 39s); 77b5acc5..5f50afd1 moves no crate, so the cell covers the tip by content",
|
||||
"in_progress": true
|
||||
},
|
||||
{
|
||||
"cell": "check:freeze",
|
||||
"status": "NOT RUN",
|
||||
"evidence": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-kaspad-check.log",
|
||||
"note": "suite GREEN (rc 0, every test passed); the cell covers its cases partially, so the case reads NOT RUN in progress with this evidence until its own acceptance is exercised; kaspad-check on release-2.0.2-node a284380b with the key-succession parent and the class v5 freeze igneum-pow (box4, 35s)",
|
||||
"in_progress": true
|
||||
},
|
||||
{
|
||||
"cell": "suite:core",
|
||||
"status": "NOT RUN",
|
||||
"evidence": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-core.log",
|
||||
"note": "suite GREEN (rc 0, every test passed); the cell covers its cases partially, so the case reads NOT RUN in progress with this evidence until its own acceptance is exercised; core on release-2.0.2-node a284380b with the key-succession parent and the class v5 freeze igneum-pow (box4, 26s)",
|
||||
"in_progress": true
|
||||
},
|
||||
{
|
||||
"cell": "suite:exec",
|
||||
"status": "NOT RUN",
|
||||
"evidence": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-exec.log",
|
||||
"note": "suite GREEN (rc 0, every test passed); the cell covers its cases partially, so the case reads NOT RUN in progress with this evidence until its own acceptance is exercised; exec on release-2.0.2-node a284380b with the key-succession parent and the class v5 freeze igneum-pow (box4, 45s)",
|
||||
"in_progress": true
|
||||
},
|
||||
{
|
||||
"cell": "suite:miner",
|
||||
"status": "NOT RUN",
|
||||
"evidence": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-miner.log",
|
||||
"note": "suite GREEN (rc 0, every test passed); the cell covers its cases partially, so the case reads NOT RUN in progress with this evidence until its own acceptance is exercised; miner on release-2.0.2-node a284380b with the key-succession parent and the class v5 freeze igneum-pow (box4, 45s)",
|
||||
"in_progress": true
|
||||
},
|
||||
{
|
||||
"cell": "suite:p2p-flows",
|
||||
"status": "NOT RUN",
|
||||
"evidence": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-p2p-flows.log",
|
||||
"note": "suite GREEN (rc 0, every test passed); the cell covers its cases partially, so the case reads NOT RUN in progress with this evidence until its own acceptance is exercised; p2p-flows on release-2.0.2-node a284380b with the key-succession parent and the class v5 freeze igneum-pow (box4, 46s)",
|
||||
"in_progress": true
|
||||
},
|
||||
{
|
||||
"cell": "suite:consensus",
|
||||
"status": "NOT RUN",
|
||||
"evidence": "build-1:/srv/artefacts/tas/202-a284380b-5f50afd1/node-a284380b/box4-consensus.log",
|
||||
"note": "suite GREEN (rc 0, every test passed); the cell covers its cases partially, so the case reads NOT RUN in progress with this evidence until its own acceptance is exercised; consensus on release-2.0.2-node a284380b with the key-succession parent and the class v5 freeze igneum-pow (box4, 45s)",
|
||||
"in_progress": true
|
||||
}
|
||||
]
|
||||
}
|
||||
|
|
@ -0,0 +1,6 @@
|
|||
{
|
||||
"run_id": "census-v6-chainseed-trace-20261008-2305",
|
||||
"manifest_sha": "1a938abe4",
|
||||
"evidence_dir": "docs/analysis/class-v6/rows/chainseed-census-1a938abe4",
|
||||
"cells": [ { "cell": "census:class-v6", "status": "PASS", "method": "native", "evidence": "docs/analysis/class-v6/rows/chainseed-census-1a938abe4.md" } ]
|
||||
}
|
||||
|
|
@ -20,7 +20,7 @@
|
|||
},
|
||||
"claim_impact": "R2-F03-R02 and R03 on one job context (the chain-seed class v6 object) across the readers that can run it tonight: node, CPU reference, CUDA and pool on three phases with the dataset-day switch at nonce 1572864; PASS only when every reader agrees bit for bit on every nonce; Metal and OpenCL are the morning's",
|
||||
"network_label": "on an island, not a network",
|
||||
"note": "The same-work run (the coordinator's ruling 22:0x UK: six readers on one context; the frozen-object context 01 is the CUDA and CPU-only row). The node1 state stream (block af89be5d..., written on build-4 at 15:34 UK, before the 20:00 UK split) keys every item; every cell carries the island label by main's 22:2x UK rule. Day D = 20730 (bytes ...2ffa50), D+1 = 20731 (...2ffb50); phase 2 switches at 1572864 (the midpoint, 32-aligned); genesis day index 20730 and dataset 2^28 words for the engine's growth rule (doublings 0 on both days). CPU reference: igneum-pow hash-bound (class-v6 1a938abe4, binary 0d3f3fca...) on build-4, 21:48 to 21:57 UK, ref-D 7c7be9b5... over [0, 1572864), ref-D1 cad7849f... over [1572864, 3145728); the hash lane's independent million-nonce reference hl-v6-all-cs.txt (b77c61d8..., 22:03 UK) equals ref-D's first million lines byte for byte. Node reader: PASS on all three phases (22:17, 22:21, 22:25 UK; 1,048,576 agree each, 0 disagree, 0 missing; the boundary switched: d28fcfd70ba70a64 before, 45905ff3d608dc12 after, both equal to the reference; the program id unchanged across the day). Reader faults recorded, not disagreements (the coordinator's ruling 22:1x UK): the node reader's first phase 1 at 22:10 UK read FAIL on every nonce with the program id equal because the standalone engine held the genesis day at 0 and the growth rule doubled the cache fourteen times (fixed at cc23f9bd: install_pow_genesis from the context and the pack, the live miner's own call); the pool lane's D1-pairing run at 22:17 UK hit the same class (its nonce 0 and 1 hashes equal the pre-fix node hashes byte for byte), fix sent 22:18 UK; the rerun on the fixed reader (pool-recheck-202 909ae65b, binary 10425cc9... on build-4, the pack geometry installed) reads PASS on every phase at 22:32, 22:38 and 22:44 UK: 1,048,576 agree each, 0 disagree, 0 missing, the boundary switched at 1572864 with one program id on both segments, 64, 128 and 64 of the sampled shares accepted through verify::check; every nonce of 3,145,728 agrees with the reference (pool/d1-pairing/pool-{1,2,3}.json, sha256 d6d6105f..., 611c7e75..., 44820ff8..., read back on build-1 at 22:45 UK; the README keeps the first run's FAIL and its cause). The pool on the 2.0.2 kit as pinned (release-2.0.2 00c7e1db, node 7cfa422a) is BLOCKED: its node's ProgramClass enum ends at V5, so a class=v6 job line is unnameable (pool/kit-row/pool-{1,2,3}.json carry that refusal in the driver's shape). CUDA: the 5090's phase 1 PASS at 22:00 and 22:24 UK (gen 6 worker 804a6f7f, 47.6 s); phases 2 and 3 running at the time of this record (the fleet lane's runner restarted at 22:22 UK for the three-context file; its stray phase 2 and 3 lines at 22:23 were the restart's artefact, overwritten by the clean sequence). Rule 24 on same-work-node cc23f9bd on build-2: crate check rc 0, igneum-miner suite 30 passed 0 failed.",
|
||||
"note": "The same-work run (the coordinator's ruling 22:0x UK: six readers on one context; the frozen-object context 01 is the CUDA and CPU-only row). The node1 state stream (block af89be5d..., written on build-4 at 15:34 UK, before the 20:00 UK split) keys every item; every cell carries the island label by main's 22:2x UK rule. Day D = 20730 (bytes ...2ffa50), D+1 = 20731 (...2ffb50); phase 2 switches at 1572864 (the midpoint, 32-aligned); genesis day index 20730 and dataset 2^28 words for the engine's growth rule (doublings 0 on both days). CPU reference: igneum-pow hash-bound (class-v6 1a938abe4, binary 0d3f3fca...) on build-4, 21:48 to 21:57 UK, ref-D 7c7be9b5... over [0, 1572864), ref-D1 cad7849f... over [1572864, 3145728); the hash lane's independent million-nonce reference hl-v6-all-cs.txt (b77c61d8..., 22:03 UK) equals ref-D's first million lines byte for byte. Node reader: PASS on all three phases (22:17, 22:21, 22:25 UK; 1,048,576 agree each, 0 disagree, 0 missing; the boundary switched: d28fcfd70ba70a64 before, 45905ff3d608dc12 after, both equal to the reference; the program id unchanged across the day). Reader faults recorded, not disagreements (the coordinator's ruling 22:1x UK): the node reader's first phase 1 at 22:10 UK read FAIL on every nonce with the program id equal because the standalone engine held the genesis day at 0 and the growth rule doubled the cache fourteen times (fixed at cc23f9bd: install_pow_genesis from the context and the pack, the live miner's own call); the pool lane's D1-pairing run at 22:17 UK hit the same class (its nonce 0 and 1 hashes equal the pre-fix node hashes byte for byte), fix sent 22:18 UK; the rerun on the fixed reader (pool-recheck-202 909ae65b, binary 10425cc9... on build-4, the pack geometry installed) reads PASS on every phase at 22:32, 22:38 and 22:44 UK: 1,048,576 agree each, 0 disagree, 0 missing, the boundary switched at 1572864 with one program id on both segments, 64, 128 and 64 of the sampled shares accepted through verify::check; every nonce of 3,145,728 agrees with the reference (pool/d1-pairing/pool-{1,2,3}.json, sha256 d6d6105f..., 611c7e75..., 44820ff8..., read back on build-1 at 22:45 UK; the README keeps the first run's FAIL and its cause). The pool on the 2.0.2 kit as pinned (release-2.0.2 00c7e1db, node 7cfa422a) is BLOCKED: its node's ProgramClass enum ends at V5, so a class=v6 job line is unnameable (pool/kit-row/pool-{1,2,3}.json carry that refusal in the driver's shape). CUDA (the gen 6 worker 804a6f7f on the pods, the driver from same-work-harness 76cb1f7c6 and later): the 5090 phase 1 PASS (22:53 UK, cuda-5090-1.json 654ddddb...) and phase 2 PASS (22:55 UK, cuda-5090-2.json 898b1cb3...: the day switch at 1572864 driven through the worker's prepare, both boundary hashes equal to the reference); the 4090 phase 1 PASS (22:57 UK, cuda-4090-1.json c91d9e16...) and phase 2 PASS (22:59 UK, cuda-4090-2.json 82a2f5e6...); the 3090 running. Two driver faults on the way, neither a disagreement: the driver's silent wait for the prepared line deadlocked against a worker that loads a finished prepare only between stdin lines (22:26 to 22:47 UK on the 5090; fixed at 76cb1f7c6: day D jobs keep flowing and the worker is polled with a line a second), and phase 3 on a worker started on the day D pack had every day D+1 job refused as a day-seed mismatch (22:56 UK, cuda-5090-3.json and cuda-4090-3.json: missing 1,048,576, the program id equal; fixed in the commit after b464021f8: the pack the worker holds is read from its ready line and every other segment is prepared); phase 3 reruns on the fixed driver. Rule 24 on same-work-node cc23f9bd on build-2: crate check rc 0, igneum-miner suite 30 passed 0 failed.",
|
||||
"cells": [
|
||||
{
|
||||
"cell": "harness:same-work",
|
||||
|
|
@ -29,8 +29,8 @@
|
|||
],
|
||||
"status": "NOT RUN",
|
||||
"method": "GPU",
|
||||
"evidence": "build-1:/srv/artefacts/tas/same-work-20261008-02/node/node-1.json; build-1:/srv/artefacts/tas/same-work-20261008-02/node/node-2.json; build-1:/srv/artefacts/tas/same-work-20261008-02/node/node-3.json; build-1:/srv/artefacts/tas/same-work-20261008-02/node/node.log; build-1:/srv/artefacts/tas/same-work-20261008-02/references/ref-D.txt; build-1:/srv/artefacts/tas/same-work-20261008-02/references/ref-D1.txt; build-1:/srv/artefacts/tas/same-work-20261008-02/job-context.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/kit-row/pool-1.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/kit-row/pool-2.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/kit-row/pool-3.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/d1-pairing/pool-1.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/d1-pairing/pool-2.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/d1-pairing/pool-3.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/d1-pairing/README.json; build-1:/srv/artefacts/tas/same-work-20261008-02/cuda-5090-1.json",
|
||||
"note": "node PASS x3; pool D1 pairing PASS x3 (every nonce, the share verifier's own path, the day switch seen); CPU reference written and cross-checked twice; CUDA 5090 phase 1 PASS, phases 2 and 3 running; pool kit row BLOCKED by its V5 pins; OpenCL queued on the project's first rig (08:30 UK at the latest, the job run-ca3-pc1-samework-opencl-7600-20261008, its report staged as opencl-rx7600.report.txt); Metal NOT RUN, 08:30 UK on the project's macOS machine if a serve worker exists by 08:00 UK; the case stays NOT RUN in progress until every reader has read",
|
||||
"evidence": "build-1:/srv/artefacts/tas/same-work-20261008-02/node/node-1.json; build-1:/srv/artefacts/tas/same-work-20261008-02/node/node-2.json; build-1:/srv/artefacts/tas/same-work-20261008-02/node/node-3.json; build-1:/srv/artefacts/tas/same-work-20261008-02/node/node.log; build-1:/srv/artefacts/tas/same-work-20261008-02/references/ref-D.txt; build-1:/srv/artefacts/tas/same-work-20261008-02/references/ref-D1.txt; build-1:/srv/artefacts/tas/same-work-20261008-02/job-context.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/kit-row/pool-1.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/kit-row/pool-2.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/kit-row/pool-3.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/d1-pairing/pool-1.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/d1-pairing/pool-2.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/d1-pairing/pool-3.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/d1-pairing/README.json; build-1:/srv/artefacts/tas/same-work-20261008-02/cuda-5090-1.json; docs/design/class-v5-stored-state.md; build-1:/srv/artefacts/tas/same-work-20261008-02/cuda-5090-2.json; build-1:/srv/artefacts/tas/same-work-20261008-02/cuda-4090-1.json; build-1:/srv/artefacts/tas/same-work-20261008-02/cuda-4090-2.json",
|
||||
"note": "node PASS x3; pool D1 pairing PASS x3 (every nonce, the share verifier's own path, the day switch seen); CPU reference written and cross-checked twice; CUDA 5090 and 4090 phases 1 and 2 PASS (the day switch seen by a GPU reader), phase 3 rerunning on the fixed driver, the 3090 running; pool kit row BLOCKED by its V5 pins; OpenCL queued on the project's first rig (08:30 UK at the latest, the job run-ca3-pc1-samework-opencl-7600-20261008, its report staged as opencl-rx7600.report.txt); Metal NOT RUN, 08:30 UK on the project's macOS machine if a serve worker exists by 08:00 UK; the case stays NOT RUN in progress until every reader has read; the design page docs/design/class-v5-stored-state.md (the node lane's fix-2 amendment, class-v5-foreign-capture-doc 4ec69ca4) moves with this row",
|
||||
"network_label": "on an island, not a network",
|
||||
"in_progress": true
|
||||
},
|
||||
|
|
@ -41,8 +41,8 @@
|
|||
],
|
||||
"status": "NOT RUN",
|
||||
"method": "GPU",
|
||||
"evidence": "build-1:/srv/artefacts/tas/same-work-20261008-02/node/node-1.json; build-1:/srv/artefacts/tas/same-work-20261008-02/node/node-2.json; build-1:/srv/artefacts/tas/same-work-20261008-02/node/node-3.json; build-1:/srv/artefacts/tas/same-work-20261008-02/node/node.log; build-1:/srv/artefacts/tas/same-work-20261008-02/references/ref-D.txt; build-1:/srv/artefacts/tas/same-work-20261008-02/references/ref-D1.txt; build-1:/srv/artefacts/tas/same-work-20261008-02/job-context.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/d1-pairing/pool-2.json",
|
||||
"note": "the dataset-day transition in phase 2: the node reader and the pool reader each switched program and dataset at nonce 1572864 with the program id unchanged and both boundary hashes equal to the reference; the CUDA phase 2 read is pending; in progress",
|
||||
"evidence": "build-1:/srv/artefacts/tas/same-work-20261008-02/node/node-1.json; build-1:/srv/artefacts/tas/same-work-20261008-02/node/node-2.json; build-1:/srv/artefacts/tas/same-work-20261008-02/node/node-3.json; build-1:/srv/artefacts/tas/same-work-20261008-02/node/node.log; build-1:/srv/artefacts/tas/same-work-20261008-02/references/ref-D.txt; build-1:/srv/artefacts/tas/same-work-20261008-02/references/ref-D1.txt; build-1:/srv/artefacts/tas/same-work-20261008-02/job-context.json; build-1:/srv/artefacts/tas/same-work-20261008-02/pool/d1-pairing/pool-2.json; build-1:/srv/artefacts/tas/same-work-20261008-02/cuda-5090-2.json; build-1:/srv/artefacts/tas/same-work-20261008-02/cuda-4090-2.json",
|
||||
"note": "the dataset-day transition in phase 2: the node reader, the pool reader, the 5090 and the 4090 each switched program and dataset at nonce 1572864 with the program id unchanged and both boundary hashes equal to the reference; the 3090's phase 2 and the OpenCL and Metal reads are pending; in progress",
|
||||
"network_label": "on an island, not a network",
|
||||
"in_progress": true
|
||||
}
|
||||
|
|
|
|||
32
tools/ci/batches/v607-floor-memory-20261008T2059Z.json
Normal file
32
tools/ci/batches/v607-floor-memory-20261008T2059Z.json
Normal file
|
|
@ -0,0 +1,32 @@
|
|||
{
|
||||
"run_id": "v607-floor-memory-20261008T2059Z",
|
||||
"manifest_sha": "c411bae9a",
|
||||
"evidence_dir": "docs/analysis/floor-memory-profile-2026-10-08.md",
|
||||
"boxes": [
|
||||
"vast:54902103",
|
||||
"vast:54906194",
|
||||
"vast:54909866",
|
||||
"vast:54909865",
|
||||
"vast:54913403",
|
||||
"build-3"
|
||||
],
|
||||
"cells": [
|
||||
{
|
||||
"cell": "rows:v607-floor-memory",
|
||||
"status": "RUNNING",
|
||||
"evidence": "docs/analysis/floor-memory-profile-2026-10-08.md",
|
||||
"note": "partial coverage of GPU-05, CAP-02 and UX-02 by the V6-07 rows (the 3060 and 4060 on the default unoverridden path, the refusal rows beside the miner); RUNNING by the map's rule since no case's coverage is full; the 4060's chain row is a FAIL at the aggregation (an 8 GB card is refused the chain before setup) and the master-host proof rows are NOT RUN (6dbd5d2f8 without the guest re-pin); island: none"
|
||||
}
|
||||
],
|
||||
"method": "GPU",
|
||||
"release_identity": {
|
||||
"commit": "c411bae9a (branch v607-floor-memory off cef5234b5)",
|
||||
"lockfile": "proving/igneum-prove/Cargo.lock at cef5234b5",
|
||||
"binary": "sp1-gpu-server db37c38b5feabdec5bbfd0298446da89f6a5d144125e399d6ab617f64f9cb61d (igneum-floor-sm8689-v607.tgz 9c02fcc6a3f045d3167c1d9cb6e22c584ea862189d11997754a57cbe1b0d355b); igneum-prove-host-0317 71bc2438856bb141; igneum-prove-host v607 66662219a3630772",
|
||||
"network_object": "none (fixture proofs on rented pods, no network)",
|
||||
"activation": "none",
|
||||
"profile_hashes": "floor patch 9098c3e5979d057031f43588668201d1f7a53ab17e980c1eaf896cd5e6f38575"
|
||||
},
|
||||
"claim_impact": "the 2.0 proving statement's memory condition per tier: a 12 GB card runs the whole segment path alone; an 8 GB card proves shards only (its aggregation fails) and is refused the chain before setup; both time-share beside the 5.5 GiB miner; the 16 GB and 24 GB tiers unmeasured tonight",
|
||||
"reviewer": ""
|
||||
}
|
||||
|
|
@ -86,7 +86,16 @@ def run(a):
|
|||
# jobs it can, never sends a job on a segment before that segment's prepared line, and when it has nothing to send it polls
|
||||
# the worker with a line a second (the worker answers "info ignored", or the prepared line first). One prepare is in flight at
|
||||
# a time (the worker refuses a second while one runs). A worker without prepare support cannot be driven across a boundary.
|
||||
held = {i: (i == 0 and not a.prepare_first) for i in range(len(seg_fields))}
|
||||
# the pack the worker holds is read from its ready line ("pack igneum-epoch/<seed>/day/<day bytes>", the CUDA and OpenCL
|
||||
# workers started on --pack); a segment on any other day is prepared, whichever phase (phase 3's worker started on the
|
||||
# day D pack refused every day D+1 job as a day-seed mismatch at 22:56 UK: missing 1,048,576, no disagreement). A ready
|
||||
# line without a pack label means the worker holds the first segment's pack unless --prepare-first says it holds none.
|
||||
rt = ready.split(); resident = next((t for t in rt[1:] if 'igneum-epoch/' in t and '/day/' in t), None)
|
||||
def holds(f):
|
||||
if a.prepare_first: return False
|
||||
if resident is None: return f is seg_fields[0][2]
|
||||
return resident.endswith('/day/' + f[1]) and ('/' + f[0] + '/') in resident
|
||||
held = {i: holds(seg_fields[i][2]) for i in range(len(seg_fields))}
|
||||
to_prepare = [i for i in range(len(seg_fields)) if not held[i]]; in_flight = None; prepare_sent = None; prepare_lines = []
|
||||
if to_prepare and ' prepare 0' in ready: p.kill(); return {'error': f'the worker has no prepare support (ready line: {ready}); a pack by prepare or a day switch needs it'}, 2
|
||||
def send_prepare():
|
||||
|
|
@ -166,7 +175,8 @@ def self_test():
|
|||
# a fake worker: hashes nonce n as (n * 0x9e3779b97f4a7c15) mod 2^64; nonce 7 wrong when WRONG=1; nonce 9 never answered when DROP=1
|
||||
w = os.path.join(d, 'worker.py')
|
||||
open(w, 'w').write('''import sys, os
|
||||
print("ready fake-gpu 0 prepare " + ("0" if os.environ.get("NOPREPARE") == "1" else "1"), flush=True)
|
||||
res0 = os.environ.get("RESIDENT", "cc" * 19)
|
||||
print("ready fake-gpu 0 pack " + ("igneum-epoch/" + "aa" * 32 + "/day/" + res0 if res0 != "none" else "none") + " prepare " + ("0" if os.environ.get("NOPREPARE") == "1" else "1"), flush=True)
|
||||
pending_prepare = None; prep_lines_left = 0
|
||||
resident = {os.environ.get("RESIDENT", "cc" * 19)} # the pack's day (the real worker starts on its --pack); another day needs prepare first, as the real worker (day seed mismatch otherwise)
|
||||
for line in sys.stdin:
|
||||
|
|
@ -214,8 +224,8 @@ for line in sys.stdin:
|
|||
r = subprocess.run([sys.executable, __file__, '--worker', sys.executable, '--worker-arg', w, '--job-context', jc, '--phase', str(phase), '--job-nonces', '8', '--out', out] + list(extra), capture_output=True, text=True, env=e, timeout=120)
|
||||
return r.returncode, (json.load(open(out)) if os.path.exists(out) else {}), r.stdout + r.stderr
|
||||
for ph, s0, pid in ((1, 0, '0x1'), (3, 32, '0x2')):
|
||||
rc, ev, o = gojc(ph, f'jc{ph}', env={'RESIDENT': 'dd' * 19} if ph == 3 else None) # phase 3's worker starts on the day D+1 pack
|
||||
if not (rc == 0 and ev.get('verdict') == 'PASS' and ev['nonces'] == {'start': s0, 'count': 16} and ev['segments'][0]['program_id'] == pid and 'boundary' not in ev): print(f"self-test failed: phase {ph} of the job context was not a PASS on its range and pack: rc={rc} {ev.get('nonces')} {ev.get('segments')} {o[-200:]}"); fails = 1
|
||||
rc, ev, o = gojc(ph, f'jc{ph}') # the worker holds the day D pack in every phase, as the pods run it; phase 3 needs the day D+1 prepare
|
||||
if not (rc == 0 and ev.get('verdict') == 'PASS' and ev['nonces'] == {'start': s0, 'count': 16} and ev['segments'][0]['program_id'] == pid and 'boundary' not in ev and (ph == 1) == ('prepare_lines' not in ev)): print(f"self-test failed: phase {ph} of the job context was not a PASS on its range and pack: rc={rc} {ev.get('nonces')} {ev.get('segments')} {o[-200:]}"); fails = 1
|
||||
rc, ev, o = gojc(2, 'jc2')
|
||||
b = ev.get('boundary', {})
|
||||
if not (rc == 0 and ev.get('verdict') == 'PASS' and ev['nonces'] == {'start': 16, 'count': 16} and b.get('nonce') == 24 and b.get('program_id_before') == '0x1' and b.get('program_id_after') == '0x2' and b.get('switched') is True and len(ev['segments']) == 2): print(f"self-test failed: phase 2 did not switch pack at the boundary nonce 24 with the boundary block: rc={rc} {b} {o[-200:]}"); fails = 1
|
||||
|
|
|
|||
|
|
@ -712,6 +712,25 @@
|
|||
"coverage": {
|
||||
"R2-F03-R02": "partial: the worker half (CPU reference, CUDA, Metal, OpenCL) on the same program packs; the node's and the pool's accepted work on the same job context are the CI steward's and the pool lane's cells; PASS only when every listed platform reads one fingerprint and the node and pool halves are green"
|
||||
}
|
||||
},
|
||||
"rows:v607-floor-memory": {
|
||||
"command": "the V6-07 lane's pod scripts (scratch v607/pod-v607.sh, ab.sh, ab2.sh, ab3.sh) on rented Vast RTX 3060 12 GB and RTX 4060 8 GB one-shots: the served 0317 host and the V6-07 host on the V6-07 floor server (igneum-floor-sm8689-v607.tgz), the fixture fees-v1-shards2.json, the ds55 kit's worker as the miner, nvidia-smi at 1 Hz; recorded in docs/analysis/floor-memory-profile-2026-10-08.md",
|
||||
"box_class": "rented pods (Vast one-shots)",
|
||||
"fixtures": [
|
||||
"F0",
|
||||
"F1"
|
||||
],
|
||||
"cases": [
|
||||
"GPU-05",
|
||||
"CAP-02",
|
||||
"UX-02"
|
||||
],
|
||||
"coverage": {
|
||||
"GPU-05": "partial: the 8 GB and 12 GB tiers' prover headroom on the 5.5 GiB dataset (free memory, the prover's peak alone, the refusal beside the miner); fragmentation, restart and next-epoch construction not run",
|
||||
"CAP-02": "partial: proving alone (shard and the whole segment path) on the 8 GB and 12 GB tiers on the default job path, the time-share refusal beside the miner with the miner unharmed, memory headroom and proof latency; wall energy only on the 3060 (the 4060 host has no power sensor); induced GPU task failure and wallet control not run; 16 GB and 24 GB tiers not run",
|
||||
"UX-02": "partial: the prover's safe refusal under memory pressure (exit 78 in one line, the miner's worker unharmed, nothing leaked after the refusal); pause, stop, power limit and the UI's crash ownership are the suite:app cell's"
|
||||
},
|
||||
"island": "none: every row proved a fixture on a rented pod; no row read a devnet"
|
||||
}
|
||||
},
|
||||
"not_run": {
|
||||
|
|
|
|||
|
|
@ -11,7 +11,8 @@ proving-v1, 272b025) and tools/proving-v1/pc2-segments.ps1 around the four binar
|
|||
offered again every pass until the segment's deadline (the 272b025 behaviour).
|
||||
3. one export (igneum_exportSegments 0..last), one fixture per block (igneum-prove-export), one host run
|
||||
(igneum-prove-host --mode chain --chain ... --save-shards [--prev]) on the patched server (HOME=/opt/igneum-floor/home,
|
||||
SP1_GPU_ELEMENT_THRESHOLD from THRESHOLD), the miner paused for the run when MINER=pause (prove-alone cards).
|
||||
IGNEUM_PROVE_WORKLOAD=chain and no threshold by default, the host's memory profile from the card's free memory;
|
||||
SP1_GPU_ELEMENT_THRESHOLD from THRESHOLD as a hand override), the miner paused for the run when MINER=pause (prove-alone cards).
|
||||
4. every shard record signed (igneum-miner sign-record) and submitted (igneum_submitProofRecord); the segment record
|
||||
(sign-segment-record, igneum_submitSegmentRecord) once every shard is accepted and the statement equals the node's.
|
||||
5. the paid state of every submitted segment polled each pass (igneum_getSegmentRecords); a state file for the
|
||||
|
|
@ -557,7 +558,10 @@ while (time.time() - t_run0) / 3600 < RUN_HOURS:
|
|||
# chain
|
||||
if MINER == "pause": miner_stop()
|
||||
kill_server()
|
||||
env = dict(os.environ, HOME=f"{FLOOR}/home", SP1_PROVER="cuda", RUST_LOG="off")
|
||||
# the default job path (V6-07, 8 October 2026): no threshold override; the host picks its memory profile for the chain
|
||||
# workload from the card's FREE memory (host/src/memory_profile.rs) and hands it to the floor server; THRESHOLD, when
|
||||
# set, is a hand override the host keeps and names in its RESULT memory_profile line
|
||||
env = dict(os.environ, HOME=f"{FLOOR}/home", SP1_PROVER="cuda", RUST_LOG="off", IGNEUM_PROVE_WORKLOAD="chain", IGNEUM_PROVE_DEVICE=str(DEV))
|
||||
if THRESHOLD: env["SP1_GPU_ELEMENT_THRESHOLD"] = THRESHOLD
|
||||
args = [HOST, "--mode", "chain", "--chain", ",".join(fixtures), "--prover", WALLET, "--save-shards", "--out", f"{d}/chain-results.json"]
|
||||
if prev_file: args += ["--prev", prev_file]
|
||||
|
|
@ -571,6 +575,10 @@ while (time.time() - t_run0) / 3600 < RUN_HOURS:
|
|||
except Exception: peak = 0
|
||||
seg["peak_mib"] = peak
|
||||
open(f"{d}/chain.log", "w").write((rr.stdout if rr else "") + "\n" + (rr.stderr if rr else "TIMEOUT"))
|
||||
if rr and rr.returncode == 78:
|
||||
# the host refused the card's free memory for the chain workload before any setup (exit 78, one line)
|
||||
line = [l for l in rr.stdout.split("\n") if "memory_profile refused" in l]
|
||||
say(f"RESULT seg {first} chain REFUSED {stamp()} {line[-1].strip() if line else 'memory profile refused'}"); seg["failed"] = "memory"; close(first, "cancelled", "memory"); time.sleep(60); continue
|
||||
if not rr or rr.returncode != 0 or not os.path.exists(f"{d}/chain-results.json"):
|
||||
say(f"RESULT seg {first} chain FAILED {stamp()} rc={rr.returncode if rr else 'timeout'} wall={seg['chain_s']} s: {((rr.stderr if rr else '') or '')[-200:].strip()}"); seg["failed"] = "chain"; close(first, "cancelled", "chain" if rr else "timeout"); continue
|
||||
res = json.load(open(f"{d}/chain-results.json"))
|
||||
|
|
|
|||
|
|
@ -22,7 +22,7 @@ fail() { echo "RESULT setup_failed $1 $(stamp)"; exit 2; }
|
|||
ARCHS="${ARCHS:-86,89,120}"; LABEL="${LABEL:-box}"; WALLET="${WALLET:-}"
|
||||
PKG_URL=https://dl.igneum.network/dl/public/igneum-hive-0.3.12.tar.gz
|
||||
PKG_SHA=7972af92e7cd9a032303eca4d95b533f53e0e68d1b9cae5bfe406a5b7c30a454
|
||||
PATCH_SHA=e81cb0d03b291f9fd4bf0a109d6da2d7c897795c9ffd7f797c0ddce723eee2b1
|
||||
PATCH_SHA=9098c3e5979d057031f43588668201d1f7a53ab17e980c1eaf896cd5e6f38575
|
||||
SEED=188.245.5.161:26611
|
||||
echo "RESULT start $(stamp) label=$LABEL archs=$ARCHS host=$(hostname) nproc=$(nproc) ram_gb=$(( $(awk '/MemTotal/{print $2}' /proc/meminfo) / 1048576 )) disk_avail=$(df -BG /root | awk 'NR==2{print $4}')"
|
||||
echo "RESULT gpu $(nvidia-smi --query-gpu=name,memory.total,driver_version,pci.bus_id,power.limit,power.min_limit,power.max_limit,clocks.max.sm --format=csv,noheader 2>&1 | tr '\n' ';')"
|
||||
|
|
|
|||
|
|
@ -199,15 +199,16 @@ diff --git a/sp1-gpu/crates/prover_components/src/builder.rs b/sp1-gpu/crates/pr
|
|||
index 5dccd9d..574d4fa 100644
|
||||
--- a/sp1-gpu/crates/prover_components/src/builder.rs
|
||||
+++ b/sp1-gpu/crates/prover_components/src/builder.rs
|
||||
@@ -23,28 +23,75 @@ use crate::{
|
||||
@@ -23,28 +23,124 @@ use crate::{
|
||||
SP1CudaProverComponents,
|
||||
};
|
||||
|
||||
+/// Igneum prover-floor patch (5 October 2026). Upstream sizes every device buffer for a 24 GB card or larger
|
||||
+/// and panics below that, whatever the shard. Here the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks a
|
||||
+/// tier, and `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly.
|
||||
+/// The proof format, the verifier and the program ids do not change: the element threshold only decides where
|
||||
+/// the executor splits shards, as upstream's own 24 GB tier already does.
|
||||
+/// Igneum prover-floor patch (5 October 2026; the memory rules of V6-07, 8 October 2026). Upstream sizes every
|
||||
+/// device buffer for a 24 GB card or larger and panics below that, whatever the shard. Here the card's FREE memory
|
||||
+/// at start (or `SP1_GPU_MEMORY_BUDGET_GB`, the host's lease) picks a tier, read once for the process, and
|
||||
+/// `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly. The proof
|
||||
+/// format, the verifier and the program ids do not change: the element threshold only decides where the executor
|
||||
+/// splits shards, as upstream's own 24 GB tier already does.
|
||||
+fn env_usize(name: &str) -> Option<usize> {
|
||||
+ std::env::var(name).ok().and_then(|s| s.parse::<usize>().ok())
|
||||
+}
|
||||
|
|
@ -216,8 +217,17 @@ index 5dccd9d..574d4fa 100644
|
|||
+ std::env::var(name).ok().and_then(|s| s.parse::<f64>().ok())
|
||||
+}
|
||||
+
|
||||
+/// The core element threshold for a memory budget in GB (upstream's own figure for the budget, +4, as it
|
||||
+/// computed it: a 32 GB card is 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16).
|
||||
+/// The recursion trace allocation under the 24 GB tier (V6-07): a recursion key or shard uses 90,177,536 elements
|
||||
+/// (35.6 M preprocessed and 54.5 M main; the floor sweep of 5 October 2026), so 2^26 + 2^25 = 100,663,296 holds it
|
||||
+/// with the stacking slack `floor_capacity` adds; the pinned host copies (four per prover) shrink with it. Upstream's
|
||||
+/// 2^27 stays on the 24 GB tier and above. Two distinct values: `floor_tests::recursion_branches_differ`.
|
||||
+pub const RECURSION_TRACE_ALLOCATION_SMALL: usize = (1 << 26) + (1 << 25);
|
||||
+
|
||||
+/// The core element threshold for a memory budget in GB (the budget is the FREE memory, +4, as upstream computed
|
||||
+/// its tiers from the total: a 32 GB card alone reads 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16, an
|
||||
+/// 8 GB card 12). Under the 18 tier the threshold is 2^26, the value every passing small-card row used (the
|
||||
+/// alone-comp-26-v1 rows of 6 October 2026 on the 3060, 3080, 4060, 4060 Ti, 4070 and 5070; the 8 October 2026
|
||||
+/// proof_alone rows on the 3060 at 7,525 MiB and the 4060 at 7,532 MiB); no small-card row passed at 2^27 here.
|
||||
+pub fn element_threshold_for_budget(gpu_memory_gb: usize, full_size_shards: bool) -> u64 {
|
||||
+ if gpu_memory_gb > 30 || (full_size_shards && gpu_memory_gb >= 24) {
|
||||
+ ELEMENT_THRESHOLD
|
||||
|
|
@ -226,31 +236,69 @@ index 5dccd9d..574d4fa 100644
|
|||
+ } else if gpu_memory_gb >= 18 {
|
||||
+ (1 << 27) + (1 << 26)
|
||||
+ } else {
|
||||
+ 1 << 27
|
||||
+ 1 << 26
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+/// The recursion trace allocation (elements) for a memory budget.
|
||||
+/// The recursion trace allocation (elements) for a memory budget: upstream's on the 24 GB tier and above, the
|
||||
+/// small constant under it.
|
||||
+pub fn recursion_trace_allocation_for_budget(gpu_memory_gb: usize) -> usize {
|
||||
+ if gpu_memory_gb >= 24 {
|
||||
+ RECURSION_TRACE_ALLOCATION
|
||||
+ } else {
|
||||
+ RECURSION_TRACE_ALLOCATION
|
||||
+ RECURSION_TRACE_ALLOCATION_SMALL
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+/// The memory budget in GB, read ONCE for the process (the core opts and the recursion prover are built at
|
||||
+/// different moments, after allocations that lower the free figure; one reading keeps both on one tier):
|
||||
+/// `SP1_GPU_MEMORY_BUDGET_GB` when set (the host's lease), else the device's FREE memory as the driver reports it
|
||||
+/// at the first call, never the total (a card beside a miner has the miner's resident set gone), +4 as upstream
|
||||
+/// computed its tiers.
|
||||
+pub fn gpu_memory_gb() -> usize {
|
||||
+ let gb = 1024.0 * 1024.0 * 1024.0;
|
||||
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
|
||||
+ Some(b) => (b.ceil() as usize) + 4,
|
||||
+ None => (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4,
|
||||
+ }
|
||||
+ static BUDGET: std::sync::OnceLock<usize> = std::sync::OnceLock::new();
|
||||
+ *BUDGET.get_or_init(|| {
|
||||
+ let gb = 1024.0 * 1024.0 * 1024.0;
|
||||
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
|
||||
+ Some(b) => (b.ceil() as usize) + 4,
|
||||
+ None => {
|
||||
+ let (free, _total) = cuda_memory_info().unwrap();
|
||||
+ (((free as f64) / gb).ceil() as usize) + 4
|
||||
+ }
|
||||
+ }
|
||||
+ })
|
||||
+}
|
||||
+
|
||||
+pub fn recursion_trace_allocation() -> usize {
|
||||
+ env_usize("SP1_GPU_RECURSION_TRACE_ALLOCATION")
|
||||
+ .unwrap_or_else(|| recursion_trace_allocation_for_budget(gpu_memory_gb()))
|
||||
+}
|
||||
+
|
||||
+#[cfg(test)]
|
||||
+mod floor_tests {
|
||||
+ use super::*;
|
||||
+
|
||||
+ #[test]
|
||||
+ fn small_tier_is_two_to_the_26() {
|
||||
+ assert_eq!(element_threshold_for_budget(12, false), 1 << 26, "an 8 GB card reads 12");
|
||||
+ assert_eq!(element_threshold_for_budget(16, false), 1 << 26, "a 12 GB card reads 16");
|
||||
+ assert_eq!(element_threshold_for_budget(17, false), 1 << 26);
|
||||
+ assert_eq!(element_threshold_for_budget(20, false), (1 << 27) + (1 << 26), "a 16 GB card reads 20");
|
||||
+ assert_eq!(element_threshold_for_budget(28, false), ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24));
|
||||
+ assert_eq!(element_threshold_for_budget(28, true), ELEMENT_THRESHOLD);
|
||||
+ assert_eq!(element_threshold_for_budget(36, false), ELEMENT_THRESHOLD);
|
||||
+ }
|
||||
+
|
||||
+ #[test]
|
||||
+ fn recursion_branches_differ() {
|
||||
+ assert_ne!(recursion_trace_allocation_for_budget(28), recursion_trace_allocation_for_budget(16));
|
||||
+ assert_eq!(recursion_trace_allocation_for_budget(28), RECURSION_TRACE_ALLOCATION);
|
||||
+ assert_eq!(recursion_trace_allocation_for_budget(16), RECURSION_TRACE_ALLOCATION_SMALL);
|
||||
+ assert_eq!(RECURSION_TRACE_ALLOCATION_SMALL, 100_663_296);
|
||||
+ assert!(RECURSION_TRACE_ALLOCATION_SMALL < RECURSION_TRACE_ALLOCATION);
|
||||
+ assert!(RECURSION_TRACE_ALLOCATION_SMALL >= 90_177_536 + (1 << 22));
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
pub fn local_gpu_opts() -> SP1CoreOpts {
|
||||
let mut opts = SP1CoreOpts::default();
|
||||
|
|
@ -266,7 +314,7 @@ index 5dccd9d..574d4fa 100644
|
|||
- if gpu_memory_gb < 24 {
|
||||
- panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB");
|
||||
- }
|
||||
+ // The card's memory plus 4, as upstream computed it (a 32 GB card reads 36), or the budget given.
|
||||
+ // The card's FREE memory plus 4, as upstream computed its tiers from the total, or the budget given.
|
||||
+ let gpu_memory_gb = gpu_memory_gb();
|
||||
|
||||
- let shard_threshold = if !opts.full_size_shards && gpu_memory_gb <= 30 {
|
||||
|
|
@ -278,17 +326,18 @@ index 5dccd9d..574d4fa 100644
|
|||
+ None => element_threshold_for_budget(gpu_memory_gb, opts.full_size_shards),
|
||||
};
|
||||
+ let height_threshold = opts.sharding_threshold.height_threshold;
|
||||
+ let (free_mib, total_mib) = cuda_memory_info().map(|(f, t)| (f >> 20, t >> 20)).unwrap_or((0, 0));
|
||||
|
||||
- tracing::debug!("Shard threshold: {shard_threshold}");
|
||||
+ eprintln!(
|
||||
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
|
||||
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} free_mib={free_mib} total_mib={total_mib} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
|
||||
+ recursion_trace_allocation(),
|
||||
+ opts.full_size_shards
|
||||
+ );
|
||||
opts.sharding_threshold.element_threshold = shard_threshold;
|
||||
|
||||
opts.global_dependencies_opt = true;
|
||||
@@ -92,7 +139,7 @@ pub async fn recursion_prover_and_verifier(
|
||||
@@ -92,7 +188,7 @@ pub async fn recursion_prover_and_verifier(
|
||||
) {
|
||||
let recursion_verifier = SP1CudaProverComponents::compress_verifier();
|
||||
(
|
||||
|
|
|
|||
|
|
@ -199,15 +199,16 @@ diff --git a/sp1-gpu/crates/prover_components/src/builder.rs b/sp1-gpu/crates/pr
|
|||
index 5dccd9d..574d4fa 100644
|
||||
--- a/sp1-gpu/crates/prover_components/src/builder.rs
|
||||
+++ b/sp1-gpu/crates/prover_components/src/builder.rs
|
||||
@@ -23,28 +23,75 @@ use crate::{
|
||||
@@ -23,28 +23,124 @@ use crate::{
|
||||
SP1CudaProverComponents,
|
||||
};
|
||||
|
||||
+/// Igneum prover-floor patch (5 October 2026). Upstream sizes every device buffer for a 24 GB card or larger
|
||||
+/// and panics below that, whatever the shard. Here the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks a
|
||||
+/// tier, and `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly.
|
||||
+/// The proof format, the verifier and the program ids do not change: the element threshold only decides where
|
||||
+/// the executor splits shards, as upstream's own 24 GB tier already does.
|
||||
+/// Igneum prover-floor patch (5 October 2026; the memory rules of V6-07, 8 October 2026). Upstream sizes every
|
||||
+/// device buffer for a 24 GB card or larger and panics below that, whatever the shard. Here the card's FREE memory
|
||||
+/// at start (or `SP1_GPU_MEMORY_BUDGET_GB`, the host's lease) picks a tier, read once for the process, and
|
||||
+/// `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly. The proof
|
||||
+/// format, the verifier and the program ids do not change: the element threshold only decides where the executor
|
||||
+/// splits shards, as upstream's own 24 GB tier already does.
|
||||
+fn env_usize(name: &str) -> Option<usize> {
|
||||
+ std::env::var(name).ok().and_then(|s| s.parse::<usize>().ok())
|
||||
+}
|
||||
|
|
@ -216,8 +217,17 @@ index 5dccd9d..574d4fa 100644
|
|||
+ std::env::var(name).ok().and_then(|s| s.parse::<f64>().ok())
|
||||
+}
|
||||
+
|
||||
+/// The core element threshold for a memory budget in GB (upstream's own figure for the budget, +4, as it
|
||||
+/// computed it: a 32 GB card is 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16).
|
||||
+/// The recursion trace allocation under the 24 GB tier (V6-07): a recursion key or shard uses 90,177,536 elements
|
||||
+/// (35.6 M preprocessed and 54.5 M main; the floor sweep of 5 October 2026), so 2^26 + 2^25 = 100,663,296 holds it
|
||||
+/// with the stacking slack `floor_capacity` adds; the pinned host copies (four per prover) shrink with it. Upstream's
|
||||
+/// 2^27 stays on the 24 GB tier and above. Two distinct values: `floor_tests::recursion_branches_differ`.
|
||||
+pub const RECURSION_TRACE_ALLOCATION_SMALL: usize = (1 << 26) + (1 << 25);
|
||||
+
|
||||
+/// The core element threshold for a memory budget in GB (the budget is the FREE memory, +4, as upstream computed
|
||||
+/// its tiers from the total: a 32 GB card alone reads 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16, an
|
||||
+/// 8 GB card 12). Under the 18 tier the threshold is 2^26, the value every passing small-card row used (the
|
||||
+/// alone-comp-26-v1 rows of 6 October 2026 on the 3060, 3080, 4060, 4060 Ti, 4070 and 5070; the 8 October 2026
|
||||
+/// proof_alone rows on the 3060 at 7,525 MiB and the 4060 at 7,532 MiB); no small-card row passed at 2^27 here.
|
||||
+pub fn element_threshold_for_budget(gpu_memory_gb: usize, full_size_shards: bool) -> u64 {
|
||||
+ if gpu_memory_gb > 30 || (full_size_shards && gpu_memory_gb >= 24) {
|
||||
+ ELEMENT_THRESHOLD
|
||||
|
|
@ -226,31 +236,69 @@ index 5dccd9d..574d4fa 100644
|
|||
+ } else if gpu_memory_gb >= 18 {
|
||||
+ (1 << 27) + (1 << 26)
|
||||
+ } else {
|
||||
+ 1 << 27
|
||||
+ 1 << 26
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+/// The recursion trace allocation (elements) for a memory budget.
|
||||
+/// The recursion trace allocation (elements) for a memory budget: upstream's on the 24 GB tier and above, the
|
||||
+/// small constant under it.
|
||||
+pub fn recursion_trace_allocation_for_budget(gpu_memory_gb: usize) -> usize {
|
||||
+ if gpu_memory_gb >= 24 {
|
||||
+ RECURSION_TRACE_ALLOCATION
|
||||
+ } else {
|
||||
+ RECURSION_TRACE_ALLOCATION
|
||||
+ RECURSION_TRACE_ALLOCATION_SMALL
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+/// The memory budget in GB, read ONCE for the process (the core opts and the recursion prover are built at
|
||||
+/// different moments, after allocations that lower the free figure; one reading keeps both on one tier):
|
||||
+/// `SP1_GPU_MEMORY_BUDGET_GB` when set (the host's lease), else the device's FREE memory as the driver reports it
|
||||
+/// at the first call, never the total (a card beside a miner has the miner's resident set gone), +4 as upstream
|
||||
+/// computed its tiers.
|
||||
+pub fn gpu_memory_gb() -> usize {
|
||||
+ let gb = 1024.0 * 1024.0 * 1024.0;
|
||||
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
|
||||
+ Some(b) => (b.ceil() as usize) + 4,
|
||||
+ None => (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4,
|
||||
+ }
|
||||
+ static BUDGET: std::sync::OnceLock<usize> = std::sync::OnceLock::new();
|
||||
+ *BUDGET.get_or_init(|| {
|
||||
+ let gb = 1024.0 * 1024.0 * 1024.0;
|
||||
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
|
||||
+ Some(b) => (b.ceil() as usize) + 4,
|
||||
+ None => {
|
||||
+ let (free, _total) = cuda_memory_info().unwrap();
|
||||
+ (((free as f64) / gb).ceil() as usize) + 4
|
||||
+ }
|
||||
+ }
|
||||
+ })
|
||||
+}
|
||||
+
|
||||
+pub fn recursion_trace_allocation() -> usize {
|
||||
+ env_usize("SP1_GPU_RECURSION_TRACE_ALLOCATION")
|
||||
+ .unwrap_or_else(|| recursion_trace_allocation_for_budget(gpu_memory_gb()))
|
||||
+}
|
||||
+
|
||||
+#[cfg(test)]
|
||||
+mod floor_tests {
|
||||
+ use super::*;
|
||||
+
|
||||
+ #[test]
|
||||
+ fn small_tier_is_two_to_the_26() {
|
||||
+ assert_eq!(element_threshold_for_budget(12, false), 1 << 26, "an 8 GB card reads 12");
|
||||
+ assert_eq!(element_threshold_for_budget(16, false), 1 << 26, "a 12 GB card reads 16");
|
||||
+ assert_eq!(element_threshold_for_budget(17, false), 1 << 26);
|
||||
+ assert_eq!(element_threshold_for_budget(20, false), (1 << 27) + (1 << 26), "a 16 GB card reads 20");
|
||||
+ assert_eq!(element_threshold_for_budget(28, false), ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24));
|
||||
+ assert_eq!(element_threshold_for_budget(28, true), ELEMENT_THRESHOLD);
|
||||
+ assert_eq!(element_threshold_for_budget(36, false), ELEMENT_THRESHOLD);
|
||||
+ }
|
||||
+
|
||||
+ #[test]
|
||||
+ fn recursion_branches_differ() {
|
||||
+ assert_ne!(recursion_trace_allocation_for_budget(28), recursion_trace_allocation_for_budget(16));
|
||||
+ assert_eq!(recursion_trace_allocation_for_budget(28), RECURSION_TRACE_ALLOCATION);
|
||||
+ assert_eq!(recursion_trace_allocation_for_budget(16), RECURSION_TRACE_ALLOCATION_SMALL);
|
||||
+ assert_eq!(RECURSION_TRACE_ALLOCATION_SMALL, 100_663_296);
|
||||
+ assert!(RECURSION_TRACE_ALLOCATION_SMALL < RECURSION_TRACE_ALLOCATION);
|
||||
+ assert!(RECURSION_TRACE_ALLOCATION_SMALL >= 90_177_536 + (1 << 22));
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
pub fn local_gpu_opts() -> SP1CoreOpts {
|
||||
let mut opts = SP1CoreOpts::default();
|
||||
|
|
@ -266,7 +314,7 @@ index 5dccd9d..574d4fa 100644
|
|||
- if gpu_memory_gb < 24 {
|
||||
- panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB");
|
||||
- }
|
||||
+ // The card's memory plus 4, as upstream computed it (a 32 GB card reads 36), or the budget given.
|
||||
+ // The card's FREE memory plus 4, as upstream computed its tiers from the total, or the budget given.
|
||||
+ let gpu_memory_gb = gpu_memory_gb();
|
||||
|
||||
- let shard_threshold = if !opts.full_size_shards && gpu_memory_gb <= 30 {
|
||||
|
|
@ -278,17 +326,18 @@ index 5dccd9d..574d4fa 100644
|
|||
+ None => element_threshold_for_budget(gpu_memory_gb, opts.full_size_shards),
|
||||
};
|
||||
+ let height_threshold = opts.sharding_threshold.height_threshold;
|
||||
+ let (free_mib, total_mib) = cuda_memory_info().map(|(f, t)| (f >> 20, t >> 20)).unwrap_or((0, 0));
|
||||
|
||||
- tracing::debug!("Shard threshold: {shard_threshold}");
|
||||
+ eprintln!(
|
||||
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
|
||||
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} free_mib={free_mib} total_mib={total_mib} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
|
||||
+ recursion_trace_allocation(),
|
||||
+ opts.full_size_shards
|
||||
+ );
|
||||
opts.sharding_threshold.element_threshold = shard_threshold;
|
||||
|
||||
opts.global_dependencies_opt = true;
|
||||
@@ -92,7 +139,7 @@ pub async fn recursion_prover_and_verifier(
|
||||
@@ -92,7 +188,7 @@ pub async fn recursion_prover_and_verifier(
|
||||
) {
|
||||
let recursion_verifier = SP1CudaProverComponents::compress_verifier();
|
||||
(
|
||||
|
|
@ -297,6 +346,86 @@ index 5dccd9d..574d4fa 100644
|
|||
.await,
|
||||
recursion_verifier,
|
||||
)
|
||||
diff --git a/sp1-gpu/crates/server/src/main.rs b/sp1-gpu/crates/server/src/main.rs
|
||||
index 65e94f5..498357c 100644
|
||||
--- a/sp1-gpu/crates/server/src/main.rs
|
||||
+++ b/sp1-gpu/crates/server/src/main.rs
|
||||
@@ -17,9 +17,55 @@ struct Args {
|
||||
version: bool,
|
||||
}
|
||||
|
||||
+/// Igneum prover-floor patch (6 October 2026, the GPU fleet's finding): a panic inside a prover task (an allocation
|
||||
+/// the card cannot meet, `cudaMallocAsync` failing and `Buffer::with_capacity_in` panicking in a tokio worker) left
|
||||
+/// the request's future waiting for ever, the client on its socket, the card at 0% for the 15 minutes until someone
|
||||
+/// killed it (the 8 GB and 10 GB cards at threshold 2^27). The server must fail the shard instead: this hook names
|
||||
+/// the stage from the panic's location and exits, so the client's proof fails at once with the server's last line.
|
||||
+fn stage_of(location: &str) -> &'static str {
|
||||
+ let l = location.to_ascii_lowercase();
|
||||
+ if l.contains("jagged_tracegen") || l.contains("/tracegen") {
|
||||
+ "trace generation"
|
||||
+ } else if l.contains("commit") || l.contains("basefold") || l.contains("merkle") {
|
||||
+ "the commit (codewords and Merkle trees)"
|
||||
+ } else if l.contains("logup_gkr") {
|
||||
+ "LogUp GKR"
|
||||
+ } else if l.contains("zerocheck") {
|
||||
+ "the zerocheck"
|
||||
+ } else if l.contains("jagged") {
|
||||
+ "the jagged sumcheck"
|
||||
+ } else if l.contains("prover_components") || l.contains("recursion") || l.contains("sp1-prover") || l.contains("sp1_prover") {
|
||||
+ "the recursion (compression)"
|
||||
+ } else if l.contains("cuda") || l.contains("slop") {
|
||||
+ "a device allocation"
|
||||
+ } else {
|
||||
+ "the prover"
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+fn install_abort_on_panic() {
|
||||
+ std::panic::set_hook(Box::new(|info| {
|
||||
+ let location = info.location().map(|l| format!("{}:{}", l.file(), l.line())).unwrap_or_else(|| "unknown".into());
|
||||
+ let message = info
|
||||
+ .payload()
|
||||
+ .downcast_ref::<&str>()
|
||||
+ .map(|s| s.to_string())
|
||||
+ .or_else(|| info.payload().downcast_ref::<String>().cloned())
|
||||
+ .unwrap_or_default();
|
||||
+ let oom = message.to_ascii_lowercase().contains("alloc") || message.contains("MemoryAllocation") || message.contains("OUT_OF_MEMORY");
|
||||
+ eprintln!(
|
||||
+ "FLOOR abort: {} failed at {location}: {message}{}; the server exits so the client's proof fails instead of waiting",
|
||||
+ stage_of(&location),
|
||||
+ if oom { " (the card's memory could not meet an allocation: lower SP1_GPU_ELEMENT_THRESHOLD one notch)" } else { "" }
|
||||
+ );
|
||||
+ std::process::exit(70);
|
||||
+ }));
|
||||
+}
|
||||
+
|
||||
#[tokio::main]
|
||||
#[allow(clippy::print_stdout)]
|
||||
async fn main() {
|
||||
+ install_abort_on_panic();
|
||||
tracing_subscriber::fmt::init();
|
||||
|
||||
let args = Args::parse();
|
||||
@@ -40,3 +86,19 @@ async fn main() {
|
||||
eprintln!("Error running server: {e}");
|
||||
}
|
||||
}
|
||||
+
|
||||
+#[cfg(test)]
|
||||
+mod floor_tests {
|
||||
+ use super::stage_of;
|
||||
+
|
||||
+ #[test]
|
||||
+ fn the_stage_is_named_from_the_panic_location() {
|
||||
+ assert_eq!(stage_of("sp1-gpu/crates/jagged_tracegen/src/lib.rs:240"), "trace generation");
|
||||
+ assert_eq!(stage_of("sp1-gpu/crates/basefold/src/fri.rs:97"), "the commit (codewords and Merkle trees)");
|
||||
+ assert_eq!(stage_of("sp1-gpu/crates/logup_gkr/src/tracegen.rs:72"), "LogUp GKR");
|
||||
+ assert_eq!(stage_of("sp1-gpu/crates/zerocheck/src/prover.rs:1163"), "the zerocheck");
|
||||
+ assert_eq!(stage_of("sp1-gpu/crates/prover_components/src/builder.rs:70"), "the recursion (compression)");
|
||||
+ assert_eq!(stage_of("sp1-gpu/crates/cuda/src/stream.rs:330"), "a device allocation");
|
||||
+ assert_eq!(stage_of("somewhere/else.rs:1"), "the prover");
|
||||
+ }
|
||||
+}
|
||||
diff --git a/sp1-gpu/crates/server/src/server.rs b/sp1-gpu/crates/server/src/server.rs
|
||||
index 4035f1f..0d0d907 100644
|
||||
--- a/sp1-gpu/crates/server/src/server.rs
|
||||
|
|
|
|||
Loading…
Reference in a new issue