diff --git a/docs/analysis/floor-memory-profile-2026-10-08.md b/docs/analysis/floor-memory-profile-2026-10-08.md new file mode 100644 index 000000000..000c61923 --- /dev/null +++ b/docs/analysis/floor-memory-profile-2026-10-08.md @@ -0,0 +1,125 @@ +# The floor's memory profile (V6-07), 8 October 2026 + +Master review R1, residual V6-07 (`docs/plans/igneum-2.0-master/evidence/04_full_system/IGNEUM_V6_Full_System_Review.md`, pages 210 to 212): the floor patch picked its limits from the card's total VRAM, not from what was free; its small-card element threshold defaulted to 2^27 where every passing row had set 2^26 by hand; its recursion-allocation budget returned one constant in both branches. The order: a pinned memory profile per workload read from free memory, with the app lane's device coordinator, and the 3060 and 4060 rows rerun on the default, unoverridden job path. This document is the record: the code (section 1), the table (section 2), the rows as the pods wrote them (section 3), the consequences per card tier beside each number (section 4), the coordinator hook and what is next (section 5). + +Branch `v607-floor-memory` off the box mirror master cef5234b5. Code commits e6abe8c2c (the patch and the host profile), 573dad0ad (the lesser of grant and free), 8c1fb754c (one floor patch), then the floors re-pinned from the rows (the commit carrying this document). Clocks below are UTC as the pods wrote them; UK time is one hour later. + +## 1. What changed + +| Item | Where | Before | After | +|---|---|---|---| +| The limits' source | `proving/prover-floor/sp1-gpu-6.8.1-floor.patch`, `builder.rs` `gpu_memory_gb()` (the fleet's copies `tools/fleet/floor.patch`, `floor-v5.patch` carry the same hunk; `box-setup.sh` re-pinned) | `cuda_memory_info().1` (the total), +4 as upstream | `cuda_memory_info().0` (the FREE memory as the driver reports it), +4, read ONCE per process in a `OnceLock` so the core opts and the recursion prover, built at different moments, sit on one tier; `SP1_GPU_MEMORY_BUDGET_GB` (the host's lease) still overrides | +| The small-card element threshold | same, `element_threshold_for_budget` | `1 << 27` under the 18 tier | `1 << 26`, the value the passing rows used (section 2 cites them) | +| The recursion-allocation budget | same, `recursion_trace_allocation_for_budget` | `RECURSION_TRACE_ALLOCATION` in both branches | upstream's 2^27 on the 24 GB tier and above; `RECURSION_TRACE_ALLOCATION_SMALL` = 2^26 + 2^25 = 100,663,296 under it (a recursion key or shard uses 90,177,536 elements, `docs/analysis/prover-floor.md` sweep 1; the patch's `floor_capacity` sizes the device buffer to the need, so the constant caps the buffer and shrinks the four pinned host copies per prover); `floor_tests::recursion_branches_differ` pins that the two branches differ and that the small one clears the measured use with a stacking height of slack; `floor_tests::small_tier_is_two_to_the_26` pins the tiers | +| The FLOOR opts line | same, `local_gpu_opts` | `gpu_memory_gb`, thresholds | adds `free_mib` and `total_mib` so a log names what the server read | +| The pinned profile per workload | `proving/igneum-prove/host/src/memory_profile.rs` (new), wired in `main.rs` before the SP1 client is built | nothing: the server guessed from the total; the fleet set `SP1_GPU_ELEMENT_THRESHOLD` by hand | one table (section 2) with tests pinning every value; the row is chosen from the engine's lease or, with no engine, from `nvidia-smi memory.free` on the device, and handed to the server by environment (`SP1_GPU_ELEMENT_THRESHOLD`, `SP1_GPU_RECURSION_TRACE_ALLOCATION`, `SP1_GPU_MEMORY_BUDGET_GB`); a hand override already in the environment is kept and named; under the workload's floor the host prints one line and exits 78 before any setup | +| The fleet's default path | `tools/fleet/box-prover.py` | `SP1_GPU_ELEMENT_THRESHOLD` from `THRESHOLD`, else the server's total-VRAM guess | `IGNEUM_PROVE_WORKLOAD=chain` and `IGNEUM_PROVE_DEVICE` on the host run, no threshold unless `THRESHOLD` is set by hand; a refusal (exit 78) closes the segment as cancelled/memory | + +Tests: `cargo test -p igneum-prove-host memory_profile` (box 3, section 6) and the patch's `floor_tests` (run on the build pod after the server build, section 6). + +## 2. The table + +The host's `TIERS` (largest first; the free-memory lines are the card tiers as the floor patch reads them, free GiB rounded up plus 4 as upstream computed its tiers from the total): + +| Tier | Free memory at start | Element threshold | Recursion trace allocation | Who lands here | +|---|---|---|---|---| +| full | 26,624 MiB and up | 2^28 + 2^27 = 402,653,184 | 2^27 = 134,217,728 | a 32 GB card alone | +| 24gb | 20,480 MiB and up | 285,212,672 (upstream's 24 GB figure) | 2^27 | a 24 GB card alone | +| 16gb | 14,336 MiB and up | 2^27 + 2^26 = 201,326,592 (the patch's own figure, unmeasured on this fixture) | 2^26 + 2^25 = 100,663,296 | a 16 GB card alone | +| small | under 14,336 MiB | 2^26 = 67,108,864 | 100,663,296 | a 12 GB or 8 GB card alone; any card beside a miner's resident set | + +Floors per workload (the least free memory the host runs in; under it the refusal, exit 78): shard 7,700 MiB, aggregate 8,500 MiB, chain 8,500 MiB. The shard floor is the small tier's measured peak (7,525 MiB on the 3060, 7,532 on the 4060, section 3) plus headroom for the driver's own context. The aggregate and chain floors are 8,500 MiB: the aggregation is the peak of a chain run, 8,306 MiB on the RTX 3060 (section 3, D2), and the RTX 4060 (7,807 MiB free) aborted at it, so an 8 GB card is refused the chain and the aggregation before any setup and proves shards only. + +The rows cited for 2^26 on small cards: `docs/analysis/prover-tiers-real-cards.md` (6 October 2026, the matrix's `alone-comp-26-v1` point, compressed, verified: RTX 3060 7.4 GB 14.4 s; RTX 3080 8.0 GB 7.1 s; RTX 4060 7.4 GB 18.4 s; RTX 4060 Ti 8 GB 7.6 GB 9.6 s; RTX 4060 Ti 16 GB 7.8 GB 11.6 s; RTX 4070 7.6 GB 12.1 s; RTX 5070 7.6 GB 4.8 s) and `docs/analysis/class-v6/coexist-rows.md` (8 October 2026, `proof_alone` at `SP1_GPU_ELEMENT_THRESHOLD=67108864`: RTX 3060 peak 7,525 MiB 13.2 s verified; RTX 4060 peak 7,532 MiB 8.2 s verified). No small-card row ever passed at 2^27 on this host path; the 10 GB tier's 2^27 reading of 7 October (8,642 MiB alone on a 3080) ran on the segment host and is not this path. + +Every value above is pinned by `memory_profile::tests::table_is_pinned`, `tiers_from_free_memory`, `refuses_under_the_floor`, `recursion_branches_differ`, `workload_from_mode` and `env_for_the_server`; a change of the table is a change of the profile and needs its rows. + +## 3. The rows, rerun on the default, unoverridden job path + +Every row is a RESULT line as the pod wrote it (the raw run logs, host logs and 1 Hz `nvidia-smi` samples are kept under the lane's scratch `v607/rows/