diff --git a/docs/bench-log.md b/docs/bench-log.md index 77d27643b..304a42264 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -2456,9 +2456,22 @@ did not fail: job `floor-pc1-hangcase` (13:06:29 to 13:08:05Z, the miners stoppe M-cycle prototype shard at 2^27 at **10,785 MiB of 12,282 in 24.8 s** (core 15.1 s, 39.7 MB; every proof verified), where the 5090's allocation had said 13,459 MiB: so a 12 GB card proves even the biggest devnet shard alone at 2^27, and the 5090's pattern overstates a 12 GB card by about 2.7 GB on big shards (0.65 GB on the v1 shard). The -known-failed cases of the gate are therefore upstream's own threshold (402,653,184; 28,307 MiB on the 5090) on that -shard on the 4070 (`floor-pc1-hangcase-2`) and the fleet's 3080 at 2^27 on a server rebuilt from v5; their lines -follow here when they land. +known-failed case of the gate, on a real card: the GPU fleet's rented 3080 10 GB (driver 570.211, sm_86), the server +rebuilt from patch v5 in 86 s (binary 5d4f92a3..., 108,212,616 bytes), the point that hung 568 s on v4 (2^27 on the +v1 shard, alone): **the host failed in 13 s wall** (setup 10.2 s, execute 0.3 s, then the abort), the server's line +in its log "FLOOR abort: a device allocation failed at slop/crates/tensor/src/inner.rs:51: called `Result::unwrap()` +on an `Err` value: AllocError { layout: Layout { size: 486586112, align: 4 } } (the card's memory could not meet an +allocation: lower SP1_GPU_ELEMENT_THRESHOLD one notch)", the tracegen line before it at 8,872 MiB in use, the +host's "Error: CudaClientError: Failed to read the response: UnexpectedEof", exit 1. The second candidate on PC 1 +(`floor-pc1-hangcase-2`, upstream's threshold 402,653,184 on the prototype shard on the 4070, 13:10:29Z) hit the +job's 5-minute cap with no result line: on that card and point the failure did not reach the hook (a C++ CUDA +exception in the sppark NTT code cannot unwind into Rust; or a stream synchronisation that never returns after a +failed launch), or the point was still running; the point's log (the restore job) says which. For a user the app's +budget covers it either way: 120 s, then the server killed and the threshold stepped down. The 4060 Ti 8 GB rows +from the fleet: alone 2^26 9.6 s at 7,740 MiB, 2^25 compressed 13.5 s at 7,676, core 2^25 5.8 s at 5,916, core 2^26 +5.0 s at 7,196, core 2^24 11.2 s at 5,404; 2^27 hung 904 s on v4; beside its miner the compressed 2^26 point hung too +(7.7 GB plus the 1.4 GB miner on 8.2 GB), so an 8 GB card proves ALONE (the tier line) and core-only beside the +miner is its open row. Consequences (the rule of 5 October 2026), as they stood after sweep 1 (superseded by the reading above for the 12 and 16 GB tiers): the shipped SP1 GPU server refuses every card under 24 GB before allocating, so 8, 12 and 16 GB NVIDIA cards cannot prove on it whatever the shard; the v1