bench-log: the known-failed case on the fleet's 3080 (v5 fails in 13 s against 568 s of hang; the FLOOR abort line), PC 1's second candidate at its cap, the 4060 Ti rows

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 13:17:03 +00:00
parent fd000e34b7
commit 500ab8c2fe

View file

@ -2456,9 +2456,22 @@ did not fail: job `floor-pc1-hangcase` (13:06:29 to 13:08:05Z, the miners stoppe
M-cycle prototype shard at 2^27 at **10,785 MiB of 12,282 in 24.8 s** (core 15.1 s, 39.7 MB; every proof verified),
where the 5090's allocation had said 13,459 MiB: so a 12 GB card proves even the biggest devnet shard alone at 2^27,
and the 5090's pattern overstates a 12 GB card by about 2.7 GB on big shards (0.65 GB on the v1 shard). The
known-failed cases of the gate are therefore upstream's own threshold (402,653,184; 28,307 MiB on the 5090) on that
shard on the 4070 (`floor-pc1-hangcase-2`) and the fleet's 3080 at 2^27 on a server rebuilt from v5; their lines
follow here when they land.
known-failed case of the gate, on a real card: the GPU fleet's rented 3080 10 GB (driver 570.211, sm_86), the server
rebuilt from patch v5 in 86 s (binary 5d4f92a3..., 108,212,616 bytes), the point that hung 568 s on v4 (2^27 on the
v1 shard, alone): **the host failed in 13 s wall** (setup 10.2 s, execute 0.3 s, then the abort), the server's line
in its log "FLOOR abort: a device allocation failed at slop/crates/tensor/src/inner.rs:51: called `Result::unwrap()`
on an `Err` value: AllocError { layout: Layout { size: 486586112, align: 4 } } (the card's memory could not meet an
allocation: lower SP1_GPU_ELEMENT_THRESHOLD one notch)", the tracegen line before it at 8,872 MiB in use, the
host's "Error: CudaClientError: Failed to read the response: UnexpectedEof", exit 1. The second candidate on PC 1
(`floor-pc1-hangcase-2`, upstream's threshold 402,653,184 on the prototype shard on the 4070, 13:10:29Z) hit the
job's 5-minute cap with no result line: on that card and point the failure did not reach the hook (a C++ CUDA
exception in the sppark NTT code cannot unwind into Rust; or a stream synchronisation that never returns after a
failed launch), or the point was still running; the point's log (the restore job) says which. For a user the app's
budget covers it either way: 120 s, then the server killed and the threshold stepped down. The 4060 Ti 8 GB rows
from the fleet: alone 2^26 9.6 s at 7,740 MiB, 2^25 compressed 13.5 s at 7,676, core 2^25 5.8 s at 5,916, core 2^26
5.0 s at 7,196, core 2^24 11.2 s at 5,404; 2^27 hung 904 s on v4; beside its miner the compressed 2^26 point hung too
(7.7 GB plus the 1.4 GB miner on 8.2 GB), so an 8 GB card proves ALONE (the tier line) and core-only beside the
miner is its open row.
Consequences (the rule of 5 October 2026), as they stood after sweep 1 (superseded by the reading above for the 12 and 16 GB tiers): the shipped SP1 GPU server refuses every
card under 24 GB before allocating, so 8, 12 and 16 GB NVIDIA cards cannot prove on it whatever the shard; the v1