bench-log: the known-failed case on the fleet's 3080 (v5 fails in 13 s against 568 s of hang; the FLOOR abort line), PC 1's second candidate at its cap, the 4060 Ti rows
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
fd000e34b7
commit
500ab8c2fe
1 changed files with 16 additions and 3 deletions
|
|
@ -2456,9 +2456,22 @@ did not fail: job `floor-pc1-hangcase` (13:06:29 to 13:08:05Z, the miners stoppe
|
|||
M-cycle prototype shard at 2^27 at **10,785 MiB of 12,282 in 24.8 s** (core 15.1 s, 39.7 MB; every proof verified),
|
||||
where the 5090's allocation had said 13,459 MiB: so a 12 GB card proves even the biggest devnet shard alone at 2^27,
|
||||
and the 5090's pattern overstates a 12 GB card by about 2.7 GB on big shards (0.65 GB on the v1 shard). The
|
||||
known-failed cases of the gate are therefore upstream's own threshold (402,653,184; 28,307 MiB on the 5090) on that
|
||||
shard on the 4070 (`floor-pc1-hangcase-2`) and the fleet's 3080 at 2^27 on a server rebuilt from v5; their lines
|
||||
follow here when they land.
|
||||
known-failed case of the gate, on a real card: the GPU fleet's rented 3080 10 GB (driver 570.211, sm_86), the server
|
||||
rebuilt from patch v5 in 86 s (binary 5d4f92a3..., 108,212,616 bytes), the point that hung 568 s on v4 (2^27 on the
|
||||
v1 shard, alone): **the host failed in 13 s wall** (setup 10.2 s, execute 0.3 s, then the abort), the server's line
|
||||
in its log "FLOOR abort: a device allocation failed at slop/crates/tensor/src/inner.rs:51: called `Result::unwrap()`
|
||||
on an `Err` value: AllocError { layout: Layout { size: 486586112, align: 4 } } (the card's memory could not meet an
|
||||
allocation: lower SP1_GPU_ELEMENT_THRESHOLD one notch)", the tracegen line before it at 8,872 MiB in use, the
|
||||
host's "Error: CudaClientError: Failed to read the response: UnexpectedEof", exit 1. The second candidate on PC 1
|
||||
(`floor-pc1-hangcase-2`, upstream's threshold 402,653,184 on the prototype shard on the 4070, 13:10:29Z) hit the
|
||||
job's 5-minute cap with no result line: on that card and point the failure did not reach the hook (a C++ CUDA
|
||||
exception in the sppark NTT code cannot unwind into Rust; or a stream synchronisation that never returns after a
|
||||
failed launch), or the point was still running; the point's log (the restore job) says which. For a user the app's
|
||||
budget covers it either way: 120 s, then the server killed and the threshold stepped down. The 4060 Ti 8 GB rows
|
||||
from the fleet: alone 2^26 9.6 s at 7,740 MiB, 2^25 compressed 13.5 s at 7,676, core 2^25 5.8 s at 5,916, core 2^26
|
||||
5.0 s at 7,196, core 2^24 11.2 s at 5,404; 2^27 hung 904 s on v4; beside its miner the compressed 2^26 point hung too
|
||||
(7.7 GB plus the 1.4 GB miner on 8.2 GB), so an 8 GB card proves ALONE (the tier line) and core-only beside the
|
||||
miner is its open row.
|
||||
|
||||
Consequences (the rule of 5 October 2026), as they stood after sweep 1 (superseded by the reading above for the 12 and 16 GB tiers): the shipped SP1 GPU server refuses every
|
||||
card under 24 GB before allocating, so 8, 12 and 16 GB NVIDIA cards cannot prove on it whatever the shard; the v1
|
||||
|
|
|
|||
Loading…
Reference in a new issue