From 6d72baf549c0d3ef861e2afecb310b7cd2a461b9 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 13:23:37 +0000 Subject: [PATCH] bench-log: prover tiers on real cards, the rented fleet (ten cards, the 3090 to follow) --- docs/bench-log.md | 33 +++++++++++++++++++++++++++++++++ 1 file changed, 33 insertions(+) diff --git a/docs/bench-log.md b/docs/bench-log.md index 85997dca..05d8c532 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1961,3 +1961,36 @@ Ten distinct programs (seed strings `igneum-devnet-v4-epoch0`, `/epoch1` .. `/ep Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. The devnet pack three times through `packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256` (the pack's two libraries, `memhard.metal` and `program.metal`): compile 79 ms, 1 ms, 1 ms (the system shader cache answers the identical source from the second run); cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU. Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: `host.c` times `clBuildProgram` only in the `prepare` path and no `prepared` line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan. + +## Prover tiers on real cards: the rented fleet, 6 October 2026 (branch gpu-fleet) + +From 11:50 UTC, Vast.ai containers (nvidia/cuda:12.8.1-devel-ubuntu24.04, the host's driver), one card each, the +0.3.12 Linux node 83089544 on the ten-field override, the 0.3.12 CUDA worker, the patched SP1 server built on each box +from `proving/prover-floor/sp1-gpu-6.8.1-floor.patch` v4 (e81cb0d0...) for the card's own arch, the cuda host with the +pinned ids 0x2b1a81cb... and 0x474678f3...; the v1 shard (`fees-v1-shards2.json` shard 0, 4,717,439 cycles); every +proof VERIFIED by the host's own SDK verifier; peak = nvidia-smi memory.used sampled once a second (a per-second loop, +not `-l 1`, which buffers and ignores SIGTERM in a container); own = peak minus the reading before the point; beside += the card's miner running (its resident set is the base). Runner `tools/fleet/box-matrix.sh`, collector +`tools/fleet/collect.py`, raw logs `~/Desktop/fleet//`, the analysis `docs/analysis/prover-tiers-real-cards.md`. + +| Card | VRAM GB | Idle MiB | Miner | Stock SP1 6.8.1 | Patched, proves alone (own) | Beside the miner (peak) | Core-only beside the miner (own) | Verdict | +|---|---|---|---|---|---|---|---|---| +| RTX 3060 | 12 | 1 | 23.78 MH/s at 103.7 W, 1.4 GB | refused: thread 'tokio-rt-worker' (48952) panicked at sp1-gpu/crates/ | 7.4 GB, 14.4 s (alone-comp-26-v1) | 8.9 GB peak, 37.5 s | 5.6 GB, 27.2 s | mines and proves | +| RTX 3080 | 10 | 11 | 40.82 MH/s at 204.9 W, 1.5 GB | refused: thread 'tokio-rt-worker' (49593) panicked at sp1-gpu/crates/ | 8.0 GB, 7.1 s (alone-comp-26-v1) | 9.2 GB peak, 25.6 s | 5.9 GB, 19.2 s | mines and proves | +| RTX 4060 Ti 16 GB | 16 | 0 | 17.58 MH/s at 72.3 W, 1.4 GB | refused: thread 'tokio-rt-worker' (49293) panicked at sp1-gpu/crates/ | 7.8 GB, 11.6 s (alone-comp-26-v1) | 9.0 GB peak, 34.6 s | 5.8 GB, 25.9 s | mines and proves | +| RTX 4060 Ti 8 GB | 8 | 0 | 19.07 MH/s at 72.6 W, 1.4 GB | refused: thread 'tokio-rt-worker' (43763) panicked at sp1-gpu/crates/ | 7.6 GB, 9.6 s (alone-comp-26-v1) | no GB peak, s | 5.8 GB, 26.3 s | mines and proves core-only | +| RTX 4060 | 8 | 2 | 17.07 MH/s at 0.0 W, 1.4 GB | refused: thread 'tokio-rt-worker' (48510) panicked at sp1-gpu/crates/ | 7.4 GB, 18.4 s (alone-comp-26-v1) | no GB peak, s | 5.6 GB, 22.1 s | mines and proves core-only | +| RTX 4070 | 12 | 9 | 24.99 MH/s at 91.1 W, 1.4 GB | refused: thread 'tokio-rt-worker' (47475) panicked at sp1-gpu/crates/ | 7.6 GB, 12.1 s (alone-comp-26-v1) | 10.1 GB peak, 27.3 s | 5.6 GB, 14.3 s | mines and proves | +| RTX 4090 | 24 | 1 | 52.25 MH/s at 183.1 W, 1.7 GB | proved 5.6 s at 17.4 GB | 7.9 GB, 6.3 s (alone-comp-26-v1) | 10.7 GB peak, 26.1 s | 6.1 GB, 10.6 s | mines and proves | +| RTX 5070 | 12 | 2 | 41.89 MH/s at 137.0 W, 2.7 GB | refused: thread 'tokio-rt-worker' (53680) panicked at sp1-gpu/crates/ | 7.6 GB, 4.8 s (alone-comp-26-v1) | 10.2 GB peak, 37.2 s | 5.8 GB, 19.8 s | mines and proves | +| RTX 5090 | 32 | 2 | 98.48 MH/s at 258.2 W, 1.8 GB | proved 8.4 s at 18.3 GB | 8.0 GB, 6.3 s (alone-comp-26-v1) | 9.9 GB peak, 10.7 s | 6.3 GB, 7.4 s | mines and proves | +| RTX A5000 | 24 | 1 | 47.7 MH/s at 222.7 W, 1.5 GB | proved 6.4 s at 17.2 GB | 7.7 GB, 8.3 s (alone-comp-26-v1) | 10.5 GB peak, 34.6 s | 6.0 GB, 18.2 s | mines and proves | + +Also measured: the stock SP1 6.8.1 server refuses every card under 24 GB at `builder.rs:38` and proves the v1 shard on +the 4090 (5.6 s, 17.4 GB), the A5000 (6.4 s, 17.2 GB) and the 5090 (8.4 s, 18.3 GB); 2^27 does not fit a 10 or 8 GB card +and the v4 server hangs at the card's limit (568 and 904 s until killed) where v5 aborts in 13 s ("FLOOR abort: a device +allocation failed at slop/crates/tensor/src/inner.rs:51 ... AllocError { size: 486586112 }", exit 70; the known-failed +case of the prover-floor gate, on the 3080); the miner beside a prover costs 1.7x (5090) to 7.7x (5070) on the proof's +time and 5 to 20% of the miner's rate; the empty-shard fixture (block-72854, a first block with a genesis witness) proves +slower than the v1 shard on every card (24 to 55 s) and is not an empty live shard; Ember's two knobs are refused in the +containers, so the ladders are baseline rows (`docs/plans/ember-tune.md`, fleet priors).