Prover tiers on real cards (eleven rented cards, patched SP1 server) and the rental cost of hash: the analysis file and its two bench-log entries, cut from gpu-fleet for the site

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 19:21:28 +00:00
parent 7d3df308fa
commit 8e6de32ea8
2 changed files with 128 additions and 0 deletions

View file

@ -0,0 +1,75 @@
# Prover tiers on real cards: the memory matrix of the patched SP1 GPU server, measured on rented GPUs
6 October 2026, from 11:50 UTC (the project lead: "rent all you need, absolute overkill", "get as many GPUs as you need to properly
test everything swiftly"). Branch `gpu-fleet`, tools in `tools/fleet/`, raw logs per instance under
`~/Desktop/fleet/<instance>/` (the sampler csv, every point's host log and results JSON, the miner log). This file
replaces the 5090-allocation rows of `docs/analysis/prover-floor.md` with the cards themselves. A row that is not here
yet is still running; the table is rewritten by `tools/fleet/collect.py` as rows land.
## Method
Every box is a Vast.ai container (nvidia/cuda:12.8.1-devel-ubuntu24.04, the host's own driver) with one card. On it:
the 0.3.12 Linux node 83089544 on the ten-field override (digest 7bd98cc4...), peered with the seed; the 0.3.12 CUDA
worker; the patched `sp1-gpu-server` built on the box from SP1 v6.8.1 (c84ada1e) with
`proving/prover-floor/sp1-gpu-6.8.1-floor.patch` v4 (sha256 e81cb0d0...) for the card's own arch (86 for Ampere, 89
for Ada, 120 for Blackwell; the 4090 box built 86,89,120); the cuda host and exporter from the prover-floor bundle
(pinned ids 0x2b1a81cb... and 0x474678f3..., `--mode id` on every box). The fixture is the v1 shard
(`proving/fixtures/fees-v1-shards2.json` shard 0, 4,717,439 cycles), the same as `prover-floor.md`.
| Row | What runs | How it is read |
|---|---|---|
| idle | nothing on the card | `nvidia-smi memory.used`, `power.draw` |
| miner | `igneum-miner mine` with the CUDA worker against the box's node, 60 s warm, 150 s sampled | the mean of the STATUS line's `now=` over the sample; watts and memory from a 1-s sampler |
| stock | the SDK's own `sp1_gpu_server_v6.8.1` (HOME=/root), `--mode compressed --shard 0` | the host's RESULT line, or the refusal in its log |
| alone | the patched server, `SP1_GPU_ELEMENT_THRESHOLD` 2^26 and 2^27 compressed, 2^25 and 2^26 `--mode core`, nothing else on the card | peak = max `memory.used` at 1 s; own = peak minus the reading before the point |
| beside | the same points with the miner running on the card (45 s warm before the first) | the same; base = the miner's resident set |
Every proof is verified by the host's own SDK verifier (the VERIFIED word in the RESULT line; the patch changes buffer
sizes and the shard split, not the circuit). The server is killed and its socket unlinked around every point. A point
whose allocation does not fit does not fail: the patched server hangs at the card's limit at 0% utilisation (the 3080
and the 4060 Ti 8 GB at 2^27, 15 minutes each until killed), so every prover run needs a wall-clock timeout.
## The table (rewritten as rows land)
| Card | VRAM GB | Idle MiB | Miner | Stock SP1 6.8.1 | Patched, proves alone (own) | Beside the miner (peak) | Core-only beside the miner (own) | Verdict |
|---|---|---|---|---|---|---|---|---|
| RTX 3060 | 12 | 1 | 23.78 MH/s at 103.7 W, 1.4 GB | refused: thread 'tokio-rt-worker' (48952) panicked at sp1-gpu/crates/ | 7.4 GB, 14.4 s (alone-comp-26-v1) | 8.9 GB peak, 37.5 s | 5.6 GB, 27.2 s | mines and proves |
| RTX 3080 | 10 | 11 | 40.82 MH/s at 204.9 W, 1.5 GB | refused: thread 'tokio-rt-worker' (49593) panicked at sp1-gpu/crates/ | 8.0 GB, 7.1 s (alone-comp-26-v1) | 9.2 GB peak, 25.6 s | 5.9 GB, 19.2 s | mines and proves |
| RTX 3090 | 24 | 1 | 37.79 MH/s at 228.8 W, 1.5 GB | not measured: the SDK's server download stalled (killed at 553 s) | 7.7 GB, 14.9 s (alone-comp-26-v1) | 9.3 GB peak, 19.9 s | 5.8 GB, 13.3 s | mines and proves |
| RTX 4060 Ti 16 GB | 16 | 0 | 17.58 MH/s at 72.3 W, 1.4 GB | refused: thread 'tokio-rt-worker' (49293) panicked at sp1-gpu/crates/ | 7.8 GB, 11.6 s (alone-comp-26-v1) | 9.0 GB peak, 34.6 s | 5.8 GB, 25.9 s | mines and proves |
| RTX 4060 Ti 8 GB | 8 | 0 | 19.07 MH/s at 72.6 W, 1.4 GB | refused: thread 'tokio-rt-worker' (43763) panicked at sp1-gpu/crates/ | 7.6 GB, 9.6 s (alone-comp-26-v1) | no GB peak, s | 5.8 GB, 26.3 s | mines and proves core-only |
| RTX 4060 | 8 | 2 | 17.07 MH/s at 0.0 W, 1.4 GB | refused: thread 'tokio-rt-worker' (48510) panicked at sp1-gpu/crates/ | 7.4 GB, 18.4 s (alone-comp-26-v1) | no GB peak, s | 5.6 GB, 22.1 s | mines and proves core-only |
| RTX 4070 | 12 | 9 | 24.99 MH/s at 91.1 W, 1.4 GB | refused: thread 'tokio-rt-worker' (47475) panicked at sp1-gpu/crates/ | 7.6 GB, 12.1 s (alone-comp-26-v1) | 10.1 GB peak, 27.3 s | 5.6 GB, 14.3 s | mines and proves |
| RTX 4090 | 24 | 1 | 52.25 MH/s at 183.1 W, 1.7 GB | proved 5.6 s at 17.4 GB | 7.9 GB, 6.3 s (alone-comp-26-v1) | 10.7 GB peak, 26.1 s | 6.1 GB, 10.6 s | mines and proves |
| RTX 5070 | 12 | 2 | 41.89 MH/s at 137.0 W, 2.7 GB | refused: thread 'tokio-rt-worker' (53680) panicked at sp1-gpu/crates/ | 7.6 GB, 4.8 s (alone-comp-26-v1) | 10.2 GB peak, 37.2 s | 5.8 GB, 19.8 s | mines and proves |
| RTX 5090 | 32 | 2 | 98.48 MH/s at 258.2 W, 1.8 GB | proved 8.4 s at 18.3 GB | 8.0 GB, 6.3 s (alone-comp-26-v1) | 9.9 GB peak, 10.7 s | 6.3 GB, 7.4 s | mines and proves |
| RTX A5000 | 24 | 1 | 47.7 MH/s at 222.7 W, 1.5 GB | proved 6.4 s at 17.2 GB | 7.7 GB, 8.3 s (alone-comp-26-v1) | 10.5 GB peak, 34.6 s | 6.0 GB, 18.2 s | mines and proves |
## What the rows say, per tier
Measured, all eleven cards complete at 13:27Z (4090, 3090, A5000, 5090, 4070, 5070, 3060, 3080, 4060 Ti 16 GB, 4060, 4060 Ti 8 GB):
| Tier | What the cards say | Consequence | What is being done |
|---|---|---|---|
| 24 GB (4090, 3090, A5000) | the STOCK server proves the v1 shard (5.6 s at 17.4 GB on the 4090, 6.4 s at 17.2 GB on the A5000; on the 3090 the SDK's server download stalled and the point was killed, not a refusal); the patched 2^26 profile does it in 7.7 to 7.9 GB (6.3 s 4090, 8.3 s A5000, 14.9 s 3090) and beside the miner the peak is 9.3 GB (3090, 19.9 s) to 10.5 to 10.7 GB (A5000, 4090; 26 to 35 s); the 3090 mines 37.8 MH/s at 229 W, the weakest MH/W of the 24 GB cards | mines and proves with 13 GB to spare; the patched profile frees 9.5 GB for nothing but a 1.1 to 1.3x slower proof, so a 24 GB card keeps upstream's threshold and the public line "24 GB: mines and proves" stands on real cards | the 24 GB rows go into the fleet night as compressed provers at the default tier |
| 12 GB (3060, 4070, 5070) | the stock server refuses (the 24 GB gate); patched 2^26 proves alone at 7.4 to 7.6 GB (14.4 s on the 3060, 12.1 s on the 4070, 4.8 s on the 5070); BESIDE THE MINER the peak is 8.9 GB (3060, 37.5 s) to 10.1 to 10.2 GB (4070, 5070; 27.3 s and 37.2 s) of 12 GB, verified; core-only beside the miner 5.6 to 5.8 GB (14.3 to 27.2 s) | a 12 GB card mines and proves compressed shards on the patched server with about 2 GB to spare before the display (Windows and a monitor take 0.5 to 1.5 GB, so a desktop 12 GB card is at the edge; a headless Linux one is fine); the 9.0 GB line of prover-floor.md is not needed for Linux headless, and core-only (5.6 to 5.8 GB) keeps 6 GB spare for a desktop | the public line becomes "12 GB: mines and proves on Linux (the patched server), proves alone on a desktop; core-only mine-and-prove on a desktop once the hand-off ships"; the 4070 and 5070 join the fleet night with the miner PAUSED per segment (prove-alone profile), the 10.2 GB beside-row is the mine-and-prove candidate for a second night |
| 16 GB (4060 Ti 16 GB) | patched 2^26 alone 7.8 GB in 11.6 s; beside the miner 9.0 GB peak of 16 GB in 34.6 s; core-only beside 5.8 GB | mines and proves with 7 GB spare, display or not; the 2^27 profile (10 GB alone) fits too | joins the fleet night mining and proving compressed |
| 10 GB (3080) | proves alone at 2^26 (7.1 s, 8,158 MiB own); BESIDE THE MINER compressed 2^26 verified at a 9,412 MiB peak of 10,240 in 25.6 s; core-only beside 2^25 at a 7,618 MiB peak in 19.2 s; 2^27 panics after 568 s | mines and proves on headless Linux with 0.8 GB spare, too close for a desktop with a display, where core-only (2.6 GB spare) is the profile | the app keeps 8 and 10 GB cards "prove alone, off by default while mining" (the prover-floor agent's provedefault rows from these numbers) |
| 8 GB (4060, 4060 Ti 8 GB) | both prove alone at 2^26 (the 4060 Ti 9.6 s at 7,740 MiB of 8,188; the 4060 16.9 s at 7,655), at 2^25 compressed (13.5 s, 7,676) and core-only at 2^25 (5.8 s, 5,916) and 2^24 (11.2 s, 5,404); BESIDE THE MINER (1.4 GB resident) compressed 2^26 does not fit on either, core-only 2^25 proves verified at a 7,123 MiB (4060, 22.1 s) and 7,352 MiB (4060 Ti, 26.3 s) peak of 8,188, core 2^24 at 6,840 to 6,867 (60 to 70 s) | an 8 GB card proves alone compressed, or mines and proves core-only at 2^25 with about 1 GB spare on headless Linux (a desktop with a display takes 2^24, 1.3 GB spare, at 60 to 70 s a shard); the 8 GB miner's hash while proving falls to 16.0 to 18.4 MH/s from 17.1 to 19.1 | the public floor becomes "8 GB proves alone; mines and proves core-only once the hand-off ships" |
| 32 GB (5090) | 98.5 MH/s at 258 W; the stock server proves in 8.4 s at 18.3 GB; patched 2^26 alone 8.0 GB in 6.3 s; beside the miner 9.9 GB peak in 10.7 s (the miner costs 1.7x here against 4x on Ada) | the strongest prover per card: beside its miner it proves a v1 shard every 11 s | the fleet night's compressed prover at upstream's tier |
| rig | one server per card at the card's profile: a 4090 rig needs 8 x 10.7 GB device memory beside its miners and about 6 GB of host RAM per server | fits any 8x 4090 rig with 64 GB of host RAM | phase 3 measures it |
| pool user | nothing changes: the pool's provers carry the proofs | | |
The empty-shard rows (the `block-72854-empty-block-first` fixture) prove SLOWER than the v1 shard on every card (24 to 50 s at 2^26 against 6 to 12 s), the opposite of PC 2's 3.3 s for its own empty shard; the fixture is a first block with a genesis-state witness, not an empty live shard, so those rows are not the "empty block" cost and are left out of the tier line.
## Ember Tune on rented cards
`nvidia-smi -pl` and `-lgc` are refused inside a Vast container (the host's driver holds the power and clock knobs), so
the two-knob ladder (`tools/fleet/box-ember.sh`) reports one baseline step per card: the untuned MH/s, W and MH/W. The
rows are in `results.json` (`tune_plan: baseline`) and in the fleet priors section of `docs/plans/ember-tune.md`; a
tuned point per model needs a bare-metal host or a VM with the driver inside.
| Card | MH/s | W | MH/W | Plan |
|---|---|---|---|---|
| RTX 4070 | 24.77 | 91.3 | 0.2713 | baseline |
| RTX 4090 | 52.24 | 179.9 | 0.2904 | baseline |

View file

@ -2544,3 +2544,56 @@ task now runs ...Programs\Igneum Miner\igneum-app.exe --power-helper"), then the
task with no prompt ("-pl 460: set to 460.00 W from 575.00 W", "-pl 160: set to 160.00 W from 100.00 W"). Task task with no prompt ("-pl 460: set to 460.00 W from 575.00 W", "-pl 160: set to 160.00 W from 100.00 W"). Task
Running, Highest, user Admin. Cards: 5090 221 W at 1845 MHz, 4070 75.8 W at 1860 MHz, 9070 XT 202 W, all mining Running, Highest, user Admin. Cards: 5090 221 W at 1845 MHz, 4070 75.8 W at 1860 MHz, 9070 XT 202 W, all mining
through the installed app. Ember closed 16:53Z. through the installed app. Ember closed 16:53Z.
## Prover tiers on real cards: the rented fleet, 6 October 2026 (branch gpu-fleet)
From 11:50 UTC, Vast.ai containers (nvidia/cuda:12.8.1-devel-ubuntu24.04, the host's driver), one card each, the
0.3.12 Linux node 83089544 on the ten-field override, the 0.3.12 CUDA worker, the patched SP1 server built on each box
from `proving/prover-floor/sp1-gpu-6.8.1-floor.patch` v4 (e81cb0d0...) for the card's own arch, the cuda host with the
pinned ids 0x2b1a81cb... and 0x474678f3...; the v1 shard (`fees-v1-shards2.json` shard 0, 4,717,439 cycles); every
proof VERIFIED by the host's own SDK verifier; peak = nvidia-smi memory.used sampled once a second (a per-second loop,
not `-l 1`, which buffers and ignores SIGTERM in a container); own = peak minus the reading before the point; beside
= the card's miner running (its resident set is the base). Runner `tools/fleet/box-matrix.sh`, collector
`tools/fleet/collect.py`, raw logs `~/Desktop/fleet/<instance>/`, the analysis `docs/analysis/prover-tiers-real-cards.md`.
| Card | VRAM GB | Idle MiB | Miner | Stock SP1 6.8.1 | Patched, proves alone (own) | Beside the miner (peak) | Core-only beside the miner (own) | Verdict |
|---|---|---|---|---|---|---|---|---|
| RTX 3060 | 12 | 1 | 23.78 MH/s at 103.7 W, 1.4 GB | refused: thread 'tokio-rt-worker' (48952) panicked at sp1-gpu/crates/ | 7.4 GB, 14.4 s (alone-comp-26-v1) | 8.9 GB peak, 37.5 s | 5.6 GB, 27.2 s | mines and proves |
| RTX 3080 | 10 | 11 | 40.82 MH/s at 204.9 W, 1.5 GB | refused: thread 'tokio-rt-worker' (49593) panicked at sp1-gpu/crates/ | 8.0 GB, 7.1 s (alone-comp-26-v1) | 9.2 GB peak, 25.6 s | 5.9 GB, 19.2 s | mines and proves |
| RTX 3090 | 24 | 1 | 37.79 MH/s at 228.8 W, 1.5 GB | not measured: the SDK's server download stalled (killed at 553 s) | 7.7 GB, 14.9 s (alone-comp-26-v1) | 9.3 GB peak, 19.9 s | 5.8 GB, 13.3 s | mines and proves |
| RTX 4060 Ti 16 GB | 16 | 0 | 17.58 MH/s at 72.3 W, 1.4 GB | refused: thread 'tokio-rt-worker' (49293) panicked at sp1-gpu/crates/ | 7.8 GB, 11.6 s (alone-comp-26-v1) | 9.0 GB peak, 34.6 s | 5.8 GB, 25.9 s | mines and proves |
| RTX 4060 Ti 8 GB | 8 | 0 | 19.07 MH/s at 72.6 W, 1.4 GB | refused: thread 'tokio-rt-worker' (43763) panicked at sp1-gpu/crates/ | 7.6 GB, 9.6 s (alone-comp-26-v1) | no GB peak, s | 5.8 GB, 26.3 s | mines and proves core-only |
| RTX 4060 | 8 | 2 | 17.07 MH/s at 0.0 W, 1.4 GB | refused: thread 'tokio-rt-worker' (48510) panicked at sp1-gpu/crates/ | 7.4 GB, 18.4 s (alone-comp-26-v1) | no GB peak, s | 5.6 GB, 22.1 s | mines and proves core-only |
| RTX 4070 | 12 | 9 | 24.99 MH/s at 91.1 W, 1.4 GB | refused: thread 'tokio-rt-worker' (47475) panicked at sp1-gpu/crates/ | 7.6 GB, 12.1 s (alone-comp-26-v1) | 10.1 GB peak, 27.3 s | 5.6 GB, 14.3 s | mines and proves |
| RTX 4090 | 24 | 1 | 52.25 MH/s at 183.1 W, 1.7 GB | proved 5.6 s at 17.4 GB | 7.9 GB, 6.3 s (alone-comp-26-v1) | 10.7 GB peak, 26.1 s | 6.1 GB, 10.6 s | mines and proves |
| RTX 5070 | 12 | 2 | 41.89 MH/s at 137.0 W, 2.7 GB | refused: thread 'tokio-rt-worker' (53680) panicked at sp1-gpu/crates/ | 7.6 GB, 4.8 s (alone-comp-26-v1) | 10.2 GB peak, 37.2 s | 5.8 GB, 19.8 s | mines and proves |
| RTX 5090 | 32 | 2 | 98.48 MH/s at 258.2 W, 1.8 GB | proved 8.4 s at 18.3 GB | 8.0 GB, 6.3 s (alone-comp-26-v1) | 9.9 GB peak, 10.7 s | 6.3 GB, 7.4 s | mines and proves |
| RTX A5000 | 24 | 1 | 47.7 MH/s at 222.7 W, 1.5 GB | proved 6.4 s at 17.2 GB | 7.7 GB, 8.3 s (alone-comp-26-v1) | 10.5 GB peak, 34.6 s | 6.0 GB, 18.2 s | mines and proves |
Also measured: the stock SP1 6.8.1 server refuses every card under 24 GB at `builder.rs:38` and proves the v1 shard on
the 4090 (5.6 s, 17.4 GB), the A5000 (6.4 s, 17.2 GB) and the 5090 (8.4 s, 18.3 GB); 2^27 does not fit a 10 or 8 GB card
and the v4 server hangs at the card's limit (568 and 904 s until killed) where v5 aborts in 13 s ("FLOOR abort: a device
allocation failed at slop/crates/tensor/src/inner.rs:51 ... AllocError { size: 486586112 }", exit 70; the known-failed
case of the prover-floor gate, on the 3080); the miner beside a prover costs 1.7x (5090) to 7.7x (5070) on the proof's
time and 5 to 20% of the miner's rate; the empty-shard fixture (block-72854, a first block with a genesis witness) proves
slower than the v1 shard on every card (24 to 55 s) and is not an empty live shard; Ember's two knobs are refused in the
containers, so the ladders are baseline rows (`docs/plans/ember-tune.md`, fleet priors).
## Rental cost of hash, 6 October 2026 (branch gpu-fleet): what a GH/s costs by the hour against the devnet
Measured on the rented fleet (RunPod community pods, list prices, 18:45Z): 38 wave pods (4090, A4000, L4, 3090, 3070, 4070 Ti,
A5000, 4000 Ada) ran 1,748 MH/s inside jobs (median pod 28 MH/s) for USD 20.44 an hour, USD 0.0117 per MH/s-hour; the 8x 4090 rig
459 MH/s at 1,636 W for USD 5.92 an hour, USD 0.0129 per MH/s-hour; a single 5090 pod 98 to 128 MH/s for USD 0.41 to 0.74 an hour.
The live devnet's difficulty read 1,156,040,186 at 1 block a second at 19:00Z, so the whole network was about 1.16 GH/s, and the
rented fleet was most of it. The live litepaper's line ("2 GH/s for USD 13/h vs 280 MH/s devnet") is corrected to: 1.75 GH/s for
USD 20 an hour on community pods, against a devnet of 1.16 GH/s.
| Buyer | What USD 20/h buys | Against the devnet (1.16 GH/s) | Against mainnet scale |
|---|---|---|---|
| Home miner, 8 to 12 GB card (11 to 28 MH/s) | nothing: the card is owned, 0.15 to 0.2 kW | one card is 1 to 2 percent of the devnet | one card is noise at a TH/s |
| Rig, 8x 4090 (459 MH/s, USD 5.92/h rented) | 1.7 rigs | one rig is 40 percent of the devnet | one rig is 0.05 percent of a TH/s |
| Renter at RunPod list prices | 1.75 GH/s while cards exist | 150 percent of the devnet: overtaken for USD 15/h | a TH/s costs USD 11,700 an hour and the market cannot supply it: asked for 20 pods of any of 8 card types at 18:59Z to 19:15Z, RunPod gave 0 ("no instances currently available") |
Consequence: the devnet's hash is rentable for the price of a dinner, so nothing on it is a security result; the counter-ASIC and
finality work is tested there for correctness, not for cost. The cost argument only starts at the TH/s scale, where the rental
market's supply (not its price) is the limit, and that number belongs in the litepaper with this caveat.