igneum/docs/analysis/prover-tiers-real-cards.md
igneum-labs 7eed16a29a Pre-public scrub, the text pass (7 October 2026, 19:5x UK): no founder name, personal login, earlier business or personal address in any tracked text file, and a gate check that keeps it so
The sweep (main's item 1): 199 tracked text files, 783 lines. The founder's full name, first name and possessive become "the founder" (sentence starts capitalised); the lowercase operating-system user name in WSL paths and commands becomes <user>; the second owner login becomes "the second owner login"; the three earlier businesses and the two other brands become "the other business", "the earlier entity", "the earlier business" and "another brand"; the Chrome profile rule names the igneum.network profile, not the profile's label. The standing commit login igneum-labs is not a founder term here: the fresh-repository step renames it in the history (docs/plans/history-rewrite.md, tools/repo/fresh-repo.sh).

The patterns never appear in plain text in the tree (a plaintext list would be the hit): tools/ci/founder-strings.b64 (perl regex, tab, a sample per row) is read by tools/ci/founder-strings-check.sh (every tracked text file, perl, known-failed first: the self-test plants each row's sample in a fixture and the hit must name the file), by tools/community/discord-hooks.mjs (the guard's founder and business rows; the test takes its fixtures from the samples) and by tools/repo/fresh-repo.sh (the business names of the rewrite rules). site/forbidden-strings.txt carries the same patterns as b64: lines, decoded case-insensitive by site/scrub.mjs and tools/ci/launch-gates-check.mjs (whose fixture now plants an encoded made-up name). The check runs in the gate's tree checks on every merge.

Not in this commit, by main's word: the 105 commit messages and 40 personal-identity commits that need the history rewrite (listed, not run), and the secrets found by gitleaks over the history (reported with owners).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 18:39:50 +00:00

10 KiB

Prover tiers on real cards: the memory matrix of the patched SP1 GPU server, measured on rented GPUs

6 October 2026, from 11:50 UTC (the founder: "rent all you need, absolute overkill", "get as many GPUs as you need to properly test everything swiftly"). Branch gpu-fleet, tools in tools/fleet/, raw logs per instance under ~/Desktop/fleet/<instance>/ (the sampler csv, every point's host log and results JSON, the miner log). This file replaces the 5090-allocation rows of docs/analysis/prover-floor.md with the cards themselves. A row that is not here yet is still running; the table is rewritten by tools/fleet/collect.py as rows land.

Method

Every box is a Vast.ai container (nvidia/cuda:12.8.1-devel-ubuntu24.04, the host's own driver) with one card. On it: the 0.3.12 Linux node 83089544 on the ten-field override (digest 7bd98cc4...), peered with the seed; the 0.3.12 CUDA worker; the patched sp1-gpu-server built on the box from SP1 v6.8.1 (c84ada1e) with proving/prover-floor/sp1-gpu-6.8.1-floor.patch v4 (sha256 e81cb0d0...) for the card's own arch (86 for Ampere, 89 for Ada, 120 for Blackwell; the 4090 box built 86,89,120); the cuda host and exporter from the prover-floor bundle (pinned ids 0x2b1a81cb... and 0x474678f3..., --mode id on every box). The fixture is the v1 shard (proving/fixtures/fees-v1-shards2.json shard 0, 4,717,439 cycles), the same as prover-floor.md.

Row What runs How it is read
idle nothing on the card nvidia-smi memory.used, power.draw
miner igneum-miner mine with the CUDA worker against the box's node, 60 s warm, 150 s sampled the mean of the STATUS line's now= over the sample; watts and memory from a 1-s sampler
stock the SDK's own sp1_gpu_server_v6.8.1 (HOME=/root), --mode compressed --shard 0 the host's RESULT line, or the refusal in its log
alone the patched server, SP1_GPU_ELEMENT_THRESHOLD 2^26 and 2^27 compressed, 2^25 and 2^26 --mode core, nothing else on the card peak = max memory.used at 1 s; own = peak minus the reading before the point
beside the same points with the miner running on the card (45 s warm before the first) the same; base = the miner's resident set

Every proof is verified by the host's own SDK verifier (the VERIFIED word in the RESULT line; the patch changes buffer sizes and the shard split, not the circuit). The server is killed and its socket unlinked around every point. A point whose allocation does not fit does not fail: the patched server hangs at the card's limit at 0% utilisation (the 3080 and the 4060 Ti 8 GB at 2^27, 15 minutes each until killed), so every prover run needs a wall-clock timeout.

The table (rewritten as rows land)

Card VRAM GB Idle MiB Miner Stock SP1 6.8.1 Patched, proves alone (own) Beside the miner (peak) Core-only beside the miner (own) Verdict
RTX 3060 12 1 23.78 MH/s at 103.7 W, 1.4 GB refused: thread 'tokio-rt-worker' (48952) panicked at sp1-gpu/crates/ 7.4 GB, 14.4 s (alone-comp-26-v1) 8.9 GB peak, 37.5 s 5.6 GB, 27.2 s mines and proves
RTX 3080 10 11 40.82 MH/s at 204.9 W, 1.5 GB refused: thread 'tokio-rt-worker' (49593) panicked at sp1-gpu/crates/ 8.0 GB, 7.1 s (alone-comp-26-v1) 9.2 GB peak, 25.6 s 5.9 GB, 19.2 s mines and proves
RTX 3090 24 1 37.79 MH/s at 228.8 W, 1.5 GB not measured: the SDK's server download stalled (killed at 553 s) 7.7 GB, 14.9 s (alone-comp-26-v1) 9.3 GB peak, 19.9 s 5.8 GB, 13.3 s mines and proves
RTX 4060 Ti 16 GB 16 0 17.58 MH/s at 72.3 W, 1.4 GB refused: thread 'tokio-rt-worker' (49293) panicked at sp1-gpu/crates/ 7.8 GB, 11.6 s (alone-comp-26-v1) 9.0 GB peak, 34.6 s 5.8 GB, 25.9 s mines and proves
RTX 4060 Ti 8 GB 8 0 19.07 MH/s at 72.6 W, 1.4 GB refused: thread 'tokio-rt-worker' (43763) panicked at sp1-gpu/crates/ 7.6 GB, 9.6 s (alone-comp-26-v1) no GB peak, s 5.8 GB, 26.3 s mines and proves core-only
RTX 4060 8 2 17.07 MH/s at 0.0 W, 1.4 GB refused: thread 'tokio-rt-worker' (48510) panicked at sp1-gpu/crates/ 7.4 GB, 18.4 s (alone-comp-26-v1) no GB peak, s 5.6 GB, 22.1 s mines and proves core-only
RTX 4070 12 9 24.99 MH/s at 91.1 W, 1.4 GB refused: thread 'tokio-rt-worker' (47475) panicked at sp1-gpu/crates/ 7.6 GB, 12.1 s (alone-comp-26-v1) 10.1 GB peak, 27.3 s 5.6 GB, 14.3 s mines and proves
RTX 4090 24 1 52.25 MH/s at 183.1 W, 1.7 GB proved 5.6 s at 17.4 GB 7.9 GB, 6.3 s (alone-comp-26-v1) 10.7 GB peak, 26.1 s 6.1 GB, 10.6 s mines and proves
RTX 5070 12 2 41.89 MH/s at 137.0 W, 2.7 GB refused: thread 'tokio-rt-worker' (53680) panicked at sp1-gpu/crates/ 7.6 GB, 4.8 s (alone-comp-26-v1) 10.2 GB peak, 37.2 s 5.8 GB, 19.8 s mines and proves
RTX 5090 32 2 98.48 MH/s at 258.2 W, 1.8 GB proved 8.4 s at 18.3 GB 8.0 GB, 6.3 s (alone-comp-26-v1) 9.9 GB peak, 10.7 s 6.3 GB, 7.4 s mines and proves
RTX A5000 24 1 47.7 MH/s at 222.7 W, 1.5 GB proved 6.4 s at 17.2 GB 7.7 GB, 8.3 s (alone-comp-26-v1) 10.5 GB peak, 34.6 s 6.0 GB, 18.2 s mines and proves

What the rows say, per tier

Measured, all eleven cards complete at 13:27Z (4090, 3090, A5000, 5090, 4070, 5070, 3060, 3080, 4060 Ti 16 GB, 4060, 4060 Ti 8 GB):

Tier What the cards say Consequence What is being done
24 GB (4090, 3090, A5000) the STOCK server proves the v1 shard (5.6 s at 17.4 GB on the 4090, 6.4 s at 17.2 GB on the A5000; on the 3090 the SDK's server download stalled and the point was killed, not a refusal); the patched 2^26 profile does it in 7.7 to 7.9 GB (6.3 s 4090, 8.3 s A5000, 14.9 s 3090) and beside the miner the peak is 9.3 GB (3090, 19.9 s) to 10.5 to 10.7 GB (A5000, 4090; 26 to 35 s); the 3090 mines 37.8 MH/s at 229 W, the weakest MH/W of the 24 GB cards mines and proves with 13 GB to spare; the patched profile frees 9.5 GB for nothing but a 1.1 to 1.3x slower proof, so a 24 GB card keeps upstream's threshold and the public line "24 GB: mines and proves" stands on real cards the 24 GB rows go into the fleet night as compressed provers at the default tier
12 GB (3060, 4070, 5070) the stock server refuses (the 24 GB gate); patched 2^26 proves alone at 7.4 to 7.6 GB (14.4 s on the 3060, 12.1 s on the 4070, 4.8 s on the 5070); BESIDE THE MINER the peak is 8.9 GB (3060, 37.5 s) to 10.1 to 10.2 GB (4070, 5070; 27.3 s and 37.2 s) of 12 GB, verified; core-only beside the miner 5.6 to 5.8 GB (14.3 to 27.2 s) a 12 GB card mines and proves compressed shards on the patched server with about 2 GB to spare before the display (Windows and a monitor take 0.5 to 1.5 GB, so a desktop 12 GB card is at the edge; a headless Linux one is fine); the 9.0 GB line of prover-floor.md is not needed for Linux headless, and core-only (5.6 to 5.8 GB) keeps 6 GB spare for a desktop the public line becomes "12 GB: mines and proves on Linux (the patched server), proves alone on a desktop; core-only mine-and-prove on a desktop once the hand-off ships"; the 4070 and 5070 join the fleet night with the miner PAUSED per segment (prove-alone profile), the 10.2 GB beside-row is the mine-and-prove candidate for a second night
16 GB (4060 Ti 16 GB) patched 2^26 alone 7.8 GB in 11.6 s; beside the miner 9.0 GB peak of 16 GB in 34.6 s; core-only beside 5.8 GB mines and proves with 7 GB spare, display or not; the 2^27 profile (10 GB alone) fits too joins the fleet night mining and proving compressed
10 GB (3080) proves alone at 2^26 (7.1 s, 8,158 MiB own); BESIDE THE MINER compressed 2^26 verified at a 9,412 MiB peak of 10,240 in 25.6 s; core-only beside 2^25 at a 7,618 MiB peak in 19.2 s; 2^27 panics after 568 s mines and proves on headless Linux with 0.8 GB spare, too close for a desktop with a display, where core-only (2.6 GB spare) is the profile the app keeps 8 and 10 GB cards "prove alone, off by default while mining" (the prover-floor agent's provedefault rows from these numbers)
8 GB (4060, 4060 Ti 8 GB) both prove alone at 2^26 (the 4060 Ti 9.6 s at 7,740 MiB of 8,188; the 4060 16.9 s at 7,655), at 2^25 compressed (13.5 s, 7,676) and core-only at 2^25 (5.8 s, 5,916) and 2^24 (11.2 s, 5,404); BESIDE THE MINER (1.4 GB resident) compressed 2^26 does not fit on either, core-only 2^25 proves verified at a 7,123 MiB (4060, 22.1 s) and 7,352 MiB (4060 Ti, 26.3 s) peak of 8,188, core 2^24 at 6,840 to 6,867 (60 to 70 s) an 8 GB card proves alone compressed, or mines and proves core-only at 2^25 with about 1 GB spare on headless Linux (a desktop with a display takes 2^24, 1.3 GB spare, at 60 to 70 s a shard); the 8 GB miner's hash while proving falls to 16.0 to 18.4 MH/s from 17.1 to 19.1 the public floor becomes "8 GB proves alone; mines and proves core-only once the hand-off ships"
32 GB (5090) 98.5 MH/s at 258 W; the stock server proves in 8.4 s at 18.3 GB; patched 2^26 alone 8.0 GB in 6.3 s; beside the miner 9.9 GB peak in 10.7 s (the miner costs 1.7x here against 4x on Ada) the strongest prover per card: beside its miner it proves a v1 shard every 11 s the fleet night's compressed prover at upstream's tier
rig one server per card at the card's profile: a 4090 rig needs 8 x 10.7 GB device memory beside its miners and about 6 GB of host RAM per server fits any 8x 4090 rig with 64 GB of host RAM phase 3 measures it
pool user nothing changes: the pool's provers carry the proofs

The empty-shard rows (the block-72854-empty-block-first fixture) prove SLOWER than the v1 shard on every card (24 to 50 s at 2^26 against 6 to 12 s), the opposite of PC 2's 3.3 s for its own empty shard; the fixture is a first block with a genesis-state witness, not an empty live shard, so those rows are not the "empty block" cost and are left out of the tier line.

Ember Tune on rented cards

nvidia-smi -pl and -lgc are refused inside a Vast container (the host's driver holds the power and clock knobs), so the two-knob ladder (tools/fleet/box-ember.sh) reports one baseline step per card: the untuned MH/s, W and MH/W. The rows are in results.json (tune_plan: baseline) and in the fleet priors section of docs/plans/ember-tune.md; a tuned point per model needs a bare-metal host or a VM with the driver inside.

Card MH/s W MH/W Plan
RTX 4070 24.77 91.3 0.2713 baseline
RTX 4090 52.24 179.9 0.2904 baseline