Prover tiers on real cards: the first four cards (4090, A5000, 4070, 5070), the method, the per-tier consequences; fleet: RunPod pods, companions pushed with every stage

This commit is contained in:
igneum-labs 2026-10-06 12:56:51 +00:00
parent 5d1fe8a697
commit b39eeff070
4 changed files with 82 additions and 6 deletions

View file

@ -0,0 +1,67 @@
# Prover tiers on real cards: the memory matrix of the patched SP1 GPU server, measured on rented GPUs
6 October 2026, from 11:50 UTC (the project lead: "rent all you need, absolute overkill", "get as many GPUs as you need to properly
test everything swiftly"). Branch `gpu-fleet`, tools in `tools/fleet/`, raw logs per instance under
`~/Desktop/fleet/<instance>/` (the sampler csv, every point's host log and results JSON, the miner log). This file
replaces the 5090-allocation rows of `docs/analysis/prover-floor.md` with the cards themselves. A row that is not here
yet is still running; the table is rewritten by `tools/fleet/collect.py` as rows land.
## Method
Every box is a Vast.ai container (nvidia/cuda:12.8.1-devel-ubuntu24.04, the host's own driver) with one card. On it:
the 0.3.12 Linux node 83089544 on the ten-field override (digest 7bd98cc4...), peered with the seed; the 0.3.12 CUDA
worker; the patched `sp1-gpu-server` built on the box from SP1 v6.8.1 (c84ada1e) with
`proving/prover-floor/sp1-gpu-6.8.1-floor.patch` v4 (sha256 e81cb0d0...) for the card's own arch (86 for Ampere, 89
for Ada, 120 for Blackwell; the 4090 box built 86,89,120); the cuda host and exporter from the prover-floor bundle
(pinned ids 0x2b1a81cb... and 0x474678f3..., `--mode id` on every box). The fixture is the v1 shard
(`proving/fixtures/fees-v1-shards2.json` shard 0, 4,717,439 cycles), the same as `prover-floor.md`.
| Row | What runs | How it is read |
|---|---|---|
| idle | nothing on the card | `nvidia-smi memory.used`, `power.draw` |
| miner | `igneum-miner mine` with the CUDA worker against the box's node, 60 s warm, 150 s sampled | the mean of the STATUS line's `now=` over the sample; watts and memory from a 1-s sampler |
| stock | the SDK's own `sp1_gpu_server_v6.8.1` (HOME=/root), `--mode compressed --shard 0` | the host's RESULT line, or the refusal in its log |
| alone | the patched server, `SP1_GPU_ELEMENT_THRESHOLD` 2^26 and 2^27 compressed, 2^25 and 2^26 `--mode core`, nothing else on the card | peak = max `memory.used` at 1 s; own = peak minus the reading before the point |
| beside | the same points with the miner running on the card (45 s warm before the first) | the same; base = the miner's resident set |
Every proof is verified by the host's own SDK verifier (the VERIFIED word in the RESULT line; the patch changes buffer
sizes and the shard split, not the circuit). The server is killed and its socket unlinked around every point. A point
whose allocation does not fit does not fail: the patched server hangs at the card's limit at 0% utilisation (the 3080
and the 4060 Ti 8 GB at 2^27, 15 minutes each until killed), so every prover run needs a wall-clock timeout.
## The table (rewritten as rows land)
| Card | VRAM GB | Idle MiB | Miner | Stock SP1 6.8.1 | Patched, proves alone (own) | Beside the miner (peak) | Core-only beside the miner (own) | Verdict |
|---|---|---|---|---|---|---|---|---|
| RTX 4070 | 12 | 9 | 24.99 MH/s at 91.1 W, 1.4 GB | refused: thread 'tokio-rt-worker' (47475) panicked at sp1-gpu/crates/ | 7.6 GB, 12.1 s (alone-comp-26-v1) | 10.1 GB peak, 27.3 s | 5.6 GB, 14.3 s | mines and proves |
| RTX 4090 | 24 | 1 | 52.25 MH/s at 183.1 W, 1.7 GB | proved 5.6 s at 17.4 GB | 7.9 GB, 6.3 s (alone-comp-26-v1) | 10.7 GB peak, 26.1 s | 6.1 GB, 10.6 s | mines and proves |
| RTX 5070 | 12 | 2 | 41.89 MH/s at 137.0 W, 2.7 GB | refused: thread 'tokio-rt-worker' (53680) panicked at sp1-gpu/crates/ | 7.6 GB, 4.8 s (alone-comp-26-v1) | 10.2 GB peak, 37.2 s | 5.8 GB, 19.8 s | mines and proves |
| RTX A5000 | 24 | 1 | 47.7 MH/s at 222.7 W, 1.5 GB | proved 6.4 s at 17.2 GB | 7.7 GB, 8.3 s (alone-comp-26-v1) | 10.5 GB peak, 34.6 s | 6.0 GB, 18.2 s | mines and proves |
## What the rows say, per tier
Measured so far (the 4090, A5000, 4070 and 5070 complete; the 3080, 4060 Ti 8 and 16 GB, 3060, 4060, 3090 and 5090 running):
| Tier | What the cards say | Consequence | What is being done |
|---|---|---|---|
| 24 GB (4090, A5000) | the STOCK server proves the v1 shard (5.6 s at 17.4 GB on the 4090, 6.4 s at 17.2 GB on the A5000), so no patch is needed to prove alone; the patched 2^26 profile does it in 7.7 to 7.9 GB (6.3 and 8.3 s) and beside the miner the peak is 10.5 to 10.7 GB at 26 to 35 s (the miner costs 4.1x on the 4090) | mines and proves with 13 GB to spare; the patched profile frees 9.5 GB for nothing but a 1.1 to 1.3x slower proof, so a 24 GB card keeps upstream's threshold and the public line "24 GB: mines and proves" stands on real cards | the 24 GB rows go into the fleet night as compressed provers at the default tier |
| 12 GB (4070, 5070) | the stock server refuses (the 24 GB gate); patched 2^26 proves alone at 7.6 GB (12.1 s on the 4070, 4.8 s on the 5070); BESIDE THE MINER the peak is 10.1 to 10.2 GB of 12 GB (27.3 s and 37.2 s), verified; core-only beside the miner 5.6 to 5.8 GB (14.3 and 19.8 s) | a 12 GB card mines and proves compressed shards on the patched server with about 2 GB to spare before the display (Windows and a monitor take 0.5 to 1.5 GB, so a desktop 12 GB card is at the edge; a headless Linux one is fine); the 9.0 GB line of prover-floor.md is not needed for Linux headless, and core-only (5.6 to 5.8 GB) keeps 6 GB spare for a desktop | the public line becomes "12 GB: mines and proves on Linux (the patched server), proves alone on a desktop; core-only mine-and-prove on a desktop once the hand-off ships"; the 4070 and 5070 join the fleet night with the miner PAUSED per segment (prove-alone profile), the 10.2 GB beside-row is the mine-and-prove candidate for a second night |
| 16 GB (4060 Ti 16 GB) | running | | |
| 10 GB (3080) and 8 GB (4060, 4060 Ti 8 GB) | the 3080 proves alone at 2^26 (7.0 to 7.2 s, 7.9 to 8.2 GB) and the 4060 Ti 8 GB too (9.6 s, 7.74 GB of 8.19); 2^27 does not fit either and the server hangs instead of failing | an 8 GB card proves alone (not beside its miner: 7.7 + 1.4 GB is over 8 GB); the public floor moves from "12 GB proves" to "8 GB proves alone, slowly"; a profile under 7.5 GB (2^25 compressed, core-only 2^25 at 5.9 GB) is the 8 GB mine-and-prove candidate | the beside-the-miner and 2^25 rows on the 3080 and 4060 Ti are running |
| 32 GB (5090) | running | | |
| rig | one server per card at the card's profile: a 4090 rig needs 8 x 10.7 GB device memory beside its miners and about 6 GB of host RAM per server | fits any 8x 4090 rig with 64 GB of host RAM | phase 3 measures it |
| pool user | nothing changes: the pool's provers carry the proofs | | |
The empty-shard rows (the `block-72854-empty-block-first` fixture) prove SLOWER than the v1 shard on every card (24 to 50 s at 2^26 against 6 to 12 s), the opposite of PC 2's 3.3 s for its own empty shard; the fixture is a first block with a genesis-state witness, not an empty live shard, so those rows are not the "empty block" cost and are left out of the tier line.
## Ember Tune on rented cards
`nvidia-smi -pl` and `-lgc` are refused inside a Vast container (the host's driver holds the power and clock knobs), so
the two-knob ladder (`tools/fleet/box-ember.sh`) reports one baseline step per card: the untuned MH/s, W and MH/W. The
rows are in `results.json` (`tune_plan: baseline`) and in the fleet priors section of `docs/plans/ember-tune.md`; a
tuned point per model needs a bare-metal host or a VM with the driver inside.
| Card | MH/s | W | MH/W | Plan |
|---|---|---|---|---|
| RTX 4070 | 24.77 | 91.3 | 0.2713 | baseline |
| RTX 4090 | 52.24 | 179.9 | 0.2904 | baseline |

View file

@ -15,8 +15,7 @@ reg = json.load(open(f"{ROOT}/boxes.json"))
def gb(mib): return round(mib / 1024, 1)
rows = {}; table = []
for iid, b in reg.items():
if not b.get("phase") == "1": continue
mp = f"{ROOT}/{iid}/matrix.json"
mp = f"{ROOT}/{iid}/matrix.json" # every box with a matrix (a box moves to phase 2 when its Ember ladder ends)
if not os.path.exists(mp): continue
m = json.load(open(mp)); pts = {r["name"]: r for r in m["rows"] if r.get("row") == "point"}; miner = next((r for r in m["rows"] if r.get("row") == "miner"), {})
total = m["total_mib"]; idle = m["idle_mib"]
@ -25,7 +24,9 @@ for iid, b in reg.items():
ml = f"{ROOT}/{iid}/miner.log"
if os.path.exists(ml):
vals = [float(l.split(" now=")[1].split()[0]) for l in open(ml, errors="replace") if "STATUS" in l and " now=" in l]
if vals: miner = dict(miner, mhs=round(sum(vals[-12:]) / len(vals[-12:]), 2), mhs_raw_lines=len(vals))
# only when the script's own row is broken (the first matrices parsed a second now= field): the miner-alone
# window is the first 21 STATUS lines (60 s warm + 150 s sample at 10 s), minus the warm-up's first 6
if vals and (not miner.get("mhs") or miner.get("mhs", 0) > 10000): w = vals[6:21]; miner = dict(miner, mhs=round(sum(w) / len(w), 2), mhs_recomputed=True)
def ok(n): p = pts.get(n); return p and p.get("verified") == "yes"
def own(n): p = pts.get(n); return p.get("own_mib") if p else None
def secs(n): p = pts.get(n); return p.get("prove_s") if p else None
@ -37,7 +38,7 @@ for iid, b in reg.items():
elif alone and core_b: verdict = "mines and proves core-only"
elif alone: verdict = "proves alone"
else: verdict = "mining only"
card = m["card"].replace("NVIDIA GeForce ", "").replace("NVIDIA ", "")
card = b.get("card") or m["card"]
r = {"card": card, "vram_gb": round(total / 1024), "mhs": miner.get("mhs"), "watts": miner.get("watts"), "miner_gb": gb(miner.get("own_mib", 0)),
"stock": "refused: " + (stock.get("err") or "")[:60] if stock.get("verified") != "yes" else f"proved {stock.get('prove_s')} s at {gb(stock.get('peak_mib', 0))} GB",
"prove_alone_gb": gb(own(alone)) if alone else None, "shard_s": secs(alone), "prove_alone_point": alone,

View file

@ -50,6 +50,13 @@ def refresh_ssh(reg):
for iid, b in reg.items():
if b.get("provider", "vast") == "vast" and iid in live:
i = live[iid]; patch(iid, ssh_host=i.get("ssh_host"), ssh_port=i.get("ssh_port"), actual_status=i.get("actual_status"), status_msg=(i.get("status_msg") or "")[:120])
if any(b.get("provider") == "runpod" and b.get("state") != "destroyed" for b in reg.values()):
import runpod
for p in runpod.pods():
pid = str(p.get("id"))
if pid in reg and reg[pid].get("state") != "destroyed":
pm = p.get("portMappings") or {}
patch(pid, ssh_host=p.get("publicIp") or None, ssh_port=pm.get("22") if p.get("publicIp") else None, actual_status=p.get("desiredStatus"), port_map=pm, status_msg="")
return load()
def ssh(b, cmd, timeout=120, capture=True):
@ -122,7 +129,8 @@ def run(script, labels, env=""):
hub = next((v for v in load().values() if v.get("hub") and v.get("hub_peer")), None)
if hub and "HUB_PEER" not in env: env = f"HUB_PEER={hub['hub_peer']} " + env
for iid, b in boxes(labels).items():
if scp(b, [os.path.join(HERE, script)], "/root/fleet/in/") != 0: print(b["label"], "scp failed"); continue
files = [os.path.join(HERE, script), os.path.join(HERE, "box-kill.sh")] + ([os.path.join(HERE, "box-prover.py")] if "prover" in script else [])
if scp(b, files, "/root/fleet/in/") != 0: print(b["label"], "scp failed"); continue
name = os.path.basename(script)
logname = {"box-matrix.sh": "matrix", "box-ember.sh": "ember", "box-prover.sh": "prover"}.get(name, name)
# rotate the stage's log before the start: an old run's end line must never be read as this run's (the 5070, 12:33Z)

View file

@ -16,7 +16,7 @@ IMAGE = "runpod/pytorch:2.8.0-py3.11-cuda12.8.1-cudnn-devel-ubuntu22.04"
PUB = open(os.path.expanduser("~/.ssh/igneum-fleet.pub")).read().strip()
def call(method, path, body=None):
req = urllib.request.Request(BASE + path, data=(json.dumps(body).encode() if body is not None else None),
headers={"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}, method=method)
headers={"Authorization": "Bearer " + KEY, "Content-Type": "application/json", "User-Agent": "curl/8.7.1", "Accept": "*/*"}, method=method) # Cloudflare 1010 refuses urllib's default agent
try:
with urllib.request.urlopen(req, timeout=90) as r:
t = r.read(); return json.loads(t) if t else {}