Prover tiers on real cards: ten cards, the 8 GB beside-the-miner rows (core-only 2^25 fits); fleet: the rig script and retry

This commit is contained in:
igneum-labs 2026-10-06 13:20:06 +00:00
parent 9c6b62ab5e
commit a1f83bf16f
3 changed files with 120 additions and 2 deletions

View file

@ -36,6 +36,7 @@ and the 4060 Ti 8 GB at 2^27, 15 minutes each until killed), so every prover run
| RTX 3060 | 12 | 1 | 23.78 MH/s at 103.7 W, 1.4 GB | refused: thread 'tokio-rt-worker' (48952) panicked at sp1-gpu/crates/ | 7.4 GB, 14.4 s (alone-comp-26-v1) | 8.9 GB peak, 37.5 s | 5.6 GB, 27.2 s | mines and proves |
| RTX 3080 | 10 | 11 | 40.82 MH/s at 204.9 W, 1.5 GB | refused: thread 'tokio-rt-worker' (49593) panicked at sp1-gpu/crates/ | 8.0 GB, 7.1 s (alone-comp-26-v1) | 9.2 GB peak, 25.6 s | 5.9 GB, 19.2 s | mines and proves |
| RTX 4060 Ti | 16 | 0 | 17.58 MH/s at 72.3 W, 1.4 GB | refused: thread 'tokio-rt-worker' (49293) panicked at sp1-gpu/crates/ | 7.8 GB, 11.6 s (alone-comp-26-v1) | 9.0 GB peak, 34.6 s | 5.8 GB, 25.9 s | mines and proves |
| RTX 4060 Ti | 8 | 0 | 19.07 MH/s at 72.6 W, 1.4 GB | refused: thread 'tokio-rt-worker' (43763) panicked at sp1-gpu/crates/ | 7.6 GB, 9.6 s (alone-comp-26-v1) | no GB peak, s | 5.8 GB, 26.3 s | mines and proves core-only |
| RTX 4070 | 12 | 9 | 24.99 MH/s at 91.1 W, 1.4 GB | refused: thread 'tokio-rt-worker' (47475) panicked at sp1-gpu/crates/ | 7.6 GB, 12.1 s (alone-comp-26-v1) | 10.1 GB peak, 27.3 s | 5.6 GB, 14.3 s | mines and proves |
| RTX 4090 | 24 | 1 | 52.25 MH/s at 183.1 W, 1.7 GB | proved 5.6 s at 17.4 GB | 7.9 GB, 6.3 s (alone-comp-26-v1) | 10.7 GB peak, 26.1 s | 6.1 GB, 10.6 s | mines and proves |
| RTX 5070 | 12 | 2 | 41.89 MH/s at 137.0 W, 2.7 GB | refused: thread 'tokio-rt-worker' (53680) panicked at sp1-gpu/crates/ | 7.6 GB, 4.8 s (alone-comp-26-v1) | 10.2 GB peak, 37.2 s | 5.8 GB, 19.8 s | mines and proves |
@ -44,7 +45,7 @@ and the 4060 Ti 8 GB at 2^27, 15 minutes each until killed), so every prover run
## What the rows say, per tier
Measured so far (eight cards complete: 4090, A5000, 5090, 4070, 5070, 3060, 3080, 4060 Ti 16 GB; the 4060, 4060 Ti 8 GB and 3090 running):
Measured (ten cards complete: 4090, A5000, 5090, 4070, 5070, 3060, 3080, 4060 Ti 16 GB, 4060, 4060 Ti 8 GB; the 3090 running):
| Tier | What the cards say | Consequence | What is being done |
|---|---|---|---|
@ -52,7 +53,7 @@ Measured so far (eight cards complete: 4090, A5000, 5090, 4070, 5070, 3060, 3080
| 12 GB (3060, 4070, 5070) | the stock server refuses (the 24 GB gate); patched 2^26 proves alone at 7.4 to 7.6 GB (14.4 s on the 3060, 12.1 s on the 4070, 4.8 s on the 5070); BESIDE THE MINER the peak is 8.9 GB (3060, 37.5 s) to 10.1 to 10.2 GB (4070, 5070; 27.3 s and 37.2 s) of 12 GB, verified; core-only beside the miner 5.6 to 5.8 GB (14.3 to 27.2 s) | a 12 GB card mines and proves compressed shards on the patched server with about 2 GB to spare before the display (Windows and a monitor take 0.5 to 1.5 GB, so a desktop 12 GB card is at the edge; a headless Linux one is fine); the 9.0 GB line of prover-floor.md is not needed for Linux headless, and core-only (5.6 to 5.8 GB) keeps 6 GB spare for a desktop | the public line becomes "12 GB: mines and proves on Linux (the patched server), proves alone on a desktop; core-only mine-and-prove on a desktop once the hand-off ships"; the 4070 and 5070 join the fleet night with the miner PAUSED per segment (prove-alone profile), the 10.2 GB beside-row is the mine-and-prove candidate for a second night |
| 16 GB (4060 Ti 16 GB) | patched 2^26 alone 7.8 GB in 11.6 s; beside the miner 9.0 GB peak of 16 GB in 34.6 s; core-only beside 5.8 GB | mines and proves with 7 GB spare, display or not; the 2^27 profile (10 GB alone) fits too | joins the fleet night mining and proving compressed |
| 10 GB (3080) | proves alone at 2^26 (7.1 s, 8,158 MiB own); BESIDE THE MINER compressed 2^26 verified at a 9,412 MiB peak of 10,240 in 25.6 s; core-only beside 2^25 at a 7,618 MiB peak in 19.2 s; 2^27 panics after 568 s | mines and proves on headless Linux with 0.8 GB spare, too close for a desktop with a display, where core-only (2.6 GB spare) is the profile | the app keeps 8 and 10 GB cards "prove alone, off by default while mining" (the prover-floor agent's provedefault rows from these numbers) |
| 8 GB (4060, 4060 Ti 8 GB) | the 4060 Ti proves alone at 2^26 (9.6 s, 7,740 MiB of 8,188), at 2^25 (13.5 s, 7,676) and core-only at 2^25 (5.8 s, 5,916) and 2^24 (11.2 s, 5,404); beside its 1.4 GB miner the 2^26 compressed point hangs (7.7 + 1.4 over 8.2 GB); the core-only beside rows are running | an 8 GB card proves alone, slowly, or mines; mine-and-prove on 8 GB is core-only at 2^25 or 2^24 (5.4 to 5.9 GB plus 1.4) once the hand-off ships | the public floor moves from "12 GB proves" to "8 GB proves alone" |
| 8 GB (4060, 4060 Ti 8 GB) | both prove alone at 2^26 (the 4060 Ti 9.6 s at 7,740 MiB of 8,188; the 4060 16.9 s at 7,655), at 2^25 compressed (13.5 s, 7,676) and core-only at 2^25 (5.8 s, 5,916) and 2^24 (11.2 s, 5,404); BESIDE THE MINER (1.4 GB resident) compressed 2^26 does not fit on either, core-only 2^25 proves verified at a 7,123 MiB (4060, 22.1 s) and 7,352 MiB (4060 Ti, 26.3 s) peak of 8,188, core 2^24 at 6,840 to 6,867 (60 to 70 s) | an 8 GB card proves alone compressed, or mines and proves core-only at 2^25 with about 1 GB spare on headless Linux (a desktop with a display takes 2^24, 1.3 GB spare, at 60 to 70 s a shard); the 8 GB miner's hash while proving falls to 16.0 to 18.4 MH/s from 17.1 to 19.1 | the public floor becomes "8 GB proves alone; mines and proves core-only once the hand-off ships" |
| 32 GB (5090) | 98.5 MH/s at 258 W; the stock server proves in 8.4 s at 18.3 GB; patched 2^26 alone 8.0 GB in 6.3 s; beside the miner 9.9 GB peak in 10.7 s (the miner costs 1.7x here against 4x on Ada) | the strongest prover per card: beside its miner it proves a v1 shard every 11 s | the fleet night's compressed prover at upstream's tier |
| rig | one server per card at the card's profile: a 4090 rig needs 8 x 10.7 GB device memory beside its miners and about 6 GB of host RAM per server | fits any 8x 4090 rig with 64 GB of host RAM | phase 3 measures it |
| pool user | nothing changes: the pool's provers carry the proofs | | |

79
tools/fleet/box-rig.sh Executable file
View file

@ -0,0 +1,79 @@
#!/usr/bin/env bash
# Phase 3 on an 8-card rig (RunPod pod, a container: no systemd, so the rig installer's unit steps are exercised with
# --preflight-only and --dry-run and the units' own scripts are run by hand). After box-setup.sh (node, workers, the
# patched server for the rig's arch, the cuda host):
# A. the rig installer: packaging/linux/install-rig.sh --preflight-only, then --dry-run with the published package
# B. eight miners on the devnet, one per card, 150 s: MH/s per card and the rig's sum, watts
# C. seven core-only provers in parallel (one sp1-gpu-server per card, CUDA_VISIBLE_DEVICES=n, socket /tmp/sp1-cuda-n),
# the v1 shard at 2^25 in a loop for PROVE_SECS: core proofs per minute for the rig and per card, host RAM
# D. the eighth card compressing the v1 shard in a loop for PROVE_SECS: compressed proofs per minute (the compression
# throughput of one card)
# E. the hand-off cost by file: a core proof (2^25, about 25 MB) copied card-to-card over the rig's disk and read back
# F. the aggregator: --mode chain over 8 block fixtures on the big card (the 8-block chain: shards plus aggregation)
# Every result a RESULT line in /root/fleet/out/rig.log; rig.json at the end.
set -uo pipefail
F=/root/fleet; OUT=$F/out; B=/opt/igneum/pkg/bin; FLOOR=/opt/igneum-floor; HOST=$FLOOR/bin/igneum-prove-host
FIX=$FLOOR/prove/proving/fixtures; V1=$FIX/fees-v1-shards2.json
LABEL="${LABEL:-rig}"; WALLET="${WALLET:-0x1919191919191919191919191919191919191919}"; PROVE_SECS="${PROVE_SECS:-300}"; MINE_SECS="${MINE_SECS:-150}"
mkdir -p $OUT $F/mine; exec >> $OUT/rig.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
N=$(nvidia-smi --query-gpu=name --format=csv,noheader | wc -l)
echo "RESULT rig_start $(stamp) label=$LABEL cards=$N $(nvidia-smi --query-gpu=name,memory.total --format=csv,noheader | head -1) host_ram_gb=$(( $(awk '/MemTotal/{print $2}' /proc/meminfo) / 1048576 )) cores=$(nproc)"
for b in igneum-prove-host igneum-prove-export; do for d in $FLOOR/target/release $FLOOR/prove/proving/igneum-prove/target/release; do [ -x $d/$b ] && ln -sfn $d/$b $FLOOR/bin/$b; done; done
# A. the installer's real parts
if [ -d $F/in/linux ]; then
bash $F/in/linux/install-rig.sh --preflight-only > $OUT/rig-preflight.log 2>&1; echo "RESULT installer_preflight exit=$? fails=$(grep -c 'FAIL' $OUT/rig-preflight.log) warns=$(grep -c 'warn' $OUT/rig-preflight.log) cards=$(grep -c '^ card' $OUT/rig-preflight.log)"
bash $F/in/linux/install-rig.sh --dry-run --yes --wallet "$WALLET" --name "$LABEL" --package-url https://dl.igneum.network/dl/public/igneum-hive-0.3.12.tar.gz --package-sha256 7972af92e7cd9a032303eca4d95b533f53e0e68d1b9cae5bfe406a5b7c30a454 --package-size 24506282 > $OUT/rig-dryrun.log 2>&1; echo "RESULT installer_dryrun exit=$? steps=$(grep -c '^\[dry-run\]' $OUT/rig-dryrun.log) manifest=$(grep -c 'signature verifies' $OUT/rig-dryrun.log) package=$(grep -c 'sha256 7972af92' $OUT/rig-dryrun.log) prover_decision=\"$(grep -o 'prover: .*' $OUT/rig-dryrun.log | head -1 | cut -c1-120)\""
fi
# B. eight miners
cd $F/mine && rm -rf packs/devnet && $B/igneum-miner export-pack grpc://127.0.0.1:26610 packs/devnet > $OUT/rig-export-pack.log 2>&1
pids=()
for d in $(seq 0 $((N-1))); do
nohup $B/igneum-miner mine grpc://127.0.0.1:26610 1 100000000 "$LABEL-gpu$d" --worker $B/igneum-worker-cuda --worker-args "--device $d --pack packs/devnet" --prepare-packs packs/prepare-$d --exit-on-seed-change --evm-address "$WALLET" --payout-label "$LABEL-gpu$d" --status-secs 10 > $OUT/rig-miner-$d.log 2>&1 &
pids+=($!)
done
cd $F; sleep 60
( while :; do nvidia-smi --query-gpu=index,power.draw,memory.used,utilization.gpu --format=csv,noheader,nounits; sleep 1; done ) > $OUT/rig-samp-mine.csv & SP=$!
sleep "$MINE_SECS"; pkill -P $SP; kill $SP 2>/dev/null
sum=0
for d in $(seq 0 $((N-1))); do
r=$(grep STATUS $OUT/rig-miner-$d.log | grep -o ' now=[0-9.]*' | tail -n 10 | cut -d= -f2 | awk '{s+=$1; n++} END {if (n) printf "%.2f", s/n; else print 0}')
w=$(awk -F', *' -v i=$d '$1==i {s+=$2; n++} END {if (n) printf "%.0f", s/n; else print 0}' $OUT/rig-samp-mine.csv)
echo "RESULT miner card=$d mhs=$r watts=$w"; sum=$(awk -v a=$sum -v b=$r 'BEGIN {print a+b}')
done
echo "RESULT rig_miners $(stamp) cards=$N sum_mhs=$sum watts=$(awk -F', *' '{s+=$2; n++} END {printf "%.0f", s/n*'$N'}' $OUT/rig-samp-mine.csv)"
for p in "${pids[@]}"; do kill $p 2>/dev/null; done; pkill -f '^/opt/igneum/pkg/bin/igneum-worker-cuda'; sleep 3
# C + D. seven core-only provers and one compressing card, in parallel
pkill -9 -x sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock
prove_loop() { # <device> <mode> <threshold> <seconds>
local d=$1 mode=$2 thr=$3 secs=$4 t0=$(date +%s) n=0 tot=0
while [ $(( $(date +%s) - t0 )) -lt $secs ]; do
env HOME=$FLOOR/home CUDA_VISIBLE_DEVICES=$d IGNEUM_CUDA_DEVICE=$d SP1_PROVER=cuda RUST_LOG=off SP1_GPU_ELEMENT_THRESHOLD=$thr timeout 600 $HOST $V1 --mode $mode --shard 0 --prover "$WALLET" --out $OUT/rig-$mode-$d.json > $OUT/rig-$mode-$d.log 2>&1
s=$(grep -E "^RESULT ($mode) shard" $OUT/rig-$mode-$d.log | tail -1 | grep -o 'prove [0-9.]* s' | grep -o '[0-9.]*' | head -1)
v=$(grep -E "^RESULT ($mode) shard" $OUT/rig-$mode-$d.log | tail -1 | grep -c VERIFIED)
[ "$v" = 1 ] && { n=$((n+1)); tot=$(awk -v a=$tot -v b=${s:-0} 'BEGIN {print a+b}'); }
echo "$(stamp) card=$d $mode n=$n last_s=${s:-fail} verified=$v" >> $OUT/rig-loop-$d.log
done
echo "RESULT prove_loop card=$d mode=$mode threshold=$thr seconds=$secs proofs=$n mean_prove_s=$(awk -v a=$tot -v n=$n 'BEGIN {if (n) printf "%.1f", a/n; else print 0}') per_min=$(awk -v n=$n -v s=$secs 'BEGIN {printf "%.2f", n*60/s}')"
}
( while :; do echo "$(date +%s) $(free -m | awk '/Mem:/{print $3}') $(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | tr '\n' ',')"; sleep 5; done ) > $OUT/rig-samp-prove.csv & SP=$!
lp=()
for d in $(seq 0 $((N-2))); do prove_loop $d core 33554432 "$PROVE_SECS" & lp+=($!); done
prove_loop $((N-1)) compressed 67108864 "$PROVE_SECS" & lp+=($!)
for p in "${lp[@]}"; do wait $p; done
pkill -P $SP; kill $SP 2>/dev/null
echo "RESULT rig_prove $(stamp) host_ram_used_mb_max=$(awk '{if ($2>m) m=$2} END {print m}' $OUT/rig-samp-prove.csv) core_cards=$((N-1)) core_per_min=$(grep 'RESULT prove_loop' $OUT/rig.log | grep 'mode=core' | grep -o 'per_min=[0-9.]*' | cut -d= -f2 | awk '{s+=$1} END {printf "%.2f", s}') compressed_per_min=$(grep 'RESULT prove_loop' $OUT/rig.log | grep 'mode=compressed' | grep -o 'per_min=[0-9.]*' | cut -d= -f2)"
pkill -9 -x sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock
# E. the hand-off cost by file
cf=$(ls -S $F/mine/*.bin $OUT/*.bin 2>/dev/null | head -1); [ -z "$cf" ] && cf=$(find / -name 'block-*-core.bin' -size +1M 2>/dev/null | head -1)
if [ -n "$cf" ]; then
sz=$(stat -c %s "$cf"); t0=$(date +%s%N); cp "$cf" /tmp/handoff.bin; sync; t1=$(date +%s%N); cat /tmp/handoff.bin > /dev/null; t2=$(date +%s%N)
echo "RESULT handoff file=$(basename $cf) bytes=$sz copy_ms=$(( (t1-t0)/1000000 )) read_ms=$(( (t2-t1)/1000000 )) loopback_tcp_ms=$( (nc -l -p 29999 > /dev/null & sleep 0.3; t=$(date +%s%N); nc -q 0 127.0.0.1 29999 < /tmp/handoff.bin; echo $(( ($(date +%s%N)-t)/1000000 )) ) 2>/dev/null)"
fi
# F. the 8-block chain on the big card
list=""; for f in block-72854-empty-block-first block-58927-empty-reward block-72803-skipped-copies block-56-transfers block-78-increment block-341-shards2 block-344-shards4 block-338-shard1; do [ -f $FIX/$f.json ] && list="$list${list:+,}$FIX/$f.json"; done
t0=$(date +%s)
env HOME=$FLOOR/home CUDA_VISIBLE_DEVICES=$((N-1)) IGNEUM_CUDA_DEVICE=$((N-1)) SP1_PROVER=cuda RUST_LOG=off timeout 1800 $HOST --mode chain --chain "$list" --prover "$WALLET" --save-shards --out $OUT/rig-chain.json > $OUT/rig-chain.log 2>&1; rc=$?
echo "RESULT chain rc=$rc wall_s=$(( $(date +%s) - t0 )) blocks=$(echo "$list" | tr ',' '\n' | wc -l) $(grep -E '^RESULT chain' $OUT/rig-chain.log | tail -1 | cut -c1-200)"
pkill -9 -x sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock
echo "RESULT rig_done $(stamp)"

38
tools/fleet/rig-retry.py Normal file
View file

@ -0,0 +1,38 @@
#!/usr/bin/env python3
"""Retries the two 8-card rigs on Vast and RunPod every ten minutes until both are held or RIG_RETRY_UNTIL passes."""
import time, sys, os, secrets, datetime
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import fleet, vast, runpod
START = ["bash", "-c", "apt-get update -qq >/dev/null 2>&1; DEBIAN_FRONTEND=noninteractive apt-get install -y -qq openssh-server >/dev/null 2>&1; mkdir -p /run/sshd /root/.ssh; echo \"$PUBLIC_KEY\" > /root/.ssh/authorized_keys; chmod 700 /root/.ssh; chmod 600 /root/.ssh/authorized_keys; sed -i 's/^#\\?PermitRootLogin.*/PermitRootLogin prohibit-password/' /etc/ssh/sshd_config; exec /usr/sbin/sshd -D"]
def log(s): print(datetime.datetime.utcnow().strftime("%H:%M:%SZ"), s, flush=True)
def have(label): return any(b["label"] == label and b.get("state") != "destroyed" for b in fleet.load().values())
def try_vast(gpu, label, archs):
offers = [o for o in vast.search(gpu, n=12, min_cores=16, gpus=8, min_disk=150) if (o.get("cpu_ram") or 0)/1024 >= 96]
offers.sort(key=lambda o: (o.get("reliability2", 0) < 0.98, o["dph_total"]))
for o in offers[:3]:
try: d = vast.call("PUT", f"/v0/asks/{o['id']}/", {"client_id": "me", "image": vast.IMAGE, "disk": 150, "label": label, "runtype": "ssh", "cancel_unavail": True})
except SystemExit as e: log(f"{label} vast: {str(e)[-60:]}"); return False
if not d.get("success"): continue
iid = d["new_contract"]; time.sleep(40)
if not any(str(i.get("id")) == str(iid) for i in vast.instances()): log(f"{label} vast {iid} cancelled"); continue
vast.ledger({"t": fleet.now(), "event": "rent", "instance": iid, "offer": o["id"], "label": label, "gpu": gpu, "num_gpus": 8, "dph": o["dph_total"], "disk": 150, "geo": o.get("geolocation")})
fleet.patch(iid, label=label, card=gpu, vram_mb=o.get("gpu_ram"), archs=archs, offer=o["id"], dph=o["dph_total"], num_gpus=8, provider="vast", phase="3", state="renting", rented_at=fleet.now(), wallet="0x"+secrets.token_hex(20), geo=o.get("geolocation"), doing=f"phase 3: the 8x {gpu.split()[-1]} rig")
log(f"{label} -> vast {iid} ${o['dph_total']:.2f}/h {o.get('geolocation')}"); return True
return False
def try_runpod(gpu, label, archs):
for cloud in ("SECURE", "COMMUNITY"):
body = {"name": label, "imageName": "nvidia/cuda:12.8.1-devel-ubuntu24.04", "gpuTypeIds": [gpu], "gpuCount": 8, "containerDiskInGb": 120, "volumeInGb": 0, "cloudType": cloud, "ports": ["22/tcp"], "env": {"PUBLIC_KEY": runpod.PUB}, "supportPublicIp": True, "computeType": "GPU", "dockerStartCmd": START}
try: d = runpod.call("POST", "/pods", body)
except SystemExit as e: continue
iid = d.get("id")
runpod.ledger({"t": fleet.now(), "event": "rent", "provider": "runpod", "instance": iid, "label": label, "gpu": gpu, "num_gpus": 8, "dph": d.get("costPerHr"), "disk": 120})
fleet.patch(iid, label=label, card=gpu.replace("NVIDIA GeForce ", ""), vram_mb=32607 if "5090" in gpu else 24564, archs=archs, dph=d.get("costPerHr") or 0, num_gpus=8, provider="runpod", phase="3", state="renting", rented_at=fleet.now(), wallet="0x"+secrets.token_hex(20), cloud=cloud, doing=f"phase 3: the 8x {gpu.split()[-1]} rig")
log(f"{label} -> runpod {iid} {cloud} ${d.get('costPerHr')}/h"); return True
return False
until = float(os.environ.get("RIG_RETRY_UNTIL", time.time() + 6 * 3600))
while time.time() < until:
for gpu, label, archs, rpg in (("RTX 4090", "rig-4090x8", "89", "NVIDIA GeForce RTX 4090"), ("RTX 5090", "rig-5090x8", "120", "NVIDIA GeForce RTX 5090")):
if have(label): continue
if not try_vast(gpu, label, archs): try_runpod(rpg, label, archs)
if have("rig-4090x8") and have("rig-5090x8"): log("both rigs held"); break
time.sleep(600)