igneum/infra/gpu-bench
igneum-josh 5d84c539d7 infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both
infra/cloud-devnet: hcloud (doctl variant) create, builder-VM provision from a git-archive source tarball,
systemd units for igneumd --devnet-suffix with a sparse --addpeer mesh and a CPU trickle miner per node,
stdlib wRPC client, experiments (latency, partition, hop, collect, observer hookup), README with the command
sequence and the Hetzner API prices of 3 Oct 2026.
infra/gpu-bench: RunPod image recipes (CUDA 12.8, ROCm), bundle, run.sh (vectors gate, 10-min raw, sweep,
inline shortcut ratio, nvcc/NVRTC/OpenCL recompile timings, results row, intake upload), bench-log template.
infra/seed-nodes: create-seed (persistent IPv4, firewall), provision on the VM, health check, addPeer from the
Mac over grpcurl, seeds.txt; igneum-seed-1 created at 188.245.5.161 (Hetzner cx23, fsn1).
docs/plans/cloud-devnet.md and docs/plans/seed-nodes.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 22:01:50 +00:00
..
results infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
.gitignore infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
clbuild-time.c infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
Dockerfile.cuda infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
Dockerfile.rocm infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
make-bundle.sh infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
make-inline.sh infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
nvrtc-time.cu infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
README.md infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
results-template.md infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
run.sh infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
upload.sh infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00

Igneum GPU bench on rented cards

One pod per card on RunPod (primary; Vast.ai variant below), each running the same bundle: the CUDA and OpenCL test harnesses from proto-cuda/ and proto-opencl/ with the three checked-in packs, built on the pod with --serve support, then a fixed sequence of measurements and one results row. Nothing is cloned on the pod; the bundle is a tarball made here. Nothing here spends until a pod is started.

The command sequence

cd infra/gpu-bench
./make-bundle.sh                                             # build/igneum-gpu-bench.tar.gz (sources + scripts, about 1 MB)
# RunPod console: Pods -> Deploy -> pick the card -> template "nvidia/cuda:12.8.1-devel-ubuntu22.04" (or the image from
#   Dockerfile.cuda), Community Cloud, tick SSH, 20 GB container disk -> Deploy. Note host and port from "Connect".
scp -P <port> build/igneum-gpu-bench.tar.gz root@<host>:/workspace/
ssh -p <port> root@<host> 'cd /workspace && tar xzf igneum-gpu-bench.tar.gz && cd bundle && LABEL=rtx4090 ./run.sh'
# about 45 minutes: build 2 to 5 min, gate, 10 min raw, sweep, inline, recompile timings; the row prints as RESULT
#   and lands in bundle/results.md and results-<stamp>/row.md; the log is uploaded under label gpubench-rtx4090
scp -P <port> -r root@<host>:/workspace/bundle/results-* results/        # keep the raw outputs here
# RunPod console: Stop and Terminate the pod (billing per minute while it runs; storage bills while stopped)
node ../../tools/logs.mjs                                    # the uploaded log, if the pod could not be reached again

Repeat for each card (the same bundle). Paste the rows into docs/bench-log.md with results-template.md. CUDA_ARCH=sm_86 overrides -arch=native if the driver in a pod is older than its toolkit.

Cards and cost (RunPod pricing page, read 3 Oct 2026, USD per hour)

Card Community Cloud Secure Cloud One 45-minute run, community Note
RTX 3060 not on RunPod's table not on RunPod's table about 0.05 to 0.10 on Vast.ai (approximate) the entry card; Vast.ai is the host for it
RTX 3090 0.22 0.50 about 0.17 Ampere sm_86, 24 GB, 6 MB L2 (approximate)
RTX 4090 0.34 0.74 about 0.26 Ada sm_89, 24 GB, 72 MB L2 (approximate)
RTX 5090 0.69 0.99 about 0.52 Blackwell sm_120, 32 GB, 96 MB L2 (bench-log); Josh's own card is the reference row
RX 7900 XTX not offered not offered n/a RunPod has no AMD consumer cards; Vast.ai lists none in general (approximate). Dockerfile.rocm and the OpenCL lane of run.sh are ready for any host that has one (or Josh's own AMD box)

Four NVIDIA cards: about USD 2 on community cloud, USD 3.50 on secure cloud, plus cents of storage. Per-minute billing; a pod left running costs the hourly rate.

Vast.ai variant: search for the card with "CUDA 12.8" and "verified" hosts, rent with the image nvidia/cuda:12.8.1-devel-ubuntu22.04 and "SSH" as the launch mode, then the same scp and ssh lines (Vast prints its own port and host). Prices on Vast.ai are per host and change by the hour; the 3060 is usually under USD 0.10.

What run.sh measures and why

Step Output Ledger
Vectors gate (--batches 3): cache check, dataset self-test, 3 warps standalone and in batch gate.txt, must say OVERALL: PASS or the run stops the cross-vendor proof per card (bench-log, "What PASS means")
Raw bench, about 10 minutes of 2^24-hash batches at 1 GiB raw.txt: Mhash/s, seconds the per-card hash rate; sustained, so thermals and neighbours show
Sweep 64, 128, 256, 512, 1024 MiB, 20 batches each sweep-*.txt, the 64 MiB over 1 GiB ratio M1: the L2 cliff per card (the 5090 gave 5.9x)
Inline-dataset shortcut (make-inline.sh: every ds[...] load becomes mh_word(ds, ...) from the cache) inline.txt: Mhash/s and the ratio to the raw rate M16: the recompute attacker's rate; the M5 Max gave 0.21. The 64 MiB-cache variant inside the 5090's L2 needs a pack exported with a smaller cache (proto-metal has no flag for it yet) and is a follow-up
Second pack igneum-hourly (closed form, 128 loads) pack2.txt comparison with the 5090 and Metal tables in the bench log
nvcc -cubin of kernel.cu x3 (the worker's prepare path, host.cu line 554), NVRTC in process x3 (nvrtc-time.cu), clBuildProgram of kernel.cl x3 (clbuild-time.c, when an OpenCL ICD is visible) nvrtc.txt, clbuild.txt, medians in the row M11, M17: the hourly program change on real drivers
--serve ready line (printf quit | worker --serve) serve-ready.txt proves the pod's binary is the worker the miner drives (ready line shows prepare 1 when nvcc is on PATH)

The inline binary fails host.cu's dataset self-test by construction (the dataset buffer holds the cache) and must pass the vectors; run.sh reads its rate and vector lines and ignores its OVERALL.

Results row (also in results-template.md): card, driver and toolkit, date, Mhash/s at 1 GiB over the 10 minutes, seconds, second-pack Mhash/s, the five sweep rates, the cliff ratio, inline Mhash/s, inline over honest, nvcc ms, NVRTC ms, OpenCL ms, vectors, build ms.

Files

run.sh (on the pod), make-bundle.sh (on the Mac), make-inline.sh, nvrtc-time.cu, clbuild-time.c, upload.sh (the intake URL and key of proto-cuda/windows-miner/upload-log.bat as a curl line), Dockerfile.cuda, Dockerfile.rocm, results-template.md, results/ (raw outputs copied back, one directory per run).

Not tested here: there is no NVIDIA or AMD card on this Mac, so run.sh, nvrtc-time.cu and clbuild-time.c were checked by read-through and bash -n only; the harness binaries they drive ran on the 5090 and the gfx1036 on 3 Oct 2026 (bench-log). The first pod run is the real test; if nvrtc-time fails to compile the kernel (a header NVRTC rejects), the nvcc column still carries the out-of-process figure the worker uses today.