igneum/infra/gpu-bench/README.md
igneum-labs d8dc8c6ba3 infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both
infra/cloud-devnet: hcloud (doctl variant) create, builder-VM provision from a git-archive source tarball,
systemd units for igneumd --devnet-suffix with a sparse --addpeer mesh and a CPU trickle miner per node,
stdlib wRPC client, experiments (latency, partition, hop, collect, observer hookup), README with the command
sequence and the Hetzner API prices of 3 Oct 2026.
infra/gpu-bench: RunPod image recipes (CUDA 12.8, ROCm), bundle, run.sh (vectors gate, 10-min raw, sweep,
inline shortcut ratio, nvcc/NVRTC/OpenCL recompile timings, results row, intake upload), bench-log template.
infra/seed-nodes: create-seed (persistent IPv4, firewall), provision on the VM, health check, addPeer from the
Mac over grpcurl, seeds.txt; igneum-seed-1 created at 188.245.5.161 (Hetzner cx23, fsn1).
docs/plans/cloud-devnet.md and docs/plans/seed-nodes.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 22:01:50 +00:00

72 lines
5.9 KiB
Markdown

# Igneum GPU bench on rented cards
One pod per card on RunPod (primary; Vast.ai variant below), each running the same bundle: the CUDA and OpenCL
test harnesses from `proto-cuda/` and `proto-opencl/` with the three checked-in packs, built on the pod with
`--serve` support, then a fixed sequence of measurements and one results row. Nothing is cloned on the pod; the
bundle is a tarball made here. Nothing here spends until a pod is started.
## The command sequence
```
cd infra/gpu-bench
./make-bundle.sh # build/igneum-gpu-bench.tar.gz (sources + scripts, about 1 MB)
# RunPod console: Pods -> Deploy -> pick the card -> template "nvidia/cuda:12.8.1-devel-ubuntu22.04" (or the image from
# Dockerfile.cuda), Community Cloud, tick SSH, 20 GB container disk -> Deploy. Note host and port from "Connect".
scp -P <port> build/igneum-gpu-bench.tar.gz root@<host>:/workspace/
ssh -p <port> root@<host> 'cd /workspace && tar xzf igneum-gpu-bench.tar.gz && cd bundle && LABEL=rtx4090 ./run.sh'
# about 45 minutes: build 2 to 5 min, gate, 10 min raw, sweep, inline, recompile timings; the row prints as RESULT
# and lands in bundle/results.md and results-<stamp>/row.md; the log is uploaded under label gpubench-rtx4090
scp -P <port> -r root@<host>:/workspace/bundle/results-* results/ # keep the raw outputs here
# RunPod console: Stop and Terminate the pod (billing per minute while it runs; storage bills while stopped)
node ../../tools/logs.mjs # the uploaded log, if the pod could not be reached again
```
Repeat for each card (the same bundle). Paste the rows into `docs/bench-log.md` with `results-template.md`.
`CUDA_ARCH=sm_86` overrides `-arch=native` if the driver in a pod is older than its toolkit.
## Cards and cost (RunPod pricing page, read 3 Oct 2026, USD per hour)
| Card | Community Cloud | Secure Cloud | One 45-minute run, community | Note |
|---|---|---|---|---|
| RTX 3060 | not on RunPod's table | not on RunPod's table | about 0.05 to 0.10 on Vast.ai (approximate) | the entry card; Vast.ai is the host for it |
| RTX 3090 | 0.22 | 0.50 | about 0.17 | Ampere sm_86, 24 GB, 6 MB L2 (approximate) |
| RTX 4090 | 0.34 | 0.74 | about 0.26 | Ada sm_89, 24 GB, 72 MB L2 (approximate) |
| RTX 5090 | 0.69 | 0.99 | about 0.52 | Blackwell sm_120, 32 GB, 96 MB L2 (bench-log); the project lead's own card is the reference row |
| RX 7900 XTX | not offered | not offered | n/a | RunPod has no AMD consumer cards; Vast.ai lists none in general (approximate). `Dockerfile.rocm` and the OpenCL lane of `run.sh` are ready for any host that has one (or the project lead's own AMD box) |
Four NVIDIA cards: about USD 2 on community cloud, USD 3.50 on secure cloud, plus cents of storage. Per-minute
billing; a pod left running costs the hourly rate.
Vast.ai variant: search for the card with "CUDA 12.8" and "verified" hosts, rent with the image
`nvidia/cuda:12.8.1-devel-ubuntu22.04` and "SSH" as the launch mode, then the same scp and ssh lines (Vast prints its
own port and host). Prices on Vast.ai are per host and change by the hour; the 3060 is usually under USD 0.10.
## What run.sh measures and why
| Step | Output | Ledger |
|---|---|---|
| Vectors gate (`--batches 3`): cache check, dataset self-test, 3 warps standalone and in batch | `gate.txt`, must say `OVERALL: PASS` or the run stops | the cross-vendor proof per card (bench-log, "What PASS means") |
| Raw bench, about 10 minutes of 2^24-hash batches at 1 GiB | `raw.txt`: Mhash/s, seconds | the per-card hash rate; sustained, so thermals and neighbours show |
| Sweep 64, 128, 256, 512, 1024 MiB, 20 batches each | `sweep-*.txt`, the 64 MiB over 1 GiB ratio | M1: the L2 cliff per card (the 5090 gave 5.9x) |
| Inline-dataset shortcut (`make-inline.sh`: every `ds[...]` load becomes `mh_word(ds, ...)` from the cache) | `inline.txt`: Mhash/s and the ratio to the raw rate | M16: the recompute attacker's rate; the M5 Max gave 0.21. The 64 MiB-cache variant inside the 5090's L2 needs a pack exported with a smaller cache (proto-metal has no flag for it yet) and is a follow-up |
| Second pack `igneum-hourly` (closed form, 128 loads) | `pack2.txt` | comparison with the 5090 and Metal tables in the bench log |
| `nvcc -cubin` of kernel.cu x3 (the worker's `prepare` path, host.cu line 554), NVRTC in process x3 (`nvrtc-time.cu`), `clBuildProgram` of kernel.cl x3 (`clbuild-time.c`, when an OpenCL ICD is visible) | `nvrtc.txt`, `clbuild.txt`, medians in the row | M11, M17: the hourly program change on real drivers |
| `--serve` ready line (`printf quit \| worker --serve`) | `serve-ready.txt` | proves the pod's binary is the worker the miner drives (ready line shows `prepare 1` when nvcc is on PATH) |
The inline binary fails host.cu's dataset self-test by construction (the dataset buffer holds the cache) and must
pass the vectors; `run.sh` reads its rate and vector lines and ignores its OVERALL.
Results row (also in `results-template.md`): card, driver and toolkit, date, Mhash/s at 1 GiB over the 10 minutes,
seconds, second-pack Mhash/s, the five sweep rates, the cliff ratio, inline Mhash/s, inline over honest, nvcc ms,
NVRTC ms, OpenCL ms, vectors, build ms.
## Files
`run.sh` (on the pod), `make-bundle.sh` (on the Mac), `make-inline.sh`, `nvrtc-time.cu`, `clbuild-time.c`,
`upload.sh` (the intake URL and key of `proto-cuda/windows-miner/upload-log.bat` as a curl line), `Dockerfile.cuda`,
`Dockerfile.rocm`, `results-template.md`, `results/` (raw outputs copied back, one directory per run).
Not tested here: there is no NVIDIA or AMD card on this Mac, so `run.sh`, `nvrtc-time.cu` and `clbuild-time.c` were
checked by read-through and `bash -n` only; the harness binaries they drive ran on the 5090 and the gfx1036 on 3 Oct
2026 (bench-log). The first pod run is the real test; if `nvrtc-time` fails to compile the kernel (a header NVRTC
rejects), the nvcc column still carries the out-of-process figure the worker uses today.