infra/cloud-devnet: hcloud (doctl variant) create, builder-VM provision from a git-archive source tarball, systemd units for igneumd --devnet-suffix with a sparse --addpeer mesh and a CPU trickle miner per node, stdlib wRPC client, experiments (latency, partition, hop, collect, observer hookup), README with the command sequence and the Hetzner API prices of 3 Oct 2026. infra/gpu-bench: RunPod image recipes (CUDA 12.8, ROCm), bundle, run.sh (vectors gate, 10-min raw, sweep, inline shortcut ratio, nvcc/NVRTC/OpenCL recompile timings, results row, intake upload), bench-log template. infra/seed-nodes: create-seed (persistent IPv4, firewall), provision on the VM, health check, addPeer from the Mac over grpcurl, seeds.txt; igneum-seed-1 created at 188.245.5.161 (Hetzner cx23, fsn1). docs/plans/cloud-devnet.md and docs/plans/seed-nodes.md. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
72 lines
5.9 KiB
Markdown
72 lines
5.9 KiB
Markdown
# Igneum GPU bench on rented cards
|
|
|
|
One pod per card on RunPod (primary; Vast.ai variant below), each running the same bundle: the CUDA and OpenCL
|
|
test harnesses from `proto-cuda/` and `proto-opencl/` with the three checked-in packs, built on the pod with
|
|
`--serve` support, then a fixed sequence of measurements and one results row. Nothing is cloned on the pod; the
|
|
bundle is a tarball made here. Nothing here spends until a pod is started.
|
|
|
|
## The command sequence
|
|
|
|
```
|
|
cd infra/gpu-bench
|
|
./make-bundle.sh # build/igneum-gpu-bench.tar.gz (sources + scripts, about 1 MB)
|
|
# RunPod console: Pods -> Deploy -> pick the card -> template "nvidia/cuda:12.8.1-devel-ubuntu22.04" (or the image from
|
|
# Dockerfile.cuda), Community Cloud, tick SSH, 20 GB container disk -> Deploy. Note host and port from "Connect".
|
|
scp -P <port> build/igneum-gpu-bench.tar.gz root@<host>:/workspace/
|
|
ssh -p <port> root@<host> 'cd /workspace && tar xzf igneum-gpu-bench.tar.gz && cd bundle && LABEL=rtx4090 ./run.sh'
|
|
# about 45 minutes: build 2 to 5 min, gate, 10 min raw, sweep, inline, recompile timings; the row prints as RESULT
|
|
# and lands in bundle/results.md and results-<stamp>/row.md; the log is uploaded under label gpubench-rtx4090
|
|
scp -P <port> -r root@<host>:/workspace/bundle/results-* results/ # keep the raw outputs here
|
|
# RunPod console: Stop and Terminate the pod (billing per minute while it runs; storage bills while stopped)
|
|
node ../../tools/logs.mjs # the uploaded log, if the pod could not be reached again
|
|
```
|
|
|
|
Repeat for each card (the same bundle). Paste the rows into `docs/bench-log.md` with `results-template.md`.
|
|
`CUDA_ARCH=sm_86` overrides `-arch=native` if the driver in a pod is older than its toolkit.
|
|
|
|
## Cards and cost (RunPod pricing page, read 3 Oct 2026, USD per hour)
|
|
|
|
| Card | Community Cloud | Secure Cloud | One 45-minute run, community | Note |
|
|
|---|---|---|---|---|
|
|
| RTX 3060 | not on RunPod's table | not on RunPod's table | about 0.05 to 0.10 on Vast.ai (approximate) | the entry card; Vast.ai is the host for it |
|
|
| RTX 3090 | 0.22 | 0.50 | about 0.17 | Ampere sm_86, 24 GB, 6 MB L2 (approximate) |
|
|
| RTX 4090 | 0.34 | 0.74 | about 0.26 | Ada sm_89, 24 GB, 72 MB L2 (approximate) |
|
|
| RTX 5090 | 0.69 | 0.99 | about 0.52 | Blackwell sm_120, 32 GB, 96 MB L2 (bench-log); the project lead's own card is the reference row |
|
|
| RX 7900 XTX | not offered | not offered | n/a | RunPod has no AMD consumer cards; Vast.ai lists none in general (approximate). `Dockerfile.rocm` and the OpenCL lane of `run.sh` are ready for any host that has one (or the project lead's own AMD box) |
|
|
|
|
Four NVIDIA cards: about USD 2 on community cloud, USD 3.50 on secure cloud, plus cents of storage. Per-minute
|
|
billing; a pod left running costs the hourly rate.
|
|
|
|
Vast.ai variant: search for the card with "CUDA 12.8" and "verified" hosts, rent with the image
|
|
`nvidia/cuda:12.8.1-devel-ubuntu22.04` and "SSH" as the launch mode, then the same scp and ssh lines (Vast prints its
|
|
own port and host). Prices on Vast.ai are per host and change by the hour; the 3060 is usually under USD 0.10.
|
|
|
|
## What run.sh measures and why
|
|
|
|
| Step | Output | Ledger |
|
|
|---|---|---|
|
|
| Vectors gate (`--batches 3`): cache check, dataset self-test, 3 warps standalone and in batch | `gate.txt`, must say `OVERALL: PASS` or the run stops | the cross-vendor proof per card (bench-log, "What PASS means") |
|
|
| Raw bench, about 10 minutes of 2^24-hash batches at 1 GiB | `raw.txt`: Mhash/s, seconds | the per-card hash rate; sustained, so thermals and neighbours show |
|
|
| Sweep 64, 128, 256, 512, 1024 MiB, 20 batches each | `sweep-*.txt`, the 64 MiB over 1 GiB ratio | M1: the L2 cliff per card (the 5090 gave 5.9x) |
|
|
| Inline-dataset shortcut (`make-inline.sh`: every `ds[...]` load becomes `mh_word(ds, ...)` from the cache) | `inline.txt`: Mhash/s and the ratio to the raw rate | M16: the recompute attacker's rate; the M5 Max gave 0.21. The 64 MiB-cache variant inside the 5090's L2 needs a pack exported with a smaller cache (proto-metal has no flag for it yet) and is a follow-up |
|
|
| Second pack `igneum-hourly` (closed form, 128 loads) | `pack2.txt` | comparison with the 5090 and Metal tables in the bench log |
|
|
| `nvcc -cubin` of kernel.cu x3 (the worker's `prepare` path, host.cu line 554), NVRTC in process x3 (`nvrtc-time.cu`), `clBuildProgram` of kernel.cl x3 (`clbuild-time.c`, when an OpenCL ICD is visible) | `nvrtc.txt`, `clbuild.txt`, medians in the row | M11, M17: the hourly program change on real drivers |
|
|
| `--serve` ready line (`printf quit \| worker --serve`) | `serve-ready.txt` | proves the pod's binary is the worker the miner drives (ready line shows `prepare 1` when nvcc is on PATH) |
|
|
|
|
The inline binary fails host.cu's dataset self-test by construction (the dataset buffer holds the cache) and must
|
|
pass the vectors; `run.sh` reads its rate and vector lines and ignores its OVERALL.
|
|
|
|
Results row (also in `results-template.md`): card, driver and toolkit, date, Mhash/s at 1 GiB over the 10 minutes,
|
|
seconds, second-pack Mhash/s, the five sweep rates, the cliff ratio, inline Mhash/s, inline over honest, nvcc ms,
|
|
NVRTC ms, OpenCL ms, vectors, build ms.
|
|
|
|
## Files
|
|
|
|
`run.sh` (on the pod), `make-bundle.sh` (on the Mac), `make-inline.sh`, `nvrtc-time.cu`, `clbuild-time.c`,
|
|
`upload.sh` (the intake URL and key of `proto-cuda/windows-miner/upload-log.bat` as a curl line), `Dockerfile.cuda`,
|
|
`Dockerfile.rocm`, `results-template.md`, `results/` (raw outputs copied back, one directory per run).
|
|
|
|
Not tested here: there is no NVIDIA or AMD card on this Mac, so `run.sh`, `nvrtc-time.cu` and `clbuild-time.c` were
|
|
checked by read-through and `bash -n` only; the harness binaries they drive ran on the 5090 and the gfx1036 on 3 Oct
|
|
2026 (bench-log). The first pod run is the real test; if `nvrtc-time` fails to compile the kernel (a header NVRTC
|
|
rejects), the nvcc column still carries the out-of-process figure the worker uses today.
|