igneum/infra/gpu-bench
igneum-labs 58cc162c4a relay: three auth tiers, signed run tasks, machine secrets, retention; clients on headers; TZ=UTC and curl -K checks (X23 X24 X25 X26 X27 X28 X29 G13 G14)
Relay (X23, X27): the intake key is its own tier (upload and file drops only, RELAY_INTAKE_COMPAT=0 closes it);
a run task needs an Ed25519 signature by the Mac run key over {to, nonce, body sha256, flags} (RELAY_RUN_PUB,
401 without) and an HMAC tag with the target's machine secret that the agent verifies before anything runs;
results and registration are bound to the machine the secret proves (403 on a forged from).
X24: every client and Mac tool sends x-relay-token as a header to /api/relay?fn=; the path token stays for the
phone page only. X25: the agent arms the logon task only for a restart a task asked for and disarms on start
and exit. X26: 30-day retention with blob deletion, feed capped at 100, the dl base as RELAY_DL_BASE held by the
agent, never in a body. X28: GET inbox never acks (POST inbox does), RELAY-REBOOT on its own line and only with a
reboot flag, 120/min and 10 failed auths/min per IP, no username or folder on register, WSL sudo scoped to
apt-get and dpkg with SETENV, no password on a command line. X29: the intake key reaches curl through -K in
upload.sh and both upload-log.bat; tools/ci/curl-header-check.sh fails the class. G14: TZ=UTC in ship-app.mjs
and publish-jobs.sh; tools/ci/commit-tz-check.sh fails the class; history-rewrite.md names the .old-2026-10-05
files as the values in the history. The handler moved to relay/lib/handler.mjs with injected sql and blobs
(relay/lib/blob.mjs holds @vercel/blob) so relay/test/handler.test.mjs drives it without a database:
47 tests across 6 suites, all green.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 18:46:13 +00:00
..
results infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
.gitignore infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
clbuild-time.c infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
Dockerfile.cuda infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
Dockerfile.rocm infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
make-bundle.sh infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
make-inline.sh infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
nvrtc-time.cu infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
README.md infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
results-template.md infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
run.sh infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both 2026-10-03 22:01:50 +00:00
upload.sh relay: three auth tiers, signed run tasks, machine secrets, retention; clients on headers; TZ=UTC and curl -K checks (X23 X24 X25 X26 X27 X28 X29 G13 G14) 2026-10-05 18:46:13 +00:00

Igneum GPU bench on rented cards

One pod per card on RunPod (primary; Vast.ai variant below), each running the same bundle: the CUDA and OpenCL test harnesses from proto-cuda/ and proto-opencl/ with the three checked-in packs, built on the pod with --serve support, then a fixed sequence of measurements and one results row. Nothing is cloned on the pod; the bundle is a tarball made here. Nothing here spends until a pod is started.

The command sequence

cd infra/gpu-bench
./make-bundle.sh                                             # build/igneum-gpu-bench.tar.gz (sources + scripts, about 1 MB)
# RunPod console: Pods -> Deploy -> pick the card -> template "nvidia/cuda:12.8.1-devel-ubuntu22.04" (or the image from
#   Dockerfile.cuda), Community Cloud, tick SSH, 20 GB container disk -> Deploy. Note host and port from "Connect".
scp -P <port> build/igneum-gpu-bench.tar.gz root@<host>:/workspace/
ssh -p <port> root@<host> 'cd /workspace && tar xzf igneum-gpu-bench.tar.gz && cd bundle && LABEL=rtx4090 ./run.sh'
# about 45 minutes: build 2 to 5 min, gate, 10 min raw, sweep, inline, recompile timings; the row prints as RESULT
#   and lands in bundle/results.md and results-<stamp>/row.md; the log is uploaded under label gpubench-rtx4090
scp -P <port> -r root@<host>:/workspace/bundle/results-* results/        # keep the raw outputs here
# RunPod console: Stop and Terminate the pod (billing per minute while it runs; storage bills while stopped)
node ../../tools/logs.mjs                                    # the uploaded log, if the pod could not be reached again

Repeat for each card (the same bundle). Paste the rows into docs/bench-log.md with results-template.md. CUDA_ARCH=sm_86 overrides -arch=native if the driver in a pod is older than its toolkit.

Cards and cost (RunPod pricing page, read 3 Oct 2026, USD per hour)

Card Community Cloud Secure Cloud One 45-minute run, community Note
RTX 3060 not on RunPod's table not on RunPod's table about 0.05 to 0.10 on Vast.ai (approximate) the entry card; Vast.ai is the host for it
RTX 3090 0.22 0.50 about 0.17 Ampere sm_86, 24 GB, 6 MB L2 (approximate)
RTX 4090 0.34 0.74 about 0.26 Ada sm_89, 24 GB, 72 MB L2 (approximate)
RTX 5090 0.69 0.99 about 0.52 Blackwell sm_120, 32 GB, 96 MB L2 (bench-log); the project lead's own card is the reference row
RX 7900 XTX not offered not offered n/a RunPod has no AMD consumer cards; Vast.ai lists none in general (approximate). Dockerfile.rocm and the OpenCL lane of run.sh are ready for any host that has one (or the project lead's own AMD box)

Four NVIDIA cards: about USD 2 on community cloud, USD 3.50 on secure cloud, plus cents of storage. Per-minute billing; a pod left running costs the hourly rate.

Vast.ai variant: search for the card with "CUDA 12.8" and "verified" hosts, rent with the image nvidia/cuda:12.8.1-devel-ubuntu22.04 and "SSH" as the launch mode, then the same scp and ssh lines (Vast prints its own port and host). Prices on Vast.ai are per host and change by the hour; the 3060 is usually under USD 0.10.

What run.sh measures and why

Step Output Ledger
Vectors gate (--batches 3): cache check, dataset self-test, 3 warps standalone and in batch gate.txt, must say OVERALL: PASS or the run stops the cross-vendor proof per card (bench-log, "What PASS means")
Raw bench, about 10 minutes of 2^24-hash batches at 1 GiB raw.txt: Mhash/s, seconds the per-card hash rate; sustained, so thermals and neighbours show
Sweep 64, 128, 256, 512, 1024 MiB, 20 batches each sweep-*.txt, the 64 MiB over 1 GiB ratio M1: the L2 cliff per card (the 5090 gave 5.9x)
Inline-dataset shortcut (make-inline.sh: every ds[...] load becomes mh_word(ds, ...) from the cache) inline.txt: Mhash/s and the ratio to the raw rate M16: the recompute attacker's rate; the M5 Max gave 0.21. The 64 MiB-cache variant inside the 5090's L2 needs a pack exported with a smaller cache (proto-metal has no flag for it yet) and is a follow-up
Second pack igneum-hourly (closed form, 128 loads) pack2.txt comparison with the 5090 and Metal tables in the bench log
nvcc -cubin of kernel.cu x3 (the worker's prepare path, host.cu line 554), NVRTC in process x3 (nvrtc-time.cu), clBuildProgram of kernel.cl x3 (clbuild-time.c, when an OpenCL ICD is visible) nvrtc.txt, clbuild.txt, medians in the row M11, M17: the hourly program change on real drivers
--serve ready line (printf quit | worker --serve) serve-ready.txt proves the pod's binary is the worker the miner drives (ready line shows prepare 1 when nvcc is on PATH)

The inline binary fails host.cu's dataset self-test by construction (the dataset buffer holds the cache) and must pass the vectors; run.sh reads its rate and vector lines and ignores its OVERALL.

Results row (also in results-template.md): card, driver and toolkit, date, Mhash/s at 1 GiB over the 10 minutes, seconds, second-pack Mhash/s, the five sweep rates, the cliff ratio, inline Mhash/s, inline over honest, nvcc ms, NVRTC ms, OpenCL ms, vectors, build ms.

Files

run.sh (on the pod), make-bundle.sh (on the Mac), make-inline.sh, nvrtc-time.cu, clbuild-time.c, upload.sh (the intake URL and key of proto-cuda/windows-miner/upload-log.bat as a curl line), Dockerfile.cuda, Dockerfile.rocm, results-template.md, results/ (raw outputs copied back, one directory per run).

Not tested here: there is no NVIDIA or AMD card on this Mac, so run.sh, nvrtc-time.cu and clbuild-time.c were checked by read-through and bash -n only; the harness binaries they drive ran on the 5090 and the gfx1036 on 3 Oct 2026 (bench-log). The first pod run is the real test; if nvrtc-time fails to compile the kernel (a header NVRTC rejects), the nvcc column still carries the out-of-process figure the worker uses today.