infra/cloud-devnet: hcloud (doctl variant) create, builder-VM provision from a git-archive source tarball, systemd units for igneumd --devnet-suffix with a sparse --addpeer mesh and a CPU trickle miner per node, stdlib wRPC client, experiments (latency, partition, hop, collect, observer hookup), README with the command sequence and the Hetzner API prices of 3 Oct 2026. infra/gpu-bench: RunPod image recipes (CUDA 12.8, ROCm), bundle, run.sh (vectors gate, 10-min raw, sweep, inline shortcut ratio, nvcc/NVRTC/OpenCL recompile timings, results row, intake upload), bench-log template. infra/seed-nodes: create-seed (persistent IPv4, firewall), provision on the VM, health check, addPeer from the Mac over grpcurl, seeds.txt; igneum-seed-1 created at 188.245.5.161 (Hetzner cx23, fsn1). docs/plans/cloud-devnet.md and docs/plans/seed-nodes.md. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
70 lines
5.8 KiB
Markdown
70 lines
5.8 KiB
Markdown
# Cloud devnet and rented-GPU benchmarks: plan
|
|
|
|
3 October 2026. Scripts in `infra/cloud-devnet/` and `infra/gpu-bench/`. Nothing in either directory spends until
|
|
the project lead answers "yes" to one command; the seed node (`docs/plans/seed-nodes.md`) is the only thing live tonight.
|
|
|
|
## Experiment 1: a 20-node Igneum devnet across regions
|
|
|
|
What: 20 small VMs (Hetzner Cloud, 4 each in Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore) running `igneumd`
|
|
on a private network (`igneum-devnet-20`, own genesis, 2^18 hashes per block) with one CPU trickle miner and one BLS
|
|
vote key per node, wired as a sparse mesh of about 4 peers each. Three measurements: inter-region RTT and block
|
|
propagation delay per node; a region cut off with iptables for 10 minutes and healed, recording the reorg depth,
|
|
the heal time and whether any checkpoint locked with two hashes; and a schedule of hash-rate steps (4x on half the
|
|
nodes, a region off) for the difficulty controller.
|
|
|
|
What it proves:
|
|
|
|
| Measurement | Closes or informs |
|
|
|---|---|
|
|
| Reorg depth distribution under real latency, and the partition's reorg depth and heal time | gate 3: the checkpoint determination depth d (spec 03 C1, placeholder 60, devnet 20, ledger F7) is set from exactly this distribution; the floor rule (lock needs 56.7% of total weight) is tested against a real partition instead of the checkpoint-level simulator (ledger F2, F3, F11; O-3.6 two certificates at one index) |
|
|
| Block propagation p50/p90/p99 between regions | the 5 s network delay bound behind GHOSTDAG k at 1 BPS; spec 03's lock latency (simulated 2.5 s median at a 2 s inter-region delay) |
|
|
| Settle time and overshoot per hash-rate step | spec 02 section 2.3: the dual-lane controller's simulator numbers (x50 settled in 62 s, /50 in 657 s) on a real network with real timestamps; on the master tree it measures Kaspa's sampled DAA instead, so build from the `difficulty` worktree for this one |
|
|
|
|
What it does not prove: nothing about GPU hash rates (CPU miners), nothing at 1,000 miners, one evening of data.
|
|
|
|
Cost (Hetzner API prices of 3 Oct 2026, net): EUR 0.76 per hour for the 20 nodes at 4 per location (EUR 476 per
|
|
month if left running), or EUR 0.47 per hour EU-heavy; the builder VM about EUR 0.03 per hour for one hour. An evening
|
|
of six hours: about EUR 5. DigitalOcean variant: USD 0.71 per hour. Time to first block after "yes": about 25 minutes
|
|
(VM creation, a 10 to 25 minute build, install).
|
|
|
|
Command sequence: `infra/cloud-devnet/README.md`, section "The command sequence". In short: `create.sh`,
|
|
`provision.sh`, `start.sh`, then `experiments/latency.sh 10`, `experiments/partition.sh sin 10`, `experiments/hop.sh`,
|
|
`experiments/collect.sh`, then `stop.sh` and `destroy.sh`.
|
|
|
|
## Experiment 2: rented GPUs
|
|
|
|
What: one RunPod pod per card (RTX 3060 if a host offers one, RTX 3090, RTX 4090, RTX 5090; an AMD RX 7900 XTX is not
|
|
offered by RunPod and Vast.ai lists no AMD consumer cards in general, approximate, so the ROCm recipe is shipped and
|
|
waits for an AMD box). Each pod gets the bundle (`proto-cuda/host.cu`, the three packs, `proto-opencl/host.c`, the
|
|
bench scripts), builds the worker with `--serve` support, and runs: the vectors gate, a 10-minute raw bench at 1 GiB
|
|
on the memory-hard pack, the dataset sweep 64 MiB to 1 GiB (the L2 cliff), the inline-dataset shortcut kernel against
|
|
the honest one (the ratio), and the hourly-recompile timing (nvcc to cubin as the worker's `prepare` path does it,
|
|
NVRTC in process, and the OpenCL compiler where an ICD is present). It prints one results row, appends it to
|
|
`results.md`, and uploads the row and the log to the intake (`/api/log`, the Windows package's key).
|
|
|
|
What it proves:
|
|
|
|
| Measurement | Closes or informs |
|
|
|---|---|
|
|
| Mhash/s at 1 GiB per card, 10 minutes sustained | the per-card hash rate table the litepaper and the miner community need; the first numbers for 3060, 3090 and 4090 (only the 5090 and the M5 Max are measured) |
|
|
| The sweep's cliff between the card's L2 and 1 GiB | ledger M1 ("the program space is tiny"): the defence is random reads over a dataset larger than any on-chip cache; the cliff per card is the measurement behind the "under 2x for a chip" target |
|
|
| Inline shortcut over honest rate | ledger M16 (a 256 MiB cache on a die): the recompute attacker's rate on three more cards than the M5 Max's 0.21; the 64 MiB-cache variant inside the 5090's L2 needs a pack exported with a smaller cache and is a follow-up |
|
|
| nvcc, NVRTC and OpenCL compile times per card | ledger M11 (hourly JIT on real rigs) and M17: whether the hourly program change costs milliseconds or seconds on real NVIDIA drivers, in process and out of process |
|
|
|
|
Cost (RunPod pricing page, 3 Oct 2026, USD per hour): RTX 3090 0.22 community / 0.50 secure, RTX 4090 0.34 / 0.74,
|
|
RTX 5090 0.69 / 0.99; the 3060 is not on RunPod's table (Vast.ai about 0.05 to 0.10, approximate). One run is about
|
|
45 minutes, so the four NVIDIA cards cost about USD 2 on community cloud or USD 3.50 on secure cloud, plus a few
|
|
cents of storage. Per-minute billing.
|
|
|
|
Command sequence: `infra/gpu-bench/README.md`. In short: `make-bundle.sh` on the Mac, start a pod from the image
|
|
recipe with SSH, `scp` the bundle, `./run.sh`, read `results.md`, paste the row into `docs/bench-log.md` with the
|
|
template.
|
|
|
|
## The go/no-go question for the project lead
|
|
|
|
Spend about EUR 5 and USD 2 to 4 tonight-to-tomorrow for (a) the first reorg-depth, propagation and partition numbers
|
|
from a real multi-region network, which gate 3 needs and the simulator cannot give, and (b) the per-card hash rates
|
|
and the three ASIC-resistance measurements (L2 cliff, inline ratio, recompile time) on 3060, 3090 and 4090, closing
|
|
the measurement half of ledger items M1, M11 and M16 for those cards? If yes: `cd infra/cloud-devnet && ./create.sh`
|
|
and `cd infra/gpu-bench && ./make-bundle.sh` are the two starting commands. If no: the scripts keep, nothing bills,
|
|
and the seed node stays at EUR 6.49 per month.
|