igneum/infra/cloud-devnet/README.md
igneum-labs 74e7c2bdf5 seed-nodes: igneum-seed-1 live and synced; relay path from the Mac; prices in the account's currency (USD)
igneum-seed-1 (Hetzner cx23, fsn1, 188.245.5.161:26611): built on the VM in 1,530 s, synced to the live devnet
(12,204 blocks, same sink as the live node) through a non-mining relay igneumd on the Mac (the live node's addPeer
RPC is refused in safe mode); the live node and the Windows PC learned the seed's address by peer exchange and dialled
it. seeds.txt written. Hetzner prices corrected to USD (pricing API currency) in the plans, READMEs and scripts;
current-generation types per location (cx23 EU, cpx22 sin, cpx21 US) in the cloud-devnet config.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 22:27:52 +00:00

120 lines
10 KiB
Markdown

# Igneum cloud devnet: 20 nodes across regions
A private Igneum devnet (`igneum-devnet-20`: own handshake magic, own genesis) on 20 small Linux VMs in five
locations, each running `igneumd` with the real lottery hash and one CPU trickle miner with its own BLS vote key.
Built for two measurements that the single-site devnet cannot give: block propagation and reorg behaviour under
real inter-region latency (gate 3, spec 03 C1 and the floor rule), and the difficulty controller under hash-rate
steps (spec 02 section 2.3). Everything here is scripts; nothing runs or spends until `create.sh` is answered "yes".
Hetzner Cloud through `hcloud` is the primary path. DigitalOcean through `doctl` is the variant for `create.sh`
(and `destroy.sh`); the rest is provider-neutral once `nodes.tsv` exists.
## The command sequence (the project lead approves the cost, then this, start to finish)
```
cd infra/cloud-devnet
brew install hcloud # once; doctl for the DigitalOcean variant
# the API token is read from ~/.config/igneum/hetzner-token (the seed-node scripts use the same file)
./create.sh # 1. shows the plan and the live prices, asks "yes", creates 20 VMs, writes nodes.tsv
./provision.sh # 2. source tarball -> builder VM -> igneumd + igneum-miner -> every node (about 20 min)
./start.sh # 3. nodes, block logs, then the 20 trickle miners
./status.sh # one line per node (blocks, blue, DAA, tips, peers, difficulty, sink, miner)
./experiments/latency.sh 10 # 4. RTT matrix, then 10 min of block propagation samples
./experiments/partition.sh sin 10 # 5. cut Singapore off for 10 min, heal, reorg depth and heal time
./experiments/hop.sh # 6. hash-rate steps for the controller (about 75 min, schedule in the script)
./experiments/collect.sh # 7. pull logs and RPC samples -> results/<date>/, summary.md, bench-log entry
./stop.sh && ./destroy.sh # 8. stop, then delete every VM (hourly billing ends)
```
Time from "yes" to the first block: VM creation 1 min, build on the builder 10 to 25 min (approximate; 500 crates
plus rocksdb), install 2 min, miners' first 256 MiB cache 10 to 30 s. Steps 4 to 6 are independent; 4 and 6 can run
at the same time, 5 should run alone.
The observer and the live page (`tools/observer`): the RPC never leaves a node's loopback, so either
`./experiments/observer.sh tunnel` (an ssh tunnel from the Mac; then run the observer here) or
`./experiments/observer.sh remote on` (installs Node 22 on node 1 and runs the observer there, keeping the Mac out of
it). With `LIVE_TABLE_PREFIX=cloud_` the record goes to `cloud_live_*` tables; without the prefix the cloud network takes
over `/live` on the site (stop the Mac's observer first). `site/api/live.mjs` reads the unprefixed tables only.
## Cost
Hetzner API prices for this project on 3 Oct 2026, net USD per month, hourly billing (the API's `price_hourly`):
| Location | Type | vCPU | RAM | Per month | Per hour | Note |
|---|---|---|---|---|---|---|
| fsn1, nbg1, hel1 | cx23 | 2 Intel | 4 GB | 6.49 | 0.0104 | the cheapest x86 type the project can create |
| sin | cpx22 | 2 AMD | 4 GB | 30.99 | 0.0497 | cx not sold outside the EU |
| ash, hil | cpx21 | 3 AMD | 4 GB | 37.49 | 0.0601 | the old cpx line, only still sold in the US; cpx11 (2 GB) is 20.49 |
| builder, any EU | cx43 | 8 Intel | 16 GB | 18.49 | 0.0296 | deleted after the build |
Plus USD 0.60 per month net per primary IPv4 (pricing API). The account bills in USD (the pricing API reports currency USD, VAT 20%); VAT is added for a private
customer; the API's gross column is net x 1.2.
| Mix | Nodes | Per month net | Per hour | An evening (6 h) |
|---|---|---|---|---|
| 4 per location (default `REGIONS`) | 8 EU, 4 sin, 4 ash, 4 hil | USD 476 | USD 0.76 | about USD 5 |
| EU-heavy (`REGIONS=hel1,fsn1,nbg1,hel1,fsn1,nbg1,hel1,ash,hil,sin`) | 14 EU, 2 sin, 2 ash, 2 hil | USD 303 | USD 0.47 | about USD 3 |
| DigitalOcean `s-2vcpu-4gb` in any region (DO pricing page, USD 24/mo, 0.036/h) | 20 | USD 480 | USD 0.71 | about USD 4.50 |
The run is meant to last one evening, so the hourly column is the one that matters: well under USD 10 all in, with
the builder. Leaving the network up costs the monthly column.
## What each script does
| Script | What |
|---|---|
| `config.sh` | every setting (provider, N, regions, types per location, suffix, genesis bits, mesh degree, miner threads, source tree) |
| `create.sh` | ssh key, firewall (22, 26611, icmp), N servers round-robin over `REGIONS`, writes `nodes.tsv` (name, index, region, ip) |
| `make-source.sh` | `git archive HEAD` of `NODE_SRC` and `igneum-pow` plus the uncommitted files (`SRC_MODE=head+dirty`, the default: on 3 Oct 2026 every worktree's branch work is uncommitted) into `build/src.tar.gz`, same layout as the Windows package's `src.zip` |
| `provision.sh` | `build`: builder VM, `builder/build-on-builder.sh` (apt deps, rustup, `cargo build --release -p kaspad -p igneum-miner --features igneum-pow`), binaries to `build/bin/`, builder deleted. `install`: binaries, `node/wrpc.py`, units and `/etc/igneum/node.env` on every node |
| `node/run-igneumd.sh` | the flag list: `--devnet --devnet-suffix=20 --override-params-file (genesis_bits) --rpclisten=127.0.0.1:26610 --rpclisten-json=127.0.0.1:28610 --listen=0.0.0.0:26611 --externalip --addpeer x2 --outpeers=0 --maxinpeers=32 --nodnsseed --disable-upnp --nologfiles --enable-unsynced-mining` |
| `node/igneum-miner.service` | `igneum-miner mine grpc://127.0.0.1:26610 1 1000000000 <node> --engine igneum-pow --payout-label <node>`: one CPU thread, its own vote key, votes on every checkpoint |
| `node/wrpc.py` | standard-library wRPC JSON client (the node's WebSocket RPC): `call`, `sample`, `watch-blocks` (blocks.tsv, chain.tsv, samples.tsv under /var/log/igneum) |
| `start.sh`, `stop.sh`, `status.sh`, `destroy.sh` | as named; `destroy.sh` deletes everything labelled `igneum=devnet` and the firewall |
| `experiments/latency.sh` | ping matrix between all nodes, then block propagation from the per-node arrival logs joined on hash (p50, p90, p99 per node and per region) |
| `experiments/partition.sh` | iptables cut of one region's p2p traffic to the rest, heal after N min, reorg depth (`getVirtualChainFromBlock` from each node's sink at heal, plus the largest `virtualChainChanged` removal), heal time (sinks converge), conflicting locks (`getFinalityCheckpoints` on both sides) |
| `experiments/hop.sh` | a schedule of miner-thread changes per node set (step up, step down, a region off), logged for the controller analysis |
| `experiments/collect.sh` | journals, block/chain/sample logs, RPC snapshots into `results/<date>/nodes/<name>/`; `analyze.py summary` and a bench-log entry template; `APPEND_BENCH_LOG=1` appends it to `docs/bench-log.md` |
| `experiments/analyze.py` | the offline analysis (RTT table, propagation join, lock conflicts, hop phases, reorg-depth histogram) |
| `experiments/observer.sh` | the observer hookup (tunnel or remote) |
## Network design choices
- Own network id. `--devnet-suffix=20` gives `igneum-devnet-20` (own handshake magic, own data directory) and the
override file sets `genesis_bits` (0x1e400000, 2^18 hashes per block: three 6-thread CPU miners on the M5 Max held
about 1 block/s at this value, fork-divergence genesis row), which recomputes the genesis hash. A node of this
network can never complete a handshake with the live devnet or the seed.
- Sparse mesh. Each node has two permanent outbound links (`--addpeer` to its ring neighbour and to the node half
way round) and `--outpeers=0`, so the connection manager does not fill its default 8 outbound slots from
exchanged addresses. Inbound links double the count: about 4 peers per node, 40 links in total. Blocks therefore
cross several hops between regions, which is what the propagation measurement needs.
- `--addpeer`, never `--connect`: `--connect` sets the inbound limit to 0 (kaspad/src/daemon.rs), the Windows node
lesson.
- RPC on loopback. Every measurement goes through ssh to `wrpc.py` on the node. The firewall opens 22, 26611 and icmp.
- One vote key per node. The miner's identity label is the node name, so weight accrues to 20 keys and the
finality rule has 20 voters; the devnet finality parameters (interval 30, dust 5, presence 20) apply.
- Which tree runs. `NODE_SRC` defaults to `vendor/igneum-node` (master worktree: finality v2 work, uncommitted). For
the controller experiment point it at `vendor/igneum-node-diff` (the `difficulty` worktree) or at a merged tree; the
binary is cached in `build/bin` and `REBUILD=1` rebuilds. `build/src.stamp` records what went in.
- Node 22 is not needed on the nodes: `wrpc.py` is standard-library Python. The observer (remote mode) is the one
thing that installs Node, on one node, only when asked.
## What the results answer
| Measurement | Where it lands | Decides |
|---|---|---|
| RTT between regions, block propagation p50/p90/p99 | `results/<date>/latency/` | the latency assumption behind GHOSTDAG k (5 s network delay bound at 1 BPS) and spec 03 C1's lock latency (simulated 2.5 s median at a 2 s inter-region delay) |
| Reorg depth distribution over the run | `summary.md` | checkpoint determination depth d (spec 03 C1, ledger F7: placeholder 60, devnet 20) |
| Partition: reorg depth, heal time, conflicting locks | `results/<date>/partition-*/partition.md` | the floor rule (lock needs 56.7% of total weight) under a real partition; ledger F2, F3, F11; O-3.6 (two certificates at one index) |
| Hop phases: settle time and overshoot per step | `hop.md` in the summary | spec 02 section 2.3's simulator claims (x50 settled 62 s, /50 657 s) on a real network with real timestamps |
Caveats, stated once: CPU hash rate only (no GPU in this network), one evening of data, clocks by chrony, 20 nodes
not 1,000. The numbers are the first real ones; they do not close gate 3 by themselves.
## Failure notes
- `create.sh` refuses to run without an active `hcloud` context or `doctl` auth. It never creates a cloud account.
- A node whose database predates a binary change must be wiped: `ssh root@<ip> 'systemctl stop igneumd; rm -rf /var/lib/igneum/*; systemctl start igneumd'`.
- If `provision.sh build` dies on the builder, the log is `build/build.log`; `KEEP_BUILDER=1` keeps the VM for a retry.
- `partition.sh heal` removes the iptables rules if a run was interrupted.
- `destroy.sh` is the only thing that stops the bill. Check the provider console afterwards.