partition.sh adapted to private-network mode: the cut sits on the region's gateway (INPUT, OUTPUT and FORWARD against the far gateways' public IPs over 26611 and the DNAT ports 27001:27099), since a per-node port-26611 rule leaves the DNAT links up. It now records the locks per side at cut, during, at heal and after convergence, the first lock after the heal from the journals, and the heal time as each minority node's first chain removal of 5+ blocks (the sink-count criterion is tip churn on a healthy network). hop.sh and partition.sh hold the Mac awake with caffeinate; analyze.py gains a 10-s hop series (difficulty, block count, 1- and 2-min rates, threads) and an overshoot table; collect.sh writes hop-series.tsv and hop.md and gzips the journals. Results 2026-10-04: partition 1 (window still filling) reorg 431/496 on the minority, 2 on the majority, healed in 10 and 14 s; partition 2 (locks active): minority locked nothing during the cut, majority locked every interval at 66.8% to 84.5% of total, 0 conflicting locks over 107 indices, healed in 11 and 15 s, first lock after heal 13 s. Hash-rate steps x1.42, x0.70, x0.75, x1.32: difficulty overshoots x1.67, x0.66, x0.53, x1.60, settle 751 s, never in 900 s, 241 s, 646 s. Failures stated in summary.md: the Mac hibernated during the hop (phase 2 ran 94 min), the first partition could not see locks, the script's heal and first-lock figures were artefacts (fixed), and another agent's v2 rollout restarted every node during the second heal. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
17 KiB
Igneum cloud devnet: 12 nodes across regions (20 planned)
A private Igneum devnet (igneum-devnet-20: own handshake magic, own genesis) on small Linux VMs in five
locations, each running igneumd with the real lottery hash and one CPU trickle miner with its own BLS vote key.
Planned for 20; the first run (4 Oct 2026) stopped at 12 because of the limits a new Hetzner account carries
(10 primary IPs, 20 shared vCPUs, 8 dedicated vCPUs, no Arm; see "Account limits" below). N=20 once the limits are raised.
Built for two measurements that the single-site devnet cannot give: block propagation and reorg behaviour under
real inter-region latency (gate 3, spec 03 C1 and the floor rule), and the difficulty controller under hash-rate
steps (spec 02 section 2.3). Everything here is scripts; nothing runs or spends until create.sh is answered "yes".
Hetzner Cloud through hcloud is the primary path. DigitalOcean through doctl is the variant for create.sh
(and destroy.sh); the rest is provider-neutral once nodes.tsv exists.
The command sequence (Josh approves the cost, then this, start to finish)
cd infra/cloud-devnet
brew install hcloud # once; doctl for the DigitalOcean variant
# the API token is read from ~/.config/igneum/hetzner-token (the seed-node scripts use the same file)
./create.sh # 1. shows the plan and the live prices, asks "yes", creates 20 VMs, writes nodes.tsv;
# private mode (default): 4 zone networks, 4 gateways, routes, NAT, port forwards, uplink check
./net/setup.sh # (re-wires the gateways and private nodes at any time; create.sh ran it already)
./provision.sh # 2. source tarball -> builder VM (behind the EU gateway) -> igneumd + igneum-miner -> every node (about 20 min)
./start.sh # 3. nodes, block logs, then the 20 trickle miners
./status.sh # one line per node (blocks, blue, DAA, tips, peers, difficulty, sink, miner)
./experiments/latency.sh 10 # 4. RTT matrix, then 10 min of block propagation samples
./experiments/partition.sh sin 10 # 5. cut Singapore off for 10 min, heal, reorg depth and heal time
./experiments/hop.sh # 6. hash-rate steps for the controller (about 75 min, schedule in the script)
./experiments/collect.sh # 7. pull logs and RPC samples -> results/<date>/, summary.md, bench-log entry
./stop.sh && ./destroy.sh # 8. stop, then delete every VM (hourly billing ends)
ssh to a node: ./net/ssh.sh igneum-07 or by hand ssh -i ~/.ssh/igneum_ed25519 -J root@<gateway public ip> root@<private ip>
(the gateway of the node's zone, nodes.tsv column 6; public nodes are reached directly at that column).
Time from "yes" to the first block: VM creation 1 min, build on the builder 10 to 25 min (approximate; 500 crates plus rocksdb), install 2 min, miners' first 256 MiB cache 10 to 30 s. Steps 4 to 6 are independent; 4 and 6 can run at the same time, 5 should run alone.
The observer and the live page (tools/observer): the RPC never leaves a node's loopback, so either
./experiments/observer.sh tunnel (an ssh tunnel from the Mac; then run the observer here) or
./experiments/observer.sh remote on (installs Node 22 on node 1 and runs the observer there, keeping the Mac out of
it). With LIVE_TABLE_PREFIX=cloud_ the record goes to cloud_live_* tables; without the prefix the cloud network takes
over /live on the site (stop the Mac's observer first). site/api/live.mjs reads the unprefixed tables only.
Account limits (what stopped the 20-node run on 4 Oct 2026)
| Limit | Value seen | How it showed | What the scripts do about it |
|---|---|---|---|
| Primary IPs (IPv4 and IPv6 count alike) | 10 | Primary IP limit exceeded on the 5th node; IPv6-only fails the same way |
private-network mode: 4 gateways with an IPv4 only, 8 private nodes, 4 networks |
| Shared vCPUs (cx, cpx) | 20, the seed node holds 2 | shared core limit exceeded on the 8th shared server |
8 shared nodes (01 to 07 and 12) |
| Dedicated vCPUs (ccx) | 8 | dedicated core limit exceeded on the 5th ccx13 |
4 ccx13 nodes (08 to 11), TYPE_BY_INDEX=8-11:ccx13 |
| Arm (cax) | none | unsupported location for server type in fsn1, hel1 and nbg1 |
not used |
Hetzner raises these from the console (Limits page) once an account has a billing history; the request is Josh's. The seed node's own primary IPs and cores are part of the same counts.
Cost
Hetzner API prices for this project on 3 Oct 2026, net USD per month, hourly billing (the API's price_hourly):
| Location | Type | vCPU | RAM | Per month | Per hour | Note |
|---|---|---|---|---|---|---|
| fsn1, nbg1, hel1 | cx23 | 2 Intel | 4 GB | 6.49 | 0.0104 | the cheapest x86 type the project can create |
| sin | cpx22 | 2 AMD | 4 GB | 30.99 | 0.0497 | cx not sold outside the EU |
| ash, hil | cpx21 | 3 AMD | 4 GB | 37.49 | 0.0601 | the old cpx line, only still sold in the US; cpx11 (2 GB) is 20.49 |
| builder, any EU | cx43 | 8 Intel | 16 GB | 18.49 | 0.0296 | deleted after the build |
Plus USD 0.60 per month net per primary IPv4 (pricing API); in private mode only the 4 gateways have one, and no IPv6 primaries (they count against the same limit). Hetzner private networks and their traffic cost nothing. The account bills in USD (the pricing API reports currency USD, VAT 20%); VAT is added for a private customer; the API's gross column is net x 1.2.
create.sh prints the plan's cost from the live API prices before it asks for "yes".
| Mix | Nodes | Per month net | Per hour | An evening (6 h) |
|---|---|---|---|---|
| As created 4 Oct 2026 (12 nodes: 6 EU cx23, 2 ash, 2 hil, 2 sin; 08 to 11 ccx13) | 12 | USD 417 | USD 0.57 | about USD 3.43 |
4 per location (N=20, all shared types; needs raised limits) |
8 EU, 4 sin, 4 ash, 4 hil | USD 476 | USD 0.76 | about USD 5 |
EU-heavy (REGIONS=hel1,fsn1,nbg1,hel1,fsn1,nbg1,hel1,ash,hil,sin) |
14 EU, 2 sin, 2 ash, 2 hil | USD 303 | USD 0.47 | about USD 3 |
DigitalOcean s-2vcpu-4gb in any region (DO pricing page, USD 24/mo, 0.036/h) |
20 | USD 480 | USD 0.71 | about USD 4.50 |
The run is meant to last one evening, so the hourly column is the one that matters: well under USD 10 all in, with the builder. Leaving the network up costs the monthly column.
What each script does
| Script | What |
|---|---|
config.sh |
every setting (provider, N, regions, types per location, suffix, genesis bits, mesh degree, miner threads, source tree) |
create.sh |
ssh key, firewall (22, 26611, 27001-27099, icmp), one private network per zone, N servers round-robin over REGIONS (zone gateways with a public IPv4, the rest private), writes nodes.tsv (name, index, region, ip, access, pub, port), then runs net/setup.sh |
net/setup.sh |
per zone: the network's 0.0.0.0/0 route to the gateway, net/gateway.sh on the gateway (ip_forward, MASQUERADE, MSS clamp, one DNAT rule per private node, persisted as igneum-nat.service), net/private-node.sh on every private node (default route, resolvers, uplink proof) |
net/ssh.sh |
./net/ssh.sh <node> [command]: ssh to any node by name, jumping through its gateway when it is private |
make-source.sh |
git archive HEAD of NODE_SRC and igneum-pow plus the uncommitted files (SRC_MODE=head+dirty, the default: on 3 Oct 2026 every worktree's branch work is uncommitted) into build/src.tar.gz, same layout as the Windows package's src.zip |
provision.sh |
build: builder VM, builder/build-on-builder.sh (apt deps, rustup, cargo build --release -p kaspad -p igneum-miner --features igneum-pow), binaries to build/bin/, builder deleted. install [node ...]: binaries (only when the node's sha256 differs), node/wrpc.py, units and /etc/igneum/node.env on every node, or on the named ones. 4 Oct 2026: with no vCPU headroom for a builder, the binaries were copied from the seed node's staged v4 build (igneum-seed-1:/opt/igneum/v4/bin, same igneum-node-v4 commit 6457ca95 and the same igneum-pow tree, sha256 checked) into build/bin/; build/bin/src.stamp says so |
node/run-igneumd.sh |
the flag list: --devnet --devnet-suffix=20 --override-params-file (genesis_bits) --rpclisten=127.0.0.1:26610 --rpclisten-json=127.0.0.1:28610 --listen=0.0.0.0:26611 --externalip --addpeer x2 --outpeers=0 --maxinpeers=32 --nodnsseed --disable-upnp --nologfiles --enable-unsynced-mining |
node/igneum-miner.service |
igneum-miner mine grpc://127.0.0.1:26610 1 1000000000 <node> --engine igneum-pow --payout-label <node>: one CPU thread, its own vote key, votes on every checkpoint |
node/wrpc.py |
standard-library wRPC JSON client (the node's WebSocket RPC): call, sample, watch-blocks (blocks.tsv, chain.tsv, samples.tsv under /var/log/igneum) |
start.sh, stop.sh, status.sh, destroy.sh |
as named; destroy.sh deletes everything labelled igneum=devnet and the firewall |
experiments/latency.sh |
ping matrix between all nodes, then block propagation from the per-node arrival logs joined on hash (p50, p90, p99 per node and per region) |
experiments/partition.sh |
iptables cut of one region's p2p traffic to the rest (private mode: on the region's gateway, INPUT, OUTPUT and FORWARD against the far gateways' public IPs over 26611 and the DNAT ports 27001:27099), heal after N min, reorg depth (getVirtualChainFromBlock from each node's sink at heal, plus the largest virtualChainChanged removal), heal time (each minority node's first chain removal of 5+ blocks after the rules come off, heal.tsv; the sink-count "converged" figure is kept but is tip churn on a healthy network), locks per side at cut, during, at heal and after convergence (locks.tsv), conflicting locks, the first lock after the heal (LOCK_WAIT, skipped while the weight window is filling) |
experiments/hop.sh |
a schedule of miner-thread changes per node set (step up, step down, a region off), logged for the controller analysis |
experiments/collect.sh |
journals, block/chain/sample logs, RPC snapshots into results/<date>/nodes/<name>/; hop-series.tsv (10-s difficulty and block-rate series over the hop run) and hop.md when hop.log exists; analyze.py summary and a bench-log entry template; APPEND_BENCH_LOG=1 appends it to docs/bench-log.md |
experiments/analyze.py |
the offline analysis (RTT table, propagation join, lock conflicts, hop phases, reorg-depth histogram) |
experiments/observer.sh |
the observer hookup (tunnel or remote) |
Network design choices
-
Private networks, four gateways (4 Oct 2026). A new Hetzner account is capped at 10 primary IPs and IPv4 and IPv6 both count (the 5th node failed with "Primary IP limit exceeded", an IPv6-only server fails the same way, and Hetzner does not raise the limit for new accounts). A Hetzner network cannot span network zones ("subnetwork zones are not aligned"), so there is one network per zone:
igneum-net-eu-central10.20.1.0/24 (hel1, fsn1),us-east10.20.2.0/24 (ash),us-west10.20.3.0/24 (hil),ap-southeast10.20.4.0/24 (sin). The lowest-index node of each zone (01 hel1, 03 ash, 04 hil, 05 sin) keeps a public IPv4 and is the zone gateway: the network's 0.0.0.0/0 route points at it (Hetzner pushes that route to every member by DHCP option 121; the gateway drops it for itself), it masquerades the zone's outbound traffic, and it forwards TCP port 27000+index to each private node's 26611. Every node is therefore dialable from everywhere: same-zone peers use the private IPs, cross-zone peers use<gateway public ip>:<27000+index>(private) or<public ip>:26611(public). The ring mesh is unchanged; a cross-zone link to a private node crosses the gateway's NAT (one extra hop inside the zone, well under 1 ms). A private node claims--externalip=<gateway>:<port>. ssh to a private node jumps through its gateway;lib/common.shresolves that fromnodes.tsvfor every script.NET_MODE=publicrestores the old layout. -
The RTT matrix pings the address a node can reach: inside a zone the private IP, across zones the public IP (for a private target, its gateway's), so a cross-zone row measures the inter-zone path to the target's zone.
-
partition.shcuts the same addresses (the gateways' public IPs across zones), since cross-zone packets never carry a far private IP. The rules sit on the region's gateway only: INPUT and OUTPUT for its own node, FORWARD for the private nodes behind it (their links to and from other zones are DNAT'd and masqueraded there), over the whole p2p port set (26611 and 27001:27099), because a cross-zone link to a private node uses its DNAT port, not 26611. The public-mode cut (per node, port 26611) would leave those links up; it is kept forNET_MODE=public. -
Own network id.
--devnet-suffix=20givesigneum-devnet-20(own handshake magic, own data directory) and the override file setsgenesis_bits(0x1e400000, 2^18 hashes per block: three 6-thread CPU miners on the M5 Max held about 1 block/s at this value, fork-divergence genesis row), which recomputes the genesis hash. A node of this network can never complete a handshake with the live devnet or the seed. -
Sparse mesh. Each node has two permanent outbound links (
--addpeerto its ring neighbour and to the node half way round) and--outpeers=0, so the connection manager does not fill its default 8 outbound slots from exchanged addresses. Inbound links double the count: about 4 peers per node, 40 links in total. Blocks therefore cross several hops between regions, which is what the propagation measurement needs. -
--addpeer, never--connect:--connectsets the inbound limit to 0 (kaspad/src/daemon.rs), the Windows node lesson. -
RPC on loopback. Every measurement goes through ssh to
wrpc.pyon the node. The firewall opens 22, 26611 and icmp. -
One vote key per node. The miner's identity label is the node name, so weight accrues to 20 keys and the finality rule has 20 voters; the devnet finality parameters (interval 30, dust 5, presence 20) apply.
-
Which tree runs.
NODE_SRCdefaults tovendor/igneum-node(master worktree: finality v2 work, uncommitted). For the controller experiment point it atvendor/igneum-node-diff(thedifficultyworktree) or at a merged tree; the binary is cached inbuild/binandREBUILD=1rebuilds.build/src.stamprecords what went in. -
Node 22 is not needed on the nodes:
wrpc.pyis standard-library Python. The observer (remote mode) is the one thing that installs Node, on one node, only when asked.
What the results answer
| Measurement | Where it lands | Decides |
|---|---|---|
| RTT between regions, block propagation p50/p90/p99 | results/<date>/latency/ |
the latency assumption behind GHOSTDAG k (5 s network delay bound at 1 BPS) and spec 03 C1's lock latency (simulated 2.5 s median at a 2 s inter-region delay) |
| Reorg depth distribution over the run | summary.md |
checkpoint determination depth d (spec 03 C1, ledger F7: placeholder 60, devnet 20) |
| Partition: reorg depth, heal time, locks per side, conflicting locks | results/<date>/partition-*/partition.md |
the floor rule (lock needs 2/3 of total weight since O-3.15, 4 Oct 2026) under a real partition; ledger F2, F3, F11; O-3.6 (two certificates at one index). No lock is possible before the weight window has filled (7,200 DAA on the devnet parameters, about 2 h at 1 BPS from genesis), so a partition that should see locks must start after that |
| Hop phases: settle time and overshoot per step | hop.md in the summary |
spec 02 section 2.3's simulator claims (x50 settled 62 s, /50 657 s) on a real network with real timestamps |
Caveats, stated once: CPU hash rate only (no GPU in this network), one evening of data, clocks by chrony, 20 nodes not 1,000. The numbers are the first real ones; they do not close gate 3 by themselves.
Failure notes
create.shrefuses to run without an activehcloudcontext ordoctlauth. It never creates a cloud account.- Every loop over
nodes.tsvreads it on fd 3 (read -u 3 ... done 3< nodes.tsv). A backgrounded ssh inside a plainwhile readloop inherits the loop's stdin and drains the file, so later nodes silently vanish from the run (seen on 4 Oct 2026: 3 of 12 installed). Keep new loops on fd 3. - A jumped
scpto a far node (hil through its gateway) ran at well under 100 KB/s on 4 Oct 2026; the install can take 10 minutes per far node.nssh/nscpretry once on a dropped connection (ssh exit 255). - A node whose database predates a binary change must be wiped:
ssh root@<ip> 'systemctl stop igneumd; rm -rf /var/lib/igneum/*; systemctl start igneumd'. - If
provision.sh builddies on the builder, the log isbuild/build.log;KEEP_BUILDER=1keeps the VM for a retry. partition.sh healremoves the iptables rules if a run was interrupted.destroy.shis the only thing that stops the bill (servers, firewall, the four networks). Check the provider console afterwards.- A private node that cannot reach apt:
./net/setup.sh nodesre-adds the default route and resolvers and proves the uplink;./net/setup.sh gatewaysrebuilds the NAT and port-forward rules (idempotent chains,iptables -t nat -S IGNEUM-DNATon a gateway shows them).