The sweep (main's item 1): 199 tracked text files, 783 lines. The founder's full name, first name and possessive become "the founder" (sentence starts capitalised); the lowercase operating-system user name in WSL paths and commands becomes <user>; the second owner login becomes "the second owner login"; the three earlier businesses and the two other brands become "the other business", "the earlier entity", "the earlier business" and "another brand"; the Chrome profile rule names the igneum.network profile, not the profile's label. The standing commit login igneum-labs is not a founder term here: the fresh-repository step renames it in the history (docs/plans/history-rewrite.md, tools/repo/fresh-repo.sh). The patterns never appear in plain text in the tree (a plaintext list would be the hit): tools/ci/founder-strings.b64 (perl regex, tab, a sample per row) is read by tools/ci/founder-strings-check.sh (every tracked text file, perl, known-failed first: the self-test plants each row's sample in a fixture and the hit must name the file), by tools/community/discord-hooks.mjs (the guard's founder and business rows; the test takes its fixtures from the samples) and by tools/repo/fresh-repo.sh (the business names of the rewrite rules). site/forbidden-strings.txt carries the same patterns as b64: lines, decoded case-insensitive by site/scrub.mjs and tools/ci/launch-gates-check.mjs (whose fixture now plants an encoded made-up name). The check runs in the gate's tree checks on every merge. Not in this commit, by main's word: the 105 commit messages and 40 personal-identity commits that need the history rewrite (listed, not run), and the secrets found by gitleaks over the history (reported with owners). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
173 lines
17 KiB
Markdown
173 lines
17 KiB
Markdown
# Igneum cloud devnet: 12 nodes across regions (20 planned)
|
|
|
|
A private Igneum devnet (`igneum-devnet-20`: own handshake magic, own genesis) on small Linux VMs in five
|
|
locations, each running `igneumd` with the real lottery hash and one CPU trickle miner with its own BLS vote key.
|
|
Planned for 20; the first run (4 Oct 2026) stopped at 12 because of the limits a new Hetzner account carries
|
|
(10 primary IPs, 20 shared vCPUs, 8 dedicated vCPUs, no Arm; see "Account limits" below). `N=20` once the limits are raised.
|
|
Built for two measurements that the single-site devnet cannot give: block propagation and reorg behaviour under
|
|
real inter-region latency (gate 3, spec 03 C1 and the floor rule), and the difficulty controller under hash-rate
|
|
steps (spec 02 section 2.3). Everything here is scripts; nothing runs or spends until `create.sh` is answered "yes".
|
|
|
|
Hetzner Cloud through `hcloud` is the primary path. DigitalOcean through `doctl` is the variant for `create.sh`
|
|
(and `destroy.sh`); the rest is provider-neutral once `nodes.tsv` exists.
|
|
|
|
## The command sequence (the founder approves the cost, then this, start to finish)
|
|
|
|
```
|
|
cd infra/cloud-devnet
|
|
brew install hcloud # once; doctl for the DigitalOcean variant
|
|
# the API token is read from ~/.config/igneum/hetzner-token (the seed-node scripts use the same file)
|
|
./create.sh # 1. shows the plan and the live prices, asks "yes", creates 20 VMs, writes nodes.tsv;
|
|
# private mode (default): 4 zone networks, 4 gateways, routes, NAT, port forwards, uplink check
|
|
./net/setup.sh # (re-wires the gateways and private nodes at any time; create.sh ran it already)
|
|
./provision.sh # 2. source tarball -> builder VM (behind the EU gateway) -> igneumd + igneum-miner -> every node (about 20 min)
|
|
./start.sh # 3. nodes, block logs, then the 20 trickle miners
|
|
./status.sh # one line per node (blocks, blue, DAA, tips, peers, difficulty, sink, miner)
|
|
./experiments/latency.sh 10 # 4. RTT matrix, then 10 min of block propagation samples
|
|
./experiments/partition.sh sin 10 # 5. cut Singapore off for 10 min, heal, reorg depth and heal time
|
|
./experiments/hop.sh # 6. hash-rate steps for the controller (about 75 min, schedule in the script)
|
|
./experiments/collect.sh # 7. pull logs and RPC samples -> results/<date>/, summary.md, bench-log entry
|
|
./stop.sh && ./destroy.sh # 8. stop, then delete every VM (hourly billing ends)
|
|
```
|
|
|
|
ssh to a node: `./net/ssh.sh igneum-07` or by hand `ssh -i ~/.ssh/igneum_ed25519 -J root@<gateway public ip> root@<private ip>`
|
|
(the gateway of the node's zone, `nodes.tsv` column 6; public nodes are reached directly at that column).
|
|
|
|
Time from "yes" to the first block: VM creation 1 min, build on the builder 10 to 25 min (approximate; 500 crates
|
|
plus rocksdb), install 2 min, miners' first 256 MiB cache 10 to 30 s. Steps 4 to 6 are independent; 4 and 6 can run
|
|
at the same time, 5 should run alone.
|
|
|
|
The observer and the live page (`tools/observer`): the RPC never leaves a node's loopback, so either
|
|
`./experiments/observer.sh tunnel` (an ssh tunnel from the Mac; then run the observer here) or
|
|
`./experiments/observer.sh remote on` (installs Node 22 on node 1 and runs the observer there, keeping the Mac out of
|
|
it). With `LIVE_TABLE_PREFIX=cloud_` the record goes to `cloud_live_*` tables; without the prefix the cloud network takes
|
|
over `/live` on the site (stop the Mac's observer first). `site/api/live.mjs` reads the unprefixed tables only.
|
|
|
|
## Account limits (what stopped the 20-node run on 4 Oct 2026)
|
|
|
|
| Limit | Value seen | How it showed | What the scripts do about it |
|
|
|---|---|---|---|
|
|
| Primary IPs (IPv4 and IPv6 count alike) | 10 | `Primary IP limit exceeded` on the 5th node; IPv6-only fails the same way | private-network mode: 4 gateways with an IPv4 only, 8 private nodes, 4 networks |
|
|
| Shared vCPUs (cx, cpx) | 20, the seed node holds 2 | `shared core limit exceeded` on the 8th shared server | 8 shared nodes (01 to 07 and 12) |
|
|
| Dedicated vCPUs (ccx) | 8 | `dedicated core limit exceeded` on the 5th ccx13 | 4 ccx13 nodes (08 to 11), `TYPE_BY_INDEX=8-11:ccx13` |
|
|
| Arm (cax) | none | `unsupported location for server type` in fsn1, hel1 and nbg1 | not used |
|
|
|
|
Hetzner raises these from the console (Limits page) once an account has a billing history; the request is the founder's.
|
|
The seed node's own primary IPs and cores are part of the same counts.
|
|
|
|
## Cost
|
|
|
|
Hetzner API prices for this project on 3 Oct 2026, net USD per month, hourly billing (the API's `price_hourly`):
|
|
|
|
| Location | Type | vCPU | RAM | Per month | Per hour | Note |
|
|
|---|---|---|---|---|---|---|
|
|
| fsn1, nbg1, hel1 | cx23 | 2 Intel | 4 GB | 6.49 | 0.0104 | the cheapest x86 type the project can create |
|
|
| sin | cpx22 | 2 AMD | 4 GB | 30.99 | 0.0497 | cx not sold outside the EU |
|
|
| ash, hil | cpx21 | 3 AMD | 4 GB | 37.49 | 0.0601 | the old cpx line, only still sold in the US; cpx11 (2 GB) is 20.49 |
|
|
| builder, any EU | cx43 | 8 Intel | 16 GB | 18.49 | 0.0296 | deleted after the build |
|
|
|
|
Plus USD 0.60 per month net per primary IPv4 (pricing API); in private mode only the 4 gateways have one, and no
|
|
IPv6 primaries (they count against the same limit). Hetzner private networks and their traffic cost nothing. The account bills in USD (the pricing API reports currency USD, VAT 20%); VAT is added for a private
|
|
customer; the API's gross column is net x 1.2.
|
|
|
|
`create.sh` prints the plan's cost from the live API prices before it asks for "yes".
|
|
|
|
| Mix | Nodes | Per month net | Per hour | An evening (6 h) |
|
|
|---|---|---|---|---|
|
|
| As created 4 Oct 2026 (12 nodes: 6 EU cx23, 2 ash, 2 hil, 2 sin; 08 to 11 ccx13) | 12 | USD 417 | USD 0.57 | about USD 3.43 |
|
|
| 4 per location (`N=20`, all shared types; needs raised limits) | 8 EU, 4 sin, 4 ash, 4 hil | USD 476 | USD 0.76 | about USD 5 |
|
|
| EU-heavy (`REGIONS=hel1,fsn1,nbg1,hel1,fsn1,nbg1,hel1,ash,hil,sin`) | 14 EU, 2 sin, 2 ash, 2 hil | USD 303 | USD 0.47 | about USD 3 |
|
|
| DigitalOcean `s-2vcpu-4gb` in any region (DO pricing page, USD 24/mo, 0.036/h) | 20 | USD 480 | USD 0.71 | about USD 4.50 |
|
|
|
|
The run is meant to last one evening, so the hourly column is the one that matters: well under USD 10 all in, with
|
|
the builder. Leaving the network up costs the monthly column.
|
|
|
|
## What each script does
|
|
|
|
| Script | What |
|
|
|---|---|
|
|
| `config.sh` | every setting (provider, N, regions, types per location, suffix, genesis bits, mesh degree, miner threads, source tree) |
|
|
| `create.sh` | ssh key, firewall (22, 26611, 27001-27099, icmp), one private network per zone, N servers round-robin over `REGIONS` (zone gateways with a public IPv4, the rest private), writes `nodes.tsv` (name, index, region, ip, access, pub, port), then runs `net/setup.sh` |
|
|
| `net/setup.sh` | per zone: the network's 0.0.0.0/0 route to the gateway, `net/gateway.sh` on the gateway (ip_forward, MASQUERADE, MSS clamp, one DNAT rule per private node, persisted as `igneum-nat.service`), `net/private-node.sh` on every private node (default route, resolvers, uplink proof) |
|
|
| `net/ssh.sh` | `./net/ssh.sh <node> [command]`: ssh to any node by name, jumping through its gateway when it is private |
|
|
| `make-source.sh` | `git archive HEAD` of `NODE_SRC` and `igneum-pow` plus the uncommitted files (`SRC_MODE=head+dirty`, the default: on 3 Oct 2026 every worktree's branch work is uncommitted) into `build/src.tar.gz`, same layout as the Windows package's `src.zip` |
|
|
| `provision.sh` | `build`: builder VM, `builder/build-on-builder.sh` (apt deps, rustup, `cargo build --release -p kaspad -p igneum-miner --features igneum-pow`), binaries to `build/bin/`, builder deleted. `install [node ...]`: binaries (only when the node's sha256 differs), `node/wrpc.py`, units and `/etc/igneum/node.env` on every node, or on the named ones. 4 Oct 2026: with no vCPU headroom for a builder, the binaries were copied from the seed node's staged v4 build (`igneum-seed-1:/opt/igneum/v4/bin`, same igneum-node-v4 commit 6457ca95 and the same igneum-pow tree, sha256 checked) into `build/bin/`; `build/bin/src.stamp` says so |
|
|
| `node/run-igneumd.sh` | the flag list: `--devnet --devnet-suffix=20 --override-params-file (genesis_bits) --rpclisten=127.0.0.1:26610 --rpclisten-json=127.0.0.1:28610 --listen=0.0.0.0:26611 --externalip --addpeer x2 --outpeers=0 --maxinpeers=32 --nodnsseed --disable-upnp --nologfiles --enable-unsynced-mining` |
|
|
| `node/igneum-miner.service` | `igneum-miner mine grpc://127.0.0.1:26610 1 1000000000 <node> --engine igneum-pow --payout-label <node>`: one CPU thread, its own vote key, votes on every checkpoint |
|
|
| `node/wrpc.py` | standard-library wRPC JSON client (the node's WebSocket RPC): `call`, `sample`, `watch-blocks` (blocks.tsv, chain.tsv, samples.tsv under /var/log/igneum) |
|
|
| `start.sh`, `stop.sh`, `status.sh`, `destroy.sh` | as named; `destroy.sh` deletes everything labelled `igneum=devnet` and the firewall |
|
|
| `experiments/latency.sh` | ping matrix between all nodes, then block propagation from the per-node arrival logs joined on hash (p50, p90, p99 per node and per region) |
|
|
| `experiments/partition.sh` | iptables cut of one region's p2p traffic to the rest (private mode: on the region's gateway, INPUT, OUTPUT and FORWARD against the far gateways' public IPs over 26611 and the DNAT ports 27001:27099), heal after N min, reorg depth (`getVirtualChainFromBlock` from each node's sink at heal, plus the largest `virtualChainChanged` removal), heal time (each minority node's first chain removal of 5+ blocks after the rules come off, `heal.tsv`; the sink-count "converged" figure is kept but is tip churn on a healthy network), locks per side at cut, during, at heal and after convergence (`locks.tsv`), conflicting locks, the first lock after the heal (`LOCK_WAIT`, skipped while the weight window is filling) |
|
|
| `experiments/hop.sh` | a schedule of miner-thread changes per node set (step up, step down, a region off), logged for the controller analysis |
|
|
| `experiments/collect.sh` | journals, block/chain/sample logs, RPC snapshots into `results/<date>/nodes/<name>/`; `hop-series.tsv` (10-s difficulty and block-rate series over the hop run) and `hop.md` when `hop.log` exists; `analyze.py summary` and a bench-log entry template; `APPEND_BENCH_LOG=1` appends it to `docs/bench-log.md` |
|
|
| `experiments/analyze.py` | the offline analysis (RTT table, propagation join, lock conflicts, hop phases, reorg-depth histogram) |
|
|
| `experiments/observer.sh` | the observer hookup (tunnel or remote) |
|
|
|
|
## Network design choices
|
|
|
|
- Private networks, four gateways (4 Oct 2026). A new Hetzner account is capped at 10 primary IPs and IPv4 and IPv6
|
|
both count (the 5th node failed with "Primary IP limit exceeded", an IPv6-only server fails the same way, and
|
|
Hetzner does not raise the limit for new accounts). A Hetzner network cannot span network zones ("subnetwork zones
|
|
are not aligned"), so there is one network per zone: `igneum-net-eu-central` 10.20.1.0/24 (hel1, fsn1), `us-east`
|
|
10.20.2.0/24 (ash), `us-west` 10.20.3.0/24 (hil), `ap-southeast` 10.20.4.0/24 (sin). The lowest-index node of each
|
|
zone (01 hel1, 03 ash, 04 hil, 05 sin) keeps a public IPv4 and is the zone gateway: the network's 0.0.0.0/0 route
|
|
points at it (Hetzner pushes that route to every member by DHCP option 121; the gateway drops it for itself), it
|
|
masquerades the zone's outbound traffic, and it forwards TCP port 27000+index to each private node's 26611. Every
|
|
node is therefore dialable from everywhere: same-zone peers use the private IPs, cross-zone peers use
|
|
`<gateway public ip>:<27000+index>` (private) or `<public ip>:26611` (public). The ring mesh is unchanged; a
|
|
cross-zone link to a private node crosses the gateway's NAT (one extra hop inside the zone, well under 1 ms). A
|
|
private node claims `--externalip=<gateway>:<port>`. ssh to a private node jumps through its gateway;
|
|
`lib/common.sh` resolves that from `nodes.tsv` for every script. `NET_MODE=public` restores the old layout.
|
|
- The RTT matrix pings the address a node can reach: inside a zone the private IP, across zones the public IP (for a
|
|
private target, its gateway's), so a cross-zone row measures the inter-zone path to the target's zone.
|
|
- `partition.sh` cuts the same addresses (the gateways' public IPs across zones), since cross-zone packets never carry
|
|
a far private IP. The rules sit on the region's gateway only: INPUT and OUTPUT for its own node, FORWARD for the
|
|
private nodes behind it (their links to and from other zones are DNAT'd and masqueraded there), over the whole p2p
|
|
port set (26611 and 27001:27099), because a cross-zone link to a private node uses its DNAT port, not 26611. The
|
|
public-mode cut (per node, port 26611) would leave those links up; it is kept for `NET_MODE=public`.
|
|
|
|
- Own network id. `--devnet-suffix=20` gives `igneum-devnet-20` (own handshake magic, own data directory) and the
|
|
override file sets `genesis_bits` (0x1e400000, 2^18 hashes per block: three 6-thread CPU miners on the M5 Max held
|
|
about 1 block/s at this value, fork-divergence genesis row), which recomputes the genesis hash. A node of this
|
|
network can never complete a handshake with the live devnet or the seed.
|
|
- Sparse mesh. Each node has two permanent outbound links (`--addpeer` to its ring neighbour and to the node half
|
|
way round) and `--outpeers=0`, so the connection manager does not fill its default 8 outbound slots from
|
|
exchanged addresses. Inbound links double the count: about 4 peers per node, 40 links in total. Blocks therefore
|
|
cross several hops between regions, which is what the propagation measurement needs.
|
|
- `--addpeer`, never `--connect`: `--connect` sets the inbound limit to 0 (kaspad/src/daemon.rs), the Windows node
|
|
lesson.
|
|
- RPC on loopback. Every measurement goes through ssh to `wrpc.py` on the node. The firewall opens 22, 26611 and icmp.
|
|
- One vote key per node. The miner's identity label is the node name, so weight accrues to 20 keys and the
|
|
finality rule has 20 voters; the devnet finality parameters (interval 30, dust 5, presence 20) apply.
|
|
- Which tree runs. `NODE_SRC` defaults to `vendor/igneum-node` (master worktree: finality v2 work, uncommitted). For
|
|
the controller experiment point it at `vendor/igneum-node-diff` (the `difficulty` worktree) or at a merged tree; the
|
|
binary is cached in `build/bin` and `REBUILD=1` rebuilds. `build/src.stamp` records what went in.
|
|
- Node 22 is not needed on the nodes: `wrpc.py` is standard-library Python. The observer (remote mode) is the one
|
|
thing that installs Node, on one node, only when asked.
|
|
|
|
## What the results answer
|
|
|
|
| Measurement | Where it lands | Decides |
|
|
|---|---|---|
|
|
| RTT between regions, block propagation p50/p90/p99 | `results/<date>/latency/` | the latency assumption behind GHOSTDAG k (5 s network delay bound at 1 BPS) and spec 03 C1's lock latency (simulated 2.5 s median at a 2 s inter-region delay) |
|
|
| Reorg depth distribution over the run | `summary.md` | checkpoint determination depth d (spec 03 C1, ledger F7: placeholder 60, devnet 20) |
|
|
| Partition: reorg depth, heal time, locks per side, conflicting locks | `results/<date>/partition-*/partition.md` | the floor rule (lock needs 2/3 of total weight since O-3.15, 4 Oct 2026) under a real partition; ledger F2, F3, F11; O-3.6 (two certificates at one index). No lock is possible before the weight window has filled (7,200 DAA on the devnet parameters, about 2 h at 1 BPS from genesis), so a partition that should see locks must start after that |
|
|
| Hop phases: settle time and overshoot per step | `hop.md` in the summary | spec 02 section 2.3's simulator claims (x50 settled 62 s, /50 657 s) on a real network with real timestamps |
|
|
|
|
Caveats, stated once: CPU hash rate only (no GPU in this network), one evening of data, clocks by chrony, 20 nodes
|
|
not 1,000. The numbers are the first real ones; they do not close gate 3 by themselves.
|
|
|
|
## Failure notes
|
|
|
|
- `create.sh` refuses to run without an active `hcloud` context or `doctl` auth. It never creates a cloud account.
|
|
- Every loop over `nodes.tsv` reads it on fd 3 (`read -u 3 ... done 3< nodes.tsv`). A backgrounded ssh inside a
|
|
plain `while read` loop inherits the loop's stdin and drains the file, so later nodes silently vanish from the run
|
|
(seen on 4 Oct 2026: 3 of 12 installed). Keep new loops on fd 3.
|
|
- A jumped `scp` to a far node (hil through its gateway) ran at well under 100 KB/s on 4 Oct 2026; the install can
|
|
take 10 minutes per far node. `nssh`/`nscp` retry once on a dropped connection (ssh exit 255).
|
|
- A node whose database predates a binary change must be wiped: `ssh root@<ip> 'systemctl stop igneumd; rm -rf /var/lib/igneum/*; systemctl start igneumd'`.
|
|
- If `provision.sh build` dies on the builder, the log is `build/build.log`; `KEEP_BUILDER=1` keeps the VM for a retry.
|
|
- `partition.sh heal` removes the iptables rules if a run was interrupted.
|
|
- `destroy.sh` is the only thing that stops the bill (servers, firewall, the four networks). Check the provider console afterwards.
|
|
- A private node that cannot reach apt: `./net/setup.sh nodes` re-adds the default route and resolvers and proves the
|
|
uplink; `./net/setup.sh gateways` rebuilds the NAT and port-forward rules (idempotent chains, `iptables -t nat -S IGNEUM-DNAT` on a gateway shows them).
|