infra: cloud devnet (20 nodes), rented GPU bench and seed node scripts; plans for both
infra/cloud-devnet: hcloud (doctl variant) create, builder-VM provision from a git-archive source tarball, systemd units for igneumd --devnet-suffix with a sparse --addpeer mesh and a CPU trickle miner per node, stdlib wRPC client, experiments (latency, partition, hop, collect, observer hookup), README with the command sequence and the Hetzner API prices of 3 Oct 2026. infra/gpu-bench: RunPod image recipes (CUDA 12.8, ROCm), bundle, run.sh (vectors gate, 10-min raw, sweep, inline shortcut ratio, nvcc/NVRTC/OpenCL recompile timings, results row, intake upload), bench-log template. infra/seed-nodes: create-seed (persistent IPv4, firewall), provision on the VM, health check, addPeer from the Mac over grpcurl, seeds.txt; igneum-seed-1 created at 188.245.5.161 (Hetzner cx23, fsn1). docs/plans/cloud-devnet.md and docs/plans/seed-nodes.md. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
22f50fb076
commit
36a4f4ca62
53 changed files with 2580 additions and 0 deletions
70
docs/plans/cloud-devnet.md
Normal file
70
docs/plans/cloud-devnet.md
Normal file
|
|
@ -0,0 +1,70 @@
|
|||
# Cloud devnet and rented-GPU benchmarks: plan
|
||||
|
||||
3 October 2026. Scripts in `infra/cloud-devnet/` and `infra/gpu-bench/`. Nothing in either directory spends until
|
||||
the project lead answers "yes" to one command; the seed node (`docs/plans/seed-nodes.md`) is the only thing live tonight.
|
||||
|
||||
## Experiment 1: a 20-node Igneum devnet across regions
|
||||
|
||||
What: 20 small VMs (Hetzner Cloud, 4 each in Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore) running `igneumd`
|
||||
on a private network (`igneum-devnet-20`, own genesis, 2^18 hashes per block) with one CPU trickle miner and one BLS
|
||||
vote key per node, wired as a sparse mesh of about 4 peers each. Three measurements: inter-region RTT and block
|
||||
propagation delay per node; a region cut off with iptables for 10 minutes and healed, recording the reorg depth,
|
||||
the heal time and whether any checkpoint locked with two hashes; and a schedule of hash-rate steps (4x on half the
|
||||
nodes, a region off) for the difficulty controller.
|
||||
|
||||
What it proves:
|
||||
|
||||
| Measurement | Closes or informs |
|
||||
|---|---|
|
||||
| Reorg depth distribution under real latency, and the partition's reorg depth and heal time | gate 3: the checkpoint determination depth d (spec 03 C1, placeholder 60, devnet 20, ledger F7) is set from exactly this distribution; the floor rule (lock needs 56.7% of total weight) is tested against a real partition instead of the checkpoint-level simulator (ledger F2, F3, F11; O-3.6 two certificates at one index) |
|
||||
| Block propagation p50/p90/p99 between regions | the 5 s network delay bound behind GHOSTDAG k at 1 BPS; spec 03's lock latency (simulated 2.5 s median at a 2 s inter-region delay) |
|
||||
| Settle time and overshoot per hash-rate step | spec 02 section 2.3: the dual-lane controller's simulator numbers (x50 settled in 62 s, /50 in 657 s) on a real network with real timestamps; on the master tree it measures Kaspa's sampled DAA instead, so build from the `difficulty` worktree for this one |
|
||||
|
||||
What it does not prove: nothing about GPU hash rates (CPU miners), nothing at 1,000 miners, one evening of data.
|
||||
|
||||
Cost (Hetzner API prices of 3 Oct 2026, net): EUR 0.76 per hour for the 20 nodes at 4 per location (EUR 476 per
|
||||
month if left running), or EUR 0.47 per hour EU-heavy; the builder VM about EUR 0.03 per hour for one hour. An evening
|
||||
of six hours: about EUR 5. DigitalOcean variant: USD 0.71 per hour. Time to first block after "yes": about 25 minutes
|
||||
(VM creation, a 10 to 25 minute build, install).
|
||||
|
||||
Command sequence: `infra/cloud-devnet/README.md`, section "The command sequence". In short: `create.sh`,
|
||||
`provision.sh`, `start.sh`, then `experiments/latency.sh 10`, `experiments/partition.sh sin 10`, `experiments/hop.sh`,
|
||||
`experiments/collect.sh`, then `stop.sh` and `destroy.sh`.
|
||||
|
||||
## Experiment 2: rented GPUs
|
||||
|
||||
What: one RunPod pod per card (RTX 3060 if a host offers one, RTX 3090, RTX 4090, RTX 5090; an AMD RX 7900 XTX is not
|
||||
offered by RunPod and Vast.ai lists no AMD consumer cards in general, approximate, so the ROCm recipe is shipped and
|
||||
waits for an AMD box). Each pod gets the bundle (`proto-cuda/host.cu`, the three packs, `proto-opencl/host.c`, the
|
||||
bench scripts), builds the worker with `--serve` support, and runs: the vectors gate, a 10-minute raw bench at 1 GiB
|
||||
on the memory-hard pack, the dataset sweep 64 MiB to 1 GiB (the L2 cliff), the inline-dataset shortcut kernel against
|
||||
the honest one (the ratio), and the hourly-recompile timing (nvcc to cubin as the worker's `prepare` path does it,
|
||||
NVRTC in process, and the OpenCL compiler where an ICD is present). It prints one results row, appends it to
|
||||
`results.md`, and uploads the row and the log to the intake (`/api/log`, the Windows package's key).
|
||||
|
||||
What it proves:
|
||||
|
||||
| Measurement | Closes or informs |
|
||||
|---|---|
|
||||
| Mhash/s at 1 GiB per card, 10 minutes sustained | the per-card hash rate table the litepaper and the miner community need; the first numbers for 3060, 3090 and 4090 (only the 5090 and the M5 Max are measured) |
|
||||
| The sweep's cliff between the card's L2 and 1 GiB | ledger M1 ("the program space is tiny"): the defence is random reads over a dataset larger than any on-chip cache; the cliff per card is the measurement behind the "under 2x for a chip" target |
|
||||
| Inline shortcut over honest rate | ledger M16 (a 256 MiB cache on a die): the recompute attacker's rate on three more cards than the M5 Max's 0.21; the 64 MiB-cache variant inside the 5090's L2 needs a pack exported with a smaller cache and is a follow-up |
|
||||
| nvcc, NVRTC and OpenCL compile times per card | ledger M11 (hourly JIT on real rigs) and M17: whether the hourly program change costs milliseconds or seconds on real NVIDIA drivers, in process and out of process |
|
||||
|
||||
Cost (RunPod pricing page, 3 Oct 2026, USD per hour): RTX 3090 0.22 community / 0.50 secure, RTX 4090 0.34 / 0.74,
|
||||
RTX 5090 0.69 / 0.99; the 3060 is not on RunPod's table (Vast.ai about 0.05 to 0.10, approximate). One run is about
|
||||
45 minutes, so the four NVIDIA cards cost about USD 2 on community cloud or USD 3.50 on secure cloud, plus a few
|
||||
cents of storage. Per-minute billing.
|
||||
|
||||
Command sequence: `infra/gpu-bench/README.md`. In short: `make-bundle.sh` on the Mac, start a pod from the image
|
||||
recipe with SSH, `scp` the bundle, `./run.sh`, read `results.md`, paste the row into `docs/bench-log.md` with the
|
||||
template.
|
||||
|
||||
## The go/no-go question for the project lead
|
||||
|
||||
Spend about EUR 5 and USD 2 to 4 tonight-to-tomorrow for (a) the first reorg-depth, propagation and partition numbers
|
||||
from a real multi-region network, which gate 3 needs and the simulator cannot give, and (b) the per-card hash rates
|
||||
and the three ASIC-resistance measurements (L2 cliff, inline ratio, recompile time) on 3060, 3090 and 4090, closing
|
||||
the measurement half of ledger items M1, M11 and M16 for those cards? If yes: `cd infra/cloud-devnet && ./create.sh`
|
||||
and `cd infra/gpu-bench && ./make-bundle.sh` are the two starting commands. If no: the scripts keep, nothing bills,
|
||||
and the seed node stays at EUR 6.49 per month.
|
||||
84
docs/plans/seed-nodes.md
Normal file
84
docs/plans/seed-nodes.md
Normal file
|
|
@ -0,0 +1,84 @@
|
|||
# Seed nodes: plan
|
||||
|
||||
3 October 2026. Scripts in `infra/seed-nodes/`. The first seed is live.
|
||||
|
||||
## The first seed
|
||||
|
||||
| Item | Value |
|
||||
|---|---|
|
||||
| Name | igneum-seed-1 |
|
||||
| Address | `188.245.5.161:26611` (Hetzner primary IPv4, auto-delete off, so a rebuilt server keeps it) |
|
||||
| Provider, location, type | Hetzner Cloud, Falkenstein (fsn1), cx23 (2 Intel vCPU, 4 GB, 40 GB), Debian 12 |
|
||||
| Cost | EUR 6.49 per month net, EUR 7.79 gross, EUR 0.0104 per hour, 20 TB traffic included (Hetzner API, 3 Oct 2026); plus the primary IPv4, about EUR 0.50 per month (approximate) |
|
||||
| Firewall | inbound tcp 26611 from anywhere, tcp 22 from anywhere (the project lead's instruction; `SSH_SOURCE=me` narrows it to this Mac's IP), icmp; RPC bound to 127.0.0.1 only |
|
||||
| Network | the shared devnet (`--devnet`, no suffix), built from `vendor/igneum-node` HEAD d62708a8 plus the uncommitted finality v2 work and `igneum-pow` HEAD, with `--features igneum-pow` |
|
||||
| Role | p2p open, no mining, address manager on (serves `RequestAddresses`), `--nodnsseed`, UPnP off, `--externalip` set |
|
||||
| How the Mac reaches it | the Mac's live node sits behind NAT (192.168.68.64), so the seed cannot dial in; the Mac dials out: `infra/seed-nodes/addpeer-from-mac.sh 188.245.5.161` (the `addPeer` RPC through grpcurl, permanent, no restart). If the live node is ever restarted by hand: `--addpeer=188.245.5.161:26611` |
|
||||
|
||||
`infra/seed-nodes/seeds.txt` holds exactly `188.245.5.161:26611`.
|
||||
|
||||
## How the seed list reaches clients
|
||||
|
||||
1. Baked into the node. `vendor/igneum-node/consensus/core/src/config/params.rs` holds the per-network seed list as
|
||||
`dns_seeders: &'static [&'static str]` on `Params`: `MAINNET_PARAMS` (line 617 on 3 Oct 2026), `TESTNET_PARAMS`
|
||||
(670), `SIMNET_PARAMS` (721), `DEVNET_PARAMS` (785), all `&[]` since the rename commit emptied Kaspa's nine mainnet
|
||||
and three testnet hostnames. The connection manager resolves each entry with `(seeder, default_p2p_port).to_socket_addrs()`
|
||||
(`components/connectionmanager/src/lib.rs`, `dns_seed_single`), so a plain IPv4 literal works as an entry with no
|
||||
DNS at all: `dns_seeders: &["188.245.5.161"]` on `DEVNET_PARAMS` is the whole change for the devnet, and the
|
||||
testnet list is the same shape with the testnet seeds. The port is the network's default p2p port (devnet 26611;
|
||||
the testnet port is still Kaspa's and must be set with the testnet genesis). The consensus engineer owns this edit.
|
||||
Two consequences for packages: a client that passes `--nodnsseed` ignores the baked list (`kaspad/src/daemon.rs`
|
||||
line 573: `dns_seeders` is emptied when `--nodnsseed` or `--connect` is given), so the Windows node package
|
||||
(`proto-cuda/windows-node/start-node.ps1`) and the cloud scripts must drop `--nodnsseed` once the list is baked;
|
||||
and the list is consulted only when the node is short of outbound peers, so a node with enough `--addpeer`
|
||||
entries never asks a seed.
|
||||
2. `SEED_PEERS` override in every package. Each launcher (Windows node, cloud devnet, seed nodes, the observer's
|
||||
helper node) reads `SEED_PEERS` (comma-separated `ip:port`) and turns every entry into `--addpeer=<entry>`; the
|
||||
baked list is the default when the variable is empty. This is what an operator uses when the baked list is stale
|
||||
between releases.
|
||||
3. DNS names only as a convenience. `seed1.igneum.network` and so on can point at the same addresses (the domains
|
||||
are on Vercel nameservers, so a record each), and `dns_seeders` accepts a hostname too; but the IPs are the source
|
||||
of truth because a DNS failure or a registrar problem must not stop bootstrapping, and because the public key of
|
||||
nothing is involved: a seed only hands out addresses, it cannot forge blocks.
|
||||
|
||||
## Public testnet seed set
|
||||
|
||||
Three to five seeds across two providers and three regions. Proposed:
|
||||
|
||||
| Seed | Provider | Location | Type | Per month net |
|
||||
|---|---|---|---|---|
|
||||
| igneum-seed-1 | Hetzner | Falkenstein (EU) | cx23 | EUR 6.49 (live) |
|
||||
| igneum-seed-2 | Hetzner | Ashburn (US east) | cpx11 (2 GB) or cpx21 (4 GB) | EUR 20.49 or 37.49 |
|
||||
| igneum-seed-3 | DigitalOcean | Singapore (sgp1) | s-2vcpu-4gb | USD 24 (DO pricing page) |
|
||||
| igneum-seed-4 (optional) | DigitalOcean | New York or Frankfurt | s-2vcpu-4gb | USD 24 |
|
||||
| igneum-seed-5 (optional) | Hetzner | Helsinki | cx23 | EUR 6.49 |
|
||||
|
||||
Three seeds: about EUR 50 per month; five: about EUR 80 (approximate, mixed currencies). The US seed is the expensive
|
||||
one because Hetzner's current cx line is EU-only. Every seed is created with `create-seed.sh` (the DigitalOcean
|
||||
variant reserves an IP in the same way) and provisioned with `provision-seed.sh`, which adds the seeds already in
|
||||
`seeds.txt` as `--addpeer` entries so the seeds form a full mesh among themselves. `health.sh` checks them all.
|
||||
|
||||
## Rotation
|
||||
|
||||
1. Add before removing: create and provision the replacement, run `health.sh` until it is synced and has peers.
|
||||
2. Bake the new list (`params.rs`) and release packages with it; keep the old address in the list for one release so
|
||||
clients on the previous build still bootstrap.
|
||||
3. Keep the old IP alive until the release after that (a Hetzner primary IP or a DO reserved IP costs under EUR 1 per
|
||||
month unattached, approximate), then delete the server and the IP, and remove the entry from `seeds.txt` and
|
||||
`seeds.tsv`.
|
||||
4. A compromised seed is the one case to remove first: delete the server, release the IP, bake and release the same
|
||||
day. The damage a bad seed can do is bounded (it hands out addresses; the node's handshake and PoW checks are
|
||||
unchanged), which is why the list may sit in a release rather than behind a signature.
|
||||
|
||||
## Health and operations
|
||||
|
||||
`health.sh` (one line per seed: p2p port reachable from the Mac, unit active, RPC answering, synced, blocks and
|
||||
headers, connected peers, known and banned addresses, disk, memory, version); `--watch` repeats every minute. The
|
||||
seed's journal: `ssh -i ~/.ssh/igneum_ed25519 root@188.245.5.161 journalctl -u igneumd -f`. Updating the binary:
|
||||
`provision-seed.sh` again (it rebuilds on the VM) or `BUILD_WHERE=bin` to push a binary built by the cloud-devnet
|
||||
builder. The database format changes with some fork commits; a seed that refuses to start after an update is wiped
|
||||
(`rm -rf /var/lib/igneum/*`) and resyncs from its peers.
|
||||
|
||||
## No-spend rule
|
||||
|
||||
Only the first seed spends tonight (approved). The remaining seeds and the 20-node network wait for the morning.
|
||||
4
infra/cloud-devnet/.gitignore
vendored
Normal file
4
infra/cloud-devnet/.gitignore
vendored
Normal file
|
|
@ -0,0 +1,4 @@
|
|||
# build artefacts and runtime state; results/ is committed on purpose
|
||||
build/
|
||||
nodes.tsv
|
||||
__pycache__/
|
||||
120
infra/cloud-devnet/README.md
Normal file
120
infra/cloud-devnet/README.md
Normal file
|
|
@ -0,0 +1,120 @@
|
|||
# Igneum cloud devnet: 20 nodes across regions
|
||||
|
||||
A private Igneum devnet (`igneum-devnet-20`: own handshake magic, own genesis) on 20 small Linux VMs in five
|
||||
locations, each running `igneumd` with the real lottery hash and one CPU trickle miner with its own BLS vote key.
|
||||
Built for two measurements that the single-site devnet cannot give: block propagation and reorg behaviour under
|
||||
real inter-region latency (gate 3, spec 03 C1 and the floor rule), and the difficulty controller under hash-rate
|
||||
steps (spec 02 section 2.3). Everything here is scripts; nothing runs or spends until `create.sh` is answered "yes".
|
||||
|
||||
Hetzner Cloud through `hcloud` is the primary path. DigitalOcean through `doctl` is the variant for `create.sh`
|
||||
(and `destroy.sh`); the rest is provider-neutral once `nodes.tsv` exists.
|
||||
|
||||
## The command sequence (the project lead approves the cost, then this, start to finish)
|
||||
|
||||
```
|
||||
cd infra/cloud-devnet
|
||||
brew install hcloud # once; doctl for the DigitalOcean variant
|
||||
# the API token is read from ~/.config/igneum/hetzner-token (the seed-node scripts use the same file)
|
||||
./create.sh # 1. shows the plan and the live prices, asks "yes", creates 20 VMs, writes nodes.tsv
|
||||
./provision.sh # 2. source tarball -> builder VM -> igneumd + igneum-miner -> every node (about 20 min)
|
||||
./start.sh # 3. nodes, block logs, then the 20 trickle miners
|
||||
./status.sh # one line per node (blocks, blue, DAA, tips, peers, difficulty, sink, miner)
|
||||
./experiments/latency.sh 10 # 4. RTT matrix, then 10 min of block propagation samples
|
||||
./experiments/partition.sh sin 10 # 5. cut Singapore off for 10 min, heal, reorg depth and heal time
|
||||
./experiments/hop.sh # 6. hash-rate steps for the controller (about 75 min, schedule in the script)
|
||||
./experiments/collect.sh # 7. pull logs and RPC samples -> results/<date>/, summary.md, bench-log entry
|
||||
./stop.sh && ./destroy.sh # 8. stop, then delete every VM (hourly billing ends)
|
||||
```
|
||||
|
||||
Time from "yes" to the first block: VM creation 1 min, build on the builder 10 to 25 min (approximate; 500 crates
|
||||
plus rocksdb), install 2 min, miners' first 256 MiB cache 10 to 30 s. Steps 4 to 6 are independent; 4 and 6 can run
|
||||
at the same time, 5 should run alone.
|
||||
|
||||
The observer and the live page (`tools/observer`): the RPC never leaves a node's loopback, so either
|
||||
`./experiments/observer.sh tunnel` (an ssh tunnel from the Mac; then run the observer here) or
|
||||
`./experiments/observer.sh remote on` (installs Node 22 on node 1 and runs the observer there, keeping the Mac out of
|
||||
it). With `LIVE_TABLE_PREFIX=cloud_` the record goes to `cloud_live_*` tables; without the prefix the cloud network takes
|
||||
over `/live` on the site (stop the Mac's observer first). `site/api/live.mjs` reads the unprefixed tables only.
|
||||
|
||||
## Cost
|
||||
|
||||
Hetzner API prices for this project on 3 Oct 2026, net EUR per month, hourly billing (the API's `price_hourly`):
|
||||
|
||||
| Location | Type | vCPU | RAM | Per month | Per hour | Note |
|
||||
|---|---|---|---|---|---|---|
|
||||
| fsn1, nbg1, hel1 | cx23 | 2 Intel | 4 GB | 6.49 | 0.0104 | the cheapest x86 type the project can create |
|
||||
| sin | cpx22 | 2 AMD | 4 GB | 30.99 | 0.0497 | cx not sold outside the EU |
|
||||
| ash, hil | cpx21 | 3 AMD | 4 GB | 37.49 | 0.0601 | the old cpx line, only still sold in the US; cpx11 (2 GB) is 20.49 |
|
||||
| builder, any EU | cx43 | 8 Intel | 16 GB | about 20 | about 0.03 | deleted after the build (approximate: cx33 is 9.99, cx43 not priced in the check) |
|
||||
|
||||
Plus about EUR 0.50 per month per primary IPv4 (approximate, Hetzner list price). VAT is added for a private
|
||||
customer; the API's gross column is net x 1.2.
|
||||
|
||||
| Mix | Nodes | Per month net | Per hour | An evening (6 h) |
|
||||
|---|---|---|---|---|
|
||||
| 4 per location (default `REGIONS`) | 8 EU, 4 sin, 4 ash, 4 hil | EUR 476 | EUR 0.76 | about EUR 5 |
|
||||
| EU-heavy (`REGIONS=hel1,fsn1,nbg1,hel1,fsn1,nbg1,hel1,ash,hil,sin`) | 14 EU, 2 sin, 2 ash, 2 hil | EUR 303 | EUR 0.47 | about EUR 3 |
|
||||
| DigitalOcean `s-2vcpu-4gb` in any region (DO pricing page, USD 24/mo, 0.036/h) | 20 | USD 480 | USD 0.71 | about USD 4.50 |
|
||||
|
||||
The run is meant to last one evening, so the hourly column is the one that matters: well under EUR 10 all in, with
|
||||
the builder. Leaving the network up costs the monthly column.
|
||||
|
||||
## What each script does
|
||||
|
||||
| Script | What |
|
||||
|---|---|
|
||||
| `config.sh` | every setting (provider, N, regions, types per location, suffix, genesis bits, mesh degree, miner threads, source tree) |
|
||||
| `create.sh` | ssh key, firewall (22, 26611, icmp), N servers round-robin over `REGIONS`, writes `nodes.tsv` (name, index, region, ip) |
|
||||
| `make-source.sh` | `git archive HEAD` of `NODE_SRC` and `igneum-pow` plus the uncommitted files (`SRC_MODE=head+dirty`, the default: on 3 Oct 2026 every worktree's branch work is uncommitted) into `build/src.tar.gz`, same layout as the Windows package's `src.zip` |
|
||||
| `provision.sh` | `build`: builder VM, `builder/build-on-builder.sh` (apt deps, rustup, `cargo build --release -p kaspad -p igneum-miner --features igneum-pow`), binaries to `build/bin/`, builder deleted. `install`: binaries, `node/wrpc.py`, units and `/etc/igneum/node.env` on every node |
|
||||
| `node/run-igneumd.sh` | the flag list: `--devnet --devnet-suffix=20 --override-params-file (genesis_bits) --rpclisten=127.0.0.1:26610 --rpclisten-json=127.0.0.1:28610 --listen=0.0.0.0:26611 --externalip --addpeer x2 --outpeers=0 --maxinpeers=32 --nodnsseed --disable-upnp --nologfiles --enable-unsynced-mining` |
|
||||
| `node/igneum-miner.service` | `igneum-miner mine grpc://127.0.0.1:26610 1 1000000000 <node> --engine igneum-pow --payout-label <node>`: one CPU thread, its own vote key, votes on every checkpoint |
|
||||
| `node/wrpc.py` | standard-library wRPC JSON client (the node's WebSocket RPC): `call`, `sample`, `watch-blocks` (blocks.tsv, chain.tsv, samples.tsv under /var/log/igneum) |
|
||||
| `start.sh`, `stop.sh`, `status.sh`, `destroy.sh` | as named; `destroy.sh` deletes everything labelled `igneum=devnet` and the firewall |
|
||||
| `experiments/latency.sh` | ping matrix between all nodes, then block propagation from the per-node arrival logs joined on hash (p50, p90, p99 per node and per region) |
|
||||
| `experiments/partition.sh` | iptables cut of one region's p2p traffic to the rest, heal after N min, reorg depth (`getVirtualChainFromBlock` from each node's sink at heal, plus the largest `virtualChainChanged` removal), heal time (sinks converge), conflicting locks (`getFinalityCheckpoints` on both sides) |
|
||||
| `experiments/hop.sh` | a schedule of miner-thread changes per node set (step up, step down, a region off), logged for the controller analysis |
|
||||
| `experiments/collect.sh` | journals, block/chain/sample logs, RPC snapshots into `results/<date>/nodes/<name>/`; `analyze.py summary` and a bench-log entry template; `APPEND_BENCH_LOG=1` appends it to `docs/bench-log.md` |
|
||||
| `experiments/analyze.py` | the offline analysis (RTT table, propagation join, lock conflicts, hop phases, reorg-depth histogram) |
|
||||
| `experiments/observer.sh` | the observer hookup (tunnel or remote) |
|
||||
|
||||
## Network design choices
|
||||
|
||||
- Own network id. `--devnet-suffix=20` gives `igneum-devnet-20` (own handshake magic, own data directory) and the
|
||||
override file sets `genesis_bits` (0x1e400000, 2^18 hashes per block: three 6-thread CPU miners on the M5 Max held
|
||||
about 1 block/s at this value, fork-divergence genesis row), which recomputes the genesis hash. A node of this
|
||||
network can never complete a handshake with the live devnet or the seed.
|
||||
- Sparse mesh. Each node has two permanent outbound links (`--addpeer` to its ring neighbour and to the node half
|
||||
way round) and `--outpeers=0`, so the connection manager does not fill its default 8 outbound slots from
|
||||
exchanged addresses. Inbound links double the count: about 4 peers per node, 40 links in total. Blocks therefore
|
||||
cross several hops between regions, which is what the propagation measurement needs.
|
||||
- `--addpeer`, never `--connect`: `--connect` sets the inbound limit to 0 (kaspad/src/daemon.rs), the Windows node
|
||||
lesson.
|
||||
- RPC on loopback. Every measurement goes through ssh to `wrpc.py` on the node. The firewall opens 22, 26611 and icmp.
|
||||
- One vote key per node. The miner's identity label is the node name, so weight accrues to 20 keys and the
|
||||
finality rule has 20 voters; the devnet finality parameters (interval 30, dust 5, presence 20) apply.
|
||||
- Which tree runs. `NODE_SRC` defaults to `vendor/igneum-node` (master worktree: finality v2 work, uncommitted). For
|
||||
the controller experiment point it at `vendor/igneum-node-diff` (the `difficulty` worktree) or at a merged tree; the
|
||||
binary is cached in `build/bin` and `REBUILD=1` rebuilds. `build/src.stamp` records what went in.
|
||||
- Node 22 is not needed on the nodes: `wrpc.py` is standard-library Python. The observer (remote mode) is the one
|
||||
thing that installs Node, on one node, only when asked.
|
||||
|
||||
## What the results answer
|
||||
|
||||
| Measurement | Where it lands | Decides |
|
||||
|---|---|---|
|
||||
| RTT between regions, block propagation p50/p90/p99 | `results/<date>/latency/` | the latency assumption behind GHOSTDAG k (5 s network delay bound at 1 BPS) and spec 03 C1's lock latency (simulated 2.5 s median at a 2 s inter-region delay) |
|
||||
| Reorg depth distribution over the run | `summary.md` | checkpoint determination depth d (spec 03 C1, ledger F7: placeholder 60, devnet 20) |
|
||||
| Partition: reorg depth, heal time, conflicting locks | `results/<date>/partition-*/partition.md` | the floor rule (lock needs 56.7% of total weight) under a real partition; ledger F2, F3, F11; O-3.6 (two certificates at one index) |
|
||||
| Hop phases: settle time and overshoot per step | `hop.md` in the summary | spec 02 section 2.3's simulator claims (x50 settled 62 s, /50 657 s) on a real network with real timestamps |
|
||||
|
||||
Caveats, stated once: CPU hash rate only (no GPU in this network), one evening of data, clocks by chrony, 20 nodes
|
||||
not 1,000. The numbers are the first real ones; they do not close gate 3 by themselves.
|
||||
|
||||
## Failure notes
|
||||
|
||||
- `create.sh` refuses to run without an active `hcloud` context or `doctl` auth. It never creates a cloud account.
|
||||
- A node whose database predates a binary change must be wiped: `ssh root@<ip> 'systemctl stop igneumd; rm -rf /var/lib/igneum/*; systemctl start igneumd'`.
|
||||
- If `provision.sh build` dies on the builder, the log is `build/build.log`; `KEEP_BUILDER=1` keeps the VM for a retry.
|
||||
- `partition.sh heal` removes the iptables rules if a run was interrupted.
|
||||
- `destroy.sh` is the only thing that stops the bill. Check the provider console afterwards.
|
||||
41
infra/cloud-devnet/builder/build-on-builder.sh
Executable file
41
infra/cloud-devnet/builder/build-on-builder.sh
Executable file
|
|
@ -0,0 +1,41 @@
|
|||
#!/usr/bin/env bash
|
||||
# Runs ON the builder VM (Debian 12 or Ubuntu 22.04 and newer, x86_64) as root. Installs the toolchain, builds
|
||||
# igneumd and igneum-miner with the real lottery hash (--features igneum-pow), leaves the binaries in /root/out/.
|
||||
# Input: /root/src.tar.gz from make-source.sh. Idempotent; a second run only rebuilds what changed.
|
||||
set -euo pipefail
|
||||
|
||||
export DEBIAN_FRONTEND=noninteractive
|
||||
log() { printf '%s %s\n' "$(date -u +%H:%M:%S)" "$*"; }
|
||||
|
||||
if ! command -v cargo >/dev/null 2>&1 && [ ! -x "$HOME/.cargo/bin/cargo" ]; then
|
||||
log "installing build dependencies"
|
||||
apt-get update -qq
|
||||
# rocksdb (librocksdb-sys) needs clang + libclang for bindgen and a C++ compiler; the gRPC and p2p crates run protoc.
|
||||
apt-get install -y -qq build-essential clang libclang-dev llvm-dev pkg-config libssl-dev protobuf-compiler libprotobuf-dev curl git >/dev/null
|
||||
log "installing rust (stable) with rustup"
|
||||
curl -sSf https://sh.rustup.rs | sh -s -- -y --profile minimal --default-toolchain stable >/dev/null
|
||||
fi
|
||||
# shellcheck disable=SC1091
|
||||
. "$HOME/.cargo/env"
|
||||
|
||||
cd /root
|
||||
if [ -f src.tar.gz ]; then
|
||||
rm -rf src; tar -xzf src.tar.gz; log "unpacked src.tar.gz"
|
||||
fi
|
||||
[ -d src/vendor/igneum-node ] || { echo "no src/vendor/igneum-node"; exit 1; }
|
||||
cd src/vendor/igneum-node
|
||||
|
||||
log "building igneumd and igneum-miner (release, --features igneum-pow); about 500 crates, 10 to 25 minutes on 8 vCPU (approximate)"
|
||||
start=$(date +%s)
|
||||
( while sleep 60; do printf '%s build running, %s s\n' "$(date -u +%H:%M:%S)" "$(( $(date +%s) - start ))"; done ) & ticker=$!
|
||||
trap 'kill $ticker 2>/dev/null || true' EXIT
|
||||
cargo build --release -p kaspad -p igneum-miner --features igneum-pow 2>&1 | tail -n 30
|
||||
kill $ticker 2>/dev/null || true
|
||||
log "build finished in $(( $(date +%s) - start )) s"
|
||||
|
||||
mkdir -p /root/out
|
||||
cp target/release/igneumd target/release/igneum-miner /root/out/
|
||||
strip /root/out/igneumd /root/out/igneum-miner 2>/dev/null || true
|
||||
/root/out/igneumd --version | head -1 | tee /root/out/version.txt
|
||||
git -C /root/src/vendor/igneum-node rev-parse --short HEAD >> /root/out/version.txt 2>/dev/null || true
|
||||
ls -la /root/out
|
||||
55
infra/cloud-devnet/config.sh
Executable file
55
infra/cloud-devnet/config.sh
Executable file
|
|
@ -0,0 +1,55 @@
|
|||
# Igneum cloud devnet: settings. Sourced by every script in this directory through lib/common.sh.
|
||||
# Override any value in the environment, for example `N=10 PROVIDER=digitalocean ./create.sh`.
|
||||
# Nothing here spends money by itself; create.sh asks before it creates anything.
|
||||
|
||||
PROVIDER="${PROVIDER:-hetzner}" # hetzner (hcloud CLI) or digitalocean (doctl CLI)
|
||||
N="${N:-20}" # node count
|
||||
PREFIX="${PREFIX:-igneum}" # server names igneum-01 .. igneum-20, the builder is igneum-builder
|
||||
|
||||
# Regions, round-robin over the node index. Hetzner locations: hel1 Helsinki, fsn1 Falkenstein, ash Ashburn (US east),
|
||||
# hil Hillsboro (US west), sin Singapore. With 20 nodes that is 4 per location: 8 in the EU, 4 US east, 4 US west, 4 Asia.
|
||||
REGIONS="${REGIONS:-hel1,fsn1,ash,hil,sin}"
|
||||
DO_REGIONS="${DO_REGIONS:-lon1,fra1,nyc3,sfo3,sgp1}"
|
||||
|
||||
# Instance size per location. The node holds up to four 256 MiB lottery caches (M15 cap), rocksdb and the CPU
|
||||
# miner's own 256 MiB cache, so 2 GB plans are too small. What this Hetzner project can create (API, 3 Oct 2026, net
|
||||
# EUR per month): cx23 (2 Intel vCPU, 4 GB) 6.49 in fsn1/nbg1/hel1 only; cpx22 (2 AMD vCPU, 4 GB) 22.99 EU, 30.99 sin;
|
||||
# cpx21 (3 vCPU, 4 GB) 37.49 in ash/hil only (the old cpx line is deprecated in the EU and Singapore since Dec 2025);
|
||||
# cpx11 (2 GB) 20.49 in ash/hil. So: cx23 in Europe, cpx22 in Singapore, cpx21 in the US. SERVER_TYPE overrides all.
|
||||
TYPE_BY_LOCATION="${TYPE_BY_LOCATION:-fsn1=cx23,nbg1=cx23,hel1=cx23,sin=cpx22,ash=cpx21,hil=cpx21}"
|
||||
SERVER_TYPE="${SERVER_TYPE:-}" # empty = pick from TYPE_BY_LOCATION
|
||||
DO_SIZE="${DO_SIZE:-s-2vcpu-4gb}"
|
||||
BUILDER_TYPE="${BUILDER_TYPE:-cx43}" # 8 Intel vCPU, 16 GB, EUR 0.03/h approximate, for the one-off node build (deleted afterwards)
|
||||
DO_BUILDER_SIZE="${DO_BUILDER_SIZE:-c-8}"
|
||||
IMAGE="${IMAGE:-debian-12}"
|
||||
DO_IMAGE="${DO_IMAGE:-debian-12-x64}"
|
||||
|
||||
# SSH. The operations key ~/.ssh/igneum_ed25519 (3 Oct 2026, uploaded to Hetzner as "igneum-ops" by the seed-node
|
||||
# scripts). If SSH_KEY_FILE is missing, create.sh generates an ed25519 pair there and uploads the public half.
|
||||
SSH_KEY_NAME="${SSH_KEY_NAME:-igneum-ops}"
|
||||
SSH_KEY_FILE="${SSH_KEY_FILE:-$HOME/.ssh/igneum_ed25519}"
|
||||
SSH_USER="${SSH_USER:-root}"
|
||||
|
||||
# The network. A suffixed devnet (igneum-devnet-20) has its own handshake magic and its own genesis (the bits
|
||||
# override recomputes the genesis hash), so it can never peer with or pollute the live devnet on the Mac.
|
||||
DEVNET_SUFFIX="${DEVNET_SUFFIX:-20}"
|
||||
# 0x1e400000 = 2^18 expected hashes per block. Three 6-thread CPU miners on the M5 Max held about 1 block/s at this
|
||||
# value (fork-divergence, genesis row); 20 one-thread cloud vCPUs are in the same range. The DAA takes over after that.
|
||||
GENESIS_BITS="${GENESIS_BITS:-0x1e400000}"
|
||||
MESH_OUT="${MESH_OUT:-2}" # outbound --addpeer links per node (ring plus a chord); inbound doubles it, about 4 peers each
|
||||
MINER_THREADS="${MINER_THREADS:-1}" # the trickle: one CPU thread per node
|
||||
MINER_STATUS_SECS="${MINER_STATUS_SECS:-60}"
|
||||
|
||||
# Source to ship. NODE_SRC is the fork worktree whose code the network runs. master (vendor/igneum-node) carries the
|
||||
# finality v2 work as uncommitted changes; the dual-lane difficulty controller is on the `difficulty` worktree
|
||||
# (vendor/igneum-node-diff). Point NODE_SRC at the tree you want measured.
|
||||
NODE_SRC="${NODE_SRC:-$REPO/vendor/igneum-node}"
|
||||
POW_SRC="${POW_SRC:-$REPO/igneum-pow}"
|
||||
# head = `git archive HEAD` only; head+dirty = HEAD plus every modified and untracked file of the worktree (target*/ excluded).
|
||||
# All seven worktrees sit at d62708a8 with their branch work uncommitted (3 Oct 2026), so head+dirty is the default.
|
||||
SRC_MODE="${SRC_MODE:-head+dirty}"
|
||||
|
||||
# Ports (consensus/core/src/network.rs, devnet). RPC stays on loopback; only p2p is open to the world.
|
||||
P2P_PORT=26611
|
||||
RPC_PORT=26610
|
||||
RPC_JSON_PORT=28610
|
||||
147
infra/cloud-devnet/create.sh
Executable file
147
infra/cloud-devnet/create.sh
Executable file
|
|
@ -0,0 +1,147 @@
|
|||
#!/usr/bin/env bash
|
||||
# Create the cloud devnet VMs: N nodes, regions round-robin, one firewall, one SSH key. Writes nodes.tsv.
|
||||
# Hetzner Cloud through `hcloud` (primary). DigitalOcean through `doctl` with PROVIDER=digitalocean.
|
||||
# Spends money from the moment the servers exist (hourly billing on both providers). It prints the plan and
|
||||
# the provider's live price and asks for "yes" first (YES=1 skips the question).
|
||||
#
|
||||
# Before the first run: `hcloud context create igneum` (paste the project's API token; the project is created by
|
||||
# the project lead in the Hetzner console, not by this script) or `doctl auth init`.
|
||||
# usage: ./create.sh create N nodes
|
||||
# ./create.sh builder create only the builder VM (provision.sh does this itself when it needs one)
|
||||
|
||||
. "$(dirname "$0")/lib/common.sh"
|
||||
|
||||
what="${1:-nodes}"
|
||||
mkdir -p "$BUILD_DIR"
|
||||
|
||||
# ---- SSH key -------------------------------------------------------------------------------------------------
|
||||
if [ ! -f "$SSH_KEY_FILE" ]; then
|
||||
log "no key at $SSH_KEY_FILE, generating an ed25519 pair (no passphrase; it only opens these test VMs)"
|
||||
ssh-keygen -q -t ed25519 -N "" -C "igneum-devnet" -f "$SSH_KEY_FILE"
|
||||
fi
|
||||
PUBKEY_FILE="$SSH_KEY_FILE.pub"
|
||||
[ -f "$PUBKEY_FILE" ] || die "missing $PUBKEY_FILE"
|
||||
|
||||
# ---- Hetzner ------------------------------------------------------------------------------------------------------
|
||||
hetzner_prepare() {
|
||||
need hcloud "brew install hcloud"
|
||||
hcloud context active >/dev/null 2>&1 || die "no active hcloud context: hcloud context create igneum"
|
||||
if ! hcloud ssh-key describe "$SSH_KEY_NAME" >/dev/null 2>&1; then
|
||||
hcloud ssh-key create --name "$SSH_KEY_NAME" --public-key-from-file "$PUBKEY_FILE" >/dev/null
|
||||
log "uploaded ssh key $SSH_KEY_NAME"
|
||||
fi
|
||||
if ! hcloud firewall describe "$PREFIX-devnet" >/dev/null 2>&1; then
|
||||
hcloud firewall create --name "$PREFIX-devnet" --label igneum=devnet >/dev/null
|
||||
hcloud firewall add-rule "$PREFIX-devnet" --direction in --protocol tcp --port 22 --source-ips 0.0.0.0/0 --source-ips ::/0 --description ssh >/dev/null
|
||||
hcloud firewall add-rule "$PREFIX-devnet" --direction in --protocol tcp --port "$P2P_PORT" --source-ips 0.0.0.0/0 --source-ips ::/0 --description igneum-p2p >/dev/null
|
||||
hcloud firewall add-rule "$PREFIX-devnet" --direction in --protocol icmp --source-ips 0.0.0.0/0 --source-ips ::/0 --description ping >/dev/null
|
||||
log "created firewall $PREFIX-devnet (in: 22, $P2P_PORT, icmp; RPC never leaves loopback)"
|
||||
fi
|
||||
}
|
||||
|
||||
hetzner_price() {
|
||||
log "live price list for $1 (per location, EUR, excl. VAT):"
|
||||
hcloud server-type describe "$1" 2>/dev/null | sed -n '/Pricings/,$p' | head -40 || true
|
||||
}
|
||||
|
||||
hetzner_create_one() { # name type location
|
||||
local name="$1" type="$2" loc="$3"
|
||||
if hcloud server describe "$name" >/dev/null 2>&1; then log "$name exists, keeping it"; return; fi
|
||||
hcloud server create --name "$name" --type "$type" --image "$IMAGE" --location "$loc" \
|
||||
--ssh-key "$SSH_KEY_NAME" --firewall "$PREFIX-devnet" --label igneum=devnet --label role="${4:-node}" >/dev/null
|
||||
log "created $name ($type, $loc)"
|
||||
}
|
||||
|
||||
hetzner_ip() { hcloud server ip "$1"; }
|
||||
|
||||
# ---- DigitalOcean --------------------------------------------------------------------------------------------------
|
||||
do_prepare() {
|
||||
need doctl "brew install doctl"
|
||||
doctl account get >/dev/null 2>&1 || die "doctl is not authenticated: doctl auth init"
|
||||
DO_KEY_ID=$(doctl compute ssh-key list --format ID,Name --no-header | awk -v n="$SSH_KEY_NAME" '$2 == n { print $1 }')
|
||||
if [ -z "$DO_KEY_ID" ]; then
|
||||
DO_KEY_ID=$(doctl compute ssh-key import "$SSH_KEY_NAME" --public-key-file "$PUBKEY_FILE" --format ID --no-header)
|
||||
log "imported ssh key $SSH_KEY_NAME ($DO_KEY_ID)"
|
||||
fi
|
||||
DO_FW_ID=$(doctl compute firewall list --format ID,Name --no-header | awk -v n="$PREFIX-devnet" '$2 == n { print $1 }')
|
||||
if [ -z "$DO_FW_ID" ]; then
|
||||
DO_FW_ID=$(doctl compute firewall create --name "$PREFIX-devnet" --tag-names "$PREFIX-devnet" \
|
||||
--inbound-rules "protocol:tcp,ports:22,address:0.0.0.0/0,address:::/0 protocol:tcp,ports:$P2P_PORT,address:0.0.0.0/0,address:::/0 protocol:icmp,address:0.0.0.0/0,address:::/0" \
|
||||
--outbound-rules "protocol:tcp,ports:all,address:0.0.0.0/0,address:::/0 protocol:udp,ports:all,address:0.0.0.0/0,address:::/0 protocol:icmp,address:0.0.0.0/0,address:::/0" \
|
||||
--format ID --no-header)
|
||||
log "created firewall $PREFIX-devnet ($DO_FW_ID), applied by tag"
|
||||
fi
|
||||
}
|
||||
|
||||
do_price() { log "live price list for $1 (USD):"; doctl compute size list --format Slug,Memory,VCPUs,Disk,PriceMonthly,PriceHourly | grep -E "^Slug|^$1 " || true; }
|
||||
|
||||
do_create_one() { # name size region
|
||||
local name="$1" size="$2" reg="$3"
|
||||
if doctl compute droplet get "$name" >/dev/null 2>&1; then log "$name exists, keeping it"; return; fi
|
||||
doctl compute droplet create "$name" --size "$size" --image "$DO_IMAGE" --region "$reg" --ssh-keys "$DO_KEY_ID" \
|
||||
--tag-names "$PREFIX-devnet,role:${4:-node}" --wait >/dev/null
|
||||
log "created $name ($size, $reg)"
|
||||
}
|
||||
|
||||
do_ip() { doctl compute droplet get "$1" --format PublicIPv4 --no-header; }
|
||||
|
||||
# ---- plan and confirm --------------------------------------------------------------------------------------------
|
||||
if [ "$PROVIDER" = digitalocean ]; then
|
||||
do_prepare; TYPE="$DO_SIZE"; BTYPE="$DO_BUILDER_SIZE"
|
||||
else
|
||||
hetzner_prepare; TYPE="${SERVER_TYPE:-by-location}"; BTYPE="$BUILDER_TYPE"
|
||||
fi
|
||||
|
||||
if [ "$what" = builder ]; then
|
||||
log "plan: 1 builder VM $BTYPE in $(region_of 1) on $PROVIDER (deleted by provision.sh after the build unless KEEP_BUILDER=1)"
|
||||
if [ "$PROVIDER" = digitalocean ]; then do_price "$BTYPE"; else hetzner_price "$BTYPE"; fi
|
||||
confirm "create the builder now (hourly billing starts)?"
|
||||
if [ "$PROVIDER" = digitalocean ]; then do_create_one "$PREFIX-builder" "$BTYPE" "$(region_of 1)" builder; ip=$(do_ip "$PREFIX-builder")
|
||||
else hetzner_create_one "$PREFIX-builder" "$BTYPE" "$(region_of 1)" builder; ip=$(hetzner_ip "$PREFIX-builder"); fi
|
||||
printf '%s\n' "$ip" > "$BUILD_DIR/builder.ip"
|
||||
log "builder $ip (saved to build/builder.ip)"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
log "plan: $N nodes ($IMAGE) on $PROVIDER, regions round-robin:"
|
||||
for i in $(seq 1 "$N"); do
|
||||
reg=$(region_of "$i"); t="$TYPE"; [ "$PROVIDER" = digitalocean ] || t=$(type_for_location "$reg")
|
||||
printf ' %s %s %s\n' "$(node_name "$i")" "$reg" "$t"
|
||||
done
|
||||
if [ "$PROVIDER" = digitalocean ]; then do_price "$TYPE"; else
|
||||
for t in $(for i in $(seq 1 "$N"); do type_for_location "$(region_of "$i")"; done | sort -u); do hetzner_price "$t"; done
|
||||
fi
|
||||
log "cost at the Hetzner API prices of 3 Oct 2026 (net EUR per month: cx23 6.49 EU, cpx22 30.99 sin, cpx21 37.49 ash/hil):"
|
||||
log " 20 nodes as 4 per location = 8 x 6.49 + 4 x 30.99 + 8 x 37.49 = EUR 476/mo, EUR 0.76/h; an evening of 6 h about EUR 5 plus the builder (about EUR 0.03/h)"
|
||||
log " EU-heavy alternative REGIONS=hel1,fsn1,nbg1,hel1,fsn1,nbg1,hel1,ash,hil,sin (14 EU, 2 each elsewhere): about EUR 303/mo, EUR 0.47/h"
|
||||
log " DigitalOcean s-2vcpu-4gb: USD 24 per node per month (USD 0.036/h, DO pricing page): 20 nodes USD 480/mo, USD 0.71/h"
|
||||
confirm "create $N servers now (hourly billing starts)?"
|
||||
|
||||
: > "$NODES_FILE.tmp"
|
||||
for i in $(seq 1 "$N"); do
|
||||
name=$(node_name "$i"); reg=$(region_of "$i")
|
||||
if [ "$PROVIDER" = digitalocean ]; then do_create_one "$name" "$TYPE" "$reg"; else hetzner_create_one "$name" "$(type_for_location "$reg")" "$reg"; fi
|
||||
done
|
||||
log "waiting 20 s for the servers to boot, then collecting addresses"
|
||||
sleep 20
|
||||
for i in $(seq 1 "$N"); do
|
||||
name=$(node_name "$i"); reg=$(region_of "$i")
|
||||
if [ "$PROVIDER" = digitalocean ]; then ip=$(do_ip "$name"); else ip=$(hetzner_ip "$name"); fi
|
||||
printf '%s\t%s\t%s\t%s\n' "$name" "$i" "$reg" "$ip" >> "$NODES_FILE.tmp"
|
||||
done
|
||||
mv "$NODES_FILE.tmp" "$NODES_FILE"
|
||||
cp "$NODES_FILE" "$BUILD_DIR/nodes-$(date -u +%Y%m%d-%H%M%S).tsv"
|
||||
log "wrote $NODES_FILE:"
|
||||
cat "$NODES_FILE"
|
||||
|
||||
log "checking ssh on every node (cloud-init can take a minute)"
|
||||
for try in 1 2 3 4 5 6; do
|
||||
bad=0
|
||||
while IFS=$'\t' read -r name idx reg ip; do
|
||||
nssh "$ip" true >/dev/null 2>&1 || { bad=$((bad + 1)); }
|
||||
done < "$NODES_FILE"
|
||||
[ "$bad" = 0 ] && break
|
||||
log "$bad nodes not reachable yet (try $try), waiting 15 s"; sleep 15
|
||||
done
|
||||
[ "$bad" = 0 ] || log "WARNING: $bad nodes still unreachable over ssh; provision.sh will retry them"
|
||||
log "done. Next: ./provision.sh"
|
||||
28
infra/cloud-devnet/destroy.sh
Executable file
28
infra/cloud-devnet/destroy.sh
Executable file
|
|
@ -0,0 +1,28 @@
|
|||
#!/usr/bin/env bash
|
||||
# Delete every VM of the devnet (nodes and builder) and the firewall. The SSH key stays on the provider (free).
|
||||
# Run experiments/collect.sh first: the logs on the VMs go with them.
|
||||
. "$(dirname "$0")/lib/common.sh"
|
||||
|
||||
if [ "$PROVIDER" = digitalocean ]; then
|
||||
need doctl "brew install doctl"
|
||||
ids=$(doctl compute droplet list --tag-name "$PREFIX-devnet" --format ID,Name --no-header)
|
||||
[ -n "$ids" ] || log "no droplets tagged $PREFIX-devnet"
|
||||
printf '%s\n' "$ids"
|
||||
confirm "delete these droplets and the firewall $PREFIX-devnet?"
|
||||
[ -n "$ids" ] && doctl compute droplet delete -f --tag-name "$PREFIX-devnet"
|
||||
doctl compute droplet delete -f "$PREFIX-builder" 2>/dev/null || true
|
||||
fw=$(doctl compute firewall list --format ID,Name --no-header | awk -v n="$PREFIX-devnet" '$2 == n { print $1 }')
|
||||
[ -n "$fw" ] && doctl compute firewall delete -f "$fw"
|
||||
else
|
||||
need hcloud "brew install hcloud"
|
||||
hcloud server list -l igneum=devnet
|
||||
confirm "delete every server labelled igneum=devnet and the firewall $PREFIX-devnet?"
|
||||
names=$(hcloud server list -l igneum=devnet -o columns=name -o noheader)
|
||||
for s in $names; do hcloud server delete "$s" >/dev/null && log "deleted $s"; done
|
||||
hcloud firewall delete "$PREFIX-devnet" >/dev/null 2>&1 && log "deleted firewall $PREFIX-devnet" || true
|
||||
fi
|
||||
if [ -f "$NODES_FILE" ]; then
|
||||
d=$(results_dir_for); cp "$NODES_FILE" "$d/nodes-destroyed-$(date -u +%H%M%S).tsv"; rm -f "$NODES_FILE" "$BUILD_DIR/builder.ip"
|
||||
fi
|
||||
printf '%s\tdestroy\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" >> "$(results_dir_for)/events.log"
|
||||
log "done. Check the provider console once: nothing labelled igneum should remain."
|
||||
308
infra/cloud-devnet/experiments/analyze.py
Normal file
308
infra/cloud-devnet/experiments/analyze.py
Normal file
|
|
@ -0,0 +1,308 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Offline analysis of a results/<date>/ directory (runs on the Mac, standard library only).
|
||||
|
||||
analyze.py rtt <dir> <nodes.tsv> region-pair RTT table from <dir>/rtt.tsv
|
||||
analyze.py propagation <dir> <nodes.tsv> <t0_ms> <t1_ms> per-node and per-region block propagation delay from
|
||||
<dir>/nodes/<name>/blocks.tsv (window t0..t1)
|
||||
analyze.py locks <partition-dir> conflicting locks between the checkpoint dumps at heal
|
||||
analyze.py hop <dir> block rate and difficulty per hop.log phase (node 1 samples)
|
||||
analyze.py summary <dir> <nodes.tsv> summary.md of everything present under <dir>
|
||||
analyze.py bench-entry <dir> <nodes.tsv> a docs/bench-log.md entry (template with the numbers filled)
|
||||
"""
|
||||
import glob, json, os, statistics, sys, time
|
||||
from collections import defaultdict
|
||||
|
||||
|
||||
def read_nodes(path):
|
||||
nodes = []
|
||||
for line in open(path):
|
||||
f = line.rstrip("\n").split("\t")
|
||||
if len(f) >= 4:
|
||||
nodes.append({"name": f[0], "index": int(f[1]), "region": f[2], "ip": f[3]})
|
||||
return nodes
|
||||
|
||||
|
||||
def pct(xs, p):
|
||||
if not xs:
|
||||
return float("nan")
|
||||
xs = sorted(xs)
|
||||
k = min(len(xs) - 1, max(0, int(round((len(xs) - 1) * p))))
|
||||
return xs[k]
|
||||
|
||||
|
||||
def fmt(x, nd=1):
|
||||
return "n/a" if x != x else f"{x:.{nd}f}"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------------------------------------------
|
||||
def rtt(d, nodes_file):
|
||||
nodes = read_nodes(nodes_file)
|
||||
by_ip = {n["ip"]: n for n in nodes}
|
||||
by_name = {n["name"]: n for n in nodes}
|
||||
pairs = defaultdict(list)
|
||||
for line in open(os.path.join(d, "rtt.tsv")):
|
||||
f = line.rstrip("\n").split("\t")
|
||||
if len(f) < 5 or f[3] == "NA":
|
||||
continue
|
||||
src, dst = by_name.get(f[0]), by_ip.get(f[1])
|
||||
if not src or not dst:
|
||||
continue
|
||||
key = tuple(sorted([src["region"], dst["region"]]))
|
||||
pairs[key].append(float(f[3]))
|
||||
regions = sorted({n["region"] for n in nodes})
|
||||
out = ["| region | " + " | ".join(regions) + " |", "|---|" + "---|" * len(regions)]
|
||||
for a in regions:
|
||||
row = [a]
|
||||
for b in regions:
|
||||
xs = pairs.get(tuple(sorted([a, b])), [])
|
||||
row.append(fmt(statistics.median(xs)) if xs else "n/a")
|
||||
out.append("| " + " | ".join(row) + " |")
|
||||
out.append("")
|
||||
out.append("Median of the per-pair average RTT in ms (ping, 5 packets). Same-region cells are the intra-location RTT.")
|
||||
return "\n".join(out)
|
||||
|
||||
|
||||
def propagation(d, nodes_file, t0, t1):
|
||||
nodes = read_nodes(nodes_file)
|
||||
t0, t1 = int(t0), int(t1)
|
||||
seen = defaultdict(dict) # hash -> node -> recv_ms
|
||||
header_ts = {}
|
||||
for n in nodes:
|
||||
p = os.path.join(d, "nodes", n["name"], "blocks.tsv")
|
||||
if not os.path.exists(p):
|
||||
continue
|
||||
for line in open(p):
|
||||
f = line.rstrip("\n").split("\t")
|
||||
if len(f) < 3:
|
||||
continue
|
||||
try:
|
||||
recv = int(f[0])
|
||||
except ValueError:
|
||||
continue
|
||||
if recv < t0 or recv > t1:
|
||||
continue
|
||||
seen[f[1]][n["name"]] = recv
|
||||
try:
|
||||
header_ts[f[1]] = int(f[2])
|
||||
except ValueError:
|
||||
pass
|
||||
total = len(seen)
|
||||
full = {h: m for h, m in seen.items() if len(m) >= max(2, len(nodes) * 0.8)}
|
||||
delays = defaultdict(list)
|
||||
region_delays = defaultdict(list)
|
||||
ts_vs_first = []
|
||||
reg = {n["name"]: n["region"] for n in nodes}
|
||||
for h, m in full.items():
|
||||
first = min(m.values())
|
||||
for name, r in m.items():
|
||||
delays[name].append(r - first)
|
||||
region_delays[reg[name]].append(r - first)
|
||||
if h in header_ts:
|
||||
ts_vs_first.append(first - header_ts[h])
|
||||
out = [f"Blocks in the window: {total}; blocks seen by at least 80% of nodes: {len(full)} (the join base).", ""]
|
||||
out.append("| node | region | blocks | p50 ms | p90 ms | max ms |")
|
||||
out.append("|---|---|---|---|---|---|")
|
||||
for n in nodes:
|
||||
xs = delays.get(n["name"], [])
|
||||
out.append(f"| {n['name']} | {n['region']} | {len(xs)} | {fmt(pct(xs, 0.5), 0)} | {fmt(pct(xs, 0.9), 0)} | {fmt(max(xs) if xs else float('nan'), 0)} |")
|
||||
out.append("")
|
||||
out.append("| region | samples | p50 ms | p90 ms | p99 ms |")
|
||||
out.append("|---|---|---|---|---|")
|
||||
for r in sorted(region_delays):
|
||||
xs = region_delays[r]
|
||||
out.append(f"| {r} | {len(xs)} | {fmt(pct(xs, 0.5), 0)} | {fmt(pct(xs, 0.9), 0)} | {fmt(pct(xs, 0.99), 0)} |")
|
||||
allx = [x for xs in delays.values() for x in xs]
|
||||
out.append("")
|
||||
out.append(f"All nodes: p50 {fmt(pct(allx, 0.5), 0)} ms, p90 {fmt(pct(allx, 0.9), 0)} ms, p99 {fmt(pct(allx, 0.99), 0)} ms, max {fmt(max(allx) if allx else float('nan'), 0)} ms "
|
||||
f"(arrival at a node minus the first arrival anywhere; 0 for the node that produced or first received the block).")
|
||||
if ts_vs_first:
|
||||
out.append(f"First arrival minus the header timestamp: median {fmt(statistics.median(ts_vs_first), 0)} ms (miner clock and template age; negative = miner clock ahead).")
|
||||
out.append("Clock caveat: chrony on every VM; per-node offsets in nodes/<name>/chrony.txt.")
|
||||
return "\n".join(out)
|
||||
|
||||
|
||||
def locks(pdir):
|
||||
dumps = {}
|
||||
for p in glob.glob(os.path.join(pdir, "checkpoints-*-at-heal.json")):
|
||||
try:
|
||||
dumps[os.path.basename(p)[12:-13]] = json.load(open(p))
|
||||
except Exception:
|
||||
continue
|
||||
locked = defaultdict(dict) # index -> hash -> [nodes]
|
||||
for node, j in dumps.items():
|
||||
for cp in (j.get("checkpoints") or []):
|
||||
if str(cp.get("state", "")).lower() == "locked":
|
||||
locked[int(cp["index"])].setdefault(cp.get("hash"), []).append(node)
|
||||
conflicts = {i: h for i, h in locked.items() if len(h) > 1}
|
||||
if not dumps:
|
||||
return "no checkpoint dumps (node without the finality layer, or the RPC was refused)"
|
||||
lines = [f"nodes dumped: {', '.join(sorted(dumps))}; locked indices seen: {len(locked)}; conflicting: {len(conflicts)}"]
|
||||
for i in sorted(conflicts):
|
||||
lines.append(f" index {i}: " + "; ".join(f"{h[:12]} by {','.join(ns)}" for h, ns in conflicts[i].items()))
|
||||
if not conflicts:
|
||||
lines.append(" none: no index locked with two different hashes on the two sides")
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def read_samples(path):
|
||||
rows = []
|
||||
for line in open(path):
|
||||
f = line.rstrip("\n").split("\t")
|
||||
if len(f) >= 9:
|
||||
try:
|
||||
rows.append({"t": int(f[0]), "sink": f[1], "blue": int(f[2] or 0), "daa": int(f[3] or 0), "tips": int(f[4]),
|
||||
"peers": int(f[5]), "difficulty": float(f[6] or 0), "headers": int(f[7] or 0), "blocks": int(f[8] or 0)})
|
||||
except ValueError:
|
||||
continue
|
||||
return rows
|
||||
|
||||
|
||||
def hop(d):
|
||||
logp = os.path.join(d, "hop.log")
|
||||
if not os.path.exists(logp):
|
||||
return "no hop.log in this directory"
|
||||
phases = []
|
||||
for line in open(logp):
|
||||
f = line.rstrip("\n").split("\t")
|
||||
if len(f) >= 3:
|
||||
phases.append({"t": int(f[0]), "sel": f[1], "threads": f[2]})
|
||||
cands = sorted(glob.glob(os.path.join(d, "nodes", "*", "samples.tsv")))
|
||||
if not cands:
|
||||
return "no samples (run collect.sh first)"
|
||||
rows = read_samples(cands[0])
|
||||
out = [f"Samples from {cands[0].split('/')[-2]}; phases from hop.log.", "",
|
||||
"| phase | selector | threads | length s | blocks/min mean | blocks/min last 3 min | difficulty start | difficulty end | settled s (3 min within 10% of 60/min) |",
|
||||
"|---|---|---|---|---|---|---|---|---|"]
|
||||
for i, ph in enumerate(phases):
|
||||
if ph["sel"] == "end":
|
||||
break
|
||||
t_end = phases[i + 1]["t"] if i + 1 < len(phases) else rows[-1]["t"]
|
||||
win = [r for r in rows if ph["t"] <= r["t"] <= t_end]
|
||||
if len(win) < 2:
|
||||
out.append(f"| {i + 1} | {ph['sel']} | {ph['threads']} | {(t_end - ph['t']) // 1000} | n/a | n/a | n/a | n/a | n/a |")
|
||||
continue
|
||||
mins = defaultdict(list)
|
||||
for r in win:
|
||||
mins[(r["t"] - ph["t"]) // 60000].append(r["blocks"])
|
||||
per_min = [(m, max(v) - min(v)) for m, v in sorted(mins.items()) if len(v) >= 2]
|
||||
rates = [x for _, x in per_min]
|
||||
settled = "never"
|
||||
for k in range(len(rates) - 2):
|
||||
if all(abs(rates[k + j] - 60) <= 6 for j in range(3)):
|
||||
settled = str(per_min[k][0] * 60); break
|
||||
mean_rate = (win[-1]["blocks"] - win[0]["blocks"]) / max(1, (win[-1]["t"] - win[0]["t"]) / 60000)
|
||||
last3 = statistics.mean(rates[-3:]) if len(rates) >= 3 else float("nan")
|
||||
out.append(f"| {i + 1} | {ph['sel']} | {ph['threads']} | {(t_end - ph['t']) // 1000} | {mean_rate:.1f} | {fmt(last3)} | {win[0]['difficulty']:.3g} | {win[-1]['difficulty']:.3g} | {settled} |")
|
||||
out.append("")
|
||||
out.append("blocks/min from the node's blockCount (every block, blue or red). Target 60. 'settled' = first minute of three in a row within 10% of target after the step (spec 02 uses a 100-block criterion on the 121-block mean rate; recompute from samples.tsv for the spec's definition).")
|
||||
return "\n".join(out)
|
||||
|
||||
|
||||
def summary(d, nodes_file):
|
||||
nodes = read_nodes(nodes_file)
|
||||
out = [f"# Cloud devnet results, {os.path.basename(d)}", ""]
|
||||
ver = os.path.join(d, "version.txt")
|
||||
if os.path.exists(ver):
|
||||
out.append("Node build: " + " ".join(open(ver).read().split()))
|
||||
stamp = os.path.join(d, "src.stamp")
|
||||
if os.path.exists(stamp):
|
||||
out.append("Source: " + "; ".join(l.strip() for l in open(stamp) if l.strip()))
|
||||
region_counts = ", ".join("%s x %d" % (r, sum(1 for n in nodes if n["region"] == r)) for r in sorted({n["region"] for n in nodes}))
|
||||
out.append(f"Nodes: {len(nodes)} ({region_counts})")
|
||||
out.append("")
|
||||
out.append("## Final state per node")
|
||||
out.append("")
|
||||
out.append("| node | region | blocks | headers | blue | daa | tips | peers | difficulty | synced | miner blocks found |")
|
||||
out.append("|---|---|---|---|---|---|---|---|---|---|---|")
|
||||
for n in nodes:
|
||||
nd = os.path.join(d, "nodes", n["name"])
|
||||
s = read_samples(os.path.join(nd, "samples.tsv")) if os.path.exists(os.path.join(nd, "samples.tsv")) else []
|
||||
last = s[-1] if s else {}
|
||||
synced = "?"
|
||||
try:
|
||||
synced = str(json.load(open(os.path.join(nd, "getInfo.json"))).get("isSynced"))
|
||||
except Exception:
|
||||
pass
|
||||
found = "?"
|
||||
ml = os.path.join(nd, "miner.log")
|
||||
if os.path.exists(ml):
|
||||
found = str(sum(1 for l in open(ml, errors="replace") if "accepted" in l.lower() and "block" in l.lower()))
|
||||
out.append(f"| {n['name']} | {n['region']} | {last.get('blocks', '?')} | {last.get('headers', '?')} | {last.get('blue', '?')} | {last.get('daa', '?')} | {last.get('tips', '?')} | {last.get('peers', '?')} | {last.get('difficulty', 0):.3g} | {synced} | {found} |")
|
||||
out.append("")
|
||||
ev = os.path.join(d, "events.log")
|
||||
if os.path.exists(ev):
|
||||
out.append("## Events"); out.append(""); out.append("```"); out.append(open(ev).read().rstrip()); out.append("```"); out.append("")
|
||||
lat = os.path.join(d, "latency")
|
||||
if os.path.isdir(lat):
|
||||
for f in ("rtt-by-region.md", "propagation.md"):
|
||||
p = os.path.join(lat, f)
|
||||
if os.path.exists(p):
|
||||
out.append(f"## Latency: {f}"); out.append(""); out.append(open(p).read().rstrip()); out.append("")
|
||||
for p in sorted(glob.glob(os.path.join(d, "partition-*", "partition.md"))):
|
||||
out.append("## " + os.path.basename(os.path.dirname(p))); out.append(""); out.append(open(p).read().rstrip()); out.append("")
|
||||
if os.path.exists(os.path.join(d, "hop.log")):
|
||||
out.append("## Hash-rate steps (hop.sh)"); out.append(""); out.append(hop(d)); out.append("")
|
||||
# reorg distribution over the whole run from chain.tsv (every virtualChainChanged removal)
|
||||
removals = []
|
||||
for n in nodes:
|
||||
p = os.path.join(d, "nodes", n["name"], "chain.tsv")
|
||||
if os.path.exists(p):
|
||||
for line in open(p):
|
||||
f = line.rstrip("\n").split("\t")
|
||||
if len(f) >= 3 and f[2].isdigit() and int(f[2]) > 0:
|
||||
removals.append(int(f[2]))
|
||||
out.append("## Reorg depth distribution (all nodes, whole run, virtualChainChanged removals > 0)")
|
||||
out.append("")
|
||||
if removals:
|
||||
hist = defaultdict(int)
|
||||
for r in removals:
|
||||
hist[r] += 1
|
||||
out.append("| depth | count |"); out.append("|---|---|")
|
||||
for k in sorted(hist):
|
||||
out.append(f"| {k} | {hist[k]} |")
|
||||
out.append("")
|
||||
out.append(f"{len(removals)} reorgs; p50 {pct(removals, 0.5)}, p99 {pct(removals, 0.99)}, max {max(removals)} (spec 03 C1: d is set from this distribution).")
|
||||
else:
|
||||
out.append("no reorgs recorded (or no chain.tsv)")
|
||||
return "\n".join(out)
|
||||
|
||||
|
||||
def bench_entry(d, nodes_file):
|
||||
nodes = read_nodes(nodes_file)
|
||||
day = os.path.basename(d)
|
||||
try:
|
||||
date_txt = time.strftime("%-d %B %Y", time.strptime(day, "%Y-%m-%d"))
|
||||
except ValueError:
|
||||
date_txt = day
|
||||
ver = " ".join(open(os.path.join(d, "version.txt")).read().split()) if os.path.exists(os.path.join(d, "version.txt")) else "<igneumd version>"
|
||||
regions = ", ".join("%s x %d" % (r, sum(1 for n in nodes if n["region"] == r)) for r in sorted({n["region"] for n in nodes}))
|
||||
lines = [f"## {date_txt}, cloud devnet: {len(nodes)} igneumd nodes across regions, CPU trickle miners, latency, partition and hash-rate steps (consensus-engineer)", "",
|
||||
f"Machines: {len(nodes)} Hetzner Cloud VMs ({regions}; type cpx21 unless noted), Debian 12, chrony. Node: {ver}, built on a builder VM from the source tarball in results/{day}/src.stamp. Network: igneum-devnet-<suffix>, genesis bits <bits>, {len(nodes)} one-thread CPU miners (igneum-miner --engine igneum-pow), one BLS vote key per node, sparse --addpeer mesh (about 4 peers each).", "",
|
||||
"Inter-region RTT (ms, median of pair averages): <paste results/" + day + "/latency/rtt-by-region.md>", "",
|
||||
"Block propagation (arrival minus first arrival anywhere): p50 <n> ms, p90 <n> ms, p99 <n> ms, max <n> ms; per region <paste the region table from propagation.md>.", "",
|
||||
"Partition <region>, <minutes> min: minority reorg depth max <n> (getVirtualChainFromBlock from the minority sink at heal), majority <n>; converged <n> s after heal; conflicting locks: <none | list>.", "",
|
||||
"Hash-rate steps (hop.sh): <paste the phase table from hop.md>: settle times per step against spec 02 section 2.3 (simulator: x50 settled 62 s, /50 657 s).", "",
|
||||
"Reorg depth distribution over the run: p50 <n>, p99 <n>, max <n> over <n> reorgs (summary.md).", "",
|
||||
"Reading: <what this says about gate 3 (checkpoint depth d, the floor rule under a real partition) and the controller under real latency>. Caveats: CPU hash rate only (no GPU), one evening, clocks by chrony.", ""]
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def main(a):
|
||||
if len(a) >= 4 and a[1] == "rtt":
|
||||
print(rtt(a[2], a[3]))
|
||||
elif len(a) >= 6 and a[1] == "propagation":
|
||||
print(propagation(a[2], a[3], a[4], a[5]))
|
||||
elif len(a) >= 3 and a[1] == "locks":
|
||||
print(locks(a[2]))
|
||||
elif len(a) >= 3 and a[1] == "hop":
|
||||
print(hop(a[2]))
|
||||
elif len(a) >= 4 and a[1] == "summary":
|
||||
print(summary(a[2], a[3]))
|
||||
elif len(a) >= 4 and a[1] == "bench-entry":
|
||||
print(bench_entry(a[2], a[3]))
|
||||
else:
|
||||
sys.stderr.write(__doc__); sys.exit(2)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main(sys.argv)
|
||||
38
infra/cloud-devnet/experiments/collect.sh
Executable file
38
infra/cloud-devnet/experiments/collect.sh
Executable file
|
|
@ -0,0 +1,38 @@
|
|||
#!/usr/bin/env bash
|
||||
# Pull logs and RPC samples from every node into results/<date>/nodes/<name>/ and write summary.md plus a
|
||||
# bench-log entry template (results/<date>/bench-log-entry.md). APPEND_BENCH_LOG=1 also appends that entry to
|
||||
# docs/bench-log.md (the log is append-only; edit the numbers in the entry file first if a run looked odd).
|
||||
# ./experiments/collect.sh [date] default today (UTC)
|
||||
. "$(dirname "$0")/../lib/common.sh"
|
||||
require_nodes
|
||||
out="$(results_dir_for "${1:-}")"; mkdir -p "$out/nodes"
|
||||
log "collecting into $out"
|
||||
cp "$NODES_FILE" "$out/nodes.tsv"
|
||||
[ -f "$BIN_DIR/version.txt" ] && cp "$BIN_DIR/version.txt" "$BIN_DIR/src.stamp" "$out/" 2>/dev/null
|
||||
|
||||
while IFS=$'\t' read -r name idx reg ip; do
|
||||
(
|
||||
d="$out/nodes/$name"; mkdir -p "$d"
|
||||
nssh "$ip" "journalctl -u igneumd --no-pager -o short-iso --since -24h | tail -c 30000000" > "$d/igneumd.log" 2>/dev/null
|
||||
nssh "$ip" "journalctl -u igneum-miner --no-pager -o short-iso --since -24h | tail -c 5000000" > "$d/miner.log" 2>/dev/null
|
||||
for f in blocks chain samples; do nscp "$SSH_USER@$ip:/var/log/igneum/$f.tsv" "$d/$f.tsv" 2>/dev/null || true; done
|
||||
nssh "$ip" "chronyc tracking 2>/dev/null; echo; free -m; echo; df -h / | tail -1; echo; cat /etc/igneum/node.env" > "$d/host.txt" 2>/dev/null
|
||||
for m in getInfo getBlockDagInfo getConnectedPeerInfo getFinalityWeights; do
|
||||
nssh "$ip" "python3 /opt/igneum/bin/wrpc.py call $m" > "$d/$m.json" 2>/dev/null || true
|
||||
done
|
||||
nssh "$ip" "python3 /opt/igneum/bin/wrpc.py call getFinalityCheckpoints '{\"last\": 50}'" > "$d/getFinalityCheckpoints.json" 2>/dev/null || true
|
||||
printf '[%s] %s blocks lines, %s chain lines, %s samples\n' "$name" "$(grep -c . "$d/blocks.tsv" 2>/dev/null || echo 0)" "$(grep -c . "$d/chain.tsv" 2>/dev/null || echo 0)" "$(grep -c . "$d/samples.tsv" 2>/dev/null || echo 0)"
|
||||
) &
|
||||
while [ "$(jobs -r | wc -l)" -ge 10 ]; do sleep 1; done
|
||||
done < "$NODES_FILE"
|
||||
wait
|
||||
|
||||
python3 "$HERE/experiments/analyze.py" summary "$out" "$NODES_FILE" | tee "$out/summary.md"
|
||||
python3 "$HERE/experiments/analyze.py" bench-entry "$out" "$NODES_FILE" > "$out/bench-log-entry.md"
|
||||
log "wrote $out/summary.md and $out/bench-log-entry.md"
|
||||
if [ "${APPEND_BENCH_LOG:-0}" = 1 ]; then
|
||||
{ printf '\n'; cat "$out/bench-log-entry.md"; } >> "$REPO/docs/bench-log.md"
|
||||
log "appended the entry to docs/bench-log.md (review it before committing)"
|
||||
else
|
||||
log "APPEND_BENCH_LOG=1 appends the entry to docs/bench-log.md"
|
||||
fi
|
||||
50
infra/cloud-devnet/experiments/hop.sh
Executable file
50
infra/cloud-devnet/experiments/hop.sh
Executable file
|
|
@ -0,0 +1,50 @@
|
|||
#!/usr/bin/env bash
|
||||
# Experiment 1c: hash-rate steps for the difficulty controller (spec 02 section 2.3; the dual-lane rule is on the
|
||||
# `difficulty` worktree, NODE_SRC=vendor/igneum-node-diff when building; on master it exercises Kaspa's sampled DAA).
|
||||
# ./experiments/hop.sh "<schedule>"
|
||||
# schedule = phases separated by ";", phase = <selector>:<threads>:<seconds>
|
||||
# selector: all | half (odd indices) | odd | even | region:<name> | <from>-<to> (indices)
|
||||
# threads: miner threads on the selected nodes (0 stops the miner); unselected nodes keep their setting
|
||||
# Default schedule (about 75 min): steady, 4x step up on half the nodes, back down, one region off, back:
|
||||
# all:1:900;half:4:900;all:1:900;region:hel1:0:600;all:1:900
|
||||
# Every phase change is logged to results/<date>/hop.log (ts_ms, selector, threads); collect.sh and analyze.py hop
|
||||
# line the difficulty samples up against it.
|
||||
. "$(dirname "$0")/../lib/common.sh"
|
||||
require_nodes
|
||||
schedule="${1:-all:1:900;half:4:900;all:1:900;region:hel1:0:600;all:1:900}"
|
||||
out="$(results_dir_for)"; logf="$out/hop.log"
|
||||
|
||||
select_nodes() {
|
||||
case "$1" in
|
||||
all) node_names ;;
|
||||
half|odd) awk -F'\t' '$2 % 2 == 1 { print $1 }' "$NODES_FILE" ;;
|
||||
even) awk -F'\t' '$2 % 2 == 0 { print $1 }' "$NODES_FILE" ;;
|
||||
region:*) nodes_in_region "${1#region:}" ;;
|
||||
*-*) awk -F'\t' -v a="${1%-*}" -v b="${1#*-}" '$2 >= a && $2 <= b { print $1 }' "$NODES_FILE" ;;
|
||||
*) die "bad selector $1" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
set_threads() { # node threads
|
||||
local ip; ip=$(node_ip "$1")
|
||||
if [ "$2" = 0 ]; then
|
||||
nssh "$ip" "systemctl stop igneum-miner"
|
||||
else
|
||||
nssh "$ip" "sed -i 's/^MINER_THREADS=.*/MINER_THREADS=$2/' /etc/igneum/node.env && systemctl restart igneum-miner"
|
||||
fi
|
||||
}
|
||||
|
||||
log "schedule: $schedule"
|
||||
IFS=';' read -r -a phases <<< "$schedule"
|
||||
for ph in "${phases[@]}"; do
|
||||
sel="${ph%%:*}"; rest="${ph#*:}"
|
||||
if [ "$sel" = region ]; then sel="region:${rest%%:*}"; rest="${rest#*:}"; fi
|
||||
threads="${rest%%:*}"; secs="${rest#*:}"
|
||||
nodes=$(select_nodes "$sel")
|
||||
log "phase: $sel -> $threads thread(s) for $secs s ($(printf '%s\n' "$nodes" | wc -l | tr -d ' ') nodes)"
|
||||
for n in $nodes; do set_threads "$n" "$threads" & done; wait
|
||||
printf '%s\t%s\t%s\t%s\n' "$(date +%s)000" "$sel" "$threads" "$secs" >> "$logf"
|
||||
sleep "$secs"
|
||||
done
|
||||
printf '%s\tend\t-\t-\n' "$(date +%s)000" >> "$logf"
|
||||
log "done; phases in $logf. Pull the samples with ./experiments/collect.sh and read results/<date>/hop.md"
|
||||
37
infra/cloud-devnet/experiments/latency.sh
Executable file
37
infra/cloud-devnet/experiments/latency.sh
Executable file
|
|
@ -0,0 +1,37 @@
|
|||
#!/usr/bin/env bash
|
||||
# Experiment 1a: inter-region RTT and block propagation delay.
|
||||
# ./experiments/latency.sh [minutes] default 10
|
||||
# 1. RTT: every node pings every other node (5 pings, 0.2 s apart) in parallel; results/<date>/latency/rtt.tsv and a
|
||||
# region-pair table.
|
||||
# 2. Propagation: the block logs run anyway (igneum-blocklog); this waits `minutes`, pulls blocks.tsv from every node
|
||||
# and joins on the block hash: first arrival anywhere = t0 of that block, per node delay = arrival - t0.
|
||||
# Clocks: chrony on every VM; the per-node offset is recorded (chronyc tracking) and the join is only as good as it.
|
||||
. "$(dirname "$0")/../lib/common.sh"
|
||||
require_nodes
|
||||
minutes="${1:-10}"
|
||||
out="$(results_dir_for)/latency"; mkdir -p "$out/nodes"
|
||||
all_ips=$(cut -f4 "$NODES_FILE" | tr '\n' ' ')
|
||||
t_start=$(date +%s)000
|
||||
|
||||
log "RTT matrix: $(node_count) nodes x $(( $(node_count) - 1 )) targets, 5 pings each"
|
||||
: > "$out/rtt.tsv"
|
||||
while IFS=$'\t' read -r name idx reg ip; do
|
||||
(
|
||||
nssh "$ip" "for t in $all_ips; do [ \"\$t\" = \"$ip\" ] && continue; r=\$(ping -c 5 -i 0.2 -W 2 \$t 2>/dev/null | awk -F'/' '/rtt|round-trip/ { print \$4\"\t\"\$5\"\t\"\$6 }'); printf '%s\t%s\t%s\n' '$name' \"\$t\" \"\${r:-NA\tNA\tNA}\"; done" 2>/dev/null
|
||||
) >> "$out/rtt.tsv" &
|
||||
done < "$NODES_FILE"
|
||||
wait
|
||||
log "rtt.tsv: $(grep -c . "$out/rtt.tsv") pairs. Region-pair medians (avg RTT ms):"
|
||||
python3 "$HERE/experiments/analyze.py" rtt "$out" "$NODES_FILE" | tee "$out/rtt-by-region.md"
|
||||
|
||||
log "propagation window: $minutes min from $(date -u +%H:%M:%S) (the miners must be running: ./start.sh miners)"
|
||||
sleep $(( minutes * 60 ))
|
||||
t_end=$(date +%s)000
|
||||
while IFS=$'\t' read -r name idx reg ip; do
|
||||
mkdir -p "$out/nodes/$name"
|
||||
nscp "$SSH_USER@$ip:/var/log/igneum/blocks.tsv" "$out/nodes/$name/blocks.tsv" 2>/dev/null || log "no blocks.tsv on $name"
|
||||
nssh "$ip" "chronyc tracking 2>/dev/null | grep -E 'System time|RMS offset'" > "$out/nodes/$name/chrony.txt" 2>/dev/null || true
|
||||
done < "$NODES_FILE"
|
||||
python3 "$HERE/experiments/analyze.py" propagation "$out" "$NODES_FILE" "$t_start" "$t_end" | tee "$out/propagation.md"
|
||||
printf '%s\tlatency %s min, window %s..%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$minutes" "$t_start" "$t_end" >> "$(results_dir_for)/events.log"
|
||||
log "written: $out/rtt.tsv, rtt-by-region.md, propagation.md"
|
||||
55
infra/cloud-devnet/experiments/observer.sh
Executable file
55
infra/cloud-devnet/experiments/observer.sh
Executable file
|
|
@ -0,0 +1,55 @@
|
|||
#!/usr/bin/env bash
|
||||
# Observer hookup: let tools/observer/observer.mjs watch one cloud node so /live on the site shows the 20-node network.
|
||||
# The RPC never leaves the node's loopback, so the observer reaches it through an ssh tunnel (from the Mac) or runs
|
||||
# on the node itself (remote mode, keeps the overloaded Mac out of it).
|
||||
#
|
||||
# ./experiments/observer.sh tunnel [node] ssh -L 28610 -> the node's wRPC JSON; prints the observer command
|
||||
# ./experiments/observer.sh remote on [node] install Node 22 on the node, copy observer.mjs and DATABASE_URL, start a unit
|
||||
# ./experiments/observer.sh remote off [node] stop and remove that unit (the DATABASE_URL file is deleted too)
|
||||
#
|
||||
# Tables: with LIVE_TABLE_PREFIX=cloud_ the observer writes cloud_live_* and the public page keeps showing the Mac's
|
||||
# devnet (site/api/live.mjs reads the unprefixed tables; it has no prefix switch, so a side-by-side page needs that
|
||||
# one-line change in site/). Without the prefix the cloud network REPLACES the Mac's devnet on /live: stop the
|
||||
# Mac's observer first, or two writers fight over live_state.
|
||||
. "$(dirname "$0")/../lib/common.sh"
|
||||
require_nodes
|
||||
mode="${1:-}"; node="${3:-${2:-}}"
|
||||
case "$mode" in
|
||||
tunnel)
|
||||
node="${2:-$(node_name 1)}"; ip=$(node_ip "$node")
|
||||
log "tunnel to $node ($ip): local ws://127.0.0.1:${LOCAL_PORT:-28620} -> node 127.0.0.1:$RPC_JSON_PORT (Ctrl+C ends it)"
|
||||
log "in another shell, from the repo root:"
|
||||
log " IGNEUM_RPC=ws://127.0.0.1:${LOCAL_PORT:-28620} LIVE_TABLE_PREFIX=cloud_ node tools/observer/observer.mjs # parallel record"
|
||||
log " IGNEUM_RPC=ws://127.0.0.1:${LOCAL_PORT:-28620} node tools/observer/observer.mjs # takes over /live (stop the Mac's observer first)"
|
||||
exec ssh "${SSH_OPTS[@]}" -N -L "${LOCAL_PORT:-28620}:127.0.0.1:$RPC_JSON_PORT" "$SSH_USER@$ip"
|
||||
;;
|
||||
remote)
|
||||
sub="${2:-}"; node="${3:-$(node_name 1)}"; ip=$(node_ip "$node")
|
||||
if [ "$sub" = on ]; then
|
||||
envf="$HOME/.config/igneum/env"; [ -f "$envf" ] || die "no $envf (DATABASE_URL)"
|
||||
dburl=$(grep -E '^DATABASE_URL=' "$envf" | head -1)
|
||||
[ -n "$dburl" ] || die "DATABASE_URL not in $envf"
|
||||
log "installing Node 22 (NodeSource apt repository) and the observer on $node"
|
||||
nssh "$ip" 'command -v node >/dev/null && node -v | grep -q "^v2[2-9]" || { curl -fsSL https://deb.nodesource.com/setup_22.x | bash - >/dev/null && apt-get install -y -qq nodejs >/dev/null; }; node -v'
|
||||
nscp "$REPO/tools/observer/observer.mjs" "$SSH_USER@$ip:/opt/igneum/observer.mjs"
|
||||
printf '%s\nIGNEUM_RPC=ws://127.0.0.1:%s\nLIVE_TABLE_PREFIX=%s\n' "$dburl" "$RPC_JSON_PORT" "${LIVE_TABLE_PREFIX:-cloud_}" | nssh "$ip" 'umask 077; cat > /etc/igneum/observer.env'
|
||||
nssh "$ip" 'cat > /etc/systemd/system/igneum-observer.service <<EOF
|
||||
[Unit]
|
||||
Description=Igneum devnet observer (writes the live tables in Neon)
|
||||
After=igneumd.service
|
||||
[Service]
|
||||
EnvironmentFile=/etc/igneum/observer.env
|
||||
ExecStart=/usr/bin/node /opt/igneum/observer.mjs
|
||||
Restart=always
|
||||
RestartSec=10
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
EOF
|
||||
systemctl daemon-reload && systemctl enable --now igneum-observer && sleep 3 && journalctl -u igneum-observer --no-pager -n 5'
|
||||
log "observer running on $node with LIVE_TABLE_PREFIX=${LIVE_TABLE_PREFIX:-cloud_} (empty prefix = it takes over /live)"
|
||||
elif [ "$sub" = off ]; then
|
||||
nssh "$ip" 'systemctl disable --now igneum-observer 2>/dev/null; rm -f /etc/systemd/system/igneum-observer.service /etc/igneum/observer.env; systemctl daemon-reload; echo removed'
|
||||
else die "usage: observer.sh remote on|off [node]"; fi
|
||||
;;
|
||||
*) die "usage: observer.sh tunnel [node] | observer.sh remote on|off [node]" ;;
|
||||
esac
|
||||
113
infra/cloud-devnet/experiments/partition.sh
Executable file
113
infra/cloud-devnet/experiments/partition.sh
Executable file
|
|
@ -0,0 +1,113 @@
|
|||
#!/usr/bin/env bash
|
||||
# Experiment 1b: cut one region off, heal it, measure the reorg depth and the heal time.
|
||||
# ./experiments/partition.sh <region> <minutes> e.g. ./experiments/partition.sh sin 10
|
||||
# ./experiments/partition.sh heal remove every partition rule (if a run was interrupted)
|
||||
#
|
||||
# Cut: on every node of the region, iptables DROP of p2p traffic (port 26611, both directions, both roles) to and
|
||||
# from every node outside the region, tagged with a comment so heal finds exactly these rules. The minority keeps
|
||||
# mining on its own tips, the majority on theirs (both sides have --enable-unsynced-mining).
|
||||
# Heal: the rules are deleted; the nodes reconnect through their --addpeer retries (backoff up to 8 minutes in
|
||||
# components/connectionmanager, so a reconnect kick restarts igneumd on the minority side when RECONNECT_KICK=1,
|
||||
# default on: a restart reconnects at once and the chain state is on disk).
|
||||
# Measured, through the loopback RPC:
|
||||
# reorg depth = removedChainBlockHashes of getVirtualChainFromBlock(start = the node's own sink at heal time),
|
||||
# per node, after convergence; plus the largest single virtualChainChanged removal in chain.tsv
|
||||
# heal time = seconds from heal until all nodes report at most 2 distinct sinks for 3 consecutive 5 s polls
|
||||
# locks = getFinalityCheckpoints on both sides at heal: any index locked with two different hashes is a
|
||||
# conflicting lock (the gate 3 question; the rule's floor predicts none)
|
||||
. "$(dirname "$0")/../lib/common.sh"
|
||||
require_nodes
|
||||
|
||||
heal_all() {
|
||||
log "removing partition rules everywhere"
|
||||
on_all "iptables -S 2>/dev/null | grep -- '--comment igneum-partition' | sed 's/^-A/-D/' | while read -r r; do iptables \$r; done; true" >/dev/null
|
||||
}
|
||||
|
||||
region="${1:-}"; minutes="${2:-10}"
|
||||
[ -n "$region" ] || die "usage: partition.sh <region> <minutes> | partition.sh heal"
|
||||
if [ "$region" = heal ]; then heal_all; exit 0; fi
|
||||
minority=$(nodes_in_region "$region"); [ -n "$minority" ] || die "no nodes in region $region (present: $(regions_present | tr '\n' ' '))"
|
||||
majority=$(node_names | grep -vxF -f <(printf '%s\n' "$minority"))
|
||||
outside_ips=$(for n in $majority; do node_ip "$n"; done | tr '\n' ' ')
|
||||
stamp=$(date -u +%Y%m%d-%H%M%S)
|
||||
out="$(results_dir_for)/partition-$region-$stamp"; mkdir -p "$out"
|
||||
log "partition: region $region ($(printf '%s\n' "$minority" | wc -l | tr -d ' ') nodes) cut off for $minutes min; majority $(printf '%s\n' "$majority" | wc -l | tr -d ' ') nodes"
|
||||
|
||||
snapshot() { # file: one sample line per node, prefixed with the node name
|
||||
: > "$1"
|
||||
while IFS=$'\t' read -r name idx reg ip; do
|
||||
( printf '%s\t%s\n' "$name" "$(nssh "$ip" 'python3 /opt/igneum/bin/wrpc.py sample' 2>/dev/null)" >> "$1" ) &
|
||||
done < "$NODES_FILE"; wait
|
||||
}
|
||||
checkpoints() { # node file
|
||||
nssh "$(node_ip "$1")" "python3 /opt/igneum/bin/wrpc.py call getFinalityCheckpoints '{\"last\": 100}'" > "$2" 2>/dev/null || echo '{}' > "$2"
|
||||
}
|
||||
|
||||
snapshot "$out/before.tsv"
|
||||
t0=$(date +%s)
|
||||
for n in $minority; do
|
||||
rules=""
|
||||
for ip in $outside_ips; do
|
||||
rules="$rules iptables -I INPUT -s $ip -p tcp --dport $P2P_PORT -m comment --comment igneum-partition -j DROP;"
|
||||
rules="$rules iptables -I INPUT -s $ip -p tcp --sport $P2P_PORT -m comment --comment igneum-partition -j DROP;"
|
||||
rules="$rules iptables -I OUTPUT -d $ip -p tcp --dport $P2P_PORT -m comment --comment igneum-partition -j DROP;"
|
||||
rules="$rules iptables -I OUTPUT -d $ip -p tcp --sport $P2P_PORT -m comment --comment igneum-partition -j DROP;"
|
||||
done
|
||||
nssh "$(node_ip "$n")" "$rules true" && log "cut $n"
|
||||
done
|
||||
printf '%s\tpartition %s cut\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$region" >> "$(results_dir_for)/events.log"
|
||||
trap 'log "interrupted: healing"; heal_all' INT TERM
|
||||
|
||||
: > "$out/during.tsv"
|
||||
for i in $(seq 1 $(( minutes * 2 ))); do
|
||||
sleep 30
|
||||
snapshot "$out/during-$i.tsv"; cat "$out/during-$i.tsv" >> "$out/during.tsv"; rm -f "$out/during-$i.tsv"
|
||||
m_sinks=$(for n in $minority; do awk -F'\t' -v n="$n" '$1 == n { print $3 }' "$out/during.tsv" | tail -1; done | sort -u | wc -l | tr -d ' ')
|
||||
M_sinks=$(for n in $majority; do awk -F'\t' -v n="$n" '$1 == n { print $3 }' "$out/during.tsv" | tail -1; done | sort -u | wc -l | tr -d ' ')
|
||||
log "t+$(( i * 30 )) s: minority distinct sinks $m_sinks, majority distinct sinks $M_sinks"
|
||||
done
|
||||
|
||||
log "heal: snapshots and checkpoints on both sides, then rules off"
|
||||
snapshot "$out/at-heal.tsv"
|
||||
for n in $minority; do checkpoints "$n" "$out/checkpoints-$n-at-heal.json"; done
|
||||
first_major=$(printf '%s\n' "$majority" | head -1); checkpoints "$first_major" "$out/checkpoints-$first_major-at-heal.json"
|
||||
heal_all
|
||||
t_heal=$(date +%s)
|
||||
trap - INT TERM
|
||||
if [ "${RECONNECT_KICK:-1}" = 1 ]; then
|
||||
for n in $minority; do nssh "$(node_ip "$n")" 'systemctl restart igneumd igneum-blocklog' && log "kicked $n (restart, reconnects at once)"; done
|
||||
fi
|
||||
printf '%s\tpartition %s healed\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$region" >> "$(results_dir_for)/events.log"
|
||||
|
||||
log "waiting for convergence (at most 2 distinct sinks for 3 polls, 5 s apart; up to 20 min)"
|
||||
ok=0; converged_at=""
|
||||
for i in $(seq 1 240); do
|
||||
sleep 5
|
||||
snapshot "$out/poll.tsv"
|
||||
d=$(cut -f3 "$out/poll.tsv" | grep -c . ); u=$(cut -f3 "$out/poll.tsv" | sort -u | grep -c .)
|
||||
if [ "$u" -le 2 ] && [ "$d" = "$(node_count)" ]; then ok=$((ok + 1)); else ok=0; fi
|
||||
[ $(( i % 6 )) = 0 ] && log "t+$(( i * 5 )) s after heal: $u distinct sinks"
|
||||
if [ "$ok" -ge 3 ]; then converged_at=$(( $(date +%s) - t_heal - 10 )); break; fi
|
||||
done
|
||||
[ -n "$converged_at" ] && log "converged $converged_at s after heal" || log "no convergence within 20 min (see poll.tsv)"
|
||||
|
||||
log "reorg depth from each node's own sink at heal (getVirtualChainFromBlock)"
|
||||
: > "$out/reorg.tsv"
|
||||
while IFS=$'\t' read -r name idx reg ip; do
|
||||
sink=$(awk -F'\t' -v n="$name" '$1 == n { print $3 }' "$out/at-heal.tsv")
|
||||
side=majority; printf '%s\n' "$minority" | grep -qx "$name" && side=minority
|
||||
r=$(nssh "$ip" "python3 /opt/igneum/bin/wrpc.py call getVirtualChainFromBlock '{\"startHash\": \"$sink\", \"includeAcceptedTransactionIds\": false}'" 2>/dev/null \
|
||||
| python3 -c 'import json,sys; j=json.load(sys.stdin); print(len(j.get("removedChainBlockHashes",[])), len(j.get("addedChainBlockHashes",[])))' 2>/dev/null || echo "NA NA")
|
||||
chainmax=$(nssh "$ip" "awk -F'\t' -v t=$(( t_heal * 1000 )) '\$1 >= t && \$3 > m { m = \$3 } END { print m + 0 }' /var/log/igneum/chain.tsv" 2>/dev/null || echo NA)
|
||||
printf '%s\t%s\t%s\t%s\t%s\n' "$name" "$side" "$sink" "$r" "$chainmax" | tr ' ' '\t' >> "$out/reorg.tsv"
|
||||
done < "$NODES_FILE"
|
||||
{
|
||||
printf '# Partition %s, %s min, %s\n\n' "$region" "$minutes" "$stamp"
|
||||
printf 'cut at %s, healed at %s, converged %s s after heal (criterion: at most 2 distinct sinks for 3 polls)\n\n' "$(date -u -r "$t0" +%H:%M:%S 2>/dev/null || date -u -d @"$t0" +%H:%M:%S)" "$(date -u -r "$t_heal" +%H:%M:%S 2>/dev/null || date -u -d @"$t_heal" +%H:%M:%S)" "${converged_at:-none}"
|
||||
printf '| node | side | sink at heal | removed (reorg depth) | added | max single removal after heal |\n|---|---|---|---|---|---|\n'
|
||||
awk -F'\t' '{ printf "| %s | %s | %s | %s | %s | %s |\n", $1, $2, substr($3, 1, 12), $4, $5, $6 }' "$out/reorg.tsv"
|
||||
printf '\nMinority reorg depth, max: %s. Majority, max: %s.\n' "$(awk -F'\t' '$2 == "minority" && $4 != "NA" && $4 > m { m = $4 } END { print m + 0 }' "$out/reorg.tsv")" "$(awk -F'\t' '$2 == "majority" && $4 != "NA" && $4 > m { m = $4 } END { print m + 0 }' "$out/reorg.tsv")"
|
||||
printf '\nConflicting locks at heal (same checkpoint index, different hash, both locked):\n'
|
||||
python3 "$HERE/experiments/analyze.py" locks "$out"
|
||||
} | tee "$out/partition.md"
|
||||
log "written: $out/partition.md (plus before/during/at-heal/poll/reorg tsv and the checkpoint dumps)"
|
||||
91
infra/cloud-devnet/lib/common.sh
Normal file
91
infra/cloud-devnet/lib/common.sh
Normal file
|
|
@ -0,0 +1,91 @@
|
|||
# Shared helpers for infra/cloud-devnet. Source this; it sources config.sh.
|
||||
# Needs bash 3.2 or newer (the Mac's /bin/bash is fine), ssh, scp, python3.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
REPO="$(cd "$HERE/../.." && pwd)"
|
||||
export REPO
|
||||
# shellcheck source=../config.sh
|
||||
. "$HERE/config.sh"
|
||||
|
||||
NODES_FILE="${NODES_FILE:-$HERE/nodes.tsv}" # name<TAB>index<TAB>region<TAB>ip, one line per node, written by create.sh
|
||||
BUILD_DIR="$HERE/build"
|
||||
BIN_DIR="$BUILD_DIR/bin"
|
||||
RESULTS_DIR="$HERE/results"
|
||||
|
||||
# Hetzner API token: the only line of ~/.config/igneum/hetzner-token (never printed, never in the repo).
|
||||
if [ -z "${HCLOUD_TOKEN:-}" ] && [ -s "${HETZNER_TOKEN_FILE:-$HOME/.config/igneum/hetzner-token}" ]; then
|
||||
HCLOUD_TOKEN="$(head -1 "${HETZNER_TOKEN_FILE:-$HOME/.config/igneum/hetzner-token}" | tr -d '[:space:]')"; export HCLOUD_TOKEN
|
||||
fi
|
||||
|
||||
log() { printf '%s %s\n' "$(date -u +%H:%M:%S)" "$*"; }
|
||||
die() { log "ERROR: $*" >&2; exit 1; }
|
||||
need() { command -v "$1" >/dev/null 2>&1 || die "$1 is not installed ($2)"; }
|
||||
|
||||
# Hetzner location name for node index i (1-based) and the DigitalOcean one.
|
||||
region_of() {
|
||||
local i="$1" list
|
||||
if [ "$PROVIDER" = digitalocean ]; then list="$DO_REGIONS"; else list="$REGIONS"; fi
|
||||
local n; n=$(printf '%s' "$list" | tr ',' '\n' | grep -c .)
|
||||
printf '%s' "$list" | tr ',' '\n' | sed -n "$(( (i - 1) % n + 1 ))p"
|
||||
}
|
||||
|
||||
node_name() { printf '%s-%02d' "$PREFIX" "$1"; }
|
||||
|
||||
# Server type for a Hetzner location: SERVER_TYPE if set, else the TYPE_BY_LOCATION map, else cx23.
|
||||
type_for_location() {
|
||||
if [ -n "$SERVER_TYPE" ]; then printf '%s' "$SERVER_TYPE"; return; fi
|
||||
local t; t=$(printf '%s' "$TYPE_BY_LOCATION" | tr ',' '\n' | awk -F= -v l="$1" '$1 == l { print $2 }')
|
||||
printf '%s' "${t:-cx23}"
|
||||
}
|
||||
|
||||
# nodes.tsv accessors
|
||||
require_nodes() { [ -s "$NODES_FILE" ] || die "no $NODES_FILE: run create.sh first"; }
|
||||
node_count() { grep -c . "$NODES_FILE"; }
|
||||
node_ip() { awk -F'\t' -v n="$1" '$1 == n || $2 == n { print $4; exit }' "$NODES_FILE"; }
|
||||
node_region() { awk -F'\t' -v n="$1" '$1 == n || $2 == n { print $3; exit }' "$NODES_FILE"; }
|
||||
node_names() { cut -f1 "$NODES_FILE"; }
|
||||
nodes_in_region() { awk -F'\t' -v r="$1" '$3 == r { print $1 }' "$NODES_FILE"; }
|
||||
regions_present() { cut -f3 "$NODES_FILE" | sort -u; }
|
||||
|
||||
# Outbound peers of node index i: the ring neighbour i+1 and chords at i + k*N/MESH_OUT (1-based, wrapping).
|
||||
# With MESH_OUT=2 that is i+1 and i+N/2; inbound links double it, so each node sees about 4 peers.
|
||||
peers_of() {
|
||||
local i="$1" n; n=$(node_count)
|
||||
local k step idx out=""
|
||||
for k in $(seq 1 "$MESH_OUT"); do
|
||||
if [ "$k" = 1 ]; then step=1; else step=$(( (k - 1) * n / MESH_OUT )); fi
|
||||
idx=$(( (i - 1 + step) % n + 1 ))
|
||||
[ "$idx" = "$i" ] && continue
|
||||
out="$out $(node_ip "$idx"):$P2P_PORT"
|
||||
done
|
||||
printf '%s' "$out" | sed 's/^ //' | tr ' ' ','
|
||||
}
|
||||
|
||||
SSH_OPTS=(-i "$SSH_KEY_FILE" -o StrictHostKeyChecking=accept-new -o UserKnownHostsFile="$HERE/build/known_hosts" -o ConnectTimeout=15 -o BatchMode=yes -o ServerAliveInterval=30)
|
||||
nssh() { local ip="$1"; shift; ssh "${SSH_OPTS[@]}" "$SSH_USER@$ip" "$@"; }
|
||||
nscp() { scp -q "${SSH_OPTS[@]}" "$@"; }
|
||||
|
||||
# Run a command on every node in parallel (PAR at a time), prefixing output with the node name.
|
||||
# usage: on_all '<shell command>' (the command runs on the node through ssh)
|
||||
on_all() {
|
||||
local cmd="$1" par="${PAR:-10}"
|
||||
mkdir -p "$BUILD_DIR"
|
||||
node_names | xargs -P "$par" -I{} sh -c '
|
||||
name="$1"; cmd="$2"; ip=$(awk -F"\t" -v n="$name" "\$1 == n { print \$4; exit }" "'"$NODES_FILE"'")
|
||||
ssh -i "'"$SSH_KEY_FILE"'" -o StrictHostKeyChecking=accept-new -o UserKnownHostsFile="'"$HERE/build/known_hosts"'" -o ConnectTimeout=15 -o BatchMode=yes "'"$SSH_USER"'@$ip" "$cmd" 2>&1 | sed "s/^/[$name] /"
|
||||
' _ {} "$cmd"
|
||||
}
|
||||
|
||||
# One RPC call on a node through the loopback wRPC JSON endpoint (node/wrpc.py is installed by provision.sh).
|
||||
rpc() { local node="$1" method="$2" params="${3:-{\}}"; nssh "$(node_ip "$node")" "python3 /opt/igneum/bin/wrpc.py call $method '$params'"; }
|
||||
|
||||
results_dir_for() { local d="$RESULTS_DIR/${1:-$(date -u +%Y-%m-%d)}"; mkdir -p "$d"; printf '%s' "$d"; }
|
||||
|
||||
genesis_bits_decimal() { printf '%d' "$GENESIS_BITS"; }
|
||||
|
||||
confirm() {
|
||||
if [ "${YES:-0}" = 1 ]; then return 0; fi
|
||||
printf '%s [type yes]: ' "$1"; read -r a; [ "$a" = yes ] || die "not confirmed"
|
||||
}
|
||||
45
infra/cloud-devnet/make-source.sh
Executable file
45
infra/cloud-devnet/make-source.sh
Executable file
|
|
@ -0,0 +1,45 @@
|
|||
#!/usr/bin/env bash
|
||||
# Export the node source as one tarball the builder VM can compile without any GitHub access:
|
||||
# src/vendor/igneum-node/ `git archive HEAD` of NODE_SRC, plus (SRC_MODE=head+dirty) every modified and untracked file
|
||||
# src/igneum-pow/ the same for POW_SRC
|
||||
# The layout keeps the fork's relative path dependency `igneum-pow = { path = "../../../../igneum-pow" }`
|
||||
# (vendor/igneum-node/consensus/pow/Cargo.toml) valid. Output: build/src.tar.gz and build/src.stamp (what went in).
|
||||
# This is the Windows package's src.zip recipe (proto-cuda/windows-node/make-package.sh) in Linux form.
|
||||
|
||||
. "$(dirname "$0")/lib/common.sh"
|
||||
|
||||
mkdir -p "$BUILD_DIR"
|
||||
stage="$BUILD_DIR/src"
|
||||
rm -rf "$stage"; mkdir -p "$stage/vendor/igneum-node" "$stage/igneum-pow"
|
||||
|
||||
export_tree() { # repo-dir dest label
|
||||
local src="$1" dst="$2" label="$3" head dirty=0
|
||||
[ -d "$src" ] || die "$label: $src does not exist"
|
||||
head=$(git -C "$src" rev-parse --short HEAD)
|
||||
git -C "$src" archive --format=tar HEAD | tar -x -C "$dst"
|
||||
if [ "$SRC_MODE" = "head+dirty" ]; then
|
||||
# modified, deleted and untracked files of the worktree (never target directories or the .git dir)
|
||||
local list; list=$(git -C "$src" ls-files -m -o --exclude-standard | grep -vE '^target' || true)
|
||||
if [ -n "$list" ]; then
|
||||
dirty=$(printf '%s\n' "$list" | grep -c .)
|
||||
printf '%s\n' "$list" | while IFS= read -r f; do
|
||||
if [ -e "$src/$f" ]; then mkdir -p "$dst/$(dirname "$f")"; cp -p "$src/$f" "$dst/$f"; else rm -f "$dst/$f"; fi
|
||||
done
|
||||
fi
|
||||
fi
|
||||
printf '%s: HEAD %s (%s), %s uncommitted files overlaid, mode %s\n' "$label" "$head" "$(git -C "$src" rev-parse --abbrev-ref HEAD)" "$dirty" "$SRC_MODE"
|
||||
}
|
||||
|
||||
{
|
||||
printf 'exported %s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
||||
export_tree "$NODE_SRC" "$stage/vendor/igneum-node" "igneum-node ($NODE_SRC)"
|
||||
export_tree "$POW_SRC" "$stage/igneum-pow" "igneum-pow ($POW_SRC)"
|
||||
} | tee "$BUILD_DIR/src.stamp"
|
||||
|
||||
# the fork's Cargo.lock is tracked, so the builder resolves the same crate versions; drop any stale target dirs
|
||||
rm -rf "$stage/vendor/igneum-node/target" "$stage/igneum-pow/target"
|
||||
tar -C "$BUILD_DIR" -czf "$BUILD_DIR/src.tar.gz" src
|
||||
rm -rf "$stage"
|
||||
log "wrote $BUILD_DIR/src.tar.gz ($(du -h "$BUILD_DIR/src.tar.gz" | cut -f1))"
|
||||
grep -q 'mode head$' "$BUILD_DIR/src.stamp" && log "NOTE: SRC_MODE=head ships HEAD only; on 3 Oct 2026 every worktree's branch work (finality v2, difficulty, hotswap) is uncommitted and would be left out"
|
||||
exit 0
|
||||
7
infra/cloud-devnet/node/blocklog.sh
Executable file
7
infra/cloud-devnet/node/blocklog.sh
Executable file
|
|
@ -0,0 +1,7 @@
|
|||
#!/usr/bin/env bash
|
||||
# Block and chain log on a node VM: one line per block the node accepts (arrival time against the header timestamp),
|
||||
# one line per virtual chain change (the reorg record), one sample line every 5 s. Written by wrpc.py into
|
||||
# /var/log/igneum/{blocks,chain,samples}.tsv, collected by experiments/collect.sh. Restarts when the node restarts.
|
||||
set -euo pipefail
|
||||
mkdir -p /var/log/igneum
|
||||
exec python3 /opt/igneum/bin/wrpc.py watch-blocks /var/log/igneum
|
||||
18
infra/cloud-devnet/node/igneum-blocklog.service
Normal file
18
infra/cloud-devnet/node/igneum-blocklog.service
Normal file
|
|
@ -0,0 +1,18 @@
|
|||
[Unit]
|
||||
Description=Igneum block arrival, chain change and sample log (wrpc.py watch-blocks)
|
||||
After=igneumd.service
|
||||
Requires=igneumd.service
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
User=igneum
|
||||
Group=igneum
|
||||
ExecStart=/opt/igneum/bin/blocklog.sh
|
||||
Restart=always
|
||||
RestartSec=5
|
||||
StandardOutput=journal
|
||||
StandardError=journal
|
||||
SyslogIdentifier=igneum-blocklog
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
24
infra/cloud-devnet/node/igneum-miner.service
Normal file
24
infra/cloud-devnet/node/igneum-miner.service
Normal file
|
|
@ -0,0 +1,24 @@
|
|||
[Unit]
|
||||
Description=Igneum CPU trickle miner (igneum-miner, real lottery hash, votes with this node's key)
|
||||
After=igneumd.service
|
||||
Requires=igneumd.service
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
User=igneum
|
||||
Group=igneum
|
||||
EnvironmentFile=/etc/igneum/node.env
|
||||
# positional: <grpc url> <threads> <seconds> <vote-key-label>. The label is the node name, so each node has its own
|
||||
# BLS vote key and its own payout address (--payout-label), deterministic across restarts.
|
||||
# 1e9 seconds = never stop on its own. MINER_THREADS=0 is handled by hop.sh (it stops the unit instead).
|
||||
ExecStart=/bin/sh -c 'exec /opt/igneum/bin/igneum-miner mine grpc://127.0.0.1:26610 "$MINER_THREADS" 1000000000 "$NODE_NAME" --engine igneum-pow --payout-label "$NODE_NAME" --label "$NODE_NAME" --status-secs "$MINER_STATUS_SECS"'
|
||||
Restart=always
|
||||
RestartSec=15
|
||||
Nice=10
|
||||
MemoryMax=1200M
|
||||
StandardOutput=journal
|
||||
StandardError=journal
|
||||
SyslogIdentifier=igneum-miner
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
22
infra/cloud-devnet/node/igneumd.service
Normal file
22
infra/cloud-devnet/node/igneumd.service
Normal file
|
|
@ -0,0 +1,22 @@
|
|||
[Unit]
|
||||
Description=Igneum devnet node (igneumd)
|
||||
After=network-online.target chrony.service
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
User=igneum
|
||||
Group=igneum
|
||||
EnvironmentFile=/etc/igneum/node.env
|
||||
ExecStart=/opt/igneum/bin/run-igneumd.sh
|
||||
Restart=always
|
||||
RestartSec=10
|
||||
LimitNOFILE=65536
|
||||
# the node keeps up to four 256 MiB lottery caches plus rocksdb; stop a runaway before the VM swaps to death
|
||||
MemoryMax=3200M
|
||||
StandardOutput=journal
|
||||
StandardError=journal
|
||||
SyslogIdentifier=igneumd
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
38
infra/cloud-devnet/node/install-node.sh
Executable file
38
infra/cloud-devnet/node/install-node.sh
Executable file
|
|
@ -0,0 +1,38 @@
|
|||
#!/usr/bin/env bash
|
||||
# Runs ON a node VM as root (Debian 12 or Ubuntu). Called by provision.sh with:
|
||||
# install-node.sh <name> <index> <devnet-suffix> <genesis-bits-decimal> <peer-list> <external-ip> <miner-threads> <status-secs>
|
||||
# Expects /opt/igneum/bin/{igneumd,igneum-miner,wrpc.py,run-igneumd.sh,blocklog.sh} and the three unit files in /root.
|
||||
set -euo pipefail
|
||||
name="$1"; idx="$2"; suffix="$3"; bits="$4"; peers="$5"; extip="$6"; threads="$7"; status_secs="${8:-60}"
|
||||
|
||||
export DEBIAN_FRONTEND=noninteractive
|
||||
if ! command -v chronyd >/dev/null 2>&1 || ! command -v iptables >/dev/null 2>&1; then
|
||||
apt-get update -qq
|
||||
apt-get install -y -qq chrony iptables python3 jq >/dev/null
|
||||
fi
|
||||
systemctl enable --now chrony >/dev/null 2>&1 || true
|
||||
|
||||
id igneum >/dev/null 2>&1 || useradd --system --home /var/lib/igneum --shell /usr/sbin/nologin igneum
|
||||
mkdir -p /var/lib/igneum /var/log/igneum /etc/igneum /opt/igneum/bin
|
||||
chmod +x /opt/igneum/bin/*
|
||||
chown -R igneum:igneum /var/lib/igneum /var/log/igneum
|
||||
|
||||
# A private devnet: own handshake magic (suffix) and own genesis (bits override recomputes the genesis hash).
|
||||
# OverrideParams is deny_unknown_fields, so only known keys go in here (consensus/core/src/config/params.rs).
|
||||
printf '{"genesis_bits": %s}\n' "$bits" > /etc/igneum/override-params.json
|
||||
|
||||
cat > /etc/igneum/node.env <<EOF
|
||||
NODE_NAME=$name
|
||||
NODE_INDEX=$idx
|
||||
DEVNET_SUFFIX=$suffix
|
||||
PEERS=$peers
|
||||
EXTERNAL_IP=$extip
|
||||
MINER_THREADS=$threads
|
||||
MINER_STATUS_SECS=$status_secs
|
||||
EXTRA_ARGS=
|
||||
EOF
|
||||
|
||||
cp /root/igneumd.service /root/igneum-miner.service /root/igneum-blocklog.service /etc/systemd/system/
|
||||
systemctl daemon-reload
|
||||
systemctl enable igneumd igneum-blocklog igneum-miner >/dev/null 2>&1
|
||||
echo "installed $name (index $idx, suffix $suffix, bits $bits, peers $peers, $threads miner thread(s)); $(/opt/igneum/bin/igneumd --version | head -1)"
|
||||
20
infra/cloud-devnet/node/run-igneumd.sh
Executable file
20
infra/cloud-devnet/node/run-igneumd.sh
Executable file
|
|
@ -0,0 +1,20 @@
|
|||
#!/usr/bin/env bash
|
||||
# igneumd launcher on a node VM. Reads /etc/igneum/node.env and builds the flag list.
|
||||
# The flags are the ones that work on the Windows node (proto-cuda/windows-node/start-node.ps1): --addpeer, never
|
||||
# --connect (--connect sets the inbound limit to 0, kaspad/src/daemon.rs). --outpeers=0 stops the connection manager
|
||||
# filling 8 outbound slots from exchanged addresses, which would turn the sparse mesh into a near-complete graph;
|
||||
# the --addpeer links are permanent connections outside that target. RPC on loopback only; p2p on every interface.
|
||||
# --enable-unsynced-mining: every node starts at genesis (timestamp 2026-10-03T00:00Z), so none is "synced" at first
|
||||
# and the node would refuse templates (rpc/service/src/service.rs).
|
||||
set -euo pipefail
|
||||
. /etc/igneum/node.env
|
||||
|
||||
args=(--devnet --devnet-suffix="$DEVNET_SUFFIX" --override-params-file=/etc/igneum/override-params.json
|
||||
--appdir=/var/lib/igneum --rpclisten=127.0.0.1:26610 --rpclisten-json=127.0.0.1:28610
|
||||
--listen=0.0.0.0:26611 --externalip="$EXTERNAL_IP"
|
||||
--nodnsseed --disable-upnp --nologfiles --enable-unsynced-mining --outpeers=0 --maxinpeers=32 --loglevel=info)
|
||||
IFS=',' read -r -a peers <<< "${PEERS:-}"
|
||||
for p in "${peers[@]}"; do [ -n "$p" ] && args+=(--addpeer="$p"); done
|
||||
# shellcheck disable=SC2206
|
||||
[ -n "${EXTRA_ARGS:-}" ] && args+=($EXTRA_ARGS)
|
||||
exec /opt/igneum/bin/igneumd "${args[@]}"
|
||||
207
infra/cloud-devnet/node/wrpc.py
Executable file
207
infra/cloud-devnet/node/wrpc.py
Executable file
|
|
@ -0,0 +1,207 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Minimal wRPC JSON client for igneumd, standard library only (any Debian python3, no websocket package).
|
||||
|
||||
The node's wRPC JSON endpoint (--rpclisten-json=127.0.0.1:28610) is a plain WebSocket. Requests are
|
||||
{"id": n, "method": "getBlockDagInfo", "params": {...}}; the answer carries the same id and either "params" (the
|
||||
result) or "error"; notifications arrive as {"method": "blockAddedNotification", "params": {...}}. Field names are
|
||||
camelCase (serde rename_all on the RPC model). This mirrors the Rpc class in tools/observer/observer.mjs.
|
||||
|
||||
wrpc.py call <method> [json-params] one call, prints the result as JSON
|
||||
wrpc.py watch-blocks <dir> subscribe and append, forever:
|
||||
<dir>/blocks.tsv recv_ms hash header_timestamp_ms blue_score daa_score parents
|
||||
<dir>/chain.tsv recv_ms added removed first_removed_hashes(up to 3, comma separated)
|
||||
<dir>/samples.tsv ts_ms sink blue_score daa_score tips peers difficulty headers blocks (every 5 s)
|
||||
wrpc.py sample one samples.tsv line to stdout
|
||||
Environment: IGNEUM_RPC (default ws://127.0.0.1:28610).
|
||||
"""
|
||||
import base64, json, os, socket, struct, sys, time, threading
|
||||
from urllib.parse import urlparse
|
||||
|
||||
RPC = os.environ.get("IGNEUM_RPC", "ws://127.0.0.1:28610")
|
||||
|
||||
|
||||
class WebSocket:
|
||||
"""RFC 6455 client: text frames, masking, ping/pong, continuation frames."""
|
||||
|
||||
def __init__(self, url, timeout=10.0):
|
||||
u = urlparse(url)
|
||||
self.sock = socket.create_connection((u.hostname, u.port or 80), timeout=timeout)
|
||||
key = base64.b64encode(os.urandom(16)).decode()
|
||||
path = u.path or "/"
|
||||
req = (f"GET {path} HTTP/1.1\r\nHost: {u.hostname}:{u.port}\r\nUpgrade: websocket\r\nConnection: Upgrade\r\n"
|
||||
f"Sec-WebSocket-Key: {key}\r\nSec-WebSocket-Version: 13\r\n\r\n")
|
||||
self.sock.sendall(req.encode())
|
||||
head = b""
|
||||
while b"\r\n\r\n" not in head:
|
||||
chunk = self.sock.recv(4096)
|
||||
if not chunk:
|
||||
raise ConnectionError("handshake: connection closed")
|
||||
head += chunk
|
||||
status, _, rest = head.partition(b"\r\n\r\n")
|
||||
if b" 101 " not in status.split(b"\r\n")[0]:
|
||||
raise ConnectionError("handshake refused: " + status.split(b"\r\n")[0].decode(errors="replace"))
|
||||
self.buf = rest
|
||||
self.lock = threading.Lock()
|
||||
|
||||
def _read(self, n):
|
||||
while len(self.buf) < n:
|
||||
chunk = self.sock.recv(65536)
|
||||
if not chunk:
|
||||
raise ConnectionError("connection closed")
|
||||
self.buf += chunk
|
||||
out, self.buf = self.buf[:n], self.buf[n:]
|
||||
return out
|
||||
|
||||
def send_text(self, text):
|
||||
data = text.encode()
|
||||
mask = os.urandom(4)
|
||||
n = len(data)
|
||||
if n < 126:
|
||||
hdr = struct.pack("!BB", 0x81, 0x80 | n)
|
||||
elif n < 65536:
|
||||
hdr = struct.pack("!BBH", 0x81, 0x80 | 126, n)
|
||||
else:
|
||||
hdr = struct.pack("!BBQ", 0x81, 0x80 | 127, n)
|
||||
masked = bytes(b ^ mask[i % 4] for i, b in enumerate(data))
|
||||
with self.lock:
|
||||
self.sock.sendall(hdr + mask + masked)
|
||||
|
||||
def _send_control(self, opcode, payload=b""):
|
||||
mask = os.urandom(4)
|
||||
with self.lock:
|
||||
self.sock.sendall(struct.pack("!BB", 0x80 | opcode, 0x80 | len(payload)) + mask + bytes(b ^ mask[i % 4] for i, b in enumerate(payload)))
|
||||
|
||||
def recv_text(self):
|
||||
"""Returns the next complete text message (handles ping and continuation)."""
|
||||
message = b""
|
||||
while True:
|
||||
b0, b1 = self._read(2)
|
||||
fin, opcode = b0 & 0x80, b0 & 0x0F
|
||||
n = b1 & 0x7F
|
||||
if n == 126:
|
||||
n = struct.unpack("!H", self._read(2))[0]
|
||||
elif n == 127:
|
||||
n = struct.unpack("!Q", self._read(8))[0]
|
||||
if b1 & 0x80:
|
||||
mask = self._read(4)
|
||||
payload = bytes(b ^ mask[i % 4] for i, b in enumerate(self._read(n)))
|
||||
else:
|
||||
payload = self._read(n)
|
||||
if opcode == 0x9:
|
||||
self._send_control(0xA, payload); continue
|
||||
if opcode == 0xA:
|
||||
continue
|
||||
if opcode == 0x8:
|
||||
raise ConnectionError("server closed the websocket")
|
||||
message += payload
|
||||
if fin:
|
||||
return message.decode()
|
||||
|
||||
def close(self):
|
||||
try:
|
||||
self._send_control(0x8)
|
||||
except Exception:
|
||||
pass
|
||||
self.sock.close()
|
||||
|
||||
|
||||
class Rpc:
|
||||
def __init__(self, url=RPC):
|
||||
self.ws = WebSocket(url)
|
||||
self.next_id = 0
|
||||
self.notifications = []
|
||||
|
||||
def call(self, method, params=None, timeout=10.0):
|
||||
self.next_id += 1
|
||||
rid = self.next_id
|
||||
self.ws.send_text(json.dumps({"id": rid, "method": method, "params": params or {}}))
|
||||
deadline = time.time() + timeout
|
||||
while time.time() < deadline:
|
||||
m = json.loads(self.ws.recv_text())
|
||||
if m.get("id") == rid:
|
||||
if "error" in m and m["error"]:
|
||||
raise RuntimeError(f"{method}: {m['error']}")
|
||||
return m.get("params")
|
||||
if m.get("method"):
|
||||
self.notifications.append(m)
|
||||
raise TimeoutError(method)
|
||||
|
||||
def next_notification(self):
|
||||
if self.notifications:
|
||||
return self.notifications.pop(0)
|
||||
while True:
|
||||
m = json.loads(self.ws.recv_text())
|
||||
if m.get("method"):
|
||||
return m
|
||||
|
||||
|
||||
def sample_line(rpc):
|
||||
dag = rpc.call("getBlockDagInfo")
|
||||
peers = rpc.call("getConnectedPeerInfo").get("peerInfo", [])
|
||||
try:
|
||||
blue = rpc.call("getSinkBlueScore").get("blueScore")
|
||||
except Exception:
|
||||
blue = ""
|
||||
return "\t".join(str(x) for x in [int(time.time() * 1000), dag.get("sink", ""), blue, dag.get("virtualDaaScore", ""),
|
||||
len(dag.get("tipHashes", []) or []), len(peers), dag.get("difficulty", ""),
|
||||
dag.get("headerCount", ""), dag.get("blockCount", "")])
|
||||
|
||||
|
||||
def watch_blocks(directory):
|
||||
os.makedirs(directory, exist_ok=True)
|
||||
while True:
|
||||
try:
|
||||
rpc = Rpc()
|
||||
rpc.call("subscribe", {"BlockAdded": {}})
|
||||
rpc.call("subscribe", {"VirtualChainChanged": {"include_accepted_transaction_ids": False}})
|
||||
sys.stdout.write(f"{time.strftime('%H:%M:%S')} subscribed on {RPC}\n"); sys.stdout.flush()
|
||||
last_sample = 0.0
|
||||
with open(os.path.join(directory, "blocks.tsv"), "a") as fb, open(os.path.join(directory, "chain.tsv"), "a") as fc, \
|
||||
open(os.path.join(directory, "samples.tsv"), "a") as fs:
|
||||
while True:
|
||||
now = time.time()
|
||||
if now - last_sample >= 5.0:
|
||||
try:
|
||||
fs.write(sample_line(rpc) + "\n"); fs.flush()
|
||||
except Exception as e:
|
||||
sys.stdout.write(f"sample failed: {e}\n"); sys.stdout.flush()
|
||||
last_sample = now
|
||||
rpc.ws.sock.settimeout(5.0)
|
||||
try:
|
||||
m = rpc.next_notification()
|
||||
except socket.timeout:
|
||||
continue
|
||||
recv_ms = int(time.time() * 1000)
|
||||
p = m.get("params") or {}
|
||||
inner = p.get("BlockAdded") or p.get("VirtualChainChanged") or p
|
||||
if m["method"] == "blockAddedNotification" and inner.get("block"):
|
||||
h = inner["block"].get("header", {})
|
||||
levels = h.get("parentsByLevel") or [] # array of arrays of hashes; level 0 = direct parents
|
||||
direct = levels[0] if levels and isinstance(levels[0], list) else []
|
||||
fb.write("\t".join(str(x) for x in [recv_ms, h.get("hash", ""), h.get("timestamp", ""), h.get("blueScore", ""),
|
||||
h.get("daaScore", ""), len(direct)]) + "\n")
|
||||
fb.flush()
|
||||
elif m["method"] == "virtualChainChangedNotification":
|
||||
added = inner.get("addedChainBlockHashes", []) or []
|
||||
removed = inner.get("removedChainBlockHashes", []) or []
|
||||
fc.write("\t".join([str(recv_ms), str(len(added)), str(len(removed)), ",".join(removed[:3])]) + "\n")
|
||||
fc.flush()
|
||||
except Exception as e:
|
||||
sys.stdout.write(f"{time.strftime('%H:%M:%S')} rpc lost ({e}); retrying in 5 s\n"); sys.stdout.flush()
|
||||
time.sleep(5)
|
||||
|
||||
|
||||
def main(argv):
|
||||
if len(argv) >= 2 and argv[1] == "call":
|
||||
params = json.loads(argv[3]) if len(argv) > 3 else {}
|
||||
print(json.dumps(Rpc().call(argv[2], params), indent=None))
|
||||
elif len(argv) >= 3 and argv[1] == "watch-blocks":
|
||||
watch_blocks(argv[2])
|
||||
elif len(argv) >= 2 and argv[1] == "sample":
|
||||
print(sample_line(Rpc()))
|
||||
else:
|
||||
sys.stderr.write(__doc__); sys.exit(2)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main(sys.argv)
|
||||
78
infra/cloud-devnet/provision.sh
Executable file
78
infra/cloud-devnet/provision.sh
Executable file
|
|
@ -0,0 +1,78 @@
|
|||
#!/usr/bin/env bash
|
||||
# Build the node once and install it on every node.
|
||||
#
|
||||
# ./provision.sh source tarball -> builder VM -> build/bin/{igneumd,igneum-miner} -> every node (systemd units)
|
||||
# ./provision.sh build only the build (creates the builder VM if build/builder.ip is missing, deletes it afterwards)
|
||||
# ./provision.sh install only the install on the nodes (needs build/bin/igneumd)
|
||||
#
|
||||
# Why a builder VM and not a cross-compile on the Mac: rustup here has only aarch64-apple-darwin and
|
||||
# x86_64-pc-windows-gnu installed and there is no x86_64 Linux linker (no musl-cross, zig, cross or docker; checked
|
||||
# 3 Oct 2026). A Linux cross-compile would need `rustup target add x86_64-unknown-linux-gnu` plus a linker
|
||||
# (`brew install filosottile/musl-cross/musl-cross` and a musl target), and the rocksdb bindgen trick of the Windows
|
||||
# cross-build again. The builder costs about EUR 0.05 for the hour (cpx41, approximate) and the Mac stays free.
|
||||
# The binaries are cached in build/bin; REBUILD=1 forces a new build.
|
||||
|
||||
. "$(dirname "$0")/lib/common.sh"
|
||||
|
||||
what="${1:-all}"
|
||||
mkdir -p "$BUILD_DIR" "$BIN_DIR"
|
||||
|
||||
# ---- build ---------------------------------------------------------------------------------------------------------
|
||||
do_build() {
|
||||
if [ -x "$BIN_DIR/igneumd" ] && [ "${REBUILD:-0}" != 1 ]; then
|
||||
log "build/bin/igneumd exists ($(cat "$BIN_DIR/version.txt" 2>/dev/null | tr '\n' ' ')); REBUILD=1 to rebuild"
|
||||
return
|
||||
fi
|
||||
"$HERE/make-source.sh"
|
||||
if [ ! -s "$BUILD_DIR/builder.ip" ]; then
|
||||
YES="${YES:-0}" "$HERE/create.sh" builder
|
||||
fi
|
||||
bip=$(cat "$BUILD_DIR/builder.ip")
|
||||
log "builder at $bip; waiting for ssh"
|
||||
for try in $(seq 1 12); do nssh "$bip" true >/dev/null 2>&1 && break; sleep 10; done
|
||||
nssh "$bip" true || die "builder $bip not reachable over ssh"
|
||||
log "uploading source ($(du -h "$BUILD_DIR/src.tar.gz" | cut -f1)) and the build script"
|
||||
nscp "$BUILD_DIR/src.tar.gz" "$HERE/builder/build-on-builder.sh" "$SSH_USER@$bip:/root/"
|
||||
nssh "$bip" "bash /root/build-on-builder.sh" | tee "$BUILD_DIR/build.log"
|
||||
nscp "$SSH_USER@$bip:/root/out/igneumd" "$SSH_USER@$bip:/root/out/igneum-miner" "$SSH_USER@$bip:/root/out/version.txt" "$BIN_DIR/"
|
||||
chmod +x "$BIN_DIR/igneumd" "$BIN_DIR/igneum-miner"
|
||||
cp "$BUILD_DIR/src.stamp" "$BIN_DIR/src.stamp"
|
||||
log "binaries in $BIN_DIR: $(cat "$BIN_DIR/version.txt" | tr '\n' ' ')"
|
||||
if [ "${KEEP_BUILDER:-0}" = 1 ]; then
|
||||
log "KEEP_BUILDER=1: builder $bip stays up (it bills by the hour)"
|
||||
else
|
||||
log "deleting the builder VM"
|
||||
if [ "$PROVIDER" = digitalocean ]; then doctl compute droplet delete -f "$PREFIX-builder"; else hcloud server delete "$PREFIX-builder" >/dev/null; fi
|
||||
rm -f "$BUILD_DIR/builder.ip"
|
||||
fi
|
||||
}
|
||||
|
||||
# ---- install ------------------------------------------------------------------------------------------------------
|
||||
do_install() {
|
||||
require_nodes
|
||||
[ -x "$BIN_DIR/igneumd" ] || die "no $BIN_DIR/igneumd: run ./provision.sh build"
|
||||
local n; n=$(node_count); local bits; bits=$(genesis_bits_decimal)
|
||||
log "installing on $n nodes (devnet suffix $DEVNET_SUFFIX, genesis bits $GENESIS_BITS = $bits, $MESH_OUT outbound peers each, $MINER_THREADS miner thread)"
|
||||
while IFS=$'\t' read -r name idx reg ip; do
|
||||
peers=$(peers_of "$idx")
|
||||
(
|
||||
for try in 1 2 3; do nssh "$ip" "mkdir -p /opt/igneum/bin /etc/igneum" && break; sleep 10; done
|
||||
nscp "$BIN_DIR/igneumd" "$BIN_DIR/igneum-miner" "$HERE/node/wrpc.py" "$HERE/node/run-igneumd.sh" "$HERE/node/blocklog.sh" "$SSH_USER@$ip:/opt/igneum/bin/"
|
||||
nscp "$HERE/node/install-node.sh" "$HERE/node/igneumd.service" "$HERE/node/igneum-miner.service" "$HERE/node/igneum-blocklog.service" "$SSH_USER@$ip:/root/"
|
||||
nssh "$ip" "bash /root/install-node.sh '$name' '$idx' '$DEVNET_SUFFIX' '$bits' '$peers' '$ip' '$MINER_THREADS' '$MINER_STATUS_SECS'" 2>&1 | sed "s/^/[$name] /"
|
||||
) &
|
||||
# at most 10 installs at once
|
||||
while [ "$(jobs -r | wc -l)" -ge 10 ]; do sleep 1; done
|
||||
done < "$NODES_FILE"
|
||||
wait
|
||||
log "installed. Peer map:"
|
||||
while IFS=$'\t' read -r name idx reg ip; do printf ' %s (%s, %s) -> %s\n' "$name" "$reg" "$ip" "$(peers_of "$idx")"; done < "$NODES_FILE"
|
||||
log "next: ./start.sh"
|
||||
}
|
||||
|
||||
case "$what" in
|
||||
all) do_build; do_install ;;
|
||||
build) do_build ;;
|
||||
install) do_install ;;
|
||||
*) die "usage: provision.sh [all|build|install]" ;;
|
||||
esac
|
||||
0
infra/cloud-devnet/results/.gitkeep
Normal file
0
infra/cloud-devnet/results/.gitkeep
Normal file
30
infra/cloud-devnet/start.sh
Executable file
30
infra/cloud-devnet/start.sh
Executable file
|
|
@ -0,0 +1,30 @@
|
|||
#!/usr/bin/env bash
|
||||
# Start the network: nodes and block logs first, miners once every RPC answers.
|
||||
# ./start.sh nodes, then miners
|
||||
# ./start.sh nodes nodes and block logs only
|
||||
# ./start.sh miners miners only
|
||||
. "$(dirname "$0")/lib/common.sh"
|
||||
require_nodes
|
||||
what="${1:-all}"
|
||||
|
||||
if [ "$what" = all ] || [ "$what" = nodes ]; then
|
||||
log "starting igneumd and igneum-blocklog on $(node_count) nodes"
|
||||
on_all 'systemctl restart igneumd && sleep 1 && systemctl restart igneum-blocklog && systemctl is-active igneumd'
|
||||
log "waiting for the RPC on every node"
|
||||
for try in $(seq 1 24); do
|
||||
bad=0
|
||||
while IFS=$'\t' read -r name idx reg ip; do
|
||||
nssh "$ip" "python3 /opt/igneum/bin/wrpc.py call getInfo >/dev/null 2>&1" || bad=$((bad + 1))
|
||||
done < "$NODES_FILE"
|
||||
[ "$bad" = 0 ] && break
|
||||
log "$bad nodes not answering yet (try $try), 5 s"; sleep 5
|
||||
done
|
||||
[ "$bad" = 0 ] || log "WARNING: $bad nodes still not answering; check: ./status.sh and journalctl -u igneumd on the node"
|
||||
fi
|
||||
|
||||
if [ "$what" = all ] || [ "$what" = miners ]; then
|
||||
log "starting the trickle miners ($MINER_THREADS thread each; the first block needs the 256 MiB cache built, about 10 to 30 s on one vCPU, approximate)"
|
||||
on_all 'systemctl restart igneum-miner && systemctl is-active igneum-miner'
|
||||
fi
|
||||
log "started at $(date -u +%Y-%m-%dT%H:%M:%SZ). Watch with ./status.sh (every 30 s: watch -n 30 ./status.sh)"
|
||||
printf '%s\tstart %s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$what" >> "$(results_dir_for)/events.log"
|
||||
15
infra/cloud-devnet/status.sh
Executable file
15
infra/cloud-devnet/status.sh
Executable file
|
|
@ -0,0 +1,15 @@
|
|||
#!/usr/bin/env bash
|
||||
# One line per node: blocks, headers, blue score, DAA, tips, peers, difficulty, sink, miner state. Reads the
|
||||
# loopback RPC through ssh (wrpc.py sample), so it costs one ssh round trip per node.
|
||||
. "$(dirname "$0")/lib/common.sh"
|
||||
require_nodes
|
||||
printf '%-10s %-5s %-15s %8s %8s %8s %8s %4s %5s %12s %-12s %s\n' node region ip blocks headers blue daa tips peers difficulty sink miner
|
||||
while IFS=$'\t' read -r name idx reg ip; do
|
||||
(
|
||||
line=$(nssh "$ip" "python3 /opt/igneum/bin/wrpc.py sample 2>/dev/null; systemctl is-active igneum-miner 2>/dev/null" 2>/dev/null | tr '\n' '\t')
|
||||
s=$(printf '%s' "$line" | cut -f1-9); miner=$(printf '%s' "$line" | cut -f10)
|
||||
IFS=$'\t' read -r ts sink blue daa tips peers diff headers blocks <<< "$s"
|
||||
printf '%-10s %-5s %-15s %8s %8s %8s %8s %4s %5s %12.0f %-12s %s\n' "$name" "$reg" "$ip" "${blocks:-?}" "${headers:-?}" "${blue:-?}" "${daa:-?}" "${tips:-?}" "${peers:-?}" "${diff:-0}" "${sink:0:12}" "${miner:-?}"
|
||||
) &
|
||||
done < "$NODES_FILE"
|
||||
wait
|
||||
14
infra/cloud-devnet/stop.sh
Executable file
14
infra/cloud-devnet/stop.sh
Executable file
|
|
@ -0,0 +1,14 @@
|
|||
#!/usr/bin/env bash
|
||||
# Stop the network. The VMs keep running (and billing); destroy.sh removes them.
|
||||
# ./stop.sh miners, block logs, nodes
|
||||
# ./stop.sh miners miners only (the chain keeps running without new blocks)
|
||||
. "$(dirname "$0")/lib/common.sh"
|
||||
require_nodes
|
||||
what="${1:-all}"
|
||||
if [ "$what" = miners ]; then
|
||||
on_all 'systemctl stop igneum-miner; systemctl is-active igneum-miner || true'
|
||||
else
|
||||
on_all 'systemctl stop igneum-miner igneum-blocklog igneumd; systemctl is-active igneumd || true'
|
||||
fi
|
||||
printf '%s\tstop %s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$what" >> "$(results_dir_for)/events.log"
|
||||
log "stopped ($what). The servers still bill by the hour: ./destroy.sh when the data is collected"
|
||||
2
infra/gpu-bench/.gitignore
vendored
Normal file
2
infra/gpu-bench/.gitignore
vendored
Normal file
|
|
@ -0,0 +1,2 @@
|
|||
# the bundle tarball; results/ is committed on purpose
|
||||
build/
|
||||
16
infra/gpu-bench/Dockerfile.cuda
Normal file
16
infra/gpu-bench/Dockerfile.cuda
Normal file
|
|
@ -0,0 +1,16 @@
|
|||
# Igneum GPU bench image, NVIDIA lane. CUDA 12.8 devel (nvcc, NVRTC, cuda.h) on Ubuntu 22.04, plus the OpenCL ICD
|
||||
# loader and headers for the optional OpenCL build-time measurement on NVIDIA, Python 3 for run.sh, curl for the
|
||||
# upload, and sshd so the bundle can be scp'd in. RunPod: create a template with this image (or use the tag below
|
||||
# directly; nvcc is in it) and tick "SSH". Blackwell (RTX 5090) needs 12.8 or newer; older toolkits do not know sm_120.
|
||||
# Build and push (optional; the public tag works as-is for RunPod with run.sh installing nothing):
|
||||
# docker build -f Dockerfile.cuda -t <you>/igneum-gpu-bench:cuda12.8 .
|
||||
FROM nvidia/cuda:12.8.1-devel-ubuntu22.04
|
||||
ENV DEBIAN_FRONTEND=noninteractive
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||
python3 curl ca-certificates openssh-server clinfo ocl-icd-opencl-dev opencl-headers perl build-essential \
|
||||
&& rm -rf /var/lib/apt/lists/* \
|
||||
&& mkdir -p /run/sshd /workspace
|
||||
# NVIDIA's OpenCL ICD is provided by the host driver through the container runtime; it appears when the pod sets
|
||||
# NVIDIA_DRIVER_CAPABILITIES=compute,utility (RunPod does). clinfo shows whether it is there; run.sh skips if not.
|
||||
WORKDIR /workspace
|
||||
CMD ["/usr/sbin/sshd", "-D"]
|
||||
15
infra/gpu-bench/Dockerfile.rocm
Normal file
15
infra/gpu-bench/Dockerfile.rocm
Normal file
|
|
@ -0,0 +1,15 @@
|
|||
# Igneum GPU bench image, AMD lane: ROCm 6 with the OpenCL runtime (rocm-opencl) and headers, so proto-opencl/host.c
|
||||
# builds with plain cc -lOpenCL and clinfo lists the gfx device. Needs /dev/kfd and /dev/dri passed in (a provider
|
||||
# that offers AMD cards does that). RunPod lists no AMD consumer cards and Vast.ai none in general (approximate,
|
||||
# 3 Oct 2026); this recipe is for whichever host has an RX 7900 XTX (gfx1100), or a datacenter MI300X as a stand-in.
|
||||
# The tag is approximate: use the newest rocm/dev-ubuntu-22.04:6.x-complete the host's kernel driver supports.
|
||||
# docker build -f Dockerfile.rocm -t <you>/igneum-gpu-bench:rocm6 .
|
||||
# docker run --device=/dev/kfd --device=/dev/dri --group-add video -it <you>/igneum-gpu-bench:rocm6
|
||||
FROM rocm/dev-ubuntu-22.04:6.3-complete
|
||||
ENV DEBIAN_FRONTEND=noninteractive
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||
python3 curl ca-certificates openssh-server clinfo ocl-icd-opencl-dev opencl-headers perl build-essential \
|
||||
&& rm -rf /var/lib/apt/lists/* \
|
||||
&& mkdir -p /run/sshd /workspace
|
||||
WORKDIR /workspace
|
||||
CMD ["/usr/sbin/sshd", "-D"]
|
||||
72
infra/gpu-bench/README.md
Normal file
72
infra/gpu-bench/README.md
Normal file
|
|
@ -0,0 +1,72 @@
|
|||
# Igneum GPU bench on rented cards
|
||||
|
||||
One pod per card on RunPod (primary; Vast.ai variant below), each running the same bundle: the CUDA and OpenCL
|
||||
test harnesses from `proto-cuda/` and `proto-opencl/` with the three checked-in packs, built on the pod with
|
||||
`--serve` support, then a fixed sequence of measurements and one results row. Nothing is cloned on the pod; the
|
||||
bundle is a tarball made here. Nothing here spends until a pod is started.
|
||||
|
||||
## The command sequence
|
||||
|
||||
```
|
||||
cd infra/gpu-bench
|
||||
./make-bundle.sh # build/igneum-gpu-bench.tar.gz (sources + scripts, about 1 MB)
|
||||
# RunPod console: Pods -> Deploy -> pick the card -> template "nvidia/cuda:12.8.1-devel-ubuntu22.04" (or the image from
|
||||
# Dockerfile.cuda), Community Cloud, tick SSH, 20 GB container disk -> Deploy. Note host and port from "Connect".
|
||||
scp -P <port> build/igneum-gpu-bench.tar.gz root@<host>:/workspace/
|
||||
ssh -p <port> root@<host> 'cd /workspace && tar xzf igneum-gpu-bench.tar.gz && cd bundle && LABEL=rtx4090 ./run.sh'
|
||||
# about 45 minutes: build 2 to 5 min, gate, 10 min raw, sweep, inline, recompile timings; the row prints as RESULT
|
||||
# and lands in bundle/results.md and results-<stamp>/row.md; the log is uploaded under label gpubench-rtx4090
|
||||
scp -P <port> -r root@<host>:/workspace/bundle/results-* results/ # keep the raw outputs here
|
||||
# RunPod console: Stop and Terminate the pod (billing per minute while it runs; storage bills while stopped)
|
||||
node ../../tools/logs.mjs # the uploaded log, if the pod could not be reached again
|
||||
```
|
||||
|
||||
Repeat for each card (the same bundle). Paste the rows into `docs/bench-log.md` with `results-template.md`.
|
||||
`CUDA_ARCH=sm_86` overrides `-arch=native` if the driver in a pod is older than its toolkit.
|
||||
|
||||
## Cards and cost (RunPod pricing page, read 3 Oct 2026, USD per hour)
|
||||
|
||||
| Card | Community Cloud | Secure Cloud | One 45-minute run, community | Note |
|
||||
|---|---|---|---|---|
|
||||
| RTX 3060 | not on RunPod's table | not on RunPod's table | about 0.05 to 0.10 on Vast.ai (approximate) | the entry card; Vast.ai is the host for it |
|
||||
| RTX 3090 | 0.22 | 0.50 | about 0.17 | Ampere sm_86, 24 GB, 6 MB L2 (approximate) |
|
||||
| RTX 4090 | 0.34 | 0.74 | about 0.26 | Ada sm_89, 24 GB, 72 MB L2 (approximate) |
|
||||
| RTX 5090 | 0.69 | 0.99 | about 0.52 | Blackwell sm_120, 32 GB, 96 MB L2 (bench-log); the project lead's own card is the reference row |
|
||||
| RX 7900 XTX | not offered | not offered | n/a | RunPod has no AMD consumer cards; Vast.ai lists none in general (approximate). `Dockerfile.rocm` and the OpenCL lane of `run.sh` are ready for any host that has one (or the project lead's own AMD box) |
|
||||
|
||||
Four NVIDIA cards: about USD 2 on community cloud, USD 3.50 on secure cloud, plus cents of storage. Per-minute
|
||||
billing; a pod left running costs the hourly rate.
|
||||
|
||||
Vast.ai variant: search for the card with "CUDA 12.8" and "verified" hosts, rent with the image
|
||||
`nvidia/cuda:12.8.1-devel-ubuntu22.04` and "SSH" as the launch mode, then the same scp and ssh lines (Vast prints its
|
||||
own port and host). Prices on Vast.ai are per host and change by the hour; the 3060 is usually under USD 0.10.
|
||||
|
||||
## What run.sh measures and why
|
||||
|
||||
| Step | Output | Ledger |
|
||||
|---|---|---|
|
||||
| Vectors gate (`--batches 3`): cache check, dataset self-test, 3 warps standalone and in batch | `gate.txt`, must say `OVERALL: PASS` or the run stops | the cross-vendor proof per card (bench-log, "What PASS means") |
|
||||
| Raw bench, about 10 minutes of 2^24-hash batches at 1 GiB | `raw.txt`: Mhash/s, seconds | the per-card hash rate; sustained, so thermals and neighbours show |
|
||||
| Sweep 64, 128, 256, 512, 1024 MiB, 20 batches each | `sweep-*.txt`, the 64 MiB over 1 GiB ratio | M1: the L2 cliff per card (the 5090 gave 5.9x) |
|
||||
| Inline-dataset shortcut (`make-inline.sh`: every `ds[...]` load becomes `mh_word(ds, ...)` from the cache) | `inline.txt`: Mhash/s and the ratio to the raw rate | M16: the recompute attacker's rate; the M5 Max gave 0.21. The 64 MiB-cache variant inside the 5090's L2 needs a pack exported with a smaller cache (proto-metal has no flag for it yet) and is a follow-up |
|
||||
| Second pack `igneum-hourly` (closed form, 128 loads) | `pack2.txt` | comparison with the 5090 and Metal tables in the bench log |
|
||||
| `nvcc -cubin` of kernel.cu x3 (the worker's `prepare` path, host.cu line 554), NVRTC in process x3 (`nvrtc-time.cu`), `clBuildProgram` of kernel.cl x3 (`clbuild-time.c`, when an OpenCL ICD is visible) | `nvrtc.txt`, `clbuild.txt`, medians in the row | M11, M17: the hourly program change on real drivers |
|
||||
| `--serve` ready line (`printf quit \| worker --serve`) | `serve-ready.txt` | proves the pod's binary is the worker the miner drives (ready line shows `prepare 1` when nvcc is on PATH) |
|
||||
|
||||
The inline binary fails host.cu's dataset self-test by construction (the dataset buffer holds the cache) and must
|
||||
pass the vectors; `run.sh` reads its rate and vector lines and ignores its OVERALL.
|
||||
|
||||
Results row (also in `results-template.md`): card, driver and toolkit, date, Mhash/s at 1 GiB over the 10 minutes,
|
||||
seconds, second-pack Mhash/s, the five sweep rates, the cliff ratio, inline Mhash/s, inline over honest, nvcc ms,
|
||||
NVRTC ms, OpenCL ms, vectors, build ms.
|
||||
|
||||
## Files
|
||||
|
||||
`run.sh` (on the pod), `make-bundle.sh` (on the Mac), `make-inline.sh`, `nvrtc-time.cu`, `clbuild-time.c`,
|
||||
`upload.sh` (the intake URL and key of `proto-cuda/windows-miner/upload-log.bat` as a curl line), `Dockerfile.cuda`,
|
||||
`Dockerfile.rocm`, `results-template.md`, `results/` (raw outputs copied back, one directory per run).
|
||||
|
||||
Not tested here: there is no NVIDIA or AMD card on this Mac, so `run.sh`, `nvrtc-time.cu` and `clbuild-time.c` were
|
||||
checked by read-through and `bash -n` only; the harness binaries they drive ran on the 5090 and the gfx1036 on 3 Oct
|
||||
2026 (bench-log). The first pod run is the real test; if `nvrtc-time` fails to compile the kernel (a header NVRTC
|
||||
rejects), the nvcc column still carries the out-of-process figure the worker uses today.
|
||||
60
infra/gpu-bench/clbuild-time.c
Normal file
60
infra/gpu-bench/clbuild-time.c
Normal file
|
|
@ -0,0 +1,60 @@
|
|||
/* Times clBuildProgram of a pack's kernel.cl N times on the first GPU device: the OpenCL half of the hourly-recompile
|
||||
* figure (ledger M11). Build options mirror proto-opencl/host.c for the local-memory exchange path
|
||||
* (-cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0).
|
||||
* cc -std=c99 -O2 -o clbuild-time clbuild-time.c -lOpenCL
|
||||
* ./clbuild-time <kernel.cl> [iterations=3] [extra build options]
|
||||
* Prints one line per iteration and "median <ms>". */
|
||||
#define CL_TARGET_OPENCL_VERSION 120
|
||||
#include <CL/cl.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <time.h>
|
||||
|
||||
static double ms(void) { struct timespec t; clock_gettime(CLOCK_MONOTONIC, &t); return t.tv_sec * 1000.0 + t.tv_nsec / 1e6; }
|
||||
static int cmp(const void* a, const void* b) { double x = *(const double*)a, y = *(const double*)b; return (x > y) - (x < y); }
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
if (argc < 2) { fprintf(stderr, "usage: clbuild-time <kernel.cl> [iterations] [extra options]\n"); return 2; }
|
||||
int iters = argc > 2 ? atoi(argv[2]) : 3;
|
||||
FILE* f = fopen(argv[1], "rb"); if (!f) { perror(argv[1]); return 2; }
|
||||
fseek(f, 0, SEEK_END); long n = ftell(f); fseek(f, 0, SEEK_SET);
|
||||
char* src = malloc(n + 1); if (fread(src, 1, n, f) != (size_t)n) { fprintf(stderr, "short read\n"); return 2; } src[n] = 0; fclose(f);
|
||||
|
||||
cl_uint np = 0; clGetPlatformIDs(0, NULL, &np); if (!np) { fprintf(stderr, "no OpenCL platform\n"); return 2; }
|
||||
cl_platform_id* ps = malloc(np * sizeof *ps); clGetPlatformIDs(np, ps, NULL);
|
||||
cl_device_id dev = NULL; char name[256] = "?", drv[128] = "?";
|
||||
for (cl_uint i = 0; i < np && !dev; ++i) {
|
||||
cl_uint nd = 0; if (clGetDeviceIDs(ps[i], CL_DEVICE_TYPE_GPU, 0, NULL, &nd) != CL_SUCCESS || !nd) continue;
|
||||
clGetDeviceIDs(ps[i], CL_DEVICE_TYPE_GPU, 1, &dev, NULL);
|
||||
}
|
||||
if (!dev) { fprintf(stderr, "no OpenCL GPU device\n"); return 2; }
|
||||
clGetDeviceInfo(dev, CL_DEVICE_NAME, sizeof name, name, NULL); clGetDeviceInfo(dev, CL_DRIVER_VERSION, sizeof drv, drv, NULL);
|
||||
cl_int err; cl_context ctx = clCreateContext(NULL, 1, &dev, NULL, NULL, &err); if (err != CL_SUCCESS) { fprintf(stderr, "clCreateContext %d\n", err); return 2; }
|
||||
char opts[1024]; snprintf(opts, sizeof opts, "-cl-std=CL1.2 -D IGNEUM_GROUP=32 -D IGNEUM_EXCHANGE=0 %s", argc > 3 ? argv[3] : "");
|
||||
printf("device %s, driver %s, options \"%s\", source %ld bytes\n", name, drv, opts, n);
|
||||
|
||||
double* times = malloc(iters * sizeof *times);
|
||||
for (int i = 0; i < iters; ++i) {
|
||||
double t0 = ms();
|
||||
const char* s = src; size_t len = (size_t)n;
|
||||
cl_program prog = clCreateProgramWithSource(ctx, 1, &s, &len, &err);
|
||||
if (err != CL_SUCCESS) { fprintf(stderr, "clCreateProgramWithSource %d\n", err); return 1; }
|
||||
err = clBuildProgram(prog, 1, &dev, opts, NULL, NULL);
|
||||
double t1 = ms();
|
||||
if (err != CL_SUCCESS) {
|
||||
size_t ls = 0; clGetProgramBuildInfo(prog, dev, CL_PROGRAM_BUILD_LOG, 0, NULL, &ls);
|
||||
char* log = malloc(ls + 1); clGetProgramBuildInfo(prog, dev, CL_PROGRAM_BUILD_LOG, ls, log, NULL); log[ls] = 0;
|
||||
fprintf(stderr, "build failed (%d):\n%s\n", err, log); return 1;
|
||||
}
|
||||
cl_kernel k = clCreateKernel(prog, "igneum_hash", &err);
|
||||
if (err != CL_SUCCESS) fprintf(stderr, "igneum_hash not found (%d)\n", err); else clReleaseKernel(k);
|
||||
clReleaseProgram(prog);
|
||||
times[i] = t1 - t0;
|
||||
printf("iteration %d: build %.1f ms\n", i + 1, times[i]);
|
||||
}
|
||||
qsort(times, iters, sizeof *times, cmp);
|
||||
printf("median %.1f ms (clBuildProgram, %d iterations)\n", times[iters / 2], iters);
|
||||
clReleaseContext(ctx);
|
||||
return 0;
|
||||
}
|
||||
23
infra/gpu-bench/make-bundle.sh
Executable file
23
infra/gpu-bench/make-bundle.sh
Executable file
|
|
@ -0,0 +1,23 @@
|
|||
#!/usr/bin/env bash
|
||||
# On the Mac: pack the sources the pod needs into build/igneum-gpu-bench.tar.gz. Nothing is cloned on the pod.
|
||||
# Contents: proto-cuda/host.cu, the three checked-in packs, proto-opencl/host.c, and this directory's scripts.
|
||||
# DEVNET_PACK=<dir> adds a pack exported from the live node (igneum-miner export-pack), which carries kernel_bound.cu
|
||||
# for a --serve job test; optional.
|
||||
set -euo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"; REPO="$(cd "$HERE/../.." && pwd)"
|
||||
mkdir -p "$HERE/build"
|
||||
stage="$HERE/build/bundle"; rm -rf "$stage"; mkdir -p "$stage/proto-cuda/packs" "$stage/proto-opencl"
|
||||
cp "$REPO/proto-cuda/host.cu" "$stage/proto-cuda/"
|
||||
for p in igneum-genesis-mh igneum-hourly igneum-genesis; do cp -R "$REPO/proto-cuda/packs/$p" "$stage/proto-cuda/packs/"; done
|
||||
[ -n "${DEVNET_PACK:-}" ] && cp -R "$DEVNET_PACK" "$stage/proto-cuda/packs/igneum-devnet"
|
||||
cp "$REPO/proto-opencl/host.c" "$stage/proto-opencl/"
|
||||
cp -R "$REPO/proto-opencl/emu" "$stage/proto-opencl/" 2>/dev/null || true # emu_opencl.h is included by kernel.cl when no OpenCL compiler is present
|
||||
cp "$HERE/run.sh" "$HERE/make-inline.sh" "$HERE/nvrtc-time.cu" "$HERE/clbuild-time.c" "$HERE/upload.sh" "$HERE/results-template.md" "$stage/"
|
||||
chmod +x "$stage"/*.sh
|
||||
git -C "$REPO" rev-parse --short HEAD > "$stage/SOURCE_COMMIT" 2>/dev/null || true
|
||||
tar -C "$HERE/build" -czf "$HERE/build/igneum-gpu-bench.tar.gz" bundle
|
||||
rm -rf "$stage"
|
||||
echo "wrote $HERE/build/igneum-gpu-bench.tar.gz ($(du -h "$HERE/build/igneum-gpu-bench.tar.gz" | cut -f1))"
|
||||
echo "on the pod (RunPod SSH, port and host from the pod's Connect panel):"
|
||||
echo " scp -P <port> $HERE/build/igneum-gpu-bench.tar.gz root@<host>:/workspace/"
|
||||
echo " ssh -p <port> root@<host> 'cd /workspace && tar xzf igneum-gpu-bench.tar.gz && cd bundle && LABEL=rtx4090 ./run.sh'"
|
||||
22
infra/gpu-bench/make-inline.sh
Executable file
22
infra/gpu-bench/make-inline.sh
Executable file
|
|
@ -0,0 +1,22 @@
|
|||
#!/usr/bin/env bash
|
||||
# Build the inline-dataset shortcut variant of a CUDA pack: the measurement of proto-metal/MEMHARD.md section 2.2
|
||||
# (ledger M16) on NVIDIA. usage: make-inline.sh <pack-dir> <out-dir>
|
||||
# Two text edits on kernel.cu, nothing else:
|
||||
# 1. every dataset load `ds[expr]` in igneum_hash becomes `mh_word(ds, expr)`: the word is recomputed from the
|
||||
# cache through memhard.h's mh_item (8 dependent cache reads and 9 mixer applications per word) instead of read
|
||||
# 2. igneum_build copies the 256 MiB cache into the head of the dataset buffer instead of building items, so the
|
||||
# `ds` pointer the hash kernel receives points at the cache (mh_word indexes words below 2^IGNEUM_CACHE_LOG2_WORDS)
|
||||
# The dataset self-test of host.cu then FAILS by construction (it reads ds words expecting items) and the vectors
|
||||
# PASS (mh_word gives the true words). The rate line is the number; run.sh checks the vector lines and ignores OVERALL.
|
||||
set -euo pipefail
|
||||
src="$1"; out="$2"
|
||||
mkdir -p "$out"
|
||||
cp "$src"/*.h "$out/"
|
||||
perl -0pe '
|
||||
s/ds\[([^\]]*)\]/mh_word(ds, $1)/g;
|
||||
s/__global__ void igneum_build\(uint32_t\* ds, const uint32_t\* cache, uint32_t nItems\) \{.*?\n\}\n/__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {\n \/\/ INLINE VARIANT (infra\/gpu-bench\/make-inline.sh): copy the cache into the dataset buffer; igneum_hash recomputes words from it\n uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;\n if (t < nItems && t < (1u << (IGNEUM_CACHE_LOG2_WORDS - 4u))) { for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)t * 16u + i] = cache[(size_t)t * 16u + i]; }\n}\n/s;
|
||||
' "$src/kernel.cu" > "$out/kernel.cu"
|
||||
n=$(grep -c 'mh_word(ds,' "$out/kernel.cu" || true)
|
||||
grep -q 'INLINE VARIANT' "$out/kernel.cu" || { echo "make-inline: igneum_build not rewritten (pack layout changed?)"; exit 1; }
|
||||
[ "$n" -gt 0 ] || { echo "make-inline: no dataset loads rewritten"; exit 1; }
|
||||
echo "make-inline: $n loads rewritten to mh_word, build kernel replaced -> $out/kernel.cu"
|
||||
82
infra/gpu-bench/nvrtc-time.cu
Normal file
82
infra/gpu-bench/nvrtc-time.cu
Normal file
|
|
@ -0,0 +1,82 @@
|
|||
// Times an in-process NVRTC compile of a pack's kernel (the device part of kernel.cu with program.h and memhard.h
|
||||
// as in-memory headers), N times, then loads the cubin through the driver API to prove it is usable.
|
||||
// This is the "hourly JIT in the miner process" figure of ledger M11 and M17; the worker today runs nvcc out of
|
||||
// process (proto-cuda/host.cu, prepareCompile), which run.sh times separately.
|
||||
// nvcc -O2 -std=c++17 -o nvrtc-time nvrtc-time.cu -lnvrtc -lcuda
|
||||
// ./nvrtc-time <pack-dir> [iterations=3]
|
||||
// Prints one line per iteration (compile ms, cubin bytes, load ms) and "median <ms>".
|
||||
#include <nvrtc.h>
|
||||
#include <cuda.h>
|
||||
#include <algorithm>
|
||||
#include <chrono>
|
||||
#include <cstdio>
|
||||
#include <cstdlib>
|
||||
#include <fstream>
|
||||
#include <sstream>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
static std::string readFile(const std::string& p) {
|
||||
std::ifstream f(p); if (!f) { std::fprintf(stderr, "cannot read %s\n", p.c_str()); std::exit(2); }
|
||||
std::stringstream ss; ss << f.rdbuf(); return ss.str();
|
||||
}
|
||||
// Drop the includes NVRTC cannot see and the host launch wrappers; the typedefs replace <cstdint>.
|
||||
static std::string deviceOnly(const std::string& src, bool cutHost) {
|
||||
std::istringstream in(src); std::string line, out = "typedef unsigned int uint32_t;\ntypedef unsigned long long uint64_t;\ntypedef unsigned char uint8_t;\ntypedef unsigned long size_t;\n";
|
||||
while (std::getline(in, line)) {
|
||||
if (cutHost && line.rfind("cudaError_t igneum_launch", 0) == 0) break;
|
||||
if (line.find("#include <cuda_runtime.h>") != std::string::npos || line.find("#include <cstdint>") != std::string::npos ||
|
||||
line.find("#include <stdint.h>") != std::string::npos || line.find("#include <stddef.h>") != std::string::npos) { out += "\n"; continue; }
|
||||
out += line + "\n";
|
||||
}
|
||||
return out;
|
||||
}
|
||||
static double ms() { return std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now().time_since_epoch()).count(); }
|
||||
|
||||
int main(int argc, char** argv) {
|
||||
if (argc < 2) { std::fprintf(stderr, "usage: nvrtc-time <pack-dir> [iterations]\n"); return 2; }
|
||||
std::string pack = argv[1]; int iters = argc > 2 ? std::atoi(argv[2]) : 3;
|
||||
std::string kernel = deviceOnly(readFile(pack + "/kernel.cu"), true);
|
||||
std::string program = deviceOnly(readFile(pack + "/program.h"), false);
|
||||
std::string memhard = deviceOnly(readFile(pack + "/memhard.h"), false);
|
||||
const char* headers[2] = { program.c_str(), memhard.c_str() };
|
||||
const char* names[2] = { "program.h", "memhard.h" };
|
||||
|
||||
if (cuInit(0) != CUDA_SUCCESS) { std::fprintf(stderr, "cuInit failed\n"); return 2; }
|
||||
CUdevice dev; cuDeviceGet(&dev, 0); int major = 0, minor = 0;
|
||||
cuDeviceGetAttribute(&major, CU_DEVICE_ATTRIBUTE_COMPUTE_CAPABILITY_MAJOR, dev);
|
||||
cuDeviceGetAttribute(&minor, CU_DEVICE_ATTRIBUTE_COMPUTE_CAPABILITY_MINOR, dev);
|
||||
CUcontext ctx; cuCtxCreate(&ctx, 0, dev);
|
||||
char arch[64]; std::snprintf(arch, sizeof arch, "--gpu-architecture=sm_%d%d", major, minor);
|
||||
const char* opts[] = { arch, "-std=c++17", "-O3", "-default-device" };
|
||||
int nvrtcMajor = 0, nvrtcMinor = 0; nvrtcVersion(&nvrtcMajor, &nvrtcMinor);
|
||||
std::printf("nvrtc %d.%d, device sm_%d%d, kernel source %zu bytes\n", nvrtcMajor, nvrtcMinor, major, minor, kernel.size());
|
||||
|
||||
std::vector<double> times;
|
||||
for (int i = 0; i < iters; ++i) {
|
||||
double t0 = ms();
|
||||
nvrtcProgram prog;
|
||||
if (nvrtcCreateProgram(&prog, kernel.c_str(), "kernel.cu", 2, headers, names) != NVRTC_SUCCESS) { std::fprintf(stderr, "nvrtcCreateProgram failed\n"); return 1; }
|
||||
nvrtcResult r = nvrtcCompileProgram(prog, 4, opts);
|
||||
double t1 = ms();
|
||||
size_t logSize = 0; nvrtcGetProgramLogSize(prog, &logSize);
|
||||
if (r != NVRTC_SUCCESS) {
|
||||
std::string log(logSize, '\0'); nvrtcGetProgramLog(prog, &log[0]);
|
||||
std::fprintf(stderr, "nvrtc compile failed:\n%s\n", log.c_str()); return 1;
|
||||
}
|
||||
size_t cubinSize = 0; nvrtcGetCUBINSize(prog, &cubinSize);
|
||||
std::vector<char> cubin(cubinSize); nvrtcGetCUBIN(prog, cubin.data());
|
||||
nvrtcDestroyProgram(&prog);
|
||||
double t2 = ms();
|
||||
CUmodule mod; CUfunction fn;
|
||||
if (cuModuleLoadData(&mod, cubin.data()) != CUDA_SUCCESS || cuModuleGetFunction(&fn, mod, "_Z11igneum_hashPKjPyjj") != CUDA_SUCCESS) {
|
||||
std::fprintf(stderr, "cubin load or igneum_hash lookup failed (name mangling differs?)\n");
|
||||
} else { cuModuleUnload(mod); }
|
||||
double t3 = ms();
|
||||
std::printf("iteration %d: compile %.1f ms, cubin %zu bytes, load %.1f ms\n", i + 1, t1 - t0, cubinSize, t3 - t2);
|
||||
times.push_back(t1 - t0);
|
||||
}
|
||||
std::sort(times.begin(), times.end());
|
||||
std::printf("median %.1f ms (nvrtc compile of the device kernel, %d iterations)\n", times[times.size() / 2], iters);
|
||||
return 0;
|
||||
}
|
||||
26
infra/gpu-bench/results-template.md
Normal file
26
infra/gpu-bench/results-template.md
Normal file
|
|
@ -0,0 +1,26 @@
|
|||
## <day> October 2026, rented GPUs: RTX 3060, 3090, 4090, 5090 on RunPod (infra/gpu-bench/run.sh; AMD RX 7900 XTX pending a host)
|
||||
|
||||
Image: nvidia/cuda:12.8.1-devel-ubuntu22.04 (RunPod, SSH), bundle of proto-cuda/host.cu plus the three packs and
|
||||
proto-opencl/host.c at commit <hash>, nvcc -arch=native, 1 warp per block. Pack igneum-genesis-mh (memory-hard, 104
|
||||
loads per hash), 1 GiB dataset, batches of 2^24; the raw column is one run sized to about 10 minutes. Sweep sizes 64
|
||||
to 1024 MiB. Inline = the shortcut kernel of make-inline.sh (every dataset load recomputed from the 256 MiB cache
|
||||
through mh_word), vectors PASS, dataset self-test FAIL by construction. Recompile = nvcc -cubin of kernel.cu out of
|
||||
process (the worker's prepare path), NVRTC in process (nvrtc-time.cu), clBuildProgram of kernel.cl where NVIDIA's
|
||||
OpenCL ICD was present. Cost: RunPod community cloud, USD <n> for the four pods (pricing page: 3090 0.22/h, 4090
|
||||
0.34/h, 5090 0.69/h; 3060 <host and price>).
|
||||
|
||||
| card | driver / toolkit | date | Mhash/s at 1 GiB (10 min) | s | igneum-hourly Mhash/s | sweep MiB:Mhash/s 64 / 128 / 256 / 512 / 1024 | 64 MiB over 1 GiB | inline Mhash/s | inline / honest | nvcc cubin ms | NVRTC ms | OpenCL build ms | vectors | build ms |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| RTX 3060 | | | | | | | | | | | | | | |
|
||||
| RTX 3090 | | | | | | | | | | | | | | |
|
||||
| RTX 4090 | | | | | | | | | | | | | | |
|
||||
| RTX 5090 | | | | | | | | | | | | | | |
|
||||
| RTX 5090 (the project lead's PC, 3 Oct 2026, for reference) | 13.4 / 12.8 | 2026-10-03 | 229 (memory-hard) | | 185.3 (closed form) | 1340 / n/a / 270 / 242 / 229 (closed form) | 5.9 | n/a | n/a | n/a | n/a | n/a | 96/96 PASS | |
|
||||
| Apple M5 Max (Metal, reference) | | 2026-10-03 | 45.2 | | 36.6 | 569 (4 MiB) / 183 / 94 / 69 / 44 | | 9.49 | 0.21 | 20 to 52 (Metal compile) | | | PASS | |
|
||||
|
||||
Reading per column: the 1 GiB rate is the mining rate of the card on today's hash; the 64 MiB over 1 GiB ratio is
|
||||
the L2 cliff (ledger M1: the on-chip-cache advantage a chip would have to buy in DRAM); inline over honest is the
|
||||
recompute attacker's rate relative to an honest miner (ledger M16; the M5 Max gives 0.21; a value above 1 on any card
|
||||
means the shortcut beats the honest kernel there); the three compile columns are the hourly program change on real
|
||||
NVIDIA drivers (ledger M11, M17; the 90 s gap of the Windows launcher was a rebuild, not a compile). The rows are
|
||||
raw bench numbers from a rented host with whatever neighbours it had; rerun before quoting a figure outside this log.
|
||||
0
infra/gpu-bench/results/.gitkeep
Normal file
0
infra/gpu-bench/results/.gitkeep
Normal file
141
infra/gpu-bench/run.sh
Executable file
141
infra/gpu-bench/run.sh
Executable file
|
|
@ -0,0 +1,141 @@
|
|||
#!/usr/bin/env bash
|
||||
# Igneum GPU benchmark on a rented card. Runs on the pod, inside the unpacked bundle (make-bundle.sh). Clones nothing:
|
||||
# the sources are in the bundle. Needs nvcc (CUDA 12.8 image) for NVIDIA, or an OpenCL ICD plus a C compiler for AMD.
|
||||
#
|
||||
# ./run.sh everything below, results in results-<stamp>/ and one row appended to results.md
|
||||
# MINUTES=10 PACK=igneum-genesis-mh PACK2=igneum-hourly LABEL=rtx4090 UPLOAD=1 ./run.sh
|
||||
#
|
||||
# Steps (NVIDIA lane): build the memory-hard pack and the closed-form second pack with --serve support (host.cu has
|
||||
# it), the inline-shortcut variant (make-inline.sh), the NVRTC timer; gate on the vectors (96/96 PASS or stop); raw
|
||||
# bench for MINUTES at 1 GiB; sweep 64, 128, 256, 512, 1024 MiB; inline run; recompile timings (nvcc -cubin as the
|
||||
# worker's prepare path does it, NVRTC in process, OpenCL build if an ICD is present); results row; upload.
|
||||
# AMD lane (ROCm image): the same through proto-opencl/host.c, minus NVRTC and the inline variant (OpenCL packs have
|
||||
# no inline kernel yet; the sed of make-inline.sh is CUDA-only).
|
||||
set -uo pipefail
|
||||
cd "$(dirname "$0")"
|
||||
MINUTES="${MINUTES:-10}"; PACK="${PACK:-igneum-genesis-mh}"; PACK2="${PACK2:-igneum-hourly}"
|
||||
LABEL="${LABEL:-}"; UPLOAD="${UPLOAD:-1}"; ITER="${ITER:-3}"
|
||||
STAMP=$(date -u +%Y%m%d-%H%M%S); OUT="results-$STAMP"; mkdir -p "$OUT"
|
||||
exec > >(tee "$OUT/run.log") 2>&1
|
||||
log() { printf '%s %s\n' "$(date -u +%H:%M:%S)" "$*"; }
|
||||
rate_of() { grep -E '^\s*(GPU|device)\b.*Mhash/s|^\s*rate' "$1" | grep -oE '[0-9]+\.[0-9]+ Mhash/s' | head -1 | cut -d' ' -f1; }
|
||||
rate_wall() { grep -oE '[0-9]+\.[0-9]+ Mhash/s' "$1" | tail -1 | cut -d' ' -f1; }
|
||||
now_ms() { python3 -c 'import time; print(int(time.time()*1000))'; }
|
||||
|
||||
PACKDIR="proto-cuda/packs/$PACK"; PACK2DIR="proto-cuda/packs/$PACK2"
|
||||
[ -d "$PACKDIR" ] || { log "no $PACKDIR in the bundle"; exit 2; }
|
||||
|
||||
if command -v nvidia-smi >/dev/null 2>&1 && nvidia-smi -L >/dev/null 2>&1; then LANE=cuda; else LANE=opencl; fi
|
||||
log "lane: $LANE; pack $PACK (memory-hard), second pack $PACK2; $MINUTES min raw bench"
|
||||
|
||||
# ---- device facts --------------------------------------------------------------------------------------------------
|
||||
if [ "$LANE" = cuda ]; then
|
||||
nvidia-smi --query-gpu=name,driver_version,memory.total,clocks.max.sm,clocks.max.mem --format=csv | tee "$OUT/device.txt"
|
||||
nvcc --version | tail -2 | tee -a "$OUT/device.txt"
|
||||
CARD=$(nvidia-smi --query-gpu=name --format=csv,noheader | head -1 | sed 's/NVIDIA //; s/GeForce //')
|
||||
DRIVER=$(nvidia-smi --query-gpu=driver_version --format=csv,noheader | head -1)
|
||||
CUDAV=$(nvcc --version | grep -oE 'release [0-9.]+' | cut -d' ' -f2)
|
||||
else
|
||||
clinfo 2>/dev/null | grep -E 'Device Name|Driver Version|Device Version|Global memory size|Max compute units' | head -12 | tee "$OUT/device.txt"
|
||||
CARD=$(clinfo 2>/dev/null | grep -m1 'Device Name' | sed 's/.*Device Name *//'); DRIVER=$(clinfo 2>/dev/null | grep -m1 'Driver Version' | sed 's/.*Version *//'); CUDAV="opencl"
|
||||
rocminfo 2>/dev/null | grep -m1 -E 'gfx[0-9]+' | tee -a "$OUT/device.txt" || true
|
||||
fi
|
||||
[ -n "$LABEL" ] || LABEL=$(printf '%s' "$CARD" | tr 'A-Z ' 'a-z-' | tr -cd 'a-z0-9-')
|
||||
log "card: $CARD, driver $DRIVER, toolkit $CUDAV, label $LABEL"
|
||||
|
||||
# ---- build ---------------------------------------------------------------------------------------------------------
|
||||
t0=$(now_ms)
|
||||
if [ "$LANE" = cuda ]; then
|
||||
ARCH="${CUDA_ARCH:-native}"
|
||||
log "nvcc -arch=$ARCH: bench-$PACK, bench-$PACK2, bench-inline, nvrtc-time"
|
||||
nvcc -O3 -std=c++17 -arch="$ARCH" -I "$PACKDIR" -o "bench-$PACK" proto-cuda/host.cu "$PACKDIR/kernel.cu" 2>&1 | tail -5 || { log "build of $PACK failed"; exit 2; }
|
||||
nvcc -O3 -std=c++17 -arch="$ARCH" -I "$PACK2DIR" -o "bench-$PACK2" proto-cuda/host.cu "$PACK2DIR/kernel.cu" 2>&1 | tail -5 || log "build of $PACK2 failed (continuing)"
|
||||
./make-inline.sh "$PACKDIR" "$OUT/inline-pack" && nvcc -O3 -std=c++17 -arch="$ARCH" -I "$OUT/inline-pack" -o bench-inline proto-cuda/host.cu "$OUT/inline-pack/kernel.cu" 2>&1 | tail -5 || log "inline variant did not build (continuing)"
|
||||
nvcc -O2 -std=c++17 -o nvrtc-time nvrtc-time.cu -lnvrtc -lcuda 2>&1 | tail -3 || log "nvrtc-time did not build (continuing)"
|
||||
WORKER="./bench-$PACK"
|
||||
else
|
||||
log "cc: bench-cl-$PACK (OpenCL at runtime), clbuild-time"
|
||||
cc -std=c99 -O2 -I "$PACKDIR" -DIGNEUM_KERNEL_PATH="\"$PACKDIR/kernel.cl\"" -o "bench-cl-$PACK" proto-opencl/host.c -lOpenCL -ldl 2>&1 | tail -5 || { log "build failed"; exit 2; }
|
||||
cc -std=c99 -O2 -I "$PACK2DIR" -DIGNEUM_KERNEL_PATH="\"$PACK2DIR/kernel.cl\"" -o "bench-cl-$PACK2" proto-opencl/host.c -lOpenCL -ldl 2>&1 | tail -5 || true
|
||||
cc -std=c99 -O2 -o clbuild-time clbuild-time.c -lOpenCL 2>&1 | tail -3 || true
|
||||
WORKER="./bench-cl-$PACK"
|
||||
fi
|
||||
BUILD_MS=$(( $(now_ms) - t0 )); log "build took $BUILD_MS ms"
|
||||
# --serve support is in host.cu and host.c (ready line "ready cuda|opencl ... prepare N"); prove it answers:
|
||||
printf 'quit\n' | timeout 120 "$WORKER" --serve 2>&1 | head -3 | tee "$OUT/serve-ready.txt" || true
|
||||
|
||||
# ---- gate: vectors -------------------------------------------------------------------------------------------------
|
||||
log "gate: vectors on $PACK (3 batches)"
|
||||
"$WORKER" --batches 3 > "$OUT/gate.txt" 2>&1
|
||||
grep -E 'verify warp|cache check|dataset self-test|OVERALL' "$OUT/gate.txt"
|
||||
if ! grep -q 'OVERALL: PASS' "$OUT/gate.txt"; then log "GATE FAIL: the card does not reproduce the Mac's vectors; stopping (send $OUT/gate.txt)"; VECTORS=FAIL; else VECTORS="96/96 PASS"; fi
|
||||
[ "$VECTORS" = FAIL ] && exit 1
|
||||
|
||||
# ---- raw bench, MINUTES at 1 GiB -------------------------------------------------------------------------------------
|
||||
r0=$(rate_of "$OUT/gate.txt"); [ -n "$r0" ] || r0=$(rate_wall "$OUT/gate.txt")
|
||||
batches=$(python3 -c "import math; r=float('${r0:-50}'); print(max(5, min(100000, int(math.ceil($MINUTES*60*r*1e6/2**24)))))")
|
||||
log "raw: $batches batches of 2^24 at about $r0 Mhash/s (about $MINUTES min)"
|
||||
t0=$(now_ms); "$WORKER" --batches "$batches" > "$OUT/raw.txt" 2>&1; RAW_S=$(( ($(now_ms) - t0) / 1000 ))
|
||||
RAW=$(rate_of "$OUT/raw.txt"); [ -n "$RAW" ] || RAW=$(rate_wall "$OUT/raw.txt")
|
||||
grep -E 'Mhash/s|OVERALL' "$OUT/raw.txt" | head -4
|
||||
log "raw: $RAW Mhash/s over $RAW_S s"
|
||||
|
||||
# ---- sweep ---------------------------------------------------------------------------------------------------------
|
||||
SWEEP=""
|
||||
for mib in 64 128 256 512 1024; do
|
||||
"$WORKER" --dataset-mib "$mib" --batches 20 > "$OUT/sweep-$mib.txt" 2>&1
|
||||
r=$(rate_of "$OUT/sweep-$mib.txt"); [ -n "$r" ] || r=$(rate_wall "$OUT/sweep-$mib.txt")
|
||||
SWEEP="$SWEEP $mib:${r:-n/a}"; log "sweep $mib MiB: ${r:-n/a} Mhash/s"
|
||||
done
|
||||
R64=$(printf '%s' "$SWEEP" | grep -oE ' 64:[0-9.n/a]+' | cut -d: -f2); R1024=$(printf '%s' "$SWEEP" | grep -oE '1024:[0-9.n/a]+' | cut -d: -f2)
|
||||
CLIFF=$(python3 -c "
|
||||
try: print('%.1f' % (float('$R64') / float('$R1024')))
|
||||
except Exception: print('n/a')")
|
||||
|
||||
# ---- second pack (closed form, for comparison with the Mac's and the 5090's tables) -----------------------------------
|
||||
R2="n/a"
|
||||
if [ -x "./bench-$PACK2" ] || [ -x "./bench-cl-$PACK2" ]; then
|
||||
W2="./bench-$PACK2"; [ -x "$W2" ] || W2="./bench-cl-$PACK2"
|
||||
"$W2" --batches 20 > "$OUT/pack2.txt" 2>&1; R2=$(rate_of "$OUT/pack2.txt"); [ -n "$R2" ] || R2=$(rate_wall "$OUT/pack2.txt")
|
||||
log "$PACK2: $R2 Mhash/s, $(grep -c 'PASS' "$OUT/pack2.txt") PASS lines"
|
||||
fi
|
||||
|
||||
# ---- inline shortcut (CUDA only) -------------------------------------------------------------------------------------
|
||||
INLINE="n/a"; RATIO="n/a"
|
||||
if [ -x ./bench-inline ]; then
|
||||
# By construction the inline binary FAILS the dataset self-test (the dataset buffer holds the cache copy) and must
|
||||
# PASS the vectors (mh_word recomputes the true words). Only the rate line and the vector lines count.
|
||||
./bench-inline --batches 20 > "$OUT/inline.txt" 2>&1 || true
|
||||
INLINE=$(rate_of "$OUT/inline.txt"); [ -n "$INLINE" ] || INLINE=$(rate_wall "$OUT/inline.txt")
|
||||
IV=$(grep -c 'verify warp.*: PASS' "$OUT/inline.txt"); log "inline: ${INLINE:-n/a} Mhash/s, $IV vector PASS lines (6 expected)"
|
||||
[ "$IV" -ge 3 ] || { log "inline vectors did not pass: ratio discarded"; INLINE="n/a(vectors)"; }
|
||||
RATIO=$(python3 -c "
|
||||
try: print('%.3f' % (float('$INLINE') / float('$RAW')))
|
||||
except Exception: print('n/a')")
|
||||
fi
|
||||
|
||||
# ---- recompile timings ---------------------------------------------------------------------------------------------
|
||||
NVCC_MS="n/a"; NVRTC_MS="n/a"; CL_MS="n/a"
|
||||
if [ "$LANE" = cuda ]; then
|
||||
ARCHSM=$(nvidia-smi --query-gpu=compute_cap --format=csv,noheader | head -1 | tr -d '.')
|
||||
xs=""; for i in $(seq 1 "$ITER"); do t0=$(now_ms); nvcc -cubin -O3 -std=c++17 -arch="sm_$ARCHSM" -allow-unsupported-compiler -I "$PACKDIR" -o "$OUT/kernel.cubin" "$PACKDIR/kernel.cu" >/dev/null 2>&1; xs="$xs $(( $(now_ms) - t0 ))"; done
|
||||
NVCC_MS=$(python3 -c "xs=sorted(int(x) for x in '$xs'.split()); print(xs[len(xs)//2] if xs else 'n/a')"); log "nvcc -cubin (the worker's prepare path): $xs ms, median $NVCC_MS"
|
||||
if [ -x ./nvrtc-time ]; then ./nvrtc-time "$PACKDIR" "$ITER" | tee "$OUT/nvrtc.txt"; NVRTC_MS=$(grep -oE 'median [0-9.]+' "$OUT/nvrtc.txt" | cut -d' ' -f2); fi
|
||||
fi
|
||||
if command -v clinfo >/dev/null 2>&1 && clinfo 2>/dev/null | grep -q 'Device Name'; then
|
||||
[ -x ./clbuild-time ] || cc -std=c99 -O2 -o clbuild-time clbuild-time.c -lOpenCL 2>/dev/null || true
|
||||
[ -x ./clbuild-time ] && { ./clbuild-time "$PACKDIR/kernel.cl" "$ITER" | tee "$OUT/clbuild.txt"; CL_MS=$(grep -oE 'median [0-9.]+' "$OUT/clbuild.txt" | cut -d' ' -f2); }
|
||||
fi
|
||||
|
||||
# ---- results row -----------------------------------------------------------------------------------------------------
|
||||
L2=$(nvidia-smi --query-gpu=name --format=csv,noheader 2>/dev/null | head -1 | grep -qE '5090' && echo 96 || echo "?")
|
||||
ROW="| $CARD | $DRIVER / $CUDAV | $(date -u +%Y-%m-%d) | $RAW | $RAW_S | $R2 | $(printf '%s' "$SWEEP" | sed 's/^ //; s/ / \/ /g') | $CLIFF | $INLINE | $RATIO | $NVCC_MS | $NVRTC_MS | $CL_MS | $VECTORS | $BUILD_MS |"
|
||||
HEAD="| card | driver / toolkit | date | Mhash/s at 1 GiB ($MINUTES min) | s | $PACK2 Mhash/s | sweep MiB:Mhash/s 64 / 128 / 256 / 512 / 1024 | 64 MiB over 1 GiB | inline Mhash/s | inline / honest | nvcc cubin ms | NVRTC ms | OpenCL build ms | vectors | build ms |"
|
||||
[ -f results.md ] || { printf '%s\n|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|\n' "$HEAD" > results.md; }
|
||||
printf '%s\n' "$ROW" >> results.md
|
||||
printf '%s\n%s\n' "$HEAD" "$ROW" > "$OUT/row.md"
|
||||
log "RESULT $ROW"
|
||||
|
||||
# ---- upload ----------------------------------------------------------------------------------------------------------
|
||||
if [ "$UPLOAD" = 1 ]; then ./upload.sh "$OUT/run.log" "gpubench-$LABEL" "gpubench-$LABEL-$STAMP" || log "upload failed (the row is in results.md and $OUT/row.md)"; fi
|
||||
log "done: $OUT/ (run.log, gate.txt, raw.txt, sweep-*.txt, inline.txt, nvrtc.txt, clbuild.txt, row.md)"
|
||||
22
infra/gpu-bench/upload.sh
Executable file
22
infra/gpu-bench/upload.sh
Executable file
|
|
@ -0,0 +1,22 @@
|
|||
#!/usr/bin/env bash
|
||||
# Upload a log (its last 256 KB) to the Igneum log intake, the same endpoint and key as
|
||||
# proto-cuda/windows-miner/upload-log.bat (the key only authorises log uploads and ships inside the packages).
|
||||
# Read back on the Mac with `node tools/logs.mjs`.
|
||||
# ./upload.sh <logfile> <label> [run_id]
|
||||
set -euo pipefail
|
||||
IGNEUM_LOG_URL="${IGNEUM_LOG_URL:-https://igneum-six.vercel.app/api/log}"
|
||||
IGNEUM_LOG_KEY="${IGNEUM_LOG_KEY:-***INTAKE-KEY-REMOVED***}"
|
||||
[ $# -ge 2 ] || { echo "usage: upload.sh <logfile> <label> [run_id]"; exit 1; }
|
||||
file="$1"; label="$2"; run_id="${3:-${IGNEUM_RUN_ID:-$label-$(date -u +%Y%m%d-%H%M)}}"
|
||||
[ -f "$file" ] || { echo "upload: file not found: $file"; exit 2; }
|
||||
body=$(mktemp)
|
||||
python3 - "$file" "$label" "$run_id" "$body" <<'EOF'
|
||||
import json, socket, sys
|
||||
path, label, run_id, out = sys.argv[1:5]
|
||||
data = open(path, 'rb').read()[-262144:]
|
||||
json.dump({"label": label, "machine": socket.gethostname(), "run_id": run_id, "lines": data.decode('utf-8', 'replace')}, open(out, 'w'))
|
||||
print(f"upload: run_id {run_id}, {len(data)} bytes")
|
||||
EOF
|
||||
curl -sS --max-time 60 -X POST "$IGNEUM_LOG_URL" -H "Content-Type: application/json" -H "x-igneum-key: $IGNEUM_LOG_KEY" --data-binary "@$body"
|
||||
echo
|
||||
rm -f "$body"
|
||||
2
infra/seed-nodes/.gitignore
vendored
Normal file
2
infra/seed-nodes/.gitignore
vendored
Normal file
|
|
@ -0,0 +1,2 @@
|
|||
# build logs, known_hosts and tarballs; seeds.txt and seeds.tsv are committed on purpose
|
||||
build/
|
||||
32
infra/seed-nodes/README.md
Normal file
32
infra/seed-nodes/README.md
Normal file
|
|
@ -0,0 +1,32 @@
|
|||
# Igneum seed nodes
|
||||
|
||||
A seed is a small VM with a fixed public IPv4 that runs `igneumd` with p2p open, no mining, RPC on loopback, and
|
||||
serves peer exchange (the address manager stays on; `--connect` is never used). `seeds.txt` lists one `<ip>:26611`
|
||||
per seed and is the file the consensus engineer bakes into the network parameters and every package reads as
|
||||
`SEED_PEERS`. The plan, the client path and the rotation procedure are in `docs/plans/seed-nodes.md`.
|
||||
|
||||
The first seed, `igneum-seed-1`, was created on 3 Oct 2026 (Hetzner cx23, Falkenstein, EUR 6.49 per month net plus
|
||||
the IPv4). It runs the shared devnet (`--devnet`, no suffix).
|
||||
|
||||
```
|
||||
cd infra/seed-nodes
|
||||
./create-seed.sh # one VM, persistent IPv4, firewall 22 + 26611 + icmp, appends seeds.tsv, rewrites seeds.txt
|
||||
./provision-seed.sh igneum-seed-1 # source tarball -> build on the VM (about an hour on 2 vCPU) -> unit igneumd, started
|
||||
./health.sh # one line per seed: p2p port, unit, RPC, synced, blocks, peers, known addresses
|
||||
./addpeer-from-mac.sh 188.245.5.161 # the Mac's live node (NAT, gRPC only) dials the seed: addPeer over grpcurl, no restart
|
||||
SEED_NAME=igneum-seed-2 SEED_LOCATION=ash SEED_TYPE=cpx11 ./create-seed.sh # the next seed; provision adds the first as --addpeer
|
||||
```
|
||||
|
||||
Settings in `config.sh`: provider, name, type, location, image, ssh key (`~/.ssh/igneum_ed25519`), `SSH_SOURCE`
|
||||
(`any`, `me` or a CIDR for port 22), `NETWORK_ARGS` (`--devnet`; a private network: `--devnet --devnet-suffix=20`),
|
||||
`SEED_PEERS`, `BUILD_WHERE` (`seed` builds on the VM; `bin` reuses `../cloud-devnet/build/bin`). The Hetzner token is
|
||||
read from `~/.config/igneum/hetzner-token` and never printed.
|
||||
|
||||
Node flags (`node/run-seed.sh`): `--devnet --appdir=/var/lib/igneum --rpclisten=127.0.0.1:26610
|
||||
--rpclisten-json=127.0.0.1:28610 --listen=0.0.0.0:26611 --externalip=<ip> --nodnsseed --disable-upnp --nologfiles
|
||||
--yes --maxinpeers=128 --outpeers=8 [--addpeer=<other seeds>]`. No `--enable-unsynced-mining` (nothing mines here) and
|
||||
no `--connect` (it would set the inbound limit to 0 and switch the address manager off).
|
||||
|
||||
Files: `create-seed.sh`, `provision-seed.sh`, `health.sh`, `addpeer-from-mac.sh`, `config.sh`, `lib.sh`,
|
||||
`node/{install-seed.sh,run-seed.sh,igneumd.service}`, `seeds.txt`, `seeds.tsv`, `build/` (logs, ignored by nothing: keep
|
||||
the build logs, they are small).
|
||||
22
infra/seed-nodes/addpeer-from-mac.sh
Executable file
22
infra/seed-nodes/addpeer-from-mac.sh
Executable file
|
|
@ -0,0 +1,22 @@
|
|||
#!/usr/bin/env bash
|
||||
# Make the Mac's live devnet node dial a seed (the Mac sits behind NAT on 192.168.68.64, so the seed cannot dial in).
|
||||
# The live node (./target/release/kaspad --devnet ... --rpclisten=0.0.0.0:26610) serves gRPC only, no wRPC JSON, so
|
||||
# this goes through grpcurl against the fork's own proto files. addPeer is an RPC call: the node is not restarted and
|
||||
# nothing on the chain changes. isPermanent=true makes the node retry the seed after a disconnect (backoff up to 8 min).
|
||||
# ./addpeer-from-mac.sh <ip[:port]> add (port defaults to 26611)
|
||||
# ./addpeer-from-mac.sh --peers list the live node's connected peers
|
||||
# ./addpeer-from-mac.sh --dag the live node's getBlockDagInfo (to compare block counts with the seed)
|
||||
# If the live node is ever restarted by hand, the permanent alternative is the flag: --addpeer=<ip>:26611
|
||||
. "$(dirname "$0")/lib.sh"
|
||||
need grpcurl "brew install grpcurl"
|
||||
RPC="${MAC_RPC:-127.0.0.1:26610}"
|
||||
PROTO="$REPO/vendor/igneum-node/rpc/grpc/core/proto"
|
||||
call() { grpcurl -plaintext -max-time 15 -import-path "$PROTO" -proto messages.proto -d "$1" "$RPC" protowire.RPC/MessageStream; }
|
||||
case "${1:-}" in
|
||||
"") die "usage: addpeer-from-mac.sh <ip[:port]> | --peers | --dag" ;;
|
||||
--peers) call '{"getConnectedPeerInfoRequest":{}}' ;;
|
||||
--dag) call '{"getBlockDagInfoRequest":{}}' ;;
|
||||
*) addr="$1"; case "$addr" in *:*) ;; *) addr="$addr:$P2P_PORT" ;; esac
|
||||
log "addPeer $addr on the live node at $RPC"
|
||||
call "{\"addPeerRequest\":{\"address\":\"$addr\",\"isPermanent\":true}}" ;;
|
||||
esac
|
||||
38
infra/seed-nodes/config.sh
Executable file
38
infra/seed-nodes/config.sh
Executable file
|
|
@ -0,0 +1,38 @@
|
|||
# Igneum seed nodes: settings. Sourced by every script in this directory. Override in the environment.
|
||||
# A seed is a small VM with a fixed public IPv4 that runs igneumd with p2p open, no mining, RPC on loopback, and
|
||||
# serves peer exchange. seeds.txt lists "<ip>:26611" per seed; seeds.tsv keeps the metadata.
|
||||
|
||||
PROVIDER="${PROVIDER:-hetzner}" # hetzner (hcloud) or digitalocean (doctl)
|
||||
SEED_NAME="${SEED_NAME:-igneum-seed-1}"
|
||||
SEED_TYPE="${SEED_TYPE:-cx23}" # 2 shared Intel vCPU, 4 GB, 40 GB: EUR 6.49/mo net, EUR 0.0104/h, in fsn1, nbg1, hel1
|
||||
# (Hetzner API, 3 Oct 2026). It is the cheapest x86 type this project can create in the
|
||||
# EU: the old cpx11/cpx21 line is deprecated there since Dec 2025 (only ash and hil keep
|
||||
# it, at EUR 20.49 and 37.49/mo), cpx12 (1 vCPU, 2 GB) costs 13.49. 4 GB is needed: the
|
||||
# node keeps up to four 256 MiB lottery caches (M15 cap) beside rocksdb, and the on-VM
|
||||
# build wants 4 GB plus swap. US seeds: cpx11 (2 GB) or cpx21; Singapore: cpx22 (30.99).
|
||||
SEED_LOCATION="${SEED_LOCATION:-fsn1}" # fsn1 Falkenstein, nbg1 Nuremberg, hel1, ash, hil, sin
|
||||
IMAGE="${IMAGE:-debian-12}"
|
||||
DO_SIZE="${DO_SIZE:-s-2vcpu-4gb}"
|
||||
DO_REGION="${DO_REGION:-fra1}"
|
||||
DO_IMAGE="${DO_IMAGE:-debian-12-x64}"
|
||||
|
||||
SSH_KEY_FILE="${SSH_KEY_FILE:-$HOME/.ssh/igneum_ed25519}" # the operations key (3 Oct 2026)
|
||||
SSH_KEY_NAME="${SSH_KEY_NAME:-igneum-ops}"
|
||||
SSH_USER="${SSH_USER:-root}"
|
||||
# Who may reach port 22. "any" = 0.0.0.0/0 (the project lead's instruction of 3 Oct 2026: SSH open); "me" = this Mac's public IP
|
||||
# at creation time (api.ipify.org); or a CIDR. Change later with: hcloud firewall replace-rules, or ./firewall.sh.
|
||||
SSH_SOURCE="${SSH_SOURCE:-any}"
|
||||
|
||||
HETZNER_TOKEN_FILE="${HETZNER_TOKEN_FILE:-$HOME/.config/igneum/hetzner-token}" # one line, the API token; never printed
|
||||
|
||||
NETWORK_ARGS="${NETWORK_ARGS:---devnet}" # the live devnet; a suffixed test network: "--devnet --devnet-suffix=20"
|
||||
SEED_PEERS="${SEED_PEERS:-}" # other seeds to --addpeer, comma separated ip:port (filled from seeds.txt by provision)
|
||||
P2P_PORT=26611
|
||||
RPC_PORT=26610
|
||||
RPC_JSON_PORT=28610
|
||||
|
||||
# The node source (shared with the 20-node network): vendor/igneum-node working tree plus igneum-pow
|
||||
NODE_SRC="${NODE_SRC:-$REPO/vendor/igneum-node}"
|
||||
POW_SRC="${POW_SRC:-$REPO/igneum-pow}"
|
||||
SRC_MODE="${SRC_MODE:-head+dirty}"
|
||||
BUILD_WHERE="${BUILD_WHERE:-seed}" # seed = build on the seed VM itself (swap added); bin = use ../cloud-devnet/build/bin
|
||||
75
infra/seed-nodes/create-seed.sh
Executable file
75
infra/seed-nodes/create-seed.sh
Executable file
|
|
@ -0,0 +1,75 @@
|
|||
#!/usr/bin/env bash
|
||||
# Create one seed VM with a fixed public IPv4. Hetzner (hcloud) primary; DigitalOcean (doctl) with PROVIDER=digitalocean.
|
||||
# ./create-seed.sh creates SEED_NAME (default igneum-seed-1) in SEED_LOCATION
|
||||
# SEED_NAME=igneum-seed-2 SEED_LOCATION=ash ./create-seed.sh
|
||||
# Firewall: inbound tcp 26611 (p2p) from anywhere, tcp 22 from SSH_SOURCE, icmp; nothing else. RPC binds to loopback
|
||||
# in the unit, so no rule is needed for it. The IPv4 is made persistent (Hetzner: primary IP auto-delete off; DO:
|
||||
# a reserved IP), so a rebuilt server keeps the address that clients have baked in.
|
||||
# Appends the seed to seeds.tsv and rewrites seeds.txt. Asks "yes" before spending (YES=1 skips).
|
||||
. "$(dirname "$0")/lib.sh"
|
||||
mkdir -p "$BUILD_DIR"
|
||||
[ -f "$SSH_KEY_FILE.pub" ] || die "no public key at $SSH_KEY_FILE.pub"
|
||||
if grep -q "^$SEED_NAME " "$SEEDS_TSV" 2>/dev/null; then die "$SEED_NAME is already in seeds.tsv"; fi
|
||||
|
||||
ssh_cidr="0.0.0.0/0"
|
||||
case "$SSH_SOURCE" in
|
||||
any) ssh_cidr="0.0.0.0/0" ;;
|
||||
me) me=$(curl -s --max-time 10 https://api.ipify.org || true); [ -n "$me" ] || die "could not learn this Mac's public IP"; ssh_cidr="$me/32" ;;
|
||||
*) ssh_cidr="$SSH_SOURCE" ;;
|
||||
esac
|
||||
|
||||
if [ "$PROVIDER" = digitalocean ]; then
|
||||
need doctl "brew install doctl"
|
||||
key_id=$(doctl compute ssh-key list --format ID,Name --no-header | awk -v n="$SSH_KEY_NAME" '$2 == n { print $1 }')
|
||||
[ -n "$key_id" ] || key_id=$(doctl compute ssh-key import "$SSH_KEY_NAME" --public-key-file "$SSH_KEY_FILE.pub" --format ID --no-header)
|
||||
log "plan: $SEED_NAME, $DO_SIZE, $DO_IMAGE, $DO_REGION, reserved IPv4, firewall 22 from $ssh_cidr + 26611 from anywhere"
|
||||
doctl compute size list --format Slug,Memory,VCPUs,Disk,PriceMonthly,PriceHourly | grep -E "^Slug|^$DO_SIZE "
|
||||
[ "${YES:-0}" = 1 ] || { printf 'create it now (billing starts) [type yes]: '; read -r a; [ "$a" = yes ] || die "not confirmed"; }
|
||||
fw=$(doctl compute firewall list --format ID,Name --no-header | awk '$2 == "igneum-seed" { print $1 }')
|
||||
[ -n "$fw" ] || fw=$(doctl compute firewall create --name igneum-seed --tag-names igneum-seed \
|
||||
--inbound-rules "protocol:tcp,ports:22,address:$ssh_cidr protocol:tcp,ports:$P2P_PORT,address:0.0.0.0/0,address:::/0 protocol:icmp,address:0.0.0.0/0,address:::/0" \
|
||||
--outbound-rules "protocol:tcp,ports:all,address:0.0.0.0/0,address:::/0 protocol:udp,ports:all,address:0.0.0.0/0,address:::/0 protocol:icmp,address:0.0.0.0/0,address:::/0" --format ID --no-header)
|
||||
did=$(doctl compute droplet create "$SEED_NAME" --size "$DO_SIZE" --image "$DO_IMAGE" --region "$DO_REGION" --ssh-keys "$key_id" --tag-names igneum-seed --wait --format ID --no-header)
|
||||
rip=$(doctl compute reserved-ip create --region "$DO_REGION" --format IP --no-header)
|
||||
doctl compute reserved-ip-action assign "$rip" "$did" >/dev/null
|
||||
ip="$rip"; loc="$DO_REGION"; typ="$DO_SIZE"
|
||||
else
|
||||
hetzner_auth
|
||||
if ! hcloud ssh-key describe "$SSH_KEY_NAME" >/dev/null 2>&1; then
|
||||
hcloud ssh-key create --name "$SSH_KEY_NAME" --public-key-from-file "$SSH_KEY_FILE.pub" >/dev/null; log "uploaded ssh key $SSH_KEY_NAME"
|
||||
fi
|
||||
if ! hcloud firewall describe igneum-seed >/dev/null 2>&1; then
|
||||
hcloud firewall create --name igneum-seed --label igneum=seed >/dev/null
|
||||
hcloud firewall add-rule igneum-seed --direction in --protocol tcp --port 22 --source-ips "$ssh_cidr" --description ssh >/dev/null
|
||||
hcloud firewall add-rule igneum-seed --direction in --protocol tcp --port "$P2P_PORT" --source-ips 0.0.0.0/0 --source-ips ::/0 --description igneum-p2p >/dev/null
|
||||
hcloud firewall add-rule igneum-seed --direction in --protocol icmp --source-ips 0.0.0.0/0 --source-ips ::/0 --description ping >/dev/null
|
||||
log "created firewall igneum-seed (in: 22 from $ssh_cidr, $P2P_PORT from anywhere, icmp)"
|
||||
fi
|
||||
log "plan: $SEED_NAME, $SEED_TYPE, $IMAGE, $SEED_LOCATION, persistent primary IPv4, firewall igneum-seed"
|
||||
hcloud server-type describe "$SEED_TYPE" -o json | python3 -c '
|
||||
import json, sys
|
||||
j = json.load(sys.stdin); loc = sys.argv[1]
|
||||
for p in j["prices"]:
|
||||
if p["location"] == loc:
|
||||
print(" %s in %s: EUR %.2f per month net (%.2f gross), EUR %.4f per hour net, %d TB traffic included (Hetzner API)" % (
|
||||
j["name"], loc, float(p["price_monthly"]["net"]), float(p["price_monthly"]["gross"]), float(p["price_hourly"]["net"]), int(p.get("included_traffic", 0)) // (1 << 40)))' "$SEED_LOCATION"
|
||||
log " plus the primary IPv4: about EUR 0.50 per month (approximate, Hetzner list price)"
|
||||
[ "${YES:-0}" = 1 ] || { printf 'create it now (billing starts) [type yes]: '; read -r a; [ "$a" = yes ] || die "not confirmed"; }
|
||||
if hcloud server describe "$SEED_NAME" >/dev/null 2>&1; then log "$SEED_NAME exists, reusing it"; else
|
||||
hcloud server create --name "$SEED_NAME" --type "$SEED_TYPE" --image "$IMAGE" --location "$SEED_LOCATION" \
|
||||
--ssh-key "$SSH_KEY_NAME" --firewall igneum-seed --label igneum=seed --label role=seed >/dev/null
|
||||
log "created $SEED_NAME"
|
||||
fi
|
||||
ip=$(hcloud server ip "$SEED_NAME")
|
||||
pip=$(hcloud primary-ip list -o noheader -o columns=id,ip | awk -v ip="$ip" '$2 == ip { print $1 }')
|
||||
if [ -n "$pip" ]; then hcloud primary-ip update "$pip" --auto-delete=false --name "$SEED_NAME-v4" >/dev/null && log "primary IPv4 $ip ($pip) set to persist (auto-delete off)"; fi
|
||||
loc="$SEED_LOCATION"; typ="$SEED_TYPE"
|
||||
fi
|
||||
|
||||
printf '%s\t%s\t%s\t%s\t%s\t%s\n' "$SEED_NAME" "$PROVIDER" "$loc" "$ip" "$typ" "$(date -u +%Y-%m-%dT%H:%M:%SZ)" >> "$SEEDS_TSV"
|
||||
write_seeds_txt
|
||||
log "seed $SEED_NAME at $ip; seeds.txt now:"; cat "$SEEDS_TXT"
|
||||
log "waiting for ssh"
|
||||
for try in $(seq 1 20); do sssh "$ip" true >/dev/null 2>&1 && break; sleep 10; done
|
||||
sssh "$ip" 'hostname; uname -m; cat /etc/debian_version' || log "WARNING: ssh not up yet; provision-seed.sh retries"
|
||||
log "next: ./provision-seed.sh $SEED_NAME"
|
||||
37
infra/seed-nodes/health.sh
Executable file
37
infra/seed-nodes/health.sh
Executable file
|
|
@ -0,0 +1,37 @@
|
|||
#!/usr/bin/env bash
|
||||
# Health check for every seed in seeds.tsv (or one name). One line per seed, exit 1 if any check fails.
|
||||
# Checks: p2p port reachable from here (nc), unit active, RPC answers, synced flag, blocks and headers, connected
|
||||
# peers, known addresses (the peer-exchange table), disk and memory. Add --watch to repeat every 60 s.
|
||||
. "$(dirname "$0")/lib.sh"
|
||||
only="${1:-}"; [ "$only" = --watch ] && only=""
|
||||
fail=0
|
||||
check_one() {
|
||||
local name="$1" ip; ip=$(seed_ip "$name")
|
||||
local port=FAIL unit=? info= dag= peers=? known=? disk=? mem=? synced=? ver=? blocks=? headers=? sink=?
|
||||
if nc -z -w 5 "$ip" "$P2P_PORT" >/dev/null 2>&1; then port=open; fi
|
||||
local raw
|
||||
raw=$(sssh "$ip" "systemctl is-active igneumd 2>/dev/null; echo '|'; python3 /opt/igneum/bin/wrpc.py call getInfo 2>/dev/null; echo '|'; python3 /opt/igneum/bin/wrpc.py call getBlockDagInfo 2>/dev/null; echo '|'; python3 /opt/igneum/bin/wrpc.py call getConnectedPeerInfo 2>/dev/null | python3 -c 'import json,sys; print(len(json.load(sys.stdin).get(\"peerInfo\",[])))' 2>/dev/null; echo '|'; python3 /opt/igneum/bin/wrpc.py call getPeerAddresses 2>/dev/null | python3 -c 'import json,sys; j=json.load(sys.stdin); print(len(j.get(\"knownAddresses\",[])), len(j.get(\"bannedAddresses\",[])))' 2>/dev/null; echo '|'; df -h / | awk 'NR==2 { print \$5 }'; echo '|'; free -m | awk 'NR==2 { print \$3 \"/\" \$2 \"MB\" }'" 2>/dev/null) || raw=""
|
||||
unit=$(printf '%s' "$raw" | awk -F'|' 'NR==1 { gsub(/\n/, "", $1); print $1 }' | tr -d '\n')
|
||||
info=$(printf '%s' "$raw" | awk 'BEGIN{RS="|"} NR==2' | tr -d '\n')
|
||||
dag=$(printf '%s' "$raw" | awk 'BEGIN{RS="|"} NR==3' | tr -d '\n')
|
||||
peers=$(printf '%s' "$raw" | awk 'BEGIN{RS="|"} NR==4' | tr -d '\n ')
|
||||
known=$(printf '%s' "$raw" | awk 'BEGIN{RS="|"} NR==5' | tr -d '\n')
|
||||
disk=$(printf '%s' "$raw" | awk 'BEGIN{RS="|"} NR==6' | tr -d '\n ')
|
||||
mem=$(printf '%s' "$raw" | awk 'BEGIN{RS="|"} NR==7' | tr -d '\n ')
|
||||
if [ -n "$info" ]; then
|
||||
synced=$(printf '%s' "$info" | python3 -c 'import json,sys; j=json.load(sys.stdin); print(j.get("isSynced"))' 2>/dev/null || echo ?)
|
||||
ver=$(printf '%s' "$info" | python3 -c 'import json,sys; j=json.load(sys.stdin); print(j.get("serverVersion"))' 2>/dev/null || echo ?)
|
||||
fi
|
||||
if [ -n "$dag" ]; then
|
||||
read -r blocks headers sink <<< "$(printf '%s' "$dag" | python3 -c 'import json,sys; j=json.load(sys.stdin); print(j.get("blockCount"), j.get("headerCount"), str(j.get("sink",""))[:12])' 2>/dev/null || echo "? ? ?")"
|
||||
fi
|
||||
local ok=OK
|
||||
if [ "$port" != open ] || [ "$unit" != active ] || [ -z "$info" ]; then ok=FAIL; fail=1; fi
|
||||
printf '%s %-14s %-15s p2p=%s unit=%s rpc=%s synced=%s blocks=%s headers=%s sink=%s peers=%s known/banned=%s disk=%s mem=%s version=%s\n' \
|
||||
"$ok" "$name" "$ip" "$port" "${unit:-?}" "$([ -n "$info" ] && echo yes || echo no)" "$synced" "${blocks:-?}" "${headers:-?}" "${sink:-?}" "$peers" "${known:-?}" "$disk" "$mem" "$ver"
|
||||
}
|
||||
run() {
|
||||
[ -s "$SEEDS_TSV" ] || die "no seeds.tsv"
|
||||
if [ -n "$only" ]; then check_one "$only"; else for n in $(seed_names); do check_one "$n"; done; fi
|
||||
}
|
||||
if [ "${1:-}" = --watch ] || [ "${2:-}" = --watch ]; then while true; do run; sleep 60; done; else run; exit $fail; fi
|
||||
30
infra/seed-nodes/lib.sh
Executable file
30
infra/seed-nodes/lib.sh
Executable file
|
|
@ -0,0 +1,30 @@
|
|||
# Shared helpers for infra/seed-nodes. Source this.
|
||||
set -euo pipefail
|
||||
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
REPO="$(cd "$HERE/../.." && pwd)"
|
||||
export REPO
|
||||
# shellcheck source=config.sh
|
||||
. "$HERE/config.sh"
|
||||
SEEDS_TXT="$HERE/seeds.txt"
|
||||
SEEDS_TSV="$HERE/seeds.tsv"
|
||||
BUILD_DIR="$HERE/build"
|
||||
|
||||
log() { printf '%s %s\n' "$(date -u +%H:%M:%S)" "$*"; }
|
||||
die() { log "ERROR: $*" >&2; exit 1; }
|
||||
need() { command -v "$1" >/dev/null 2>&1 || die "$1 is not installed ($2)"; }
|
||||
|
||||
hetzner_auth() {
|
||||
need hcloud "brew install hcloud"
|
||||
[ -s "$HETZNER_TOKEN_FILE" ] || die "no token at $HETZNER_TOKEN_FILE"
|
||||
HCLOUD_TOKEN="$(head -1 "$HETZNER_TOKEN_FILE" | tr -d '[:space:]')"; export HCLOUD_TOKEN
|
||||
}
|
||||
|
||||
SSH_OPTS=(-i "$SSH_KEY_FILE" -o StrictHostKeyChecking=accept-new -o UserKnownHostsFile="$HERE/build/known_hosts" -o ConnectTimeout=15 -o BatchMode=yes -o ServerAliveInterval=30)
|
||||
sssh() { local ip="$1"; shift; ssh "${SSH_OPTS[@]}" "$SSH_USER@$ip" "$@"; }
|
||||
sscp() { scp -q "${SSH_OPTS[@]}" "$@"; }
|
||||
|
||||
seed_ip() { awk -F'\t' -v n="$1" '$1 == n { print $4; exit }' "$SEEDS_TSV" 2>/dev/null; }
|
||||
seed_names() { cut -f1 "$SEEDS_TSV" 2>/dev/null | grep . || true; }
|
||||
|
||||
# seeds.txt is derived from seeds.tsv: one "<ip>:26611" per seed, nothing else, so a script or a package can read it as is.
|
||||
write_seeds_txt() { awk -F'\t' -v p="$P2P_PORT" '{ print $4 ":" p }' "$SEEDS_TSV" > "$SEEDS_TXT"; }
|
||||
21
infra/seed-nodes/node/igneumd.service
Normal file
21
infra/seed-nodes/node/igneumd.service
Normal file
|
|
@ -0,0 +1,21 @@
|
|||
[Unit]
|
||||
Description=Igneum seed node (igneumd, p2p open, no mining)
|
||||
After=network-online.target chrony.service
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
User=igneum
|
||||
Group=igneum
|
||||
EnvironmentFile=/etc/igneum/seed.env
|
||||
ExecStart=/opt/igneum/bin/run-seed.sh
|
||||
Restart=always
|
||||
RestartSec=10
|
||||
LimitNOFILE=65536
|
||||
MemoryMax=3200M
|
||||
StandardOutput=journal
|
||||
StandardError=journal
|
||||
SyslogIdentifier=igneumd
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
25
infra/seed-nodes/node/install-seed.sh
Executable file
25
infra/seed-nodes/node/install-seed.sh
Executable file
|
|
@ -0,0 +1,25 @@
|
|||
#!/usr/bin/env bash
|
||||
# Runs ON the seed VM as root: install-seed.sh <name> <external-ip> "<network args>" "<peer list>"
|
||||
# Expects /opt/igneum/bin/{igneumd,igneum-miner,wrpc.py,run-seed.sh} and /root/igneumd.service.
|
||||
set -euo pipefail
|
||||
name="$1"; extip="$2"; netargs="$3"; peers="${4:-}"
|
||||
export DEBIAN_FRONTEND=noninteractive
|
||||
command -v chronyd >/dev/null 2>&1 || { apt-get update -qq; apt-get install -y -qq chrony python3 >/dev/null; }
|
||||
systemctl enable --now chrony >/dev/null 2>&1 || true
|
||||
id igneum >/dev/null 2>&1 || useradd --system --home /var/lib/igneum --shell /usr/sbin/nologin igneum
|
||||
mkdir -p /var/lib/igneum /etc/igneum /opt/igneum/bin
|
||||
chmod +x /opt/igneum/bin/*
|
||||
chown -R igneum:igneum /var/lib/igneum
|
||||
cat > /etc/igneum/seed.env <<EOF
|
||||
SEED_NAME=$name
|
||||
EXTERNAL_IP=$extip
|
||||
NETWORK_ARGS=$netargs
|
||||
SEED_PEERS=$peers
|
||||
EXTRA_ARGS=
|
||||
EOF
|
||||
cp /root/igneumd.service /etc/systemd/system/igneumd.service
|
||||
systemctl daemon-reload
|
||||
systemctl enable igneumd >/dev/null 2>&1
|
||||
systemctl restart igneumd
|
||||
sleep 3
|
||||
systemctl is-active igneumd && journalctl -u igneumd --no-pager -n 5 -o cat
|
||||
17
infra/seed-nodes/node/run-seed.sh
Executable file
17
infra/seed-nodes/node/run-seed.sh
Executable file
|
|
@ -0,0 +1,17 @@
|
|||
#!/usr/bin/env bash
|
||||
# igneumd launcher on a seed VM. Reads /etc/igneum/seed.env.
|
||||
# No mining, no --connect (the address manager stays on and answers RequestAddresses, which is what a seed is for),
|
||||
# --nodnsseed (a seed never asks another seed list), UPnP off, RPC on loopback only, p2p on every interface,
|
||||
# --externalip so peers learn the fixed address, generous inbound limit. --yes answers the database-version prompt
|
||||
# so the unit never waits on a terminal. Other seeds come in through --addpeer (SEED_PEERS).
|
||||
set -euo pipefail
|
||||
. /etc/igneum/seed.env
|
||||
# shellcheck disable=SC2206
|
||||
args=($NETWORK_ARGS --appdir=/var/lib/igneum --rpclisten=127.0.0.1:26610 --rpclisten-json=127.0.0.1:28610
|
||||
--listen=0.0.0.0:26611 --externalip="$EXTERNAL_IP" --nodnsseed --disable-upnp --nologfiles --yes
|
||||
--maxinpeers=128 --outpeers=8 --loglevel=info)
|
||||
IFS=',' read -r -a peers <<< "${SEED_PEERS:-}"
|
||||
for p in "${peers[@]}"; do [ -n "$p" ] && args+=(--addpeer="$p"); done
|
||||
# shellcheck disable=SC2206
|
||||
[ -n "${EXTRA_ARGS:-}" ] && args+=($EXTRA_ARGS)
|
||||
exec /opt/igneum/bin/igneumd "${args[@]}"
|
||||
39
infra/seed-nodes/provision-seed.sh
Executable file
39
infra/seed-nodes/provision-seed.sh
Executable file
|
|
@ -0,0 +1,39 @@
|
|||
#!/usr/bin/env bash
|
||||
# Install igneumd on a seed and start it.
|
||||
# ./provision-seed.sh <seed-name> build on the seed VM from the source tarball (BUILD_WHERE=seed, default)
|
||||
# BUILD_WHERE=bin ./provision-seed.sh <name> use ../cloud-devnet/build/bin/igneumd (built by the 20-node route)
|
||||
# Same binary route as the 20-node network: ../cloud-devnet/make-source.sh exports vendor/igneum-node (HEAD plus
|
||||
# uncommitted work) and igneum-pow; builder/build-on-builder.sh compiles `-p kaspad -p igneum-miner --features
|
||||
# igneum-pow`. On a 4 GB seed the build gets a 6 GB swap file and 2 jobs; expect about an hour (approximate).
|
||||
# Other seeds already in seeds.txt become --addpeer entries, so the seeds form a full mesh among themselves.
|
||||
. "$(dirname "$0")/lib.sh"
|
||||
name="${1:-$SEED_NAME}"; ip=$(seed_ip "$name"); [ -n "$ip" ] || die "$name not in seeds.tsv"
|
||||
mkdir -p "$BUILD_DIR"
|
||||
CD="$HERE/../cloud-devnet"
|
||||
|
||||
for try in $(seq 1 20); do sssh "$ip" true >/dev/null 2>&1 && break; log "waiting for ssh ($try)"; sleep 10; done
|
||||
sssh "$ip" true || die "ssh to $ip failed"
|
||||
|
||||
if [ "$BUILD_WHERE" = bin ]; then
|
||||
[ -x "$CD/build/bin/igneumd" ] || die "no $CD/build/bin/igneumd (run ../cloud-devnet/provision.sh build)"
|
||||
log "uploading prebuilt binaries"
|
||||
sssh "$ip" 'mkdir -p /opt/igneum/bin'
|
||||
sscp "$CD/build/bin/igneumd" "$CD/build/bin/igneum-miner" "$SSH_USER@$ip:/opt/igneum/bin/"
|
||||
else
|
||||
NODE_SRC="$NODE_SRC" POW_SRC="$POW_SRC" SRC_MODE="$SRC_MODE" "$CD/make-source.sh"
|
||||
cp "$CD/build/src.stamp" "$BUILD_DIR/src.stamp"
|
||||
log "uploading the source tarball ($(du -h "$CD/build/src.tar.gz" | cut -f1)) and the build script"
|
||||
sscp "$CD/build/src.tar.gz" "$CD/builder/build-on-builder.sh" "$SSH_USER@$ip:/root/"
|
||||
log "swap and build on the seed (log: $BUILD_DIR/build-$name.log)"
|
||||
sssh "$ip" 'if [ ! -f /swapfile ]; then fallocate -l 6G /swapfile && chmod 600 /swapfile && mkswap /swapfile >/dev/null && swapon /swapfile && echo "/swapfile none swap sw 0 0" >> /etc/fstab; fi; free -m | head -3'
|
||||
sssh "$ip" 'export CARGO_BUILD_JOBS=2; bash /root/build-on-builder.sh' 2>&1 | tee "$BUILD_DIR/build-$name.log" | grep -E 'build running|build finished|error|igneumd|installing' || true
|
||||
sssh "$ip" 'mkdir -p /opt/igneum/bin && cp /root/out/igneumd /root/out/igneum-miner /opt/igneum/bin/ && /opt/igneum/bin/igneumd --version | head -1' || die "build did not produce igneumd (see $BUILD_DIR/build-$name.log)"
|
||||
fi
|
||||
|
||||
peers=$(awk -F'\t' -v n="$name" -v p="$P2P_PORT" '$1 != n { print $4 ":" p }' "$SEEDS_TSV" | paste -sd, -)
|
||||
[ -n "$SEED_PEERS" ] && peers="${peers:+$peers,}$SEED_PEERS"
|
||||
sscp "$CD/node/wrpc.py" "$HERE/node/run-seed.sh" "$SSH_USER@$ip:/opt/igneum/bin/"
|
||||
sscp "$HERE/node/install-seed.sh" "$HERE/node/igneumd.service" "$SSH_USER@$ip:/root/"
|
||||
sssh "$ip" "bash /root/install-seed.sh '$name' '$ip' '$NETWORK_ARGS' '$peers'"
|
||||
log "started. Health in 30 s:"; sleep 30
|
||||
"$HERE/health.sh" "$name" || true
|
||||
1
infra/seed-nodes/seeds.tsv
Normal file
1
infra/seed-nodes/seeds.tsv
Normal file
|
|
@ -0,0 +1 @@
|
|||
igneum-seed-1 hetzner fsn1 188.245.5.161 cx23 2026-10-03T21:51:56Z
|
||||
|
1
infra/seed-nodes/seeds.txt
Normal file
1
infra/seed-nodes/seeds.txt
Normal file
|
|
@ -0,0 +1 @@
|
|||
188.245.5.161:26611
|
||||
Loading…
Reference in a new issue