Merge gpu-fleet: the rented fleet's tooling (box library, standing supervisor and library, the 10 percent and 75 percent gates, the deploy gate, publish and canary scripts, Devnet 2 scripts and gate reader, the wind-down) and the night's analyses (prover tiers, block rate, the fleet night rows 1 to 18, the bench-log entries)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 23:23:46 +00:00
commit 4dd30d4bec
61 changed files with 3956 additions and 0 deletions

View file

@ -2597,3 +2597,39 @@ USD 20 an hour on community pods, against a devnet of 1.16 GH/s.
Consequence: the devnet's hash is rentable for the price of a dinner, so nothing on it is a security result; the counter-ASIC and
finality work is tested there for correctness, not for cost. The cost argument only starts at the TH/s scale, where the rental
market's supply (not its price) is the limit, and that number belongs in the litepaper with this caveat.
## Block rate on Devnet 2, 6 October 2026 (branch gpu-fleet): 10 blocks per second against 1 on 42 rented cards
Run A (10 blocks/s profile, star topology, 65 min): 4.87 DAG blocks/s, 1.09 blue blocks/s, 77.6 percent red, tips 250 to 660,
max reorg 55, difficulty easing all hour (6,719 to 1,307), the exec follower at 0.05 blocks/s (lag 18,901 at the end). Run B
(1 block/s, same boxes, 30 min): 1.41 blocks/s over the window with the join burst, 1.0 blocks/s and under 2 percent red from
minute six, tips 1 to 3, difficulty settled in six minutes (453 to 482 M), the exec follower at 0.46 blocks/s. The network
lane's read: run A's reds came from node throughput (61 to 345 ms CPU per accepted block at mergeset 8 to 200), not from the
star. Per tier the blue rate decides the payout interval (1.09 against 1.19 blue/s: a 4070 at 10 TH/s waits about three days
for a paying block either way), so the higher rate buys the solo miner nothing until the node processes a block in under 50 ms
at mergeset 248. Recommendation (the lane's): 1 block/s for the testnet and the launch, 10 behind three measured gates.
Full tables and sources: `docs/analysis/block-rate-devnet2.md`, rows in `~/Desktop/fleet/bps/{A,B}.jsonl`.
## The rented fleet is the devnet's finality, 6 October 2026 (branch gpu-fleet)
Measured at 21:57Z from the hub's last 2,000 blocks: the 38 wave pods held 77.8 percent of the voter weight (mean 2.05 percent
a pod), the 14 standing boxes most of the rest, the hands and the hub the remainder; the last lock signed 93.5 percent of the
active voters and 89.9 percent of the frozen table (53 of 84 voters on the first certificate). Earlier the same evening the
fleet removed 13 miners' GPUs inside three minutes and finality paused for two hours five minutes (18:39:36Z to 20:44:44Z,
42.7 percent of the table gone with earlier leavers; rule v3 holds a full window). From that came the 10 percent rule (never
remove more than 10 percent of the live devnet's weight in an hour, `tools/fleet/lib/standing.py weight_check`) and the wave's
wind-down by hourly slices (`tools/fleet/winddown.py`: slice 1 at 21:58Z took 12 pods and the 8x rig at 8.6 percent of weight).
When the wave is gone the 14 standing boxes hold about 95 percent of the weight, so from then until public hash arrives the
fleet alone is the devnet's finality: a home miner's lock lands only while the fleet is up. What holds it up: every standing
box runs under `box-standing.sh`, which restarts a dead node within one of its 60-second passes (the hub's three deaths
tonight: 63 s, 41 s and 56 s to the restart line), restarts the miner with the node, prunes the prover's exports and trims
the node log, and runs the exec recovery recipe when the state layer reads zero; `lib/standing.py loop` re-rents a dead host
in the same shape and reports a box behind its wanted binary.
| Tier | What it means |
|---|---|
| Home miner | your lock depends on 14 rented cards staying up and mining; a finality pause is not your node's fault and nothing you can fix; the rule above is what keeps it from recurring on the fleet's side |
| Rig | the same, and a rig that leaves is itself a weight removal: at 459 MH/s on tonight's devnet it is about 20 percent of the weight, over the hour's budget by itself |
| Pool | a pool node is one voter carrying its members' whole weight; a pool restart is the largest single removal on the network and must be sliced like the fleet's |
| The network | finality by miner weight is only as steady as the miners' uptime; until public hash dwarfs the fleet, the fleet's supervisor is a consensus component |

View file

@ -326,3 +326,42 @@ Earlier (shipped in 0.3.12 and 0.3.13 unless marked):
- The confirm plan's neighbour is one step; a second neighbour (the other knob) would cost 75 s more and catch a
prior that is wrong on both knobs.
- Intel: no knob yet; the row says measure only.
## 7a. One administrator approval, ever (0.3.13; the project lead, 6 October 2026, 11:50 UTC)
What 0.3.12 does: Power control on raises one prompt and sets every cap in that step; every later cap (an app start, a
reboot, a slider move) and every tune's helper is another elevated launch, so another prompt. Not "once, ever".
What `src/powertask.rs` does: the first approval's elevated step also registers a per-user Windows scheduled task,
`Igneum Power Helper` (principal = the signed-in user, interactive logon, RunLevel Highest, no trigger, hidden, one
hour limit, new starts ignored while one runs), whose action is the app's own exe in the install folder with
`--power-helper`. A task the user owns is started by the user's unelevated engine with `Start-ScheduledTask`, no
prompt, and runs elevated. Every later cap and every tune's helper starts the task and writes the command file
`<app data>/app/sweep/cmd.txt` (`<seq> dev <n>`, `<seq> pl <W>`, `<seq> lgc <MHz>`, `<seq> rgc`, `quit`). The task
survives app restarts, updates (the per-user installer replaces the exe in place; the task's action path is the
install folder) and reboots. Power control off starts the task once and sends `remove`: the helper unregisters the
task (elevated) and exits; nothing is left behind. Linux keeps pkexec per step; macOS has no cap.
Threat note: the helper runs only fixed verbs with digit-only arguments through `Command::new(nvidia-smi).args`
(the driver's own path, never PATH, never a shell); a line that is anything else is ignored; the sequence must rise
(a stale file runs nothing); an attacker running as the user gains the power limit and clock cap of the user's own
NVIDIA cards inside the driver's ranges, which the same user could set with one approved prompt anyway; no file,
process, registry key or other binary is reachable through it. Tests: `powertask::tests` (the parser refuses every
non-digit or extra argument, the arguments reach nvidia-smi as a list, the registration is per-user, highest,
trigger-less and quote-safe, a stale command file runs nothing).
## Fleet priors from rented cards (6 October 2026, branch gpu-fleet)
Measured by `tools/fleet/box-ember.sh` on Vast.ai containers, the 0.3.12 CUDA worker against the box's devnet node, 15 s settle and 60 s hold per step. `nvidia-smi -pl` and `-lgc` are refused inside the containers (the host's driver holds the knobs), so every ladder is its baseline step only: the untuned point per model, as `plan: baseline` records in `relay/lib/ember.mjs` terms (they summarise beside a prior, never set one). A tuned prior per model needs bare metal or a VM with the driver inside.
| Card | Driver | MH/s | W | MH/W | Power default W | Clock max MHz | Plan |
|---|---|---|---|---|---|---|---|
| RTX 3060 | 580.126.20 | 24.59 | 106.1 | 0.2318 | 170.0 | 2100.0 | baseline |
| RTX 4060 Ti | | 17.64 | 73.5 | 0.24 | | | baseline |
| RTX 4070 | 580.159.04 | 24.77 | 91.3 | 0.2713 | 200.0 | 3120.0 | baseline |
| RTX 4090 | 595.71.05 | 52.24 | 179.9 | 0.2904 | 450.0 | 3105.0 | baseline |
| RTX A5000 | 580.82.09 | 47.6 | 222.3 | 0.2141 | 230.0 | 2100.0 | baseline |
| RTX 5090 | 580.173.02 | 98.12 | 258.2 | 0.38 | 575.0 | 3090.0 | baseline |
| RTX 5070 | 595.84 | 56.83 | 136.4 | 0.4166 | 250.0 | 3135.0 | baseline |
The TUNE records themselves (`ember.json` per instance under `~/Desktop/fleet/<instance>/`) carry the step, the mean core and memory clock and the hottest reading; `relay/lib/ember.mjs parseRecords` reads the `TUNE {...}` line each ladder prints.

65
docs/plans/gpu-fleet.md Normal file
View file

@ -0,0 +1,65 @@
# The rented GPU fleet: the standing rule and the one-shot rule
the project lead's ruling, 6 October 2026, 19:50 UK ("keep rented cards up"), written as the fleet agent runs it. Tooling in
`tools/fleet/` on branch `gpu-fleet`; every box operation goes through `tools/fleet/lib/` (the box library) and the
box-side scripts it ships.
## Two kinds of box
| Kind | Rule | Examples |
|---|---|---|
| Standing | stays up, never destroyed on a job's end; replaced in the same shape when its host dies; node under a supervisor; version follows the live manifest; never joins an experiment | the 13 live-devnet boxes of 6 October (USD 3.50/h), the Devnet 2 set (seed, 3 miners, 2 provers) |
| One-shot | rented for one measurement, destroyed the minute its measurement is in (a running meter is a bug) | the memory-matrix cards, the 8x rigs, the wave pods, the RISC Zero box |
The standing set's size and cost: 16 on the live devnet (a mix: 5090, 4090s, 3090s, 4070, 3080, A5000, an L4, at least
two AMD when Vast reopens) plus 6 on Devnet 2, about USD 8 an hour, USD 200 a day, inside the daily budget (USD 250 to
1,000). Tonight's 13 cost USD 3.50 an hour (USD 84 a day); the L4, the AMD pair and the Devnet 2 six are owed to the
roster as providers free cards (RunPod gave no pod of 8 card types from 18:59Z, Vast refuses every new rent).
## What a standing box carries
| Piece | Where | What it does |
|---|---|---|
| `box-standing.sh` (the supervisor) | `/root/fleet/in/`, started once by `lib/standing.py install` | every 60 s restarts the node, the miner loop (CUDA worker, pack re-exported on seed change) and the prover when gone; every 10 min writes a `RESULT standing` line (uptime, node pid and binary, blocks, DAA, peers, synced, exec tip, miner rate, prover) to `/root/fleet/out/standing.log`; runs the shipper's recovery recipe (`box-exec-snapshot.sh`) when consensus is synced above 1,000 blocks and the exec tip reads 0 |
| `/root/fleet/standing.node` | the box | the node binary the supervisor runs; the supervisor adopts whatever node already runs when the file is absent (a canary's newer binary is never downgraded) and restarts the node only when a publish rewrites the pointer |
| the registry row | `~/Desktop/fleet/boxes.json` | `standing: true`, `role: live|dn2`, `standing_since`, `node_sha16_wanted` after a publish |
| `lib/standing.py` | the Mac | `roster` (role, card, price, uptime), `install`, `check` (ssh alive, supervisor up, synced, exec moving, prover, node binary against the wanted sha), `update` (the manifest's hive package onto the box, pointer rewritten), `rerent` (same card and provider, label suffixed `-r<n>`, old row marked destroyed), `loop` (check every 10 min, re-rent after two dead checks, report what is behind) |
| the fleet page | `dl.igneum.network/fleet-22adafa34bc2/` | a `standing` block (count, USD/h, USD/day, roles, the rule) and `standing`, `role`, `uptime_h` on every box row |
A publish on the standing set is a script, not the loop: `tools/fleet/publish-0315.py` is the first (the Linux igneumd,
igneum-miner and the generator-4 workers over the package's, the pointer rewritten, the read-back table of version line,
digest and worker sha per box). The loop reports a box behind its wanted sha; it does not move binaries by itself.
## Ports
A standing box should carry a mapped p2p port whenever the provider allows it (RunPod: request `26611/tcp` at rent;
Vast: none of tonight's offers mapped one), so that the live devnet stops being a star in which every rented node dials
only the two hand nodes and the hub: the 18:39Z to 19:4xZ finality pause coincided with the hands being down, and run A
(10 blocks/s on a star, 77 to 84 percent red blocks) showed the same weakness on Devnet 2. Tonight's inbound-capable
nodes were igneum-build-1 (26611) and dn2-seed's one mapped port; a true mesh needs pods rented with a p2p port.
## Clock
- 19:41Z: the 13 live boxes converted (supervisor started, every node synced, every prover up).
- 20:00Z: the 10 percent rule, after the finality pause the fleet caused (`lib/standing.py weight_check`).
- 20:42Z to 20:48Z: the disk class (the prover's exports) found on the hub's third death; the supervisor prunes, `disk-sweep.py` watches, box-prover.py caps its own directory.
- 20:58Z: Devnet 2 standing at 1 block/s: the seed on igneum-build-1 (188.40.146.49:26611), miners dn2-seed, dn2-1, dn2-2, dn2-3, provers on dn2-1 and dn2-2 (the floor host copied from p1-4090); the gate owed until the rehearsal chain ends and dn2-seed's mapped port is free.
- 21:44Z to 21:53Z: publish 1 of 0.3.15 on the 14 live standing boxes (`publish-0315.py`: node 7f0bde70 = f1ea7a38, miner, generator-4 workers; the thirteen-field object, digest b18ed271).
- 22:18Z to 22:24Z: the forced poisoned-peer confirmation on the live network (relay on 6615571c over a 1026 datadir, a 713ef876 miner into it, the f1ea7a38 target refusing six relayed 1026 blocks with the old rule's line and never itself rejected).
- 22:49Z: publish 2 (the sixteen-field object: exec_restart_state_root, class v4 floor 831,600, window 86,400) on the 14 live standing boxes by `~/Desktop/fleet/move-16.py` (file to /root/fleet/override.json, the previous kept as override-13.json, node restarted by the supervisor, miner with it); expected digest eada4bda; `~/Desktop/fleet/expected-digest` rewritten so the five-minute check posts on any box left on b18ed271. The Devnet 2 five stay on their own object (4a0b8726).
- Owed: the L4 and the AMD pair when a provider reopens; the ports on re-rent.
## The 10 percent rule (6 October 2026, 20:00Z, from the finality pause)
What happened: the class v4 rehearsal job took the GPUs of 13 live-devnet miners between 18:27Z and 18:30Z (their
live nodes stayed up and synced, but a voter's weight is its blue blocks, and a miner that stops mining stops being
a voter), and with seven earlier leavers that was 42.7 percent of the frozen voter table; finality rule v3 then holds
the pause for one full window, and the first lock after 18:39:36Z is expected at about 20:40Z. The fleet's read that
"the boxes never left the devnet" was true of the nodes and wrong about the weight.
The rule: never remove more than 10 percent of the live devnet's 30-day weight in any hour. Weight is counted by blue
blocks per key over the window, read from the hub; any experiment that borrows live miners does it in slices with an
hour between slices; a standing box's miner is never stopped (or its GPU shared with a second worker) by a job
without that check. `lib/standing.py weight_check <labels>` is the gate: it reads the hub's blue blocks per payout key
over the window, sums the share of the labels asked for, and refuses the job when the share since the last hour's
removals exceeds 10 percent; every fleet job that touches a standing box's miner calls it first.

View file

@ -221,3 +221,34 @@ the project lead, 5 October 2026: "fix everything else in the numbers tonight".
| The chosen combination | the defaults | the defaults stay: batch-log2 22 and SP1's default knobs. The one knob that moves the prover (2^16) costs a fifth of the hash rate all the time for a prover that is busy a few seconds a minute on the devnet; it is the project lead's trade, not a default (below) |
Reading. The per-block aggregation is 2.1 s and a block 4.1 s on a 5090 that only proves, 9.7 and 17.5 s on one that also mines; no knob, fold or stream on tonight's SP1 changes the first pair, and only the miner's kernel length changes the second, at 1 MH/s per 0.37 s of block time. So "under 3 s a block" and "under 1.5x" are met on a card that is not mining and are not reachable on one that is. What that means per tier: a 5090 that mines and proves delivers a proven empty block every 17.5 s (6 cards for 1 block/s), the same card proving only every 4.1 s (2 cards, plus the shard work of full blocks: the fleet table above), and a batch fold of the aggregator (a new pinned guest) would bring the proving-only card to about 2.7 s a block and the mining one to about 12 s. What is being done: the app and host defaults are left as measured; the plan's open decision for the project lead is whether a card that holds a shard assignment should drop to 2^16 for the proof's minute (1.6x faster proof, 19% of its hash rate for that minute) or whether proving-only cards carry the aggregation (the clean 2.1 s), and the batch fold goes on the next pin's list. The state class found on the way (`/api/state` answering `{}` once `paid_wei` passes u64::MAX, fixed on the app branch at 6714a45) is in the bench log with the rest.
## The fleet night (6 October 2026, from 11:50 UTC)
The rented fleet (branch `gpu-fleet`, `tools/fleet/`, raw logs under `~/Desktop/fleet/<instance>/`): 11 cards for the
memory matrix (`docs/analysis/prover-tiers-real-cards.md`), 15 extras for the prover night, a p2p hub on RunPod with its
port public, the proving agent's three RISC Zero boxes. Every box: the 0.3.12 Linux node 83089544 on the ten-field
override (digest 7bd98cc4...), the 0.3.12 workers, the patched SP1 server, the cuda host, a throwaway payout address
made on the box, a per-card key label. The prover is `tools/fleet/box-prover.py`, a port of `app/igneum-app/src/prover.rs`
(whole-segment claiming, FNV spread over the fleet, `--mode chain --save-shards [--prev]`, sign and submit per shard and
per segment, held fresh records offered again every pass).
| Row | What was found | Per tier | What is being done |
|---|---|---|---|
| 1, 13:03Z: a node that joined the devnet today never executes the chain | On every fleet node `eth_blockNumber` reads 0x0 and `igneum_getProvingStatus` tipDaa 0, v1 active=false, an hour after consensus synced (blocks = headers, synced=true, blocks flowing). At exec debug the follower says `cannot find header edc4fa84... (genesis)` on a proof-synced node and `the queried hash does not have retention root on its chain` on a node started over a copy of the observer's full datadir, archival or not: the exec state is memory only (`igneum/exec/src/service.rs`: `_db_dir` unused, `IgneumDb::genesis()` at start, the follower walks the virtual chain from genesis), and once the devnet's pruning point left genesis (today) no node can start that walk. The observer's own exec read tipDaa 0 at 12:05Z (before the fleet touched it) and node 1's reads tipDaa 0, both restarted for publish 2 at about 11:37Z: the hand nodes' exec layer has been dead since, and with it every work list the provers read | Home miner (any card, any OS): a 0.3.12 install today mines but never proves, its wallet and the explorer against its own node read zero. Rig: the same. Pool user: nothing visible until the pool's own node restarts. The devnet: proving v1 produced nothing after the publish-2 restarts; the fleet night's "before the switch" window cannot exist on 0.3.12 | The coordinator's decision (13:15Z): no genesis restart; the proving agent builds the follower rebuild (walk the stored blocks where the full history is on disk, persist the exec state), tested on a copy of node 1's datadir tonight, then the hand nodes swap binaries and the PCs take a node-only 0.3.13. The fleet keeps the hub and the observer's full-history tarball (`/root/fleet/share/observer-datadir.tgz` on the hub, sha256 e67cc649...) for the phase-2 nodes, keeps the fifteen extras building until 14:15Z, and runs phase 1 and phase 3 meanwhile |
| 2, 12:24Z: the seed drops every new node every 30 s | `P2P, route error: incoming route capacity for message type IgneumFinality has been reached (peer: 188.245.5.161:26611)` then `P2P Connected to outgoing peer 188.245.5.161:26611` at :06, :36, :06 on every fleet node; IBD through the headers proof restarts at each drop, so 4 of 11 nodes had 0 blocks after 35 minutes while the seed served 26 nodes at once; `IBD with peer 188.245.5.161:26611 completed with error: peer connection is closed` | Every joiner with the seed as its only peer syncs in pieces; a rig the same once; pools unaffected | The hub (a RunPod 4090 with 26611 public) is every fleet node's second peer; the hub itself drew 23 fleet peers within minutes through the seed's address exchange. For the node: the IgneumFinality route's capacity against the per-checkpoint burst |
| 3, 12:38Z: the patched server hangs instead of failing when a profile does not fit | The 3080 (10 GB) and the 4060 Ti 8 GB at 2^27 (a 10.3 GB allocation): the server holds the card's limit at 0% for 568 and 904 s until killed; patch v5 (prover-floor) turns it into `FLOOR abort: a device allocation failed at slop/crates/tensor/src/inner.rs:51 ... AllocError { size: 486586112 }` and exit 70 in 13 s on the 3080 (the known-failed case of its gate) | Every prover run needs a wall-clock timeout (the rig unit has one; the app's prover and this fleet's loop have one) | v5 is the server the fleet ships from here |
| 4, 15:37Z to 15:41Z: the 0.3.13 swap on the fleet | The hands switched to the 0.3.13 node (bb43e9a8) with the thirteen-field override (`exec_restart_number` 27276, `exec_restart_hash` bb45cf0d..., `exec_restart_trust_daa` 200000; digest b18ed271...) at 15:29Z; the fleet's boxes followed in two steps (`tools/fleet/box-node-swap.sh`, `swap.py`): the binary with the ten-field file (digest 7bd98cc4 kept, the exec layer blocked by design), then the file. Replay from the restart to an executed tip at the chain tip, per box: hub (RunPod 4090) 88 s, 3080 87 s, 4070-1 113 s, 4090-3 138 s, the 8x 4090 rig 138 s, A5000 163 s, 3090-4 188 s; a second set 15:45Z to 15:57Z (3090, 5090, 3090-1, 3090-2, 3090-3, 4090-1b) about 10 to 12 min each including a 438 MB datadir pull. Seven boxes failed the first file pass on a partial pre-pull of the full-history tarball (the swap now checks its sha256) | A joiner on 0.3.13 with a full-history datadir executes within 1.5 to 3 minutes; a joiner without one still cannot (the fresh-join path, a snapshot from a peer, is the next item); every tier | The swap and the replay times are the fleet's measurement for the 0.3.13 release note |
| 5, 15:42:57Z: DAA 198,000, the fresh-record rule armed; 15:43:42Z: a 229-block reorg reset every executing node | The hub's first `PoW accepted ... daa 198000` at 15:42:57Z. One-block selected-chain reorgs at 15:42:32, :43, :53 and 15:43:05Z (heights 135,065 to 135,088), then at 15:43:42Z `selected-chain reorg: 229 chain blocks removed, unwinding to height 134884` (our tip blue score 194,395, the last removed 194,117), `reorg deeper than the snapshot ring; replaying from genesis`, genesis executed, then `exec not synced: the executor is at chain block 0 and the bodies below this node's retention root are gone`: executedTip 0, persistedTip 135,028, blocked null. The same on every fleet box that was executing and on the observer and the seed (the shipper's reading). The chain ran two-sided for about a minute after the switch | Every node operator whose node was executing at 15:43Z (home miner, rig, pool) read a zero wallet and an empty work list until a restart through the exec-restart path; the miners of the 229 losing blocks lost those rewards; a deep reorg after a snapshot ring on a pruned node is the class: the fallback must be the exec-restart point, not genesis (the proving agent's item) | The fleet restarted its boxes through the exec-restart path (90 to 190 s each, the night loop `tools/fleet/night.py` re-runs it on any box whose executed tip falls to 0 for 150 s) and the provers started on the executed tip from 15:56Z |
| 6, 16:38:45Z: the first segment records after the fresh-rule switch, accepted and paid | The fleet's RTX 5090 (Vast, a restart-path 0.3.13 node, the segment host from the proving-v1 bundle, the prover loop `tools/fleet/box-prover.py` exporting from the exec restart block) claimed segment 137142..137149 at 16:35:30Z, cut 8 fixtures, ran `--mode chain --save-shards` (8 shards and 8 aggregations on a card that also mines), had 8 of 8 shard records accepted and the fresh segment record accepted at 16:38:45Z: 195.8 s claim to acceptance, aggregator share 1.5076 IGN; a 4090 (137102..137109) followed at 16:39:51Z in 211.2 s, 1.6150 IGN. The hub, a node on the hands' real state, read paidSegments 2, paidSegmentWei 3.12 IGN, pool 2 entries 2 verified, paidShards 1,482 to 1,516 at 16:42Z: carried and paid | A 24 or 32 GB card that mines earns a segment's aggregator share about every 200 s on top of its 8 shards' 90%; six such cards cover about a quarter of the chain's segments (1.8 of 7.5 a minute); 47 mining 24 GB cards or 6 proving-only ones cover it (the fleet table's arithmetic, now with a measured 200 s) | The six restart-path boxes prove through the night; the hourly rows carry segments per hour and the chain's coverage |
| 7, 16:50Z: one exec state, and the export gap on a snapshot-recovered node | The state roots at 27,276 (0xed27bb2d...), 130,272 (0xf0a762da...) and 130,273 (0x2e22e029...) are identical on a restart-path node and on a snapshot-recovered node (export segment root and eth_getBlockByNumber agree), so the two recovery paths of 16:00Z to 16:25Z are one chain state and records from either side are valid on the other (row 6 confirms it). But on a snapshot-recovered node every `igneum_exportSegments` (from 0, 27,276 or 130,272) carries 41 to 42 accounts, the exporter's port state root differs from the node's at the first segment, and from 0 the 27,276 segments below the restart carry zero roots: the prover kit cannot cut against the hands' kind of node. The fleet's first report of this (16:52Z) called it two states; the roots corrected it 10 minutes later | An operator whose node recovered through the snapshot cannot run a prover until the export carries the state (the 0.3.14 account-dump export, the proving agent's item); a node that joined through the exec-restart path proves today | The nine snapshot boxes mine without provers tonight (each prover loop had pulled a 100 MB export every 6 s for nothing); the roots above are the 0.3.14 pin's numbers |
| 8, 17:45Z to 18:45Z: finality at 78 to 93 voters (the 38-pod wave, 1,748 MH/s for USD 20.44/h) | 116 checkpoints determined, 110 locked in the hour; lock delay p50 1.30 s, p90 1.54 s, max 16.6 s (n 107); the first certificate carries p50 59 voters (max 80) and the fold lifts it to 86 of 93; lock share p50 90.1% of active voters (min 76.9%), 77.2% of all (min 67.0%); 1.65 certificate replacements per checkpoint (max 5). The detector's dry run on the window: no alert, no event, the 5090 band 108/115/128 MH/s (n 1,555). The first hour of the wave (16:50Z to 17:45Z) is void: `pgrep -f igneumd-0313` in box-wave.sh matched its own launching shell, so no wave node started for 55 minutes (about USD 20 of pods); the fix is the anchored pattern and the CI check `tools/ci/pgrep-self-match-check.sh` | Home miner: a lock lands about every second block at 1 block/s, 0.2 s later at 93 voters than at 56; what to watch past 100 voters is the certificate's size (the fold to 86 signatures), not the vote count. Rig: the same. Pool: a pool's one node votes once for all its members; its weight is the day's blue blocks, so a pool of 38 pods is one voter with 40% of the weight | the rows are the fleet night's finality baseline for the 100-voter question; the certificate-size series continues in the hour marks |
| 9, 18:14Z: pool-v0 under load, 10 members, 0 shares | igneum-pool (PPLNS, port 4463) with 10 wave pods as members: every member's GPU shares were refused `WORKER MISMATCH` because the pool's job carried no program class and the member's worker hashed class v2 while the devnet's templates are class v3; 0 shares accepted in 25 minutes, the pool's stats page alive. The pool agent rebased the pool-mode miner onto 0.3.14 (pool-v0-rebase c4c92e8: the job carries program_class and era_seed, the member re-checks shares with the template's class); the rerun with both new binaries is scheduled for tonight after the block-rate runs | Pool user: on the old pool binary every share is wasted work; on the rebased pair a member without a node takes the day and the dataset size from the pool's seeds line. Home miner and rig: unaffected (solo mining never touched the pool) | the rerun's counts (shares accepted per member inside the first vardiff interval, mismatches 0, rejected 0) go to main |
| 10, 17:25Z to 17:38Z: the 0.3.14 canary on the live devnet | the release's Linux igneumd (4c6b129d) on five restart-path boxes for ten minutes: 0 new rejected blocks on every box, exec state roots equal to the hub's at a common height, a segment record paid on the new binary, every node on the new version: PASS on all four criteria; the first run's FAIL was the gate counting the hub (whose trait is about 3 rejects a minute) and a `cut -c1-140` that truncated the digest line before the match | Operator on the restart path: the swap is a binary change with the data dir kept, seconds of downtime, no replay. Snapshot-path operator: the same binary, the snapshot recipe unchanged | the canary form (canary-next.sh) is the fleet's release gate; the 0.3.15 run is in flight on four boxes at 19:41Z |
| 11, 18:31Z to 19:35Z: the class v4 rehearsal on igneum-devnet-400 (16 nodes, 38 pods joining, one stale 0.3.13 box) | The flip by miner signal landed at DAA 1,200 (19:00:5xZ) with the identical line on every node ("epoch 2 ... seed block f0606e20..., threshold 9500 bps, 599 of 599 blue blocks"), one program id per epoch across boxes (epoch 2 0x24304f0788ea9408, epoch 3 0xcc266b4f5dbc3447, epoch 4 0x634018bab5e5f283; the Mac's igneum-pow show agrees and the v3 id differs), 0 PoW rejected all run, the stale box refused by digest at every connect (12 by 19:46Z) and refusing the sixteen-field file at parse ("unknown field program_class_v4_activation_daa"), the floor at 2,400 crossed with v4 in force and nothing to print. Two findings outside the protocol: the plan's CPU engine (5 kH/s a box) cannot move a chain whose genesis difficulty is 2^27 (the CUDA worker fixed it: 14 to 98 MH/s a box), and the 0.3.14 package's igneum-worker-cuda refuses a generator-4 pack ("not a generator version this worker runs (2 or 3)"), so the chain stood at DAA 1,202 from 19:00:5xZ until the release tree's generator-4 Linux worker (97e036e2) went on at 19:15Z; a worker restarted on its old --pack after --exit-on-seed-change also answers every job "epoch seed mismatch" until the pack is re-exported | Home miner, any card: at a class flip the old worker stops dead; the flip is safe only when the generator-4 worker ships in the hive, Mac and Windows packages before the signal window closes, and the app re-exports the pack at every seed change. Rig: the same, times eight. Pool: the pool's node decides the class; a member on an old worker mismatches every share | the cut's rule (publish 2 gated on every worker, not every node) is with the shipper; P1 PASSED is on master |
| 12, 18:39:40Z onward: the live devnet's finality paused at 93 voters | On every node read the last lock is checkpoint 6842 at 18:39:36Z; checkpoints keep being determined and no certificate is received or built anywhere after 18:39:40Z. The fleet's first reading (the hands down 18:2xZ to 19:42Z, the star topology) was wrong: the observer rows show 13 fleet keys stopped mining the live devnet between 18:27Z and 18:30Z when the rehearsal job took their GPUs (their live nodes stayed up and synced, but a voter's weight is its blue blocks), and with seven earlier leavers that was 42.7 percent of the frozen voter table; rule v3 then holds the pause for one full window, the first lock expected about 20:40Z | Home miner: a voter that stops mining stops counting within the window, and 10 percent of the weight leaving in an hour is the most the table absorbs without a pause. Rig and pool: a pool is one voter with its members' whole weight; its restart is the biggest single removal on the network | the standing-fleet rule from it (6 October 2026, 20:00Z): never remove more than 10 percent of the live devnet's 30-day weight in any hour; `lib/standing.py weight_check` gates every job that stops or shares a standing miner |
| 13, 19:14Z to 20:30Z: the 0.3.15 canary, FAIL on the first binary, the retry confounded | 713ef876 on four live boxes: every block a 0.3.15 node mined or relayed carried version 1026 (the class v4 signal bit stamped from its first block, no window set) and every 0.3.14 node answered "wrong block version: got 1026 but expected 2" and disconnected it (the hub: 45 such rejects in the first 14 minutes, 468 by 20:29Z), so a 0.3.15 node that fell behind could not re-sync (p2-3090-1: connected and dropped every 30 s for 32 minutes) and every 0.3.15 miner lost every block it found; FAIL, publish 1 held. Three more findings on the way back: (a) a 0.3.14 node whose datadir holds 1026 blocks keeps relaying them and stays a disconnected peer after the rollback, so the four canary datadirs are poisoned until wiped; (b) a pruned 0.3.14 node dies ("consensus/src/processes/sync/mod.rs:87 KeyNotFound(GhostdagCompact/0/<genesis>)") when a peer syncing a gap asks below its retention (the hub twice, 19:58Z and 20:01Z, restarted by the standing supervisor); (c) a standalone igneum-miner keeps a dead template subscription after its node restarts (templates frozen, fetch_errors climbing, no submits). The shipper's 7961c5f1 gates the stamp on publish 2's object and fixes the serving side of (b); its retry on the poisoned boxes was confounded by (a) and (c), the clean retry runs on two untouched boxes | Home miner: an update that stamps a new block version before the network accepts it is a silent death (the app shows hashing, nothing is paid); the fix is the version gate on the object plus a datadir that never held a bad block. Rig: the same, times eight. Pool: a pool node on the bad version drops every member's share from the network's view | the canary form stays the release gate; the fresh-join line from a wiped datadir is the last read |
| 14, 20:00Z: the standing fleet | the project lead's ruling (19:50 UK): rented cards stay up and are never destroyed on a job's end. 13 live-devnet boxes converted at 19:41Z (USD 3.50/h, USD 84/day), each under box-standing.sh (node, miner and prover restarted when gone, the recovery recipe on a dead exec, a status line every 10 min), lib/standing.py on the Mac (roster, check, update, re-rent in the same shape, the 10 percent weight gate), a standing block on the fleet page; the Devnet 2 six and the L4 and AMD cards owed as providers free them. The supervisor's own two faults tonight (it matched any igneumd, so it mistook the rehearsal node for the live one and restarted a canary box's dead node on the 0.3.15 file) are fixed in db58804 and 9a294b4 | Operator of a standing box: the node comes back within a minute of dying, the miner with it, and no job takes its GPU without the weight gate | docs/plans/gpu-fleet.md carries the rule |
| 15, 19:58Z, 20:01Z, 20:42Z: the hub's live node died three times | Twice on a peer's sync request below its retention ("consensus/src/processes/sync/mod.rs:87 KeyNotFound(GhostdagCompact/0/<genesis>)", while the rolled-back p2-3090-1 synced a 32-minute gap against it; the serving-side fix is in the 0.3.15 node), once on a full disk ("header_processor/processor.rs:534 IO error: No space left on device"): the prover's segment exports under /root/fleet/out/segs (50 to 500 MB a segment, never pruned by box-prover.py) had filled the hub's 60 GB and 3 to 46 GB on every standing box since 11:50Z; two boxes stood at 100 percent. The standing supervisor restarted the node each time (19:59:11Z, 20:02:36Z, 20:44:38Z; the third after 38 GB were freed by hand) and now prunes exports older than 20 minutes and trims the node log every ten minutes (353cc5f); a disk sweep every five minutes writes disk per box to the fleet page and posts one #incidents line per box per hour at 85 percent (disk-sweep.py); the exporter-side cap is the proving lane's | Home miner: the hub is one of three public peers a fresh node dials; a dead hub means a slower first join and nothing lost; a home node's own disk is not at risk (the app's prover does not write exports). Rig: the same. Pool: a pool node that exports segments for its provers has the same disk clock | the three deaths' lines are in the hub's node.log (preserved on the box) and the finality pulls under the scratchpad |
| 16, 19:17Z to 20:58Z: the block rate on Devnet 2, 10 blocks a second against 1 | Run A (the fork's 10 blocks/s profile, fresh genesis, 42 cards, the seed on igneum-build-1, every miner dialling the seed only): 4.87 DAG blocks/s but 1.09 blue blocks/s, 77.6 percent red, tips 250 to 660, difficulty easing all hour, the exec follower at 0.05 blocks/s. Run B (1 block/s, same boxes): 1.0 blocks/s and under 2 percent red from minute six, tips 1 to 3, difficulty settled in six minutes, the follower at 0.46 blocks/s. The network lane's read: the reds came from node throughput (61 to 345 ms of CPU per accepted block), not the star | Home miner: the payout interval follows the blue rate, which the profile did not move (1.09 against 1.19 blue/s), so a 4070 at a 10 TH/s network waits about three days for a paying block at either rate; the pool, not the block rate, is the small card's shorter wait. Rig: 4 to 5 hours at 10 TH/s either way. Everyone: a wallet or prover on a 10 blocks/s chain would read state hours behind within the first hour at tonight's follower rate | `docs/analysis/block-rate-devnet2.md`: 1 block/s for the testnet and the launch, 10 behind three measured gates |
| 17, 21:52Z: two miners on one GPU after a restart | The read-back after publish 1 showed miner_up=2 on p1-4090, p1-a5000 and p2-4090-3: the publish killed the miner once, the supervisor's miner loop restarted it within ten seconds, and the supervisor's main loop, reading "no miner" in the same gap, started a second loop; two igneum-miner processes then shared one GPU at half rate each. The supervisor's miner loop now kills any other miner and worker before it starts its own (409bc3d); redeployed on the 14 standing boxes at 21:55Z | Home miner: the app owns one miner per GPU and never sees this; a hand-run box with two miner loops halves its rate silently and the only sign is the STATUS line's MH/s. Rig: times eight. Pool: a member with two miners doubles its share submissions at half the rate each, the pool sees one member with a jittery rate | one miner per GPU is now the supervisor's invariant, not the operator's care |
| 18, 21:43Z to 21:56Z: two provers on one Devnet 2 box killed each other's GPU server | Repeated by-hand relaunches of box-prover.py on dn2-1 and dn2-2 left two instances on a box: each relaunch's "kill the server, remove the socket" took the other instance's SP1 GPU server away mid-proof ("CudaClientError: early eof" on every chain step from 21:43Z) and both wrote the same prover-state.json.tmp, so one lost the file (FileNotFoundError at the os.replace). box-prover.py now holds a pid file under its out directory and a second instance exits at once, and each process writes its own tmp (c29c6b9); the Devnet 2 provers relaunched clean at 21:54Z | Home miner: the app's prover is one process by construction. Operator of a hand-run prover box: start the prover once; a second start now refuses with the first's pid instead of taking its GPU server down | the Devnet 2 gate's "paid segments" read waits on the first record from these provers |

7
tools/ci/README.md Normal file
View file

@ -0,0 +1,7 @@
# CI checks
| Check | What it fails | Since |
|---|---|---|
| `pgrep-self-match-check.sh` | a `pgrep -f` / `pkill -f` with a bare literal pattern, or `ps \| grep <word>` without a bracket or `grep -v grep`, in tools/, relay/playbooks/, infra/ or packaging/: the pattern matches the shell that runs it (the wave script of 6 October 2026 never started a node on 38 cards because `pgrep -f igneumd-0313` saw the launching shell; a kill file killed its caller the same day). Anchor to the executable's path, bracket the first letter, or use `-x`. Owed (allow-listed, finished measurements): `tools/prover-floor/pc2-*.ps1`, `tools/proving-v1/pc2-*.ps1` (their `pkill -f sp1-gpu-server` becomes `pkill -x`), `tools/repo/fresh-repo.sh:223`, `tools/observer/autosync.sh:22`, and the live devnet's operational scripts the fleet agent does not own: `infra/devnet/restart-hand-nodes.sh`, `infra/seed-nodes/addpeer-from-mac.sh`, `relay/playbooks/shard-test.ps1`, `tools/ci/fixtures/bash-body-ok.ps1`, `tools/ci/pgrep-self-match-check.sh`, `tools/ci/prover-socket-check.sh`, `tools/exec-attacks/net.sh` (their owners anchor the pattern when next touched; the hand-node and seed scripts run tonight and were not edited blind) | 6 October 2026, branch gpu-fleet |
| kill by exact command or pid file (owed as a check) | 6 October 2026, 21:09Z: a Mac-side `pkill -f <log file name>` matched nothing (the log name was a redirect, not part of the command line), the roll-everything script lived on and wiped a box it had been told to hold. Rule: a job is stopped by its pid file (`tools/fleet/fleet-bg.sh start|stop <name>`) or by a pattern anchored on its exact command line (`^python3 -u /root/fleet/in/box-prover.py`), never by a word that may or may not appear in it. The check that flags a `pkill -f`/`pgrep -f` whose literal is a path or a name that never starts a command line is owed to the CI lane |

View file

@ -0,0 +1,53 @@
#!/usr/bin/env bash
# The self-matching process-pattern class (6 October 2026). Three times in one day a script matched its own shell:
# the shipper's recovery at 16:1xZ, the fleet's wave script at 16:25Z (`pgrep -f igneumd-0313 || start the node` matched
# the launching shell's command line, which carried the file name, so no wave pod ever started its node and 38 cards
# hashed against nothing for an hour), and the fleet's Devnet 2 kill step at 17:18Z (`pkill -f '^bash in/box-dn2.sh'`
# inside a file the same script called killed the caller). Rule: a `pgrep -f`, `pkill -f` or `ps ... | grep` whose
# pattern is a literal word matches every process whose command line carries that word, including the shell that
# runs the pattern and any ssh command that carries the script's text; the pattern must therefore exclude itself:
# anchored to the executable's path (`'^/opt/igneum/pkg/bin/igneumd'`), the bracket form (`'[i]gneumd'`), or
# `pgrep -x <name>` / `pkill -x <name>` on the binary name (15 characters at most). This check fails CI when a script
# under tools/, relay/playbooks/, infra/ or packaging/ runs pgrep -f / pkill -f with a bare literal pattern (no `^`,
# no bracket, no `$`), or pipes `ps` into `grep <word>` without a bracket or a `grep -v grep`.
# With file arguments it checks those files only; --self-test runs the two fixtures.
set -euo pipefail
cd "$(dirname "$0")/../.."
fail=0
bad_pattern() { # the pattern text between the quotes after -f; prints 1 when it is a bare literal
local p="$1"
[[ "$p" == ^* || "$p" == *'['* || "$p" == *'$' || "$p" == '$'* ]] && return 1
return 0
}
check_file() {
local f="$1" n=0
while IFS= read -r line; do
n=$((n + 1))
[[ "$line" =~ ^[[:space:]]*# ]] && continue
# pgrep -f / pkill -f with a quoted or bare pattern
while read -r pat; do
[ -z "$pat" ] && continue
if bad_pattern "$pat"; then echo "pgrep-self-match: $f:$n: p(grep|kill) -f with the bare pattern '$pat' matches the shell that runs it; anchor it (^/path), bracket it ([x]rest) or use -x"; fail=1; fi
done < <(printf '%s\n' "$line" | grep -oE "p(grep|kill)( -[0-9A-Za-z]+)* -f(a|c|l)? +(\"[^\"]*\"|'[^']*'|[^ |;)]+)" | sed -E "s/^p(grep|kill)( -[0-9A-Za-z]+)* -f[acl]* +//; s/^[\"']//; s/[\"']$//")
# ps | grep word
if printf '%s\n' "$line" | grep -qE 'ps [^|]*\| *grep ' && ! printf '%s\n' "$line" | grep -qE "grep +(-[a-zA-Z]+ +)*['\"]?\[" && ! printf '%s\n' "$line" | grep -q 'grep -v grep'; then
echo "pgrep-self-match: $f:$n: ps | grep without a bracket pattern or 'grep -v grep' matches the grep itself"; fail=1
fi
done < "$f"
}
if [ "${1:-}" = "--self-test" ]; then
t="$(mktemp -d)"
printf 'pgrep -f igneumd-0313 >/dev/null || start\npkill -f "bash in/box-x.sh"\nps aux | grep igneumd\n' > "$t/bad.sh"
printf "pgrep -f '^/opt/igneum/pkg/bin/igneumd' || start\npkill -x igneum-miner\npkill -f '[i]gneumd-0313'\nps aux | grep '[i]gneumd'\nps -eo cmd | grep igneumd | grep -v grep\n" > "$t/good.sh"
fail=0; check_file "$t/bad.sh"; [ "$fail" = 1 ] || { echo "pgrep-self-match: self-test FAILED: the bad fixture passed"; exit 1; }
fail=0; check_file "$t/good.sh"; [ "$fail" = 0 ] || { echo "pgrep-self-match: self-test FAILED: the good fixture was flagged"; exit 1; }
echo "pgrep-self-match: self-test ok (the bad fixture fails, the good one passes)"; exit 0
fi
# Owed, not exempt: the PC 2 playbooks of 5 and 6 October use `pkill -f sp1-gpu-server` (the prover-socket check's own
# required line) inside a WSL `bash -c` whose command line carries the word, so the pkill kills that shell too when it
# runs first; they are finished measurements and get `pkill -x sp1-gpu-server` when next touched (tools/ci/README.md).
ALLOW='^(tools/prover-floor/pc2-.*\.ps1|tools/proving-v1/pc2-.*\.ps1|tools/repo/fresh-repo\.sh|tools/observer/autosync\.sh|infra/devnet/restart\-hand\-nodes\.sh|infra/seed\-nodes/addpeer\-from\-mac\.sh|relay/playbooks/shard\-test\.ps1|tools/ci/fixtures/bash\-body\-ok\.ps1|tools/ci/pgrep\-self\-match\-check\.sh|tools/ci/prover\-socket\-check\.sh|tools/exec\-attacks/net\.sh)$'
list_files() { if [ $# -gt 0 ]; then printf '%s\n' "$@"; else git ls-files 'tools/**' 'relay/playbooks/**' 'infra/**' 'packaging/**' | grep -E '\.(sh|bash|ps1|mjs|py)$'; fi; }
while IFS= read -r f; do [ -f "$f" ] || continue; [[ "$f" =~ $ALLOW ]] && continue; check_file "$f"; done < <(list_files "$@")
[ "$fail" = 0 ] && echo "pgrep-self-match: no script matches its own shell"
exit $fail

61
tools/fleet/autorun.py Normal file
View file

@ -0,0 +1,61 @@
#!/usr/bin/env python3
"""The fleet's orchestrator loop (runs on the Mac in the background, one pass a minute): advances every live box
through its stages without a human, pulls the results of each finished stage into ~/Desktop/fleet/<instance>/, runs the
collector after a matrix or an Ember ladder lands, and refreshes the fleet page every 5 minutes.
phase 1: setup_done -> box-matrix.sh -> matrix_done -> box-ember.sh -> ember_done -> box-prover.sh (joins phase 2)
phase 2: setup_done -> box-prover.sh
setup_failed: retried once (the toolchain CDN class), then marked failed and left for a human
Writes ~/Desktop/fleet/autorun.log; every change is a line there and in the page log.
"""
import json, os, sys, time, subprocess, datetime
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import fleet
ROOT = fleet.ROOT; LOG = os.path.join(ROOT, "autorun.log")
def now(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def log(s, page=False):
line = f"{now()} {s}"; open(LOG, "a").write(line + "\n"); print(line, flush=True)
if page: subprocess.run([sys.executable, os.path.join(fleet.HERE, "page.py"), "log", s])
def probe(b):
rc, out, err = fleet.ssh(b, "tail -1 /root/fleet/setup.log 2>/dev/null | grep -o '^RESULT setup_[a-z]*'; grep -c . /root/fleet/out/matrix.log 2>/dev/null; grep -o '^RESULT matrix_[a-z]*' /root/fleet/out/matrix.log 2>/dev/null | tail -1; grep -c . /root/fleet/out/ember.log 2>/dev/null; grep -o '^RESULT ember_done' /root/fleet/out/ember.log 2>/dev/null | tail -1; pgrep -c -f '^(bash in/box-|python3 -u /root/fleet/in/box-prover)'; grep -E '^RESULT (point|miner|step|choice|claim|submitted|paid|segment_refused|seg [0-9]+ (chain|shards))' /root/fleet/out/matrix.log /root/fleet/out/ember.log /root/fleet/out/prover.log 2>/dev/null | tail -1 | cut -c1-200", timeout=40)
if rc != 0 and not out.strip(): return None
l = out.split("\n")
g = lambda i: l[i].strip() if i < len(l) else ""
return {"setup": g(0), "matrix_lines": g(1), "matrix": g(2), "ember_lines": g(3), "ember": g(4), "running": g(5), "last": g(6)}
def start(b, script, env=""):
fleet.run(script, [b["label"]]) if not env else fleet.run_env(script, [b["label"]], env)
last_publish = 0; retried = set()
while True:
reg = fleet.load(); changed = False
live = [(iid, b) for iid, b in reg.items() if b.get("state") not in ("destroyed", "failed") and b.get("ssh_ok") and b.get("phase") != "5"]
from concurrent.futures import ThreadPoolExecutor
with ThreadPoolExecutor(max_workers=16) as ex: probes = dict(zip([i for i, _ in live], ex.map(lambda x: probe(x[1]), live)))
for iid, b in live:
p = probes.get(iid)
if p is None: continue
stage = b.get("stage", "setup")
if p["last"]: fleet.patch(iid, last_line=p["last"])
if p["running"] not in ("0", ""): # a stage script is running on the box: never start another (the 5070 ran two matrices at once, 12:23Z)
if stage == "setup" and p["setup"] == "RESULT setup_done": fleet.patch(iid, stage="matrix" if b["phase"] == "1" else "prover", state="running")
continue
if stage == "setup":
if p["setup"] == "RESULT setup_done":
nxt = "matrix" if b["phase"] == "1" else "prover"
fleet.patch(iid, stage=nxt, state="running", doing=("phase 1 matrix: idle, miner, stock server, patched server alone and beside the miner" if nxt == "matrix" else "phase 2: node + miner + segment prover on the devnet"))
fleet.run("box-matrix.sh" if nxt == "matrix" else "box-prover.sh", [b["label"]]); log(f"{b['label']}: setup done, {nxt} started", page=True); changed = True
elif p["setup"] == "RESULT setup_failed":
if iid not in retried: retried.add(iid); fleet.setup([b["label"]]); log(f"{b['label']}: setup failed once, retried")
else: fleet.patch(iid, state="failed", doing="setup failed twice; see setup.log"); log(f"{b['label']}: setup failed twice", page=True)
elif stage == "matrix":
if p["matrix"] == "RESULT matrix_done":
fleet.pull([b["label"]]); subprocess.run([sys.executable, os.path.join(fleet.HERE, "collect.py")], capture_output=True)
fleet.patch(iid, stage="ember", doing="Ember two-knob ladder: 6 power steps, 4 clock caps, 75 s each"); fleet.run("box-ember.sh", [b["label"]]); log(f"{b['label']}: matrix done ({p['last'][:80]}), Ember ladder started", page=True); changed = True
elif p["matrix"] == "RESULT matrix_failed":
fleet.pull([b["label"]]); fleet.patch(iid, stage="ember", doing="matrix failed (see matrix.log); Ember ladder started"); fleet.run("box-ember.sh", [b["label"]]); log(f"{b['label']}: matrix FAILED ({p['last'][:80]}), Ember started", page=True)
elif stage == "ember":
if p["ember"] == "RESULT ember_done" or (p["ember_lines"].isdigit() and int(p["ember_lines"]) > 0 and "ember_failed" in p["last"]):
fleet.pull([b["label"]]); subprocess.run([sys.executable, os.path.join(fleet.HERE, "collect.py")], capture_output=True)
fleet.patch(iid, stage="prover", phase="2", doing="phase 2: node + miner + segment prover on the devnet"); fleet.run("box-prover.sh", [b["label"]]); log(f"{b['label']}: Ember done ({p['last'][:80]}), prover started", page=True); changed = True
if time.time() - last_publish > 300 or changed:
subprocess.run([sys.executable, os.path.join(fleet.HERE, "page.py"), "publish"], capture_output=True); last_publish = time.time()
time.sleep(60)

38
tools/fleet/box-datadir.sh Executable file
View file

@ -0,0 +1,38 @@
#!/usr/bin/env bash
# The joiner stop-gap of 6 October 2026 (docs/analysis/prover-tiers-real-cards.md, the fleet night): a node that synced
# through the headers proof holds no genesis header, so the in-memory exec state never rebuilds and eth_blockNumber
# stays 0. This swaps the box's consensus datadir for a copy of the observer's full-history datadir (pulled from the
# hub over the throwaway fleet-internal key), restarts the node with the hub and the seed as peers and the proof
# verifier, and reports the follower's replay: eth_blockNumber every 15 s until it reaches the node's DAA tip.
set -uo pipefail
F=/root/fleet; OUT=$F/out; B=/opt/igneum/pkg/bin; FLOOR=/opt/igneum-floor; HOST=$FLOOR/bin/igneum-prove-host
HUB_SSH="${HUB_SSH:-213.173.107.74}"; HUB_PORT="${HUB_PORT:-16515}"; HUB_PEER="${HUB_PEER:-213.173.107.74:16516}"
mkdir -p $OUT; exec >> $OUT/datadir.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
echo "RESULT datadir_start $(stamp)"
t0=$(date +%s)
if [ ! -s $F/observer-datadir.tgz ]; then
scp -i $F/in/fleet-internal -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -P "$HUB_PORT" "root@$HUB_SSH:/root/fleet/share/observer-datadir.tgz" $F/observer-datadir.tgz || { echo "RESULT datadir_failed pull"; exit 2; }
fi
echo "RESULT datadir_pulled $(stamp) bytes=$(stat -c %s $F/observer-datadir.tgz) s=$(( $(date +%s) - t0 ))"
bash $F/in/box-kill.sh >/dev/null 2>&1 # every stage process; never the node (next line)
pkill -x igneumd; sleep 4; pkill -9 -x igneumd 2>/dev/null; sleep 1
rm -rf $F/node.proof && mv $F/node $F/node.proof && mkdir -p $F/node && tar -C $F/node -xzf $F/observer-datadir.tgz || { echo "RESULT datadir_failed untar"; exit 2; }
echo "RESULT datadir_swapped $(stamp) $(du -sh $F/node | cut -f1)"
IGNEUM_PROOF_VERIFIER=$HOST nohup $B/igneumd --devnet --appdir=$F/node --rpclisten=127.0.0.1:26610 --evm-rpclisten=127.0.0.1:26790 --listen=0.0.0.0:26611 \
--addpeer=$HUB_PEER --addpeer=188.245.5.161:26611 --override-params-file=$F/override.json --nodnsseed --disable-upnp --nologfiles --yes >> $F/node.log 2>&1 &
t1=$(date +%s)
bn() { curl -s -m 8 -X POST -H 'Content-Type: application/json' --data '{"jsonrpc":"2.0","id":1,"method":"eth_blockNumber","params":[]}' http://127.0.0.1:26790/ | grep -o '"result":"[^"]*"' | cut -d'"' -f4; }
for i in $(seq 1 240); do
sleep 15
h="$(bn)"; w="$($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1)"
n=$(( ${h:-0x0} ))
echo "RESULT replay $(stamp) s=$(( $(date +%s) - t1 )) evm_block=$n $w"
if [ "$n" -gt 0 ] && [[ "$w" == *synced=true* ]]; then
st="$(curl -s -m 8 -X POST -H 'Content-Type: application/json' --data '{"jsonrpc":"2.0","id":1,"method":"igneum_getProvingStatus","params":[]}' http://127.0.0.1:26790/ | python3 -c 'import sys,json; d=json.load(sys.stdin).get("result",{}); print("tipDaa", int(d.get("tipDaa","0x0"),16), "active", d.get("v1",{}).get("active"), "fresh", d.get("v1",{}).get("freshRuleActive"))' 2>/dev/null)"
daa=$(printf '%s' "$w" | grep -o 'daa=[0-9]*' | cut -d= -f2)
tip=$(printf '%s' "$st" | awk '{print $2}')
if [ -n "$tip" ] && [ "$tip" -ge $(( ${daa:-0} - 20 )) ]; then echo "RESULT datadir_done $(stamp) replay_s=$(( $(date +%s) - t1 )) total_s=$(( $(date +%s) - t0 )) $st"; exit 0; fi
fi
done
echo "RESULT datadir_failed replay did not reach the tip in 60 min"

36
tools/fleet/box-dn2.sh Executable file
View file

@ -0,0 +1,36 @@
#!/usr/bin/env bash
# A Devnet 2 box (6 October 2026, the project lead's standing structure: the rented fleet is the staging chain every release and
# activation crosses before the live devnet). The node runs igneum-devnet-2 (--devnet --devnet-suffix=2: own handshake
# magic, own data directory, a live-devnet peer refuses it at the handshake) with /root/fleet/dn2-override.json (its own
# genesis bits, every activation at a low DAA, NO exec-restart fields: a fresh chain executes from genesis), peered with
# the Devnet 2 seed; one miner on card 0 with its vote key; PROVER=1 adds the segment prover loop (box-prover.py with
# EXPORT_FROM=0 and the Devnet 2 chain name). NODE_BIN names the igneumd to run (the gate swaps it).
set -uo pipefail
F=/root/fleet; OUT=$F/out; mkdir -p $F/in $OUT $F/dn2 $F/mine/packs; exec >> $OUT/dn2.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
LABEL="${LABEL:-dn2}"; WALLET="${WALLET:-0x1919191919191919191919191919191919191919}"; SEED="${SEED:-}"; NODE_BIN="${NODE_BIN:-$F/in/igneumd-0313}"; # UNSYNCED=1 only for a fresh genesis (a node that mines while unsynced forks at every restart: Devnet 2 reorgs after the 22:1xZ restarts)
PROVER="${PROVER:-0}"; BPS="${BPS:-}"; UNSYNCED="${UNSYNCED:-}"; RPC_PORT="${RPC_PORT:-26610}"; P2P_PORT="${P2P_PORT:-26611}"; EVM_PORT="${EVM_PORT:-26790}" # other ports on a box whose live node holds 26610/26611 (the wave pods in run A) # BPS=10|5 -> the devnet2-bps fork's IGNEUMD_DEVNET_BPS profile (block-rate run A); unset = 1 block/s
echo "RESULT dn2_start $(stamp) label=$LABEL node=$(sha256sum $NODE_BIN | cut -c1-16) seed=${SEED:-none} prover=$PROVER"
command -v curl >/dev/null || { apt-get update -qq >/dev/null 2>&1; apt-get install -y -qq curl ca-certificates python3 >/dev/null 2>&1; }
if [ ! -x /opt/igneum/pkg/bin/igneum-miner ]; then
read -r PKG_PATH PKG_SHA PKG_VER <<< "$(curl -fsSL -m 30 https://dl.igneum.network/dl/public/igneum-downloads.json | python3 -c 'import sys,json; d=json.load(sys.stdin)["files"]["miner-hive"]; print(d["path"], d["sha256"], d["version"])')"
curl -fsSL -o $F/pkg.tgz "https://dl.igneum.network$PKG_PATH" && echo "$PKG_SHA $F/pkg.tgz" | sha256sum -c - >/dev/null && mkdir -p /opt/igneum/pkg && tar -C /opt/igneum/pkg --strip-components=1 -xzf $F/pkg.tgz || { echo "RESULT dn2_failed package"; exit 2; }
fi
B=/opt/igneum/pkg/bin; cp $F/in/dn2-override.json $F/dn2-override.json
pkill -9 -f '^/root/fleet/in/igneumd-(0313|[0-9a-f]{16}) ' 2>/dev/null; pkill -9 -f "^/opt/igneum/pkg/bin/igneum-miner mine grpc://127.0.0.1:$RPC_PORT " 2>/dev/null; sleep 2 # the Devnet 2 node only, never igneumd-v4 (the rehearsal node beside it, 19:00Z)
[ "${FRESH:-0}" = 1 ] && { rm -rf $F/dn2; mkdir -p $F/dn2; [ -s $F/dn2-node.log ] && mv $F/dn2-node.log $F/dn2-node.prev.log; echo "RESULT dn2_fresh $(stamp) appdir wiped for a new genesis"; } # never the script's own pattern (17:18Z: dn2-kill.sh killed its caller)
PEER=""; for sd in ${SEED//,/ }; do PEER="$PEER --addpeer=$sd"; done # SEED may be a comma list (run A2: two relays per box)
# the statement's program ids (proving agent, 22:1xZ): on 4c6b129d they come from the environment or the verifier host, else zero, and a zero id refuses every segment record
env ${BPS:+IGNEUMD_DEVNET_BPS=$BPS} IGNEUM_PROOF_PROGRAM_IDS="${PROGRAM_IDS:-0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a,0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896}" IGNEUM_PROOF_VERIFIER=/opt/igneum-floor/bin/igneum-prove-host nohup $NODE_BIN --devnet --devnet-suffix=2 --appdir=$F/dn2 --rpclisten=0.0.0.0:$RPC_PORT --evm-rpclisten=127.0.0.1:$EVM_PORT --listen=0.0.0.0:$P2P_PORT $PEER --override-params-file=$F/dn2-override.json --nodnsseed --disable-upnp --nologfiles --yes ${UNSYNCED:+--enable-unsynced-mining} ${NODE_EXTRA:-} >> $F/dn2-node.log 2>&1 &
# (--enable-unsynced-mining: a fresh chain's nodes start unsynced and must mine anyway, Reject(IsInIBD) on the seed at 16:52Z; a comment put inside this line at 17:00Z swallowed the redirect and the ampersand, so the node ran in the foreground and the script never reached the miner)
sleep 10
echo "RESULT dn2_node $(stamp) pid=$(pgrep -f '^/root/fleet/in/igneumd-(0313|[0-9a-f]{16}) ' | head -1) version=$($NODE_BIN --version 2>&1 | head -1) digest=$(grep -o 'digest: [0-9a-f]*' $F/dn2-node.log | tail -1 | awk '{print substr($2,1,16)}') network=$(grep -oiE 'igneum-devnet-2[^ ,]*' $F/dn2-node.log | head -1) genesis=$(grep -oiE 'genesis [0-9a-f]{16}' $F/dn2-node.log | head -1)"
for i in $(seq 1 30); do w="$($B/igneum-miner watch 1 grpc://127.0.0.1:$RPC_PORT 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1)"; [ -n "$w" ] && break; sleep 5; done
echo "RESULT dn2_watch $(stamp) $(printf '%s' "$w" | sed -E 's/difficulty=[0-9.]* sink=[0-9a-f]* //')"
cd $F/mine && rm -rf packs/dn2 && $B/igneum-miner export-pack grpc://127.0.0.1:$RPC_PORT packs/dn2 > $OUT/dn2-export-pack.log 2>&1
( while :; do $B/igneum-miner mine grpc://127.0.0.1:$RPC_PORT 1 100000000 "$LABEL" --worker $B/igneum-worker-cuda --worker-args "--device 0 --pack packs/dn2" --prepare-packs packs/dn2-prepare --exit-on-seed-change --evm-address "$WALLET" --payout-label "$LABEL" --status-secs 30 >> $OUT/dn2-miner.log 2>&1; rc=$?; echo "RESULT dn2_miner_exit $(stamp) rc=$rc" >> $OUT/dn2.log; [ $rc = 42 ] && { rm -rf packs/dn2; $B/igneum-miner export-pack grpc://127.0.0.1:$RPC_PORT packs/dn2 >> $OUT/dn2-export-pack.log 2>&1; } || sleep 10; done ) &
echo "RESULT dn2_miner_started $(stamp)"
if [ "$PROVER" = 1 ] && [ -x /opt/igneum-floor/bin/igneum-prove-host ]; then
cd $F && LABEL="$LABEL" WALLET="$WALLET" EXPORT_FROM=0 CHAIN_NAME=igneum-devnet-2 MINER=none RUN_HOURS=48 setsid nohup python3 -u $F/in/box-prover.py </dev/null >> $OUT/dn2-prover-launch.log 2>&1 &
echo "RESULT dn2_prover_started $(stamp)"
fi

86
tools/fleet/box-ember.sh Executable file
View file

@ -0,0 +1,86 @@
#!/usr/bin/env bash
# Ember Tune's two-knob ladder on a rented NVIDIA card (docs/plans/ember-tune.md, branch ember-tune): the miner runs
# throughout; the power ladder 100, 90, 80, 70, 60, 50% of the default limit at the unlocked clock (clamped at the
# card's reported minimum), then the clock ladder 90, 80, 70, 60% of the maximum graphics clock at the power the
# first ladder chose; 15 s settle and 60 s hold per step; the choice is the best MH/W among the steps whose rate is
# within 1% of the fastest. Linux root: nvidia-smi -pl and -lgc, no prompt. Every step prints a RESULT line; the
# ladder and the choice go to /root/fleet/out/ember.json as a TUNE-shaped record (relay/lib/ember.mjs parseRecords).
set -uo pipefail
F=/root/fleet; OUT=$F/out; LOG=$OUT/ember.log; B=/opt/igneum/pkg/bin
LABEL="${LABEL:-box}"; WALLET="${WALLET:-0x1919191919191919191919191919191919191919}"
SETTLE="${SETTLE:-15}"; HOLD="${HOLD:-60}"
mkdir -p $OUT $F/mine/packs
exec > >(tee -a $LOG) 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
say() { echo "$(stamp) $*"; }
q() { nvidia-smi --query-gpu="$1" --format=csv,noheader,nounits -i 0 2>/dev/null | head -1 | tr -d ' '; }
NAME="$(q name)"; DRV="$(q driver_version)"; PDEF="$(q power.default_limit)"; PMIN="$(q power.min_limit)"; PMAX="$(q power.max_limit)"; CMAX="$(q clocks.max.graphics)"
pkill -x sp1-gpu-server 2>/dev/null; rm -f /tmp/sp1-cuda-*.sock
echo "RESULT start $(stamp) card=$NAME driver=$DRV power_default_w=$PDEF min_w=$PMIN max_w=$PMAX clock_max_mhz=$CMAX"
# can we set anything?
nvidia-smi -i 0 -pl "$PDEF" >/dev/null 2>&1 && PL_OK=1 || PL_OK=0
nvidia-smi -i 0 -lgc 0,"$CMAX" >/dev/null 2>&1 && LGC_OK=1 || LGC_OK=0
nvidia-smi -i 0 -rgc >/dev/null 2>&1
echo "RESULT knobs power_limit_settable=$PL_OK clock_cap_settable=$LGC_OK"
cd $F/mine
rm -rf packs/devnet; $B/igneum-miner export-pack grpc://127.0.0.1:26610 packs/devnet > $OUT/ember-export.log 2>&1
nohup $B/igneum-miner mine grpc://127.0.0.1:26610 1 100000000 "$LABEL" --worker $B/igneum-worker-cuda --worker-args "--device 0 --pack packs/devnet" \
--prepare-packs packs/prepare --exit-on-seed-change --evm-address "$WALLET" --payout-label "$LABEL" --status-secs 10 > $OUT/ember-miner.log 2>&1 &
MPID=$!; cd $F
cleanup() { nvidia-smi -i 0 -rgc >/dev/null 2>&1; [ "$PL_OK" = 1 ] && nvidia-smi -i 0 -pl "$PDEF" >/dev/null 2>&1; kill $MPID 2>/dev/null; pkill -f '^/opt/igneum/pkg/bin/igneum-worker-cuda' 2>/dev/null; }
trap cleanup EXIT
say "miner warming 90 s"; sleep 90
STEPS=$OUT/ember-steps.jsonl; : > $STEPS
step() { # <label> <power_pct> <limit_w> <clock_cap_mhz or 0>
local lab="$1" pct="$2" lim="$3" cap="$4"
if [ "$PL_OK" = 1 ]; then nvidia-smi -i 0 -pl "$lim" >/dev/null 2>&1 || say "could not set -pl $lim"; fi
if [ "$cap" != 0 ]; then nvidia-smi -i 0 -lgc 0,"$cap" >/dev/null 2>&1 || say "could not set -lgc $cap"; else nvidia-smi -i 0 -rgc >/dev/null 2>&1; fi
sleep "$SETTLE"
local n0; n0="$(grep -c 'STATUS' $OUT/ember-miner.log)"
( while :; do nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,temperature.gpu --format=csv,noheader,nounits -i 0 2>/dev/null; sleep 1; done ) > $OUT/ember-samp-$lab.csv & local sp=$!
sleep "$HOLD"; pkill -P $sp 2>/dev/null; kill $sp 2>/dev/null; sleep 1; kill -9 $sp 2>/dev/null
local n1; n1="$(grep -c 'STATUS' $OUT/ember-miner.log)"
local mhs; mhs="$(grep STATUS $OUT/ember-miner.log | grep -o ' now=[0-9.]*' | tail -n $((n1 - n0 > 0 ? n1 - n0 : 1)) | cut -d= -f2 | awk '{s+=$1; n++} END {if (n) printf "%.2f", s/n; else print 0}')"
local w gclk mclk tmax; read -r w gclk mclk tmax <<< "$(awk -F', *' '{w+=$1; g+=$2; m+=$3; if ($4+0 > t) t=$4+0; n++} END {if (n) printf "%.1f %.0f %.0f %d", w/n, g/n, m/n, t; else print "0 0 0 0"}' $OUT/ember-samp-$lab.csv)"
local eff; eff="$(awk -v a="$mhs" -v b="$w" 'BEGIN {if (b > 0) printf "%.4f", a / b; else print 0}')"
echo "RESULT step $lab power_pct=$pct limit_w=$lim clock_cap_mhz=$cap watts=$w mhs=$mhs mhw=$eff gclk=$gclk mclk=$mclk tmax=$tmax"
printf '%s\n' "{\"label\":\"$lab\",\"clock_mhz\":$cap,\"power_pct\":$pct,\"limit_w\":$lim,\"watts\":$w,\"mhs\":$mhs,\"eff\":$eff,\"gclk\":$gclk,\"mclk\":$mclk,\"tmax\":$tmax,\"faults\":0,\"mark\":\"ok\"}" >> $STEPS
}
# the power ladder
last=""
for pct in 100 90 80 70 60 50; do
lim="$(awk -v d="$PDEF" -v p="$pct" -v mn="$PMIN" 'BEGIN {l = d * p / 100; if (l < mn) l = mn; printf "%d", l}')"
[ "$lim" = "$last" ] && { say "power $pct% clamps to the same $lim W; skipped"; continue; }
last="$lim"; step "p$pct" "$pct" "$lim" 0
[ "$PL_OK" = 0 ] && { say "power limit not settable here; one baseline step only"; break; }
done
# choose the power point: best eff within 1% of the top rate
read -r BEST_PCT BEST_LIM <<< "$(python3 - $STEPS <<'PY'
import json, sys
s = [json.loads(l) for l in open(sys.argv[1]) if l.strip()]
top = max(x["mhs"] for x in s); ok = [x for x in s if x["mhs"] >= top * 0.99]
b = max(ok, key=lambda x: x["eff"]); print(b["power_pct"], b["limit_w"])
PY
)"
echo "RESULT power_choice pct=$BEST_PCT limit_w=$BEST_LIM"
if [ "$LGC_OK" = 1 ]; then
for cpct in 90 80 70 60; do
cap="$(awk -v c="$CMAX" -v p="$cpct" 'BEGIN {printf "%d", int(c * p / 100 / 10) * 10}')"
step "c$cpct" "$BEST_PCT" "$BEST_LIM" "$cap"
done
fi
python3 - $STEPS $OUT/ember.json "$NAME" "$DRV" "$LABEL" "$PDEF" "$CMAX" "$PL_OK" "$LGC_OK" <<'PY'
import json, sys, datetime
steps = [json.loads(l) for l in open(sys.argv[1]) if l.strip()]
name, drv, label, pdef, cmax, pl, lgc = sys.argv[3:]
top = max(x["mhs"] for x in steps); ok = [x for x in steps if x["mhs"] >= top * 0.99]
chosen = max(ok, key=lambda x: x["eff"]); base = steps[0]
rec = {"ts": datetime.datetime.utcnow().strftime("%Y-%m-%dT%H:%M:%SZ"), "machine": "fleet-" + label, "app": "fleet-ladder", "os": "linux", "card": name, "vendor": "nvidia",
"driver": drv, "driver_major": drv.split(".")[0], "class": "v3", "key": f"nvidia|{name}|{drv.split('.')[0]}|v3",
"plan": "full" if pl == "1" else "baseline", "steps": steps, "chosen": chosen, "before": base, "eff": chosen["eff"], "mhs": chosen["mhs"], "watts": chosen["watts"],
"power_default_w": float(pdef or 0), "clock_max_mhz": float(cmax or 0), "knobs": {"power": pl == "1", "clock": lgc == "1"}}
json.dump(rec, open(sys.argv[2], "w"), indent=1)
print("TUNE " + json.dumps(rec))
print(f"RESULT choice clock_cap_mhz={chosen['clock_mhz']} power_pct={chosen['power_pct']} limit_w={chosen['limit_w']} watts={chosen['watts']} mhs={chosen['mhs']} mhw={chosen['eff']} baseline_mhs={base['mhs']} baseline_watts={base['watts']} baseline_mhw={base['eff']} steps={len(steps)}")
PY
echo "RESULT ember_done $(stamp)"

View file

@ -0,0 +1,34 @@
#!/usr/bin/env bash
# The shipper's exec recovery (6 October 2026, 16:20Z): a box whose exec layer sits at tip 0 "from snapshot" (the genesis
# replay after the 15:43Z reorg wrote a tip-0 exec-snapshot.bin, which loads first and is served to peers) takes node 1's
# exported snapshot (tip chain block 130,272, sha256 ac101f13...): stop the node, delete the data dir's exec-snapshot.bin
# and .prev.bin, start the node as before plus --igneum-exec-snapshot=<file>,0x<sha256>; then eth_blockNumber climbs
# from 130,272 to the sink within a minute. SNAP = the file's path on this box.
set -uo pipefail
F=/root/fleet; OUT=$F/out; B=/opt/igneum/pkg/bin; FLOOR=/opt/igneum-floor; HOST=$FLOOR/bin/igneum-prove-host
SNAP="${SNAP:-$F/node1-copy-snapshot.bin}"; SNAP_SHA=ac101f13576179fd7d7f5e8ee902c9a7b6cc47730e3a3c069f389f0ca46d9221
HUB_SSH="${HUB_SSH:-213.173.107.74}"; HUB_PORT="${HUB_PORT:-16515}"; HUB_PEER="${HUB_PEER:-213.173.107.74:16516}"
mkdir -p $OUT; exec >> $OUT/exec-snapshot.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
rpc() { curl -s -m 8 -X POST -H 'Content-Type: application/json' --data "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"$1\",\"params\":[]}" http://127.0.0.1:26790/; }
echo "RESULT snap_start $(stamp)"
[ -s "$SNAP" ] || scp -i $F/in/fleet-internal -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -P "$HUB_PORT" "root@$HUB_SSH:/root/fleet/share/node1-copy-snapshot.bin" "$SNAP" || { echo "RESULT snap_failed pull"; exit 2; }
[ "$(sha256sum "$SNAP" | cut -c1-64)" = "$SNAP_SHA" ] || { echo "RESULT snap_failed sha256 $(sha256sum "$SNAP" | cut -c1-16)"; exit 2; }
chmod 644 "$SNAP"
bash $F/in/box-kill.sh >/dev/null 2>&1
pkill -f '^/opt/igneum/pkg/bin/igneumd-0313'; pkill -x igneumd; sleep 4; pkill -9 -f '^/opt/igneum/pkg/bin/igneumd-0313' 2>/dev/null; sleep 1
rm -f $F/node/igneum-devnet/datadir/evm/exec-snapshot.bin $F/node/igneum-devnet/datadir/evm/exec-snapshot.prev.bin
echo "RESULT snap_files_removed $(stamp) left=$(ls $F/node/igneum-devnet/datadir/evm/ 2>/dev/null | tr '\n' ' ')"
IGNEUM_PROOF_VERIFIER=$HOST nohup $B/igneumd-0313 --devnet --appdir=$F/node --rpclisten=127.0.0.1:26610 --evm-rpclisten=127.0.0.1:26790 --listen=0.0.0.0:26611 \
--addpeer=$HUB_PEER --addpeer=188.245.5.161:26611 --override-params-file=$F/override.json --nodnsseed --disable-upnp --nologfiles --yes "--igneum-exec-snapshot=$SNAP,0x$SNAP_SHA" >> $F/node.log 2>&1 &
t1=$(date +%s); sleep 15
echo "RESULT node_started $(stamp) pid=$(pgrep -f '^/opt/igneum/pkg/bin/igneumd-0313' | head -1) digest=$(grep -o 'digest: [0-9a-f]*' $F/node.log | tail -1 | awk '{print substr($2,1,16)}') snapshot_line=\"$(grep -E 'igneum-exec\].*snapshot' $F/node.log | tail -1 | cut -c1-140)\""
for i in $(seq 1 40); do
sleep 10; ex="$(rpc igneum_getExecStatus | python3 -c 'import sys,json; r=sys.stdin.read(); d=json.loads(r).get("result",{}) if r.strip() else {}; print(int(d.get("executedTip","0x0"),16), "blocked" if d.get("blocked") else "ok", str(d.get("startedFrom",""))[:30])' 2>/dev/null)"
w="$($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'daa=[0-9]*' | tail -1)"
echo "RESULT climb $(stamp) s=$(( $(date +%s) - t1 )) exec=$ex $w"
n=$(printf '%s' "$ex" | awk '{print $1}'); st="$(rpc igneum_getProvingStatus | python3 -c 'import sys,json; r=sys.stdin.read(); d=json.loads(r).get("result",{}) if r.strip() else {}; print(int(d.get("tipDaa","0x0"),16), d.get("paidShards"))' 2>/dev/null)"
tip=$(printf '%s' "$st" | awk '{print $1}'); daa=${w#daa=}
if [ "${n:-0}" -gt 130272 ] && [ -n "$tip" ] && [ "$tip" -ge $(( ${daa:-0} - 20 )) ]; then echo "RESULT snap_done $(stamp) s=$(( $(date +%s) - t1 )) exec=$n tipDaa=$tip paidShards=$(printf '%s' "$st" | awk '{print $2}')"; exit 0; fi
done
echo "RESULT snap_failed no climb in 400 s: exec=$ex"

28
tools/fleet/box-floor-v5.sh Executable file
View file

@ -0,0 +1,28 @@
#!/usr/bin/env bash
# The prover-floor agent's known-failed case (6 October 2026, 13:05Z): rebuild sp1-gpu-server from patch v5 (the panic
# hook that turns a failed device allocation into "FLOOR abort ..." and exit 70) and re-run the 2^27 compressed point
# alone on a card it cannot fit; report the host's failing line, the server's FLOOR abort line and the seconds.
set -uo pipefail
F=/root/fleet; OUT=$F/out; FLOOR=/opt/igneum-floor; SRC=$FLOOR/sp1; HOST=$FLOOR/bin/igneum-prove-host
export PATH="$HOME/.cargo/bin:$FLOOR/go/bin:$PATH" RUSTUP_TOOLCHAIN=stable GOPATH=$FLOOR/gopath GOCACHE=$FLOOR/gocache GOFLAGS=-mod=mod CUDA_ARCHS="${ARCHS:-86}" CARGO_TARGET_DIR=$FLOOR/target
exec >> $OUT/floor-v5.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
echo "RESULT v5_start $(stamp) patch_sha256=$(sha256sum $F/in/floor-v5.patch | cut -c1-16)"
bash $F/in/box-kill.sh >/dev/null 2>&1
cd "$SRC" && git checkout -q -- . && git clean -qfd sp1-gpu/crates >/dev/null 2>&1
git apply $F/in/floor-v5.patch || { echo "RESULT v5_failed patch does not apply"; exit 2; }
touch sp1-gpu/crates/prover_components/src/builder.rs sp1-gpu/crates/jagged_tracegen/src/lib.rs sp1-gpu/crates/server/src/server.rs sp1-gpu/crates/cuda/src/task.rs sp1-gpu/crates/server/src/main.rs
echo "RESULT v5_patched $(git diff --stat | tail -1)"
t0=$(date +%s)
cargo build --release --bin sp1-gpu-server -j $(( $(nproc) > 16 ? 16 : $(nproc) )) > $FLOOR/logs/build-server-v5.log 2>&1; rc=$?
echo "RESULT v5_build_exit $rc time_s=$(( $(date +%s) - t0 ))"
[ $rc -eq 0 ] || { grep -n -A6 '^error' $FLOOR/logs/build-server-v5.log | head -30; echo "RESULT v5_failed build"; exit 2; }
mkdir -p $FLOOR/home-v5/.sp1/bin && cp $FLOOR/target/release/sp1-gpu-server $FLOOR/home-v5/.sp1/bin/ && chmod +x $FLOOR/home-v5/.sp1/bin/sp1-gpu-server
echo "RESULT v5_server sha256=$(sha256sum $FLOOR/home-v5/.sp1/bin/sp1-gpu-server | cut -c1-16) bytes=$(stat -c %s $FLOOR/home-v5/.sp1/bin/sp1-gpu-server)"
pkill -9 -x sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock; sleep 2
t1=$(date +%s)
env HOME=$FLOOR/home-v5 SP1_PROVER=cuda RUST_LOG=off SP1_GPU_FLOOR_LOG=1 SP1_GPU_ELEMENT_THRESHOLD=134217728 timeout 900 $HOST $FLOOR/prove/proving/fixtures/fees-v1-shards2.json --mode compressed --shard 0 --out $OUT/v5-point.json > $OUT/v5-point.log 2>&1; rc=$?
echo "RESULT v5_point rc=$rc wall_s=$(( $(date +%s) - t1 ))"
grep -n -E 'FLOOR abort|panicked|Error|error|RESULT' $OUT/v5-point.log | head -12 | cut -c1-300
pkill -9 -x sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock
echo "RESULT v5_done $(stamp)"

15
tools/fleet/box-kill.sh Executable file
View file

@ -0,0 +1,15 @@
#!/usr/bin/env bash
# Stops every stage process on a box (the matrix, Ember, the prover loop, their miners, workers, samplers and SP1
# servers); never the node. Run as a FILE (bash in/box-kill.sh) so that no pattern below can match the shell that runs
# it (an inline `pkill -f 'igneum-prove-host'` killed its own ssh shell on 6 October 2026, 12:38Z).
for i in 1 2; do
pkill -9 -f '^bash in/box-matrix.sh' ; pkill -9 -f '^bash in/box-ember.sh'; pkill -9 -f '^bash in/box-prover.sh'; pkill -9 -f '^python3 -u /root/fleet/in/box-prover.py'
pkill -9 -x igneum-miner; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-worker-cuda'; pkill -9 -x sp1-gpu-server; pkill -9 -f '^/opt/igneum-floor/(bin|target|prove)/.*igneum-prove-(host|export)' # pkill -x cannot match a name over 15 characters (igneum-worker-cuda, igneum-prove-host)
pkill -9 -f '^nvidia-smi --query'
sleep 2
done
rm -f /tmp/sp1-cuda-*.sock
nvidia-smi -i 0 -rgc >/dev/null 2>&1
d="$(nvidia-smi --query-gpu=power.default_limit --format=csv,noheader,nounits -i 0 | tr -d ' ' | cut -d. -f1)"; [ -n "$d" ] && nvidia-smi -i 0 -pl "$d" >/dev/null 2>&1
cd /root/fleet/out 2>/dev/null && for f in matrix.log ember.log prover.log prover-launch.log rows.jsonl; do [ -f "$f" ] && mv "$f" "$f.$(date +%s).old"; done
echo "count=$(pgrep -c -f '^bash in/box-|^python3 -u /root/fleet/in/box-prover')+$(pgrep -c -x igneum-miner)+$(pgrep -c -f '^/opt/igneum/pkg/bin/igneum-worker-cuda')+$(pgrep -c -x sp1-gpu-server)+$(pgrep -c -f '^/opt/igneum-floor/(bin|target|prove)/.*igneum-prove-host') mem=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits -i 0 | tr -d ' ')"

127
tools/fleet/box-matrix.sh Executable file
View file

@ -0,0 +1,127 @@
#!/usr/bin/env bash
# Phase 1 on one rented card: the memory matrix of docs/analysis/prover-floor.md on the real card. Every point is one
# igneum-prove-host run against a fresh sp1-gpu-server (killed and its socket unlinked around every point), with an
# nvidia-smi sampler at 1 s (memory.used, power.draw, utilization); peak = max memory.used, base = the reading just
# before the point (the card's idle, or the miner's resident set when the miner runs beside the prover), own = peak
# minus base. Every proof is verified by the host's own (unpatched) verifier: the VERIFIED word in the RESULT line.
# Writes /root/fleet/out/matrix.json (card facts, the miner row, every point) and the raw logs beside it; the last log
# line is "RESULT matrix_done" or "RESULT matrix_failed <why>".
set -uo pipefail
F=/root/fleet; OUT=$F/out; LOG=$OUT/matrix.log; B=/opt/igneum/pkg/bin; FLOOR=/opt/igneum-floor
for b in igneum-prove-host igneum-prove-export; do for d in $FLOOR/target/release $FLOOR/prove/proving/igneum-prove/target/release /opt/igneum-segal/proving/igneum-prove/target/release; do [ -x $d/$b ] && ln -sfn $d/$b $FLOOR/bin/$b; done; done # the last one that exists wins: the segment host (--mode chain) over the floor host
HOST=$FLOOR/bin/igneum-prove-host; FIX=$FLOOR/prove/proving/fixtures
V1=$FIX/fees-v1-shards2.json; EMPTY=$FIX/block-72854-empty-block-first.json
LABEL="${LABEL:-box}"; WALLET="${WALLET:-0x1919191919191919191919191919191919191919}"
MINER_SECS="${MINER_SECS:-150}"; POINT_TIMEOUT="${POINT_TIMEOUT:-1500}"
mkdir -p $OUT $F/mine/packs
exec > >(tee -a $LOG) 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
say() { echo "$(stamp) $*"; }
q() { nvidia-smi --query-gpu="$1" --format=csv,noheader,nounits -i 0 2>/dev/null | head -1 | tr -d ' '; }
[ -x "$HOST" ] || { echo "RESULT matrix_failed no host"; exit 2; }
[ -x "$FLOOR/bin/sp1-gpu-server" ] || { echo "RESULT matrix_failed no patched server"; exit 2; }
NAME="$(q name)"; TOTAL="$(q memory.total)"; DRV="$(q driver_version)"; PLIM="$(q power.limit)"
echo "RESULT start $(stamp) label=$LABEL card=$NAME total_mib=$TOTAL driver=$DRV power_limit_w=$PLIM"
ROWS=$OUT/rows.jsonl; : > $ROWS
# the sampler: one csv per point, "ts,mem_mib,power_w,util_pct"
SAMP=""
# one nvidia-smi call a second in a loop: `nvidia-smi -l 1` buffers its file output in 4 KB chunks and ignored SIGTERM
# on the 3080 box (12:19Z), which hung the first matrix in sampler_stop
sampler_start() { ( while :; do nvidia-smi --query-gpu=timestamp,memory.used,power.draw,utilization.gpu --format=csv,noheader,nounits -i 0 2>/dev/null; sleep 1; done ) > "$1" & SAMP=$!; }
sampler_stop() { [ -n "$SAMP" ] && { pkill -P $SAMP 2>/dev/null; kill $SAMP 2>/dev/null; sleep 1; kill -9 $SAMP 2>/dev/null; }; SAMP=""; }
peak_of() { awk -F', *' 'NR>0 {if ($2+0 > m) m=$2+0} END {print m+0}' "$1"; }
mean_col() { awk -F', *' -v c="$2" '{s+=$c; n++} END {if (n) printf "%.1f", s/n; else print 0}' "$1"; }
kill_server() { pkill -x sp1-gpu-server 2>/dev/null; sleep 2; pkill -9 -x sp1-gpu-server 2>/dev/null; rm -f /tmp/sp1-cuda-*.sock; }
idle_mib() { sleep 3; q memory.used; }
# 1. idle
kill_server
sleep 5
IDLE_MIB="$(q memory.used)"; IDLE_W="$(q power.draw)"
echo "RESULT idle mem_mib=$IDLE_MIB power_w=$IDLE_W"
# 2. the node's state (the miner mines IBD templates when not synced; the rate is the same, the pack may move)
NODE_LINE="$($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1)"
waited=0; while [[ "$NODE_LINE" != *synced=true* && $waited -lt 900 ]]; do sleep 30; waited=$((waited+30)); NODE_LINE="$($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1)"; done
echo "RESULT node $NODE_LINE waited_s=$waited"
# 3. the miner alone: rate, watts, working set
MPID=""
miner_start() {
cd $F/mine
rm -rf packs/devnet; $B/igneum-miner export-pack grpc://127.0.0.1:26610 packs/devnet > $OUT/export-pack.log 2>&1 || say "export-pack failed (see export-pack.log)"
nohup $B/igneum-miner mine grpc://127.0.0.1:26610 1 100000000 "$LABEL" --worker $B/igneum-worker-cuda --worker-args "--device 0 --pack packs/devnet" \
--prepare-packs packs/prepare --exit-on-seed-change --evm-address "$WALLET" --payout-label "$LABEL" --status-secs 10 > $OUT/miner.log 2>&1 &
MPID=$!; cd $F
}
miner_stop() { [ -n "$MPID" ] && { kill $MPID 2>/dev/null; sleep 2; kill -9 $MPID 2>/dev/null; }; pkill -f '^/opt/igneum/pkg/bin/igneum-worker-cuda' 2>/dev/null; pkill -x igneum-miner 2>/dev/null; MPID=""; sleep 3; }
miner_rate() { # mean of the STATUS now= values in the last N lines
grep STATUS $OUT/miner.log | grep -o ' now=[0-9.]*' | tail -n "${1:-10}" | cut -d= -f2 | awk '{s+=$1; n++} END {if (n) printf "%.2f", s/n; else print 0}'
}
miner_start
say "miner warming 60 s"; sleep 60
sampler_start $OUT/samp-miner.csv; sleep "$MINER_SECS"; sampler_stop
MINER_MHS="$(miner_rate 12)"; MINER_W="$(mean_col $OUT/samp-miner.csv 3)"; MINER_UTIL="$(mean_col $OUT/samp-miner.csv 4)"; MINER_MIB="$(peak_of $OUT/samp-miner.csv)"
MINER_ACC="$(grep -o 'accepted=[0-9]*' $OUT/miner.log | tail -1 | cut -d= -f2)"
MINER_STATUS="$(grep STATUS $OUT/miner.log | tail -1 | cut -c1-200)"
echo "RESULT miner mhs=$MINER_MHS watts=$MINER_W util=$MINER_UTIL mem_mib=$MINER_MIB own_mib=$((MINER_MIB - IDLE_MIB)) accepted=${MINER_ACC:-?} status=\"$MINER_STATUS\""
printf '%s\n' "{\"row\":\"miner\",\"mhs\":$MINER_MHS,\"watts\":$MINER_W,\"util\":$MINER_UTIL,\"mem_mib\":$MINER_MIB,\"own_mib\":$((MINER_MIB - IDLE_MIB)),\"accepted\":\"${MINER_ACC:-}\"}" >> $ROWS
miner_stop
# 4. one point: name, home (stock or patched), mode, threshold env, fixture, beside
point() {
local name="$1" home="$2" mode="$3" thr="$4" fixture="$5" beside="$6"
kill_server; sleep 2
local base; base="$(q memory.used)"
local env=(HOME="$home" SP1_PROVER=cuda RUST_LOG=off SP1_GPU_FLOOR_LOG=1)
[ -n "$thr" ] && env+=(SP1_GPU_ELEMENT_THRESHOLD="$thr")
local log=$OUT/point-$name.log res=$OUT/point-$name.json samp=$OUT/samp-$name.csv
say "point $name: mode $mode threshold ${thr:-default} fixture $(basename $fixture) beside_miner=$beside base_mib=$base"
sampler_start "$samp"
local t0=$(date +%s)
env "${env[@]}" timeout "$POINT_TIMEOUT" "$HOST" "$fixture" --mode "$mode" --shard 0 --prover "$WALLET" --out "$res" > "$log" 2>&1
local rc=$?; local wall=$(( $(date +%s) - t0 ))
sampler_stop; kill_server
local peak; peak="$(peak_of "$samp")"
local rline; rline="$(grep -E "^RESULT (core|compressed) shard" "$log" | tail -1)"
local secs; secs="$(printf '%s' "$rline" | grep -o 'prove [0-9.]* s' | grep -o '[0-9.]*' | head -1)"
local bytes; bytes="$(printf '%s' "$rline" | grep -o 'proof [0-9]* bytes' | grep -o '[0-9]*' | head -1)"
local ver="no"; printf '%s' "$rline" | grep -q 'VERIFIED' && ver="yes"; printf '%s' "$rline" | grep -q 'VERIFY FAILED' && ver="FAILED"
local err; err="$(grep -m1 -E 'Unsupported GPU memory|out of memory|OutOfMemory|CUDA_ERROR|panicked|Error:|error:' "$log" | cut -c1-200 | tr '"' "'")"
local setup; setup="$(grep -o 'RESULT setup: [0-9.]* s' "$log" | grep -o '[0-9.]* s' | head -1)"
echo "RESULT point name=$name mode=$mode thr=${thr:-default} beside=$beside rc=$rc wall_s=$wall prove_s=${secs:-} proof_bytes=${bytes:-} verified=$ver peak_mib=$peak base_mib=$base own_mib=$((peak - base)) setup=\"${setup:-}\" err=\"${err:-}\""
printf '%s\n' "{\"row\":\"point\",\"name\":\"$name\",\"home\":\"$home\",\"mode\":\"$mode\",\"threshold\":\"${thr:-default}\",\"fixture\":\"$(basename $fixture)\",\"beside_miner\":$beside,\"rc\":$rc,\"wall_s\":$wall,\"prove_s\":${secs:-null},\"proof_bytes\":${bytes:-null},\"verified\":\"$ver\",\"peak_mib\":$peak,\"base_mib\":$base,\"own_mib\":$((peak - base)),\"err\":\"${err:-}\",\"result\":\"$(printf '%s' "$rline" | cut -c1-160 | tr '"' "'")\"}" >> $ROWS
}
STOCK=/root; PATCHED=$FLOOR/home
T25=33554432; T26=67108864; T27=134217728; T24=16777216
SMALL=0; [ "$TOTAL" -lt 11000 ] && SMALL=1
# 5. the stock server, alone (the SDK downloads it into /root/.sp1/bin on the first run)
point stock-comp-v1 $STOCK compressed "" $V1 false
# 6. the patched server, alone
point alone-comp-26-v1 $PATCHED compressed $T26 $V1 false
point alone-comp-27-v1 $PATCHED compressed $T27 $V1 false
point alone-comp-26-empty $PATCHED compressed $T26 $EMPTY false
point alone-core-25-v1 $PATCHED core $T25 $V1 false
point alone-core-26-v1 $PATCHED core $T26 $V1 false
if [ $SMALL = 1 ]; then point alone-comp-25-v1 $PATCHED compressed $T25 $V1 false; point alone-core-24-v1 $PATCHED core $T24 $V1 false; fi
# 7. beside the miner
miner_start; say "miner warming 45 s for the beside rows"; sleep 45
MINER_RES="$(q memory.used)"; echo "RESULT miner_resident mem_mib=$MINER_RES mhs=$(miner_rate 4)"
point miner-comp-26-v1 $PATCHED compressed $T26 $V1 true
[ "$TOTAL" -ge 15000 ] && point miner-comp-27-v1 $PATCHED compressed $T27 $V1 true
point miner-core-25-v1 $PATCHED core $T25 $V1 true
point miner-core-26-v1 $PATCHED core $T26 $V1 true
if [ $SMALL = 1 ]; then point miner-core-24-v1 $PATCHED core $T24 $V1 true; fi
MINER_MHS_BESIDE="$(miner_rate 30)"
miner_stop
echo "RESULT miner_beside mhs_mean_during_points=$MINER_MHS_BESIDE"
python3 - "$OUT" "$LABEL" "$NAME" "$TOTAL" "$DRV" "$IDLE_MIB" "$IDLE_W" "$PLIM" <<'PY'
import json, sys
out, label, name, total, drv, idle, idle_w, plim = sys.argv[1:]
rows = [json.loads(l) for l in open(f"{out}/rows.jsonl") if l.strip()]
json.dump({"label": label, "card": name, "total_mib": int(total), "driver": drv, "idle_mib": int(idle), "idle_w": float(idle_w) if idle_w.replace(".", "").isdigit() else 0, "power_limit_w": float(plim) if plim.replace(".", "").isdigit() else 0, "rows": rows}, open(f"{out}/matrix.json", "w"), indent=1)
PY
echo "RESULT matrix_done $(stamp) points=$(grep -c '"row":"point"' $ROWS)"

60
tools/fleet/box-node-swap.sh Executable file
View file

@ -0,0 +1,60 @@
#!/usr/bin/env bash
# The 0.3.13 node swap for a phase-2 box (6 October 2026, the shipper's two-word procedure).
# STEP=binary (publish 1): the fixed igneumd (/root/fleet/in/igneumd-0313, sha256 NODE_SHA256) replaces the running
# node over the observer's full-history datadir (pulled from the hub if not in place), the TEN-field file kept;
# the digest must read EXPECT_DIGEST (7bd98cc4...); the exec layer stays blocked by design; one status read, done.
# STEP=file (publish 2): ov13.json becomes the override, the node restarts, the digest must read b18ed271...;
# then eth_blockNumber is polled every 15 s until it reaches the tip (the replay time).
set -uo pipefail
F=/root/fleet; OUT=$F/out; B=/opt/igneum/pkg/bin; FLOOR=/opt/igneum-floor; HOST=$FLOOR/bin/igneum-prove-host
HUB_SSH="${HUB_SSH:-213.173.107.74}"; HUB_PORT="${HUB_PORT:-16515}"; HUB_PEER="${HUB_PEER:-213.173.107.74:16516}"
NODE_SHA256="${NODE_SHA256:?the fixed igneumd sha256}"
STEP="${STEP:-binary}"
EXPECT_DIGEST="${EXPECT_DIGEST:-7bd98cc4118616455709d5e32a30b799e6e67caa42d2b5d09875cd49848a7ed7}"
mkdir -p $OUT; exec >> $OUT/node-swap.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
rpc() { curl -s -m 8 -X POST -H 'Content-Type: application/json' --data "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"$1\",\"params\":[]}" http://127.0.0.1:26790/; }
echo "RESULT swap_start $(stamp) step=$STEP"
echo "$NODE_SHA256 $F/in/igneumd-0313" | sha256sum -c - >/dev/null || { echo "RESULT swap_failed sha256 mismatch: $(sha256sum $F/in/igneumd-0313 | cut -c1-16)"; exit 2; }
chmod +x $F/in/igneumd-0313
t0=$(date +%s)
bash $F/in/box-kill.sh >/dev/null 2>&1 # every stage process; the node is stopped on the next line
pkill -x igneumd; pkill -f '^/opt/igneum/pkg/bin/igneumd-0313'; sleep 4; pkill -9 -x igneumd 2>/dev/null; pkill -9 -f '^/opt/igneum/pkg/bin/igneumd-0313' 2>/dev/null; sleep 1
TGZ_SHA=e67cc6493cd86dc0a993c3e1005b1df43b661fc4973ef7f0095f602f1536cd85
if [ ! -f $F/node/.full-history ]; then
[ "$(sha256sum $F/observer-datadir.tgz 2>/dev/null | cut -c1-64)" = "$TGZ_SHA" ] || rm -f $F/observer-datadir.tgz # a partial pre-pull is not a tarball
[ -s $F/observer-datadir.tgz ] || scp -i $F/in/fleet-internal -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -P "$HUB_PORT" "root@$HUB_SSH:/root/fleet/share/observer-datadir.tgz" $F/observer-datadir.tgz || { echo "RESULT swap_failed pull"; exit 2; }
[ "$(sha256sum $F/observer-datadir.tgz | cut -c1-64)" = "$TGZ_SHA" ] || { echo "RESULT swap_failed tarball sha256 $(sha256sum $F/observer-datadir.tgz | cut -c1-16)"; exit 2; }
rm -rf $F/node.proof && mv $F/node $F/node.proof && mkdir -p $F/node && tar -C $F/node -xzf $F/observer-datadir.tgz 2>/dev/null && touch $F/node/.full-history || { echo "RESULT swap_failed untar"; exit 2; }
echo "RESULT datadir_in_place $(stamp) s=$(( $(date +%s) - t0 )) $(du -sh $F/node | cut -f1)"
fi
if [ "$STEP" = file ]; then
cp $F/override.json $F/override-10.json; cp $F/in/ov13.json $F/override.json
echo "RESULT override_swapped $(stamp) fields=$(python3 -c 'import json,sys; print(len(json.load(open(sys.argv[1]))))' $F/override.json)"
fi
cp $F/in/igneumd-0313 $B/igneumd-0313
IGNEUM_PROOF_VERIFIER=$HOST nohup $B/igneumd-0313 --devnet --appdir=$F/node --rpclisten=127.0.0.1:26610 --evm-rpclisten=127.0.0.1:26790 --listen=0.0.0.0:26611 \
--addpeer=$HUB_PEER --addpeer=188.245.5.161:26611 --override-params-file=$F/override.json --nodnsseed --disable-upnp --nologfiles --yes ${NODE_EXTRA:-} >> $F/node.log 2>&1 &
t1=$(date +%s); sleep 12
d="$(grep -o 'digest: [0-9a-f]*' $F/node.log | tail -1 | awk '{print $2}')"
ok=DIGEST_MISMATCH; [ "$d" = "$EXPECT_DIGEST" ] && ok=digest_ok
echo "RESULT node_started $(stamp) step=$STEP pid=$(pgrep -f '^/opt/igneum/pkg/bin/igneumd-0313' | head -1) version=$($B/igneumd-0313 --version 2>&1 | head -1) digest=${d:0:16} expected=${EXPECT_DIGEST:0:16} $ok exec_restart=\"$(grep -o 'Exec restart from the override file.*' $F/node.log | tail -1 | cut -c1-120)\""
[ "$ok" = digest_ok ] || { echo "RESULT swap_failed digest mismatch"; exit 2; }
watch() { $B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1; }
if [ "$STEP" = binary ]; then
sleep 45
echo "RESULT exec_status $(rpc igneum_getExecStatus | cut -c1-240)"
echo "RESULT swap_done $(stamp) step=binary $(watch)"; exit 0
fi
for i in $(seq 1 240); do
sleep 15
h="$(rpc eth_blockNumber | grep -o '"result":"[^"]*"' | cut -d'"' -f4)"; n=$(( ${h:-0x0} )); w="$(watch)"
echo "RESULT replay $(stamp) s=$(( $(date +%s) - t1 )) evm_block=$n $w"
if [ "$n" -gt 0 ] && [[ "$w" == *synced=true* ]]; then
echo "RESULT exec_status $(rpc igneum_getExecStatus | cut -c1-240)"
st="$(rpc igneum_getProvingStatus | python3 -c 'import sys,json; d=json.load(sys.stdin).get("result",{}); print("tipDaa", int(d.get("tipDaa","0x0"),16), "active", d.get("v1",{}).get("active"), "fresh", d.get("v1",{}).get("freshRuleActive"))' 2>/dev/null)"
daa=$(printf '%s' "$w" | grep -o 'daa=[0-9]*' | cut -d= -f2); tip=$(printf '%s' "$st" | awk '{print $2}')
if [ -n "$tip" ] && [ "$tip" -ge $(( ${daa:-0} - 20 )) ]; then echo "RESULT swap_done $(stamp) step=file replay_s=$(( $(date +%s) - t1 )) $st"; exit 0; fi
fi
done
echo "RESULT swap_failed replay did not reach the tip in 60 min"

24
tools/fleet/box-pool.sh Executable file
View file

@ -0,0 +1,24 @@
#!/usr/bin/env bash
# The pool host (6 October 2026): the 0.3.13 node on the thirteen-field file peered with the hub and the seed, then
# pool-v0 (igneum-pool, built on the hub) against it on 0.0.0.0:4463 with a throwaway coinbase address, dry-run off,
# payouts at the default minimum. RESULT lines in /root/fleet/out/pool.log; the pool's own log in pool-run.log.
set -uo pipefail
F=/root/fleet; OUT=$F/out; mkdir -p $F/in $OUT $F/node $F/pool; exec >> $OUT/pool.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
HUB_PEER="${HUB_PEER:-213.173.107.74:16516}"; WALLET="${WALLET:-0x1919191919191919191919191919191919191919}"
echo "RESULT pool_host_start $(stamp)"
command -v curl >/dev/null || { apt-get update -qq >/dev/null 2>&1; apt-get install -y -qq curl ca-certificates python3 >/dev/null 2>&1; }
if [ ! -x /opt/igneum/pkg/bin/igneum-miner ]; then
read -r PKG_PATH PKG_SHA PKG_VER <<< "$(curl -fsSL -m 30 https://dl.igneum.network/dl/public/igneum-downloads.json | python3 -c 'import sys,json; d=json.load(sys.stdin)["files"]["miner-hive"]; print(d["path"], d["sha256"], d["version"])')"
curl -fsSL -o $F/pkg.tgz "https://dl.igneum.network$PKG_PATH" && echo "$PKG_SHA $F/pkg.tgz" | sha256sum -c - >/dev/null && mkdir -p /opt/igneum/pkg && tar -C /opt/igneum/pkg --strip-components=1 -xzf $F/pkg.tgz || { echo "RESULT pool_host_failed package"; exit 2; }
fi
B=/opt/igneum/pkg/bin; cp $F/in/igneumd-0313 $B/igneumd-0313; chmod +x $B/igneumd-0313; cp $F/in/ov13.json $F/override.json
pgrep -f '^/opt/igneum/pkg/bin/igneumd-0313' >/dev/null || nohup $B/igneumd-0313 --devnet --appdir=$F/node --rpclisten=0.0.0.0:26610 --evm-rpclisten=127.0.0.1:26790 --listen=0.0.0.0:26611 --addpeer=$HUB_PEER --addpeer=188.245.5.161:26611 --override-params-file=$F/override.json --nodnsseed --disable-upnp --nologfiles --yes > $F/node.log 2>&1 &
sleep 10; echo "RESULT node $(stamp) digest=$(grep -o 'digest: [0-9a-f]*' $F/node.log | tail -1 | awk '{print substr($2,1,16)}')"
for i in $(seq 1 90); do w="$($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1)"; [[ "$w" == *synced=true* ]] && break; sleep 10; done
echo "RESULT synced $(stamp) after $((i*10)) s"
for i in $(seq 1 120); do [ -x $F/in/igneum-pool ] && break; sleep 15; done
[ -x $F/in/igneum-pool ] || { echo "RESULT pool_host_failed no pool binary after 30 min"; exit 2; }
cp $F/in/igneum-pool $F/pool/; cd $F/pool
echo "RESULT pool_help $($F/pool/igneum-pool --help 2>&1 | tr '\n' ' ' | cut -c1-600)"
echo "RESULT pool_binary_ready $(stamp)"

294
tools/fleet/box-prover.py Normal file
View file

@ -0,0 +1,294 @@
#!/usr/bin/env python3
"""The fleet's segment-aligned prover for a Linux box (phase 2): a port of app/igneum-app/src/prover.rs (branch
proving-v1, 272b025) and tools/proving-v1/pc2-segments.ps1 around the four binaries. One pass every 15 s:
1. igneum_getProvingStatus (start, n, unproven, tip) and igneum_getAssignedShards [[keyHash], 600], grouped into
whole untouched segments (every block present, every shard open, unpaid, not in this pool) whose deadline
(last DAA + unproven) is at least 240 DAA (or 1.5x the last segment's time) past the tip; ranked by FNV-1a of
(first, key hash) so a fleet of provers spreads over the segments instead of racing for one.
2. igneum_getSegmentStatement [first]: executed and pending; chained with --prev when the previous segment's proof is
in this pool (igneum_getSegmentProofBytes), else fresh. A fresh record refused with "pending until" is held and
offered again every pass until the segment's deadline (the 272b025 behaviour).
3. one export (igneum_exportSegments 0..last), one fixture per block (igneum-prove-export), one host run
(igneum-prove-host --mode chain --chain ... --save-shards [--prev]) on the patched server (HOME=/opt/igneum-floor/home,
SP1_GPU_ELEMENT_THRESHOLD from THRESHOLD), the miner paused for the run when MINER=pause (prove-alone cards).
4. every shard record signed (igneum-miner sign-record) and submitted (igneum_submitProofRecord); the segment record
(sign-segment-record, igneum_submitSegmentRecord) once every shard is accepted and the statement equals the node's.
5. the paid state of every submitted segment polled each pass (igneum_getSegmentRecords); a state file for the
collector: /root/fleet/out/prover-state.json; every event a RESULT line in /root/fleet/out/prover.log.
Env: LABEL (the key label, kept for the box's life), WALLET (payout), THRESHOLD (element threshold or empty),
MINER (keep|pause), RUN_HOURS (default 9).
"""
import json, os, sys, time, subprocess, datetime, binascii, urllib.request, signal
# a rig runs one loop per card under /root/fleet/card<n>/ (FLEET_CARD)
F = "/root/fleet"; CARD = os.environ.get("FLEET_CARD"); OUT = f"{F}/card{CARD}/out" if CARD else f"{F}/out"; B = "/opt/igneum/pkg/bin"; FLOOR = "/opt/igneum-floor"
HOST = f"{FLOOR}/bin/igneum-prove-host"; EXPORT = f"{FLOOR}/bin/igneum-prove-export"
EVM = "http://127.0.0.1:26790"; GRPC = "grpc://127.0.0.1:26610"; CHAIN = os.environ.get("CHAIN_NAME", "igneum-devnet") # Devnet 2 signs for igneum-devnet-2
LABEL = os.environ.get("LABEL", "box"); WALLET = os.environ.get("WALLET", "0x" + "19" * 20)
THRESHOLD = os.environ.get("THRESHOLD", ""); MINER = os.environ.get("MINER", "keep"); RUN_HOURS = float(os.environ.get("RUN_HOURS", "9"))
EXPORT_FROM = int(os.environ.get("EXPORT_FROM", "27276"))
# the export directory's own rule (6 October 2026, 20:4xZ, after the hub's node died on a full disk): a size cap and an age cap
# on OUT/segs, each segment's export deleted the moment its record is accepted or paid, and a disk-free check before each
# export that skips with a logged line under 10 percent free. Defaults fit a 100 GB box for a week.
SEGS_CAP_GB = float(os.environ.get("SEGS_CAP_GB", "20")); SEGS_MAX_AGE_H = float(os.environ.get("SEGS_MAX_AGE_H", str(7 * 24))); DISK_MIN_FREE_PCT = float(os.environ.get("DISK_MIN_FREE_PCT", "10"))
import shutil
def seg_dir_size(d):
return sum(os.path.getsize(os.path.join(r, f)) for r, _, fs in os.walk(d) for f in fs if os.path.exists(os.path.join(r, f)))
def drop_export(first, why):
d = f"{OUT}/segs/seg-{first}"
if os.path.isdir(d): sz = seg_dir_size(d); shutil.rmtree(d, ignore_errors=True); say(f"RESULT export_dropped {stamp()} segment {first} ({sz/1e6:.0f} MB): {why}")
def prune_exports():
root = f"{OUT}/segs"
if not os.path.isdir(root): return
entries = sorted((os.path.join(root, n) for n in os.listdir(root)), key=lambda x: os.path.getmtime(x))
now_ = time.time(); total = sum(seg_dir_size(e) if os.path.isdir(e) else os.path.getsize(e) for e in entries)
for e in entries:
age_h = (now_ - os.path.getmtime(e)) / 3600
if age_h > SEGS_MAX_AGE_H or total > SEGS_CAP_GB * 1e9:
sz = seg_dir_size(e) if os.path.isdir(e) else os.path.getsize(e); shutil.rmtree(e, ignore_errors=True) if os.path.isdir(e) else os.remove(e); total -= sz
say(f"RESULT export_pruned {stamp()} {os.path.basename(e)} ({sz/1e6:.0f} MB, {age_h:.1f} h old): dir {total/1e9:.1f} GB against the {SEGS_CAP_GB:.0f} GB cap, {SEGS_MAX_AGE_H:.0f} h age cap")
def disk_free_pct(path="/"):
st = os.statvfs(path); return 100.0 * st.f_bavail / max(st.f_blocks, 1) # the devnet's exec restart block (ov13.json exec_restart_number) when the node does not report one
MINE = f"{F}/card{CARD}/mine" if CARD else f"{F}/mine"
os.makedirs(f"{OUT}/segs", exist_ok=True); os.makedirs(f"{MINE}/packs", exist_ok=True)
DEV = os.environ.get("IGNEUM_CUDA_DEVICE", "0")
LOG = open(f"{OUT}/prover.log", "a")
def stamp(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def say(s): LOG.write(f"{s}\n"); LOG.flush(); print(s, flush=True)
def hexi(v): return int(v, 16) if isinstance(v, str) and v.startswith("0x") else int(v or 0)
def rpc(method, params, timeout=60):
body = json.dumps({"jsonrpc": "2.0", "id": 1, "method": method, "params": params}).encode()
try:
with urllib.request.urlopen(urllib.request.Request(EVM, data=body, headers={"Content-Type": "application/json"}), timeout=timeout) as r:
d = json.loads(r.read())
if d.get("error"): say(f"RESULT rpc_error {stamp()} {method}: {str(d['error'])[:200]}"); return None
return d.get("result")
except Exception as e:
say(f"RESULT rpc_fail {stamp()} {method}: {str(e)[:120]}"); return None
def fnv1a(s):
h = 0xcbf29ce484222325
for c in s.encode(): h ^= c; h = (h * 0x100000001b3) & 0xffffffffffffffff
return h
def kill_server():
if CARD: subprocess.run(f"rm -f /tmp/sp1-cuda-{DEV}.sock", shell=True) # a rig: never another card's server
else: subprocess.run("pkill -x sp1-gpu-server; sleep 1; rm -f /tmp/sp1-cuda-*.sock", shell=True)
MPROC = None
def miner_start():
global MPROC
if MINER == "none": return # the box's own miner loop runs outside (Devnet 2 boxes)
if MPROC and MPROC.poll() is None: return
subprocess.run(f"cd {MINE} && rm -rf packs/devnet && {B}/igneum-miner export-pack {GRPC} packs/devnet > {OUT}/prover-export-pack.log 2>&1", shell=True)
MPROC = subprocess.Popen([f"{B}/igneum-miner", "mine", GRPC, "1", "100000000", LABEL, "--worker", f"{B}/igneum-worker-cuda", "--worker-args", f"--device {DEV} --pack packs/devnet",
"--prepare-packs", "packs/prepare", "--exit-on-seed-change", "--evm-address", WALLET, "--payout-label", LABEL, "--status-secs", "30"],
cwd=MINE, stdout=open(f"{OUT}/prover-miner.log", "a"), stderr=subprocess.STDOUT)
say(f"RESULT miner_start {stamp()} pid={MPROC.pid}")
def miner_stop():
global MPROC
if MPROC: MPROC.terminate(); time.sleep(2); MPROC.kill(); MPROC = None
subprocess.run(f"pkill -f '^/opt/igneum/pkg/bin/igneum-worker-cuda --device {DEV} '", shell=True)
def miner_rate(n=6):
try:
vals = [float(l.split(" now=")[1].split()[0]) for l in open(f"{OUT}/prover-miner.log").read().split("\n") if "STATUS" in l and " now=" in l][-n:]
return round(sum(vals) / len(vals), 2) if vals else 0
except Exception: return 0
# the key
kh = subprocess.run([f"{B}/igneum-miner", "key-hash", LABEL], capture_output=True, text=True).stdout.strip().split("\n")[-1].strip()
if len(kh) == 64: kh = "0x" + kh
if len(kh) != 66: say(f"RESULT prover_failed key-hash gave '{kh[:40]}'"); sys.exit(2)
say(f"RESULT start {stamp()} label={LABEL} key={kh[:18]} wallet={WALLET[:10]} threshold={THRESHOLD or 'default'} miner={MINER} host={os.path.exists(HOST)}")
# wait for the node to be synced
for _ in range(120):
w = subprocess.run(f"{B}/igneum-miner watch 1 {GRPC} 2>/dev/null | grep -o 'synced=[a-z]*' | tail -1", shell=True, capture_output=True, text=True).stdout.strip()
if w == "synced=true": break
time.sleep(15)
say(f"RESULT node {stamp()} {w}")
miner_start()
state = {"passes": 0, "claimed": 0, "submitted": 0, "paid": 0, "paid_wei": 0, "shards_accepted": 0, "shards_refused": 0, "segment_refused": 0, "held": 0,
"last_segment_s": 0, "segments": [], "started": stamp(), "label": LABEL, "wallet": WALLET, "key": kh}
attempted = set(); submitted = {}; held = {} # held: first -> {body file, deadline, last}
last_seg_secs = 0; t_run0 = time.time()
def save_state():
tmp = f"{OUT}/prover-state.json.{os.getpid()}.tmp" # one tmp per process: two instances racing on one name lost the file (21:4xZ)
json.dump(state, open(tmp, "w"), indent=1); os.replace(tmp, f"{OUT}/prover-state.json")
# one prover per box: a pid file under OUT; a second instance exits at once instead of killing the first's GPU server
_pidf = f"{OUT}/prover.pid"
try:
_old = int(open(_pidf).read().strip()); os.kill(_old, 0); print(f"RESULT refused {stamp()} another box-prover.py runs as pid {_old}; exiting", flush=True); sys.exit(3)
except (FileNotFoundError, ValueError, ProcessLookupError): pass
open(_pidf, "w").write(str(os.getpid()))
def submit(method, record, proof_file):
proof = "0x" + binascii.hexlify(open(proof_file, "rb").read()).decode()
return rpc(method, [{"record": record, "proof": proof}], timeout=180)
def candidates(st):
start = hexi(st["v1"]["start"]); n = max(1, hexi(st["v1"]["segmentBlocks"])); unproven = hexi(st["v1"]["unprovenDaa"]); tip = hexi(st["tipDaa"])
work = rpc("igneum_getAssignedShards", [[kh], 600]) or []
by = {}
for w in work:
num = hexi(w.get("number"))
if num < start: continue
by.setdefault(num, []).append(w)
need = max(240, int(last_seg_secs * 1.5) + 1)
segs = []; seen = set()
for num in sorted(by):
k = (num - start) // n; first = start + k * n; last = first + n - 1
if first in seen or first in attempted: continue
seen.add(first)
whole = True; shards = []; last_daa = 0
for b in range(first, last + 1):
if b not in by: whole = False; break
es = {int(e.get("shard", 0)): e for e in by[b]}
for si in sorted(es):
e = es[si]
if not e.get("open") or e.get("paid") is not None or (e.get("pool") and len(e["pool"]) > 0): whole = False; break
shards.append({"number": b, "hash": e.get("hash"), "shard": si})
if b == last: last_daa = hexi(e.get("daaScore"))
if not whole: break
if whole and shards:
deadline = last_daa + unproven
if deadline >= tip + 1 + need: segs.append({"first": first, "last": last, "last_daa": last_daa, "deadline": deadline, "shards": shards, "margin": deadline - tip - 1})
segs.sort(key=lambda s: fnv1a(f"{s['first']}:{kh}"))
return segs, start, n, tip, len(work)
while (time.time() - t_run0) / 3600 < RUN_HOURS:
state["passes"] += 1; p = state["passes"]
if MINER != "none" and MPROC and MPROC.poll() is not None and MINER == "keep": say(f"RESULT miner_exit {stamp()} rc={MPROC.returncode}; restarting"); MPROC = None; miner_start()
st = rpc("igneum_getProvingStatus", [])
if not st or not st.get("v1") or not st["v1"].get("active"): say(f"RESULT pass {p} {stamp()} v1 not active or no status"); time.sleep(20); continue
tip = hexi(st["tipDaa"])
# paid state
for first in list(submitted):
rec = rpc("igneum_getSegmentRecords", [hex(first)])
if rec and rec.get("paid"):
wei = hexi(rec["paid"]["wei"]); state["paid"] += 1; state["paid_wei"] += wei; drop_export(first, "record paid")
say(f"RESULT paid {stamp()} segment {first}..{submitted[first]['last']} wei={wei} ({wei/1e18:.4f} IGN) carrier={hexi(rec['paid'].get('carrierNumber'))} after {int(time.time()-submitted[first]['at'])} s")
for s in state["segments"]:
if s["first"] == first: s["paid_wei"] = wei; s["paid_at"] = stamp()
del submitted[first]
elif rec is not None and tip > submitted[first]["deadline"] + 50:
say(f"RESULT unpaid {stamp()} segment {first} past its deadline unpaid; carried={len(rec.get('carried') or [])} pool={len(rec.get('pool') or [])}"); del submitted[first]
# held fresh records offered again
for first in list(held):
h = held[first]
if tip > h["deadline"]: say(f"RESULT held_expired {stamp()} segment {first} deadline passed"); del held[first]; continue
rr = submit("igneum_submitSegmentRecord", h["record"], h["proof_file"])
if rr and rr.get("accepted"):
state["submitted"] += 1; submitted[first] = {"last": h["last"], "at": time.time(), "deadline": h["deadline"]}; drop_export(first, "record accepted"); say(f"RESULT submitted {stamp()} segment {first}..{h['last']} record accepted on retry (held {int(time.time()-h['since'])} s)"); del held[first]
state["held"] = len(held)
cands, start, n, tip, entries = candidates(st)
save_state()
if not cands:
if p % 4 == 1: say(f"RESULT pass {p} {stamp()} no whole segment inside the margin (worklist {entries} entries, tip {tip}, mhs {miner_rate()}); waiting")
time.sleep(15); continue
picked = None; prev_file = None; expected = ""
for c in cands[:3]:
stmt = rpc("igneum_getSegmentStatement", [hex(c["first"])])
if not stmt or not stmt.get("executed") or (stmt.get("status") or {}).get("status") != "pending":
attempted.add(c["first"]); say(f"RESULT skip {stamp()} segment {c['first']}: executed={stmt and stmt.get('executed')} status={(stmt or {}).get('status')}"); continue
if stmt.get("previous") is None:
if not st["v1"].get("freshRuleActive") and c["first"] >= start + n:
pr = rpc("igneum_getSegmentRecords", [hex(c["first"] - n)])
waiting = any(e.get("verified") and e.get("includedIn") is None for e in (pr or {}).get("pool") or [])
if waiting or (pr and pr.get("paid")): say(f"RESULT skip {stamp()} segment {c['first']}: previous has a record waiting or paid, a fresh chain would be refused (fresh rule off)"); continue
picked = c; expected = stmt.get("publicValuesFresh") or ""; break
if not stmt["previous"].get("proofInPool"): say(f"RESULT skip {stamp()} segment {c['first']}: previous paid, its proof not in this pool"); continue
got = rpc("igneum_getSegmentProofBytes", [stmt["previous"]["first"], stmt["previous"]["keyHash"]])
if not got or not got.get("proof"): continue
prev_file = f"{OUT}/segs/prev-{c['first']}.bin"; h = got["proof"]; open(prev_file, "wb").write(binascii.unhexlify(h[2:] if h.startswith("0x") else h))
picked = c; expected = stmt.get("publicValuesContinuing") or ""; break
if not picked: say(f"RESULT pass {p} {stamp()} {len(cands)} candidates, none usable; waiting"); time.sleep(15); continue
first, last = picked["first"], picked["last"]; attempted.add(first); state["claimed"] += 1
say(f"RESULT claim {stamp()} segment {first}..{last} ({len(picked['shards'])} shards, {'continuing' if prev_file else 'fresh'}) margin={picked['margin']} tip={tip} candidates={len(cands)} rank_by=fnv")
seg = {"first": first, "last": last, "claimed_at": stamp(), "shards": len(picked["shards"]), "fresh": prev_file is None}; state["segments"].append(seg)
prune_exports(); free = disk_free_pct()
if free < DISK_MIN_FREE_PCT: say(f"RESULT skip {stamp()} segment {first}: disk {free:.1f}% free is under the {DISK_MIN_FREE_PCT:.0f}% floor, no export"); time.sleep(60); continue
d = f"{OUT}/segs/seg-{first}"; os.makedirs(d, exist_ok=True); t_seg0 = time.time()
# export
# a node whose EVM restarted at a chain block (0.3.13's exec restart rule) exports from that block, not genesis: the
# blocks below it are not executed and the exporter refuses their zero state roots (6 October 2026, 16:02Z)
ex = rpc("igneum_getExecStatus", []) or {}
start_blk = hexi(ex.get("restartNumber") or ex.get("execRestartNumber") or ex.get("startedAt") or 0)
if not start_blk:
sf = str(ex.get("startedFrom", ""))
import re as _re; m = _re.search(r"chain block (\d+)", sf); start_blk = int(m.group(1)) if m else 0
if not start_blk and EXPORT_FROM: start_blk = EXPORT_FROM
t = time.time(); body = json.dumps({"jsonrpc": "2.0", "id": 1, "method": "igneum_exportSegments", "params": [hex(start_blk), hex(last)]})
seg["export_from"] = start_blk
r = subprocess.run(["curl", "-s", "-m", "600", "-X", "POST", EVM, "-H", "Content-Type: application/json", "--data-binary", body, "-o", f"{d}/seq.json"])
try: json.dump(json.load(open(f"{d}/seq.json"))["result"], open(f"{d}/export.json", "w"))
except Exception as e: say(f"RESULT seg {first} export FAILED {str(e)[:100]}"); continue
seg["export_s"] = round(time.time() - t, 1); os.remove(f"{d}/seq.json")
# cut
t = time.time(); fixtures = []
ok = True
for b in range(first, last + 1):
rr = subprocess.run([EXPORT, f"{d}/export.json", str(b), f"{d}/block-{b}.json", "--source", f"fleet {LABEL} live devnet, segment-aligned prover"], capture_output=True, text=True, timeout=600)
if rr.returncode != 0: say(f"RESULT seg {first} cut {b} FAILED: {(rr.stdout + rr.stderr)[-200:]}"); ok = False; break
fixtures.append(f"{d}/block-{b}.json")
if not ok: continue
seg["cut_s"] = round(time.time() - t, 1)
# chain
if MINER == "pause": miner_stop()
kill_server()
env = dict(os.environ, HOME=f"{FLOOR}/home", SP1_PROVER="cuda", RUST_LOG="off")
if THRESHOLD: env["SP1_GPU_ELEMENT_THRESHOLD"] = THRESHOLD
args = [HOST, "--mode", "chain", "--chain", ",".join(fixtures), "--prover", WALLET, "--save-shards", "--out", f"{d}/chain-results.json"]
if prev_file: args += ["--prev", prev_file]
samp = subprocess.Popen(f"while :; do nvidia-smi -i {DEV} --query-gpu=memory.used,utilization.gpu,power.draw --format=csv,noheader,nounits; sleep 1; done", shell=True, stdout=open(f"{d}/smi.csv", "w"), stderr=subprocess.DEVNULL, start_new_session=True)
t = time.time()
try: rr = subprocess.run(args, env=env, capture_output=True, text=True, timeout=3600)
except subprocess.TimeoutExpired: rr = None
seg["chain_s"] = round(time.time() - t, 1); os.killpg(samp.pid, signal.SIGKILL); kill_server()
if MINER == "pause": miner_start()
try: peak = max(float(l.split(",")[0]) for l in open(f"{d}/smi.csv") if l.strip())
except Exception: peak = 0
seg["peak_mib"] = peak
open(f"{d}/chain.log", "w").write((rr.stdout if rr else "") + "\n" + (rr.stderr if rr else "TIMEOUT"))
if not rr or rr.returncode != 0 or not os.path.exists(f"{d}/chain-results.json"):
say(f"RESULT seg {first} chain FAILED {stamp()} rc={rr.returncode if rr else 'timeout'} wall={seg['chain_s']} s: {((rr.stderr if rr else '') or '')[-200:].strip()}"); seg["failed"] = "chain"; continue
res = json.load(open(f"{d}/chain-results.json"))
recs = [s for blk in res.get("blocks", []) for s in blk.get("shard_records", [])]
say(f"RESULT seg {first} chain {stamp()} {len(recs)} shard records, chain_len {res.get('segment_chain_len')}, proof {res.get('segment_proof_bytes')} bytes, shards {res.get('shard_prove_seconds_total', 0):.1f} s, aggregation {res.get('aggregate_prove_seconds_total', 0):.1f} s, wall {seg['chain_s']} s, peak {peak:.0f} MiB")
# shard records
ok_shards = 0
for rcd in recs:
sg = subprocess.run([f"{B}/igneum-miner", "sign-record", LABEL, CHAIN, rcd["block_hash"], str(rcd["number"]), str(rcd["shard"]), WALLET, rcd["statement"], rcd["proof_sha256"]], capture_output=True, text=True).stdout.strip().split("\n")[-1]
try: record = json.loads(sg).get("record")
except Exception: record = None
if not record: say(f"RESULT seg {first} shard {rcd['number']}/{rcd['shard']} sign FAILED: {sg[:120]}"); state["shards_refused"] += 1; continue
reply = submit("igneum_submitProofRecord", record, rcd["proof_file"])
if reply and reply.get("accepted"): ok_shards += 1; state["shards_accepted"] += 1
else: state["shards_refused"] += 1; say(f"RESULT seg {first} shard {rcd['number']}/{rcd['shard']} refused: {(reply or {}).get('reason', reply)}")
seg["shards_accepted"] = ok_shards
say(f"RESULT seg {first} shards {stamp()} accepted {ok_shards} of {len(recs)}")
if ok_shards != len(recs): seg["failed"] = "shards"; continue
pv = res.get("segment_public_values", "")
strip = lambda h: (h[2:] if h.startswith("0x") else h); strip2 = lambda h: (strip(h)[:472] + strip(h)[536:]) if len(strip(h)) == 680 else strip(h)
if strip2(pv) != strip2(expected):
a, b = strip2(pv), strip2(expected); off = next((i for i in range(min(len(a), len(b))) if a[i] != b[i]), min(len(a), len(b)))
say(f"RESULT seg {first} FAILED: statement differs from the node's at hex offset {off} (lengths {len(a)} vs {len(b)}); ours ...{a[max(0,off-8):off+56]} node ...{b[max(0,off-8):off+56]}"); seg["failed"] = "statement"; continue
last_hash = next(s["hash"] for s in picked["shards"] if s["number"] == last)
sg = subprocess.run([f"{B}/igneum-miner", "sign-segment-record", LABEL, CHAIN, str(first), str(last), last_hash, WALLET, pv, res["segment_proof_sha256"]], capture_output=True, text=True).stdout.strip().split("\n")[-1]
try: record = json.loads(sg).get("record")
except Exception: record = None
if not record: say(f"RESULT seg {first} segment sign FAILED: {sg[:120]}"); seg["failed"] = "sign"; continue
reply = submit("igneum_submitSegmentRecord", record, res["segment_proof_file"])
seg_s = round(time.time() - t_seg0, 1); seg["end_to_end_s"] = seg_s
if reply and reply.get("accepted"):
state["submitted"] += 1; last_seg_secs = seg_s; state["last_segment_s"] = seg_s
stmt2 = rpc("igneum_getSegmentStatement", [hex(first)]) or {}
submitted[first] = {"last": last, "at": time.time(), "deadline": picked["deadline"]}; seg["submitted_at"] = stamp()
say(f"RESULT submitted {stamp()} segment {first}..{last} record accepted (new={reply.get('new')}, chain_len {res.get('segment_chain_len')}), aggregator share {hexi(stmt2.get('aggregatorWei'))/1e18:.4f} IGN, end to end {seg_s} s, mhs {miner_rate()}")
else:
reason = str((reply or {}).get("reason", reply))[:200]; state["segment_refused"] += 1; seg["refused"] = reason
say(f"RESULT segment_refused {stamp()} segment {first}..{last}: {reason}; end to end {seg_s} s")
if "pending until" in reason or "does not chain" in reason:
held[first] = {"record": record, "proof_file": res["segment_proof_file"], "last": last, "deadline": picked["deadline"], "since": time.time()}; say(f"RESULT held {stamp()} segment {first} held for retry until DAA {picked['deadline']}")
for b in range(first, last + 1):
try: os.remove(f"{d}/block-{b}.json")
except OSError: pass
try: os.remove(f"{d}/export.json")
except OSError: pass
save_state()
miner_stop(); kill_server(); save_state()
say(f"RESULT summary {stamp()} passes={state['passes']} claimed={state['claimed']} submitted={state['submitted']} paid={state['paid']} paid_wei={state['paid_wei']} shards_accepted={state['shards_accepted']} shards_refused={state['shards_refused']} segment_refused={state['segment_refused']}")
say(f"RESULT prover_done {stamp()}")

25
tools/fleet/box-prover.sh Executable file
View file

@ -0,0 +1,25 @@
#!/usr/bin/env bash
# Phase 2 launcher: restarts the box's node with IGNEUM_PROOF_VERIFIER (so its pool verifies records and its templates
# carry them, the proving agent's note of 6 October 2026), waits for it to be synced again (the data dir is kept, so
# seconds), then runs box-prover.py with the card's profile: THRESHOLD and MINER from the card's memory unless given.
set -uo pipefail
F=/root/fleet; OUT=$F/out; B=/opt/igneum/pkg/bin; FLOOR=/opt/igneum-floor; for b in igneum-prove-host igneum-prove-export; do for d in $FLOOR/target/release $FLOOR/prove/proving/igneum-prove/target/release /opt/igneum-segal/proving/igneum-prove/target/release; do [ -x $d/$b ] && ln -sfn $d/$b $FLOOR/bin/$b; done; done # the last one that exists wins: the segment host (--mode chain) over the floor host
HOST=$FLOOR/bin/igneum-prove-host
mkdir -p $OUT; exec >> $OUT/prover-launch.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
TOTAL="$(nvidia-smi --query-gpu=memory.total --format=csv,noheader,nounits -i 0 | tr -d ' ')"
if [ -z "${THRESHOLD:-}" ]; then
if [ "$TOTAL" -ge 23000 ]; then THRESHOLD=""; elif [ "$TOTAL" -ge 15000 ]; then THRESHOLD=134217728; else THRESHOLD=67108864; fi
fi
if [ -z "${MINER:-}" ]; then if [ "$TOTAL" -ge 15000 ]; then MINER=keep; else MINER=pause; fi; fi
echo "RESULT launch $(stamp) total_mib=$TOTAL threshold=${THRESHOLD:-default} miner=$MINER"
pkill -x igneum-miner 2>/dev/null; pkill -f '^/opt/igneum/pkg/bin/igneum-worker-cuda' 2>/dev/null; pkill -x sp1-gpu-server 2>/dev/null; rm -f /tmp/sp1-cuda-*.sock
pkill -x igneumd; sleep 4; pkill -9 -x igneumd 2>/dev/null; sleep 1
HUBARG=""; [ -n "${HUB_PEER:-}" ] && HUBARG="--addpeer=$HUB_PEER" # the fleet's hub box (a second peer beside the seed, 12:35Z)
NODE_BIN=$B/igneumd; [ -x $B/igneumd-0313 ] && NODE_BIN=$B/igneumd-0313 # the 0.3.13 swap's binary once box-node-swap.sh has run
IGNEUM_PROOF_VERIFIER=$HOST nohup $NODE_BIN --devnet --appdir=$F/node --rpclisten=127.0.0.1:26610 --evm-rpclisten=127.0.0.1:26790 --listen=0.0.0.0:26611 \
--addpeer=188.245.5.161:26611 $HUBARG --override-params-file=$F/override.json --nodnsseed --disable-upnp --nologfiles --yes >> $F/node.log 2>&1 &
sleep 10
echo "RESULT node_restarted $(stamp) pid=$(pgrep -x igneumd | head -1) verifier=$(grep -c 'IGNEUM_PROOF_VERIFIER\|proof verifier' $F/node.log)"
export THRESHOLD MINER
exec python3 -u $F/in/box-prover.py

48
tools/fleet/box-rehearsal.sh Executable file
View file

@ -0,0 +1,48 @@
#!/usr/bin/env bash
# The class v4 rehearsal node on a fleet box (docs/plans/counter-asic-3-rehearsal.md, 6 October 2026, the re-cut 1-block/s
# object): igneum-devnet-400 from its own genesis in a FRESH appdir, beside the box's live-devnet node (other ports).
# MODE=seed the fork's igneumd, --listen on P2P_LISTEN (the seed box's public port), --utxoindex, no miner
# MODE=miner the fork's igneumd --connect=$SEED, then the miner (igneum-miner-v4 --no-vote): ENGINE=cpu is --engine igneum-pow
# (5 kH/s a box, useless at genesis difficulty 2^27); ENGINE=cuda (default) is the app's shape, the package's
# igneum-worker-cuda through in/dn400-worker.sh (GPU IGNEUM_CUDA_DEVICE, packs under the dn400 appdir). MINER_ONLY=1 leaves the node up and restarts
# the miner and the sampler (run rehearsal-kill.sh first: it stops the loops, the miner and the worker, not the node).
# MODE=stale 0.3.13's igneumd (NODE_BIN) with the same file, --connect=$SEED, no miner: refused at the handshake
# (0.3.13 refuses the v4 fields at parse; OVERRIDE_SRC=<file without them> gives the handshake refusal)
# A 5-minute sampler writes RESULT lines to /root/fleet/out/dn400.log: the watch line, the node's digest, switch and
# signal lines, PoW rejections, the miner's seed and program-id lines and its rejected= count.
set -uo pipefail
F=/root/fleet; OUT=$F/out; R=$F/in; mkdir -p $OUT; exec >> $OUT/dn400.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
MODE="${MODE:?seed|miner|stale}"; LABEL="${LABEL:-box}"; SEED="${SEED:-}"; NODE_BIN="${NODE_BIN:-$R/igneumd-v4}"; MINER_BIN="$R/igneum-miner-v4"
P2P_LISTEN="${P2P_LISTEN:-0.0.0.0:16411}"; RPC=127.0.0.1:16410; JSON=127.0.0.1:16412; EVM=127.0.0.1:16413; ENGINE="${ENGINE:-cuda}"; MINER_ONLY="${MINER_ONLY:-0}"; OVERRIDE_SRC="${OVERRIDE_SRC:-$R/rehearsal-1bps.json}"; APP=$F/dn400; LOG=$F/dn400-node.log
if [ "$MINER_ONLY" != 1 ]; then
pkill -9 -f "^$R/igneumd-v4 " 2>/dev/null; pkill -9 -f "^$R/igneumd-0313 .*devnet-suffix=400" 2>/dev/null; pkill -9 -f "^$MINER_BIN " 2>/dev/null; sleep 2
rm -rf $APP; mkdir -p $APP
cp "$OVERRIDE_SRC" $F/dn400-override.json
echo "RESULT dn400_start $(stamp) mode=$MODE label=$LABEL node=$(sha256sum $NODE_BIN | cut -c1-16) object=$(sha256sum $F/dn400-override.json | cut -c1-16) seed=${SEED:-none}"
PEER=""; [ -n "$SEED" ] && PEER="--connect=$SEED"
EXTRA=""; [ "$MODE" = seed ] && EXTRA="--utxoindex"
nohup $NODE_BIN --devnet --devnet-suffix=400 --appdir=$APP --rpclisten=$RPC --rpclisten-json=$JSON --evm-rpclisten=$EVM --listen=$P2P_LISTEN $PEER --override-params-file=$F/dn400-override.json $EXTRA --enable-unsynced-mining --nodnsseed --disable-upnp --nologfiles --yes >> $LOG 2>&1 &
sleep 10
echo "RESULT dn400_node $(stamp) pid=$(pgrep -f "^$NODE_BIN " | head -1) version=$($NODE_BIN --version 2>&1 | head -1) digest=$(grep -o 'digest: [0-9a-f]*' $LOG | tail -1 | awk '{print $2}') switch=\"$(grep -iE 'class v4' $LOG | head -2 | sed -E 's/^[0-9: .+-]*//' | tr '\n' ';' | cut -c1-300)\""
else
echo "RESULT dn400_miner_restart $(stamp) mode=$MODE label=$LABEL engine=$ENGINE node_pid=$(pgrep -f "^$NODE_BIN " | head -1)"
fi
if [ "$MODE" = miner ]; then
if [ "$ENGINE" = cpu ]; then MARGS="--engine igneum-pow"; else
mkdir -p $APP/packs/prepare; [ -s $APP/packs/live/program.h ] || $MINER_BIN export-pack grpc://$RPC $APP/packs/live >> $OUT/dn400-miner.log 2>&1
echo "RESULT dn400_pack $(stamp) $(grep -hoE 'IGNEUM_(PROGRAM_CLASS|PROGRAM_ID)[^\n]{0,40}' $APP/packs/live/program.h 2>/dev/null | tr '\n' ' ' | cut -c1-120)"
MARGS="--worker $R/dn400-worker.sh --prepare-packs $APP/packs/prepare --exit-on-seed-change"; fi
( while :; do $MINER_BIN mine grpc://$RPC 1 100000000 "$LABEL" $MARGS --no-vote --payout-label "$LABEL" --status-secs 30 >> $OUT/dn400-miner.log 2>&1; rc=$?; echo "RESULT dn400_miner_exit $(stamp) rc=$rc" >> $OUT/dn400.log
# the worker restarts on --pack: re-export it from the node so it holds the CURRENT epoch's program (19:00Z: a worker
# restarted on the epoch-0 pack answered every epoch-2 job "epoch seed mismatch" and the chain stood at DAA 1,202)
[ "$ENGINE" = cpu ] || { rm -rf $APP/packs/live; $MINER_BIN export-pack grpc://$RPC $APP/packs/live >> $OUT/dn400-miner.log 2>&1; echo "RESULT dn400_pack $(stamp) $(grep -hoE 'IGNEUM_(PROGRAM_CLASS|PROGRAM_ID)[^\n]{0,40}' $APP/packs/live/program.h 2>/dev/null | tr '\n' ' ' | cut -c1-120)" >> $OUT/dn400.log; }
sleep 3; done ) &
echo "RESULT dn400_miner_started $(stamp)"
fi
( while :; do
sleep 300
w="$($MINER_BIN watch 1 grpc://$RPC 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | sed -E 's/difficulty=[0-9.]* sink=//' | tail -1)"
echo "RESULT sample $(stamp) $w rejected_node=$(grep -c 'PoW rejected' $LOG) rejected_miner=$(grep STATUS $OUT/dn400-miner.log 2>/dev/null | tail -1 | grep -oE 'rejected=[0-9]+' | head -1) signal=\"$(grep -E 'Program class v4 by miner signal' $LOG | tail -1 | sed -E 's/^[0-9: .+-]*\[INFO \] //' | cut -c1-220)\" ids=\"$(grep -oE 'epoch seed [0-9a-f]{16}.*(class v[34]|program id [0-9a-fx]{18})' $OUT/dn400-miner.log 2>/dev/null | tail -2 | tr '\n' ';' | cut -c1-240)\""
done ) &
echo "RESULT dn400_sampler_started $(stamp)"

12
tools/fleet/box-rig-prover.sh Executable file
View file

@ -0,0 +1,12 @@
#!/usr/bin/env bash
# The rig as N provers for the fleet night: one box-prover.py per card with its own key label, payout, output dir and
# IGNEUM_CUDA_DEVICE (distinct sockets), the miners kept on every card (24 GB cards keep mining while proving).
set -uo pipefail
F=/root/fleet; N=$(nvidia-smi --query-gpu=name --format=csv,noheader | wc -l)
LABEL="${LABEL:-rig}"; WALLET="${WALLET:-0x1919191919191919191919191919191919191919}"
for d in $(seq 0 $((N-1))); do
mkdir -p $F/card$d/out $F/card$d/mine/packs
sleep 20 # the SDK starts one sp1-gpu-server per host and waits a fixed time for its socket; eight at once overran it (15:18Z)
( cd $F/card$d && LABEL="$LABEL-gpu$d" WALLET="$WALLET" THRESHOLD="${THRESHOLD:-}" MINER="${MINER:-keep}" RUN_HOURS="${RUN_HOURS:-9}" IGNEUM_CUDA_DEVICE=$d FLEET_CARD=$d setsid nohup python3 -u $F/in/box-prover.py </dev/null > $F/card$d/out/launch.log 2>&1 & )
done
echo "RESULT rig_provers_started $(date -u +%Y-%m-%dT%H:%M:%SZ) cards=$N"

84
tools/fleet/box-rig.sh Executable file
View file

@ -0,0 +1,84 @@
#!/usr/bin/env bash
# Phase 3 on an 8-card rig (RunPod pod, a container: no systemd, so the rig installer's unit steps are exercised with
# --preflight-only and --dry-run and the units' own scripts are run by hand). After box-setup.sh (node, workers, the
# patched server for the rig's arch, the cuda host):
# A. the rig installer: packaging/linux/install-rig.sh --preflight-only, then --dry-run with the published package
# B. eight miners on the devnet, one per card, 150 s: MH/s per card and the rig's sum, watts
# C. seven core-only provers in parallel (one sp1-gpu-server per card, CUDA_VISIBLE_DEVICES=n, socket /tmp/sp1-cuda-n),
# the v1 shard at 2^25 in a loop for PROVE_SECS: core proofs per minute for the rig and per card, host RAM
# D. the eighth card compressing the v1 shard in a loop for PROVE_SECS: compressed proofs per minute (the compression
# throughput of one card)
# E. the hand-off cost by file: a core proof (2^25, about 25 MB) copied card-to-card over the rig's disk and read back
# F. the aggregator: --mode chain over 8 block fixtures on the big card (the 8-block chain: shards plus aggregation)
# Every result a RESULT line in /root/fleet/out/rig.log; rig.json at the end.
set -uo pipefail
F=/root/fleet; OUT=$F/out; B=/opt/igneum/pkg/bin; FLOOR=/opt/igneum-floor; HOST=$FLOOR/bin/igneum-prove-host
FIX=$FLOOR/prove/proving/fixtures; V1=$FIX/fees-v1-shards2.json
LABEL="${LABEL:-rig}"; WALLET="${WALLET:-0x1919191919191919191919191919191919191919}"; PROVE_SECS="${PROVE_SECS:-300}"; MINE_SECS="${MINE_SECS:-150}"
mkdir -p $OUT $F/mine; exec >> $OUT/rig.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
N=$(nvidia-smi --query-gpu=name --format=csv,noheader | wc -l)
echo "RESULT rig_start $(stamp) label=$LABEL cards=$N $(nvidia-smi --query-gpu=name,memory.total --format=csv,noheader | head -1) host_ram_gb=$(( $(awk '/MemTotal/{print $2}' /proc/meminfo) / 1048576 )) cores=$(nproc)"
for b in igneum-prove-host igneum-prove-export; do for d in $FLOOR/target/release $FLOOR/prove/proving/igneum-prove/target/release; do [ -x $d/$b ] && ln -sfn $d/$b $FLOOR/bin/$b; done; done
# A. the installer's real parts
if [ -d $F/in/linux ] && [ -z "${PROVE_ONLY:-}" ]; then
bash $F/in/linux/install-rig.sh --preflight-only > $OUT/rig-preflight.log 2>&1; echo "RESULT installer_preflight exit=$? fails=$(grep -c 'FAIL' $OUT/rig-preflight.log) warns=$(grep -c 'warn' $OUT/rig-preflight.log) cards=$(grep -c '^ card' $OUT/rig-preflight.log)"
bash $F/in/linux/install-rig.sh --dry-run --yes --wallet "$WALLET" --name "$LABEL" --package-url https://dl.igneum.network/dl/public/igneum-hive-0.3.12.tar.gz --package-sha256 7972af92e7cd9a032303eca4d95b533f53e0e68d1b9cae5bfe406a5b7c30a454 --package-size 24506282 > $OUT/rig-dryrun.log 2>&1; echo "RESULT installer_dryrun exit=$? steps=$(grep -c '^\[dry-run\]' $OUT/rig-dryrun.log) manifest=$(grep -c 'signature verifies' $OUT/rig-dryrun.log) package=$(grep -c 'sha256 7972af92' $OUT/rig-dryrun.log) prover_decision=\"$(grep -o 'prover: .*' $OUT/rig-dryrun.log | head -1 | cut -c1-120)\""
fi
# B. eight miners
if [ -z "${PROVE_ONLY:-}" ]; then
cd $F/mine && rm -rf packs/devnet && $B/igneum-miner export-pack grpc://127.0.0.1:26610 packs/devnet > $OUT/rig-export-pack.log 2>&1
pids=()
for d in $(seq 0 $((N-1))); do
nohup $B/igneum-miner mine grpc://127.0.0.1:26610 1 100000000 "$LABEL-gpu$d" --worker $B/igneum-worker-cuda --worker-args "--device $d --pack packs/devnet" --prepare-packs packs/prepare-$d --exit-on-seed-change --evm-address "$WALLET" --payout-label "$LABEL-gpu$d" --status-secs 10 > $OUT/rig-miner-$d.log 2>&1 &
pids+=($!)
done
cd $F; sleep 60
( while :; do nvidia-smi --query-gpu=index,power.draw,memory.used,utilization.gpu --format=csv,noheader,nounits; sleep 1; done ) > $OUT/rig-samp-mine.csv & SP=$!
sleep "$MINE_SECS"; pkill -P $SP; kill $SP 2>/dev/null
sum=0
for d in $(seq 0 $((N-1))); do
r=$(grep STATUS $OUT/rig-miner-$d.log | grep -o ' now=[0-9.]*' | tail -n 10 | cut -d= -f2 | awk '{s+=$1; n++} END {if (n) printf "%.2f", s/n; else print 0}')
w=$(awk -F', *' -v i=$d '$1==i {s+=$2; n++} END {if (n) printf "%.0f", s/n; else print 0}' $OUT/rig-samp-mine.csv)
echo "RESULT miner card=$d mhs=$r watts=$w"; sum=$(awk -v a=$sum -v b=$r 'BEGIN {print a+b}')
done
echo "RESULT rig_miners $(stamp) cards=$N sum_mhs=$sum watts=$(awk -F', *' '{s+=$2; n++} END {printf "%.0f", s/n*'$N'}' $OUT/rig-samp-mine.csv)"
for p in "${pids[@]}"; do kill $p 2>/dev/null; done; pkill -f '^/opt/igneum/pkg/bin/igneum-worker-cuda'; sleep 3
fi
# C + D. seven core-only provers and one compressing card, in parallel
pkill -9 -x sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock
prove_loop() { # <device> <mode> <threshold> <seconds>
local d=$1 mode=$2 thr=$3 secs=$4 t0=$(date +%s) n=0 tot=0
while [ $(( $(date +%s) - t0 )) -lt $secs ]; do
env HOME=$FLOOR/home IGNEUM_CUDA_DEVICE=$d SP1_PROVER=cuda RUST_LOG=off SP1_GPU_ELEMENT_THRESHOLD=$thr timeout 600 $HOST $V1 --mode $mode --shard 0 --prover "$WALLET" --out $OUT/rig-$mode-$d.json > $OUT/rig-$mode-$d.log 2>&1
s=$(grep -E "^RESULT ($mode) shard" $OUT/rig-$mode-$d.log | tail -1 | grep -o 'prove [0-9.]* s' | grep -o '[0-9.]*' | head -1)
v=$(grep -E "^RESULT ($mode) shard" $OUT/rig-$mode-$d.log | tail -1 | grep -c VERIFIED)
[ "$v" = 1 ] && { n=$((n+1)); tot=$(awk -v a=$tot -v b=${s:-0} 'BEGIN {print a+b}'); }
echo "$(stamp) card=$d $mode n=$n last_s=${s:-fail} verified=$v" >> $OUT/rig-loop-$d.log
done
echo "RESULT prove_loop card=$d mode=$mode threshold=$thr seconds=$secs proofs=$n mean_prove_s=$(awk -v a=$tot -v n=$n 'BEGIN {if (n) printf "%.1f", a/n; else print 0}') per_min=$(awk -v n=$n -v s=$secs 'BEGIN {printf "%.2f", n*60/s}')"
}
( while :; do echo "$(date +%s) $(free -m | awk '/Mem:/{print $3}') $(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | tr '\n' ',')"; sleep 5; done ) > $OUT/rig-samp-prove.csv & SP=$!
lp=()
for d in $(seq 0 $((N-2))); do prove_loop $d core 33554432 "$PROVE_SECS" & lp+=($!); done
prove_loop $((N-1)) compressed 67108864 "$PROVE_SECS" & lp+=($!)
for p in "${lp[@]}"; do wait $p; done
pkill -P $SP; kill $SP 2>/dev/null
echo "RESULT rig_prove $(stamp) host_ram_used_mb_max=$(awk '{if ($2>m) m=$2} END {print m}' $OUT/rig-samp-prove.csv) core_cards=$((N-1)) core_per_min=$(grep 'RESULT prove_loop' $OUT/rig.log | grep 'mode=core' | grep -o 'per_min=[0-9.]*' | cut -d= -f2 | awk '{s+=$1} END {printf "%.2f", s}') compressed_per_min=$(grep 'RESULT prove_loop' $OUT/rig.log | grep 'mode=compressed' | grep -o 'per_min=[0-9.]*' | cut -d= -f2)"
pkill -9 -x sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock
# E. the hand-off cost by file
cf=$(ls -S $F/mine/*.bin $OUT/*.bin 2>/dev/null | head -1); [ -z "$cf" ] && cf=$(find / -name 'block-*-core.bin' -size +1M 2>/dev/null | head -1)
if [ -n "$cf" ]; then
sz=$(stat -c %s "$cf"); t0=$(date +%s%N); cp "$cf" /tmp/handoff.bin; sync; t1=$(date +%s%N); cat /tmp/handoff.bin > /dev/null; t2=$(date +%s%N)
echo "RESULT handoff file=$(basename $cf) bytes=$sz copy_ms=$(( (t1-t0)/1000000 )) read_ms=$(( (t2-t1)/1000000 )) loopback_tcp_ms=$( (nc -l -p 29999 > /dev/null & sleep 0.3; t=$(date +%s%N); nc -q 0 127.0.0.1 29999 < /tmp/handoff.bin; echo $(( ($(date +%s%N)-t)/1000000 )) ) 2>/dev/null)"
fi
# F. the 8-block chain on the big card: --mode chain needs CONSECUTIVE live blocks (the chain rule checks number and parent hash), so this
# runs only when RIG_CHAIN_LIST names eight consecutive fixtures cut from the box's node (after the exec layer executes); else skipped
list="${RIG_CHAIN_LIST:-}"
if [ -n "$list" ]; then
t0=$(date +%s)
env HOME=$FLOOR/home IGNEUM_CUDA_DEVICE=$((N-1)) SP1_PROVER=cuda RUST_LOG=off timeout 1800 $HOST --mode chain --chain "$list" --prover "$WALLET" --save-shards --out $OUT/rig-chain.json > $OUT/rig-chain.log 2>&1; rc=$?
echo "RESULT chain rc=$rc wall_s=$(( $(date +%s) - t0 )) blocks=$(echo "$list" | tr ',' '\n' | wc -l) $(grep -E '^RESULT chain' $OUT/rig-chain.log | tail -1 | cut -c1-200)"
else echo "RESULT chain skipped: no consecutive live fixtures yet (needs the exec layer)"; fi
pkill -9 -x sp1-gpu-server; rm -f /tmp/sp1-cuda-*.sock
echo "RESULT rig_done $(stamp)"

19
tools/fleet/box-segal-host.sh Executable file
View file

@ -0,0 +1,19 @@
#!/usr/bin/env bash
# Builds the segment-aligned prover host (the proving-v1 bundle igneum-prove-wsl2-segal.zip: --mode chain, --save-shards,
# --prev) with the cuda feature into /opt/igneum-segal, and points /opt/igneum-floor/bin/igneum-prove-host at it. The
# prover-floor bundle's host (08:28Z) predates --mode chain (16:20Z, the fleet's chain runs failed on its usage line).
set -uo pipefail
F=/root/fleet; OUT=$F/out; S=/opt/igneum-segal; exec >> $OUT/segal-build.log 2>&1
export PATH="$HOME/.cargo/bin:$PATH" RUSTUP_TOOLCHAIN=stable; unset CARGO_TARGET_DIR
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
echo "RESULT segal_start $(stamp)"
rm -rf $S && mkdir -p $S && unzip -q -o $F/in/igneum-prove-wsl2-segal.zip -d $F/unz-segal && cp -r $F/unz-segal/igneum-prove-wsl2/package/. $S/ && find $S -type f -exec touch {} +
# the warm target of the floor host saves most of the build (the same crates)
[ -d /opt/igneum-floor/prove/proving/igneum-prove/target ] && cp -r /opt/igneum-floor/prove/proving/igneum-prove/target $S/proving/igneum-prove/target 2>/dev/null
cd $S/proving/igneum-prove && t0=$(date +%s)
cargo build --release -p igneum-prove-export -p igneum-prove-host --features igneum-prove-host/cuda -j $(( $(nproc) > 16 ? 16 : $(nproc) )) > $OUT/segal-cargo.log 2>&1; rc=$?
echo "RESULT segal_build_exit $rc time_s=$(( $(date +%s) - t0 ))"
[ $rc -eq 0 ] || { grep -n -A6 '^error' $OUT/segal-cargo.log | head -30; echo "RESULT segal_failed"; exit 2; }
ln -sfn $S/proving/igneum-prove/target/release/igneum-prove-host /opt/igneum-floor/bin/igneum-prove-host
ln -sfn $S/proving/igneum-prove/target/release/igneum-prove-export /opt/igneum-floor/bin/igneum-prove-export
echo "RESULT segal_done $(stamp) host=$(sha256sum /opt/igneum-floor/bin/igneum-prove-host | cut -c1-16) ids=$(/opt/igneum-floor/bin/igneum-prove-host --mode id 2>&1 | grep -o '0x[0-9a-f]*' | head -2 | tr '\n' ' ') chain_mode=$(/opt/igneum-floor/bin/igneum-prove-host 2>&1 | grep -c chain)"

114
tools/fleet/box-setup.sh Executable file
View file

@ -0,0 +1,114 @@
#!/usr/bin/env bash
# Igneum GPU fleet, on-box setup (6 October 2026). Runs as root inside a rented Vast or RunPod container built from
# nvidia/cuda:12.8.1-devel-ubuntu24.04 (or RunPod's CUDA 12.8 image). Idempotent; every step prints a RESULT line to
# /root/fleet/setup.log and the last line is "RESULT setup_done" or "RESULT setup_failed <step>".
#
# What it does, in order: tools (apt), the 0.3.12 Linux package (sha256 checked) unpacked to /opt/igneum/pkg, the node
# started at once on the devnet with the ten-field override (so it syncs while the builds run), rustup, Go 1.27.1 (pinned
# sha256), SP1 v6.8.1 cloned and patched with proving/prover-floor/sp1-gpu-6.8.1-floor.patch, the patched server built
# with CUDA_ARCHS=$ARCHS into /opt/igneum-floor (the SDK's stock server stays under /root/.sp1 for the stock rows), and
# igneum-prove-host + igneum-prove-export built with the cuda feature from the prover-floor host bundle.
#
# Inputs in /root/fleet/in: igneum-prove-wsl2-floor.zip (the host sources, pinned ELFs, fixtures), floor.patch,
# override.json. Env: ARCHS (CUDA arch list, default 86,89,120), LABEL (this box's name), WALLET (0x payout, throwaway).
set -uo pipefail
export DEBIAN_FRONTEND=noninteractive
F=/root/fleet; IN=$F/in; OUT=$F/out; LOG=$F/setup.log
mkdir -p $F $IN $OUT /opt/igneum /opt/igneum-floor/bin /opt/igneum-floor/home/.sp1/bin /opt/igneum-floor/logs
exec > >(tee -a $LOG) 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
say() { echo "$(stamp) $*"; }
fail() { echo "RESULT setup_failed $1 $(stamp)"; exit 2; }
ARCHS="${ARCHS:-86,89,120}"; LABEL="${LABEL:-box}"; WALLET="${WALLET:-}"
PKG_URL=https://dl.igneum.network/dl/public/igneum-hive-0.3.12.tar.gz
PKG_SHA=7972af92e7cd9a032303eca4d95b533f53e0e68d1b9cae5bfe406a5b7c30a454
PATCH_SHA=e81cb0d03b291f9fd4bf0a109d6da2d7c897795c9ffd7f797c0ddce723eee2b1
SEED=188.245.5.161:26611
echo "RESULT start $(stamp) label=$LABEL archs=$ARCHS host=$(hostname) nproc=$(nproc) ram_gb=$(( $(awk '/MemTotal/{print $2}' /proc/meminfo) / 1048576 )) disk_avail=$(df -BG /root | awk 'NR==2{print $4}')"
echo "RESULT gpu $(nvidia-smi --query-gpu=name,memory.total,driver_version,pci.bus_id,power.limit,power.min_limit,power.max_limit,clocks.max.sm --format=csv,noheader 2>&1 | tr '\n' ';')"
echo "RESULT os $(. /etc/os-release; echo "$ID $VERSION_ID") glibc $(ldd --version | head -1 | awk '{print $NF}') nvcc $(nvcc --version 2>/dev/null | grep -o 'release [0-9.]*' || echo none)"
# 1. tools
if ! command -v protoc >/dev/null 2>&1 || ! command -v cmake >/dev/null 2>&1 || ! command -v clang >/dev/null 2>&1; then
say "apt"
apt-get update -qq >/dev/null 2>&1
apt-get install -y -qq build-essential clang cmake pkg-config libssl-dev git curl ca-certificates python3 wget unzip protobuf-compiler jq rsync bc pciutils >/dev/null 2>&1 || fail apt
fi
echo "RESULT tools protoc=$(protoc --version 2>/dev/null) cmake=$(cmake --version | head -1) clang=$(clang --version | head -1 | cut -c1-40)"
# 2. the 0.3.12 Linux package and the node
if [ ! -x /opt/igneum/pkg/bin/igneumd ]; then
say "package"
curl -fsSL -o $F/pkg.tgz "$PKG_URL" || fail package_download
echo "$PKG_SHA $F/pkg.tgz" | sha256sum -c - >/dev/null || fail package_sha256
rm -rf /opt/igneum/pkg.new && mkdir -p /opt/igneum/pkg.new && tar -C /opt/igneum/pkg.new --strip-components=1 -xzf $F/pkg.tgz && mv /opt/igneum/pkg.new /opt/igneum/pkg
fi
B=/opt/igneum/pkg/bin
echo "RESULT package igneumd=$(sha256sum $B/igneumd | cut -c1-16) miner=$(sha256sum $B/igneum-miner | cut -c1-16) cuda_worker=$(sha256sum $B/igneum-worker-cuda | cut -c1-16) version=$($B/igneumd --version 2>&1 | head -1)"
cp $IN/override.json $F/override.json
if [ "${NO_NODE:-0}" = 1 ]; then echo "RESULT node skipped (NO_NODE=1: a Devnet 2 box runs its own chain's node)"
elif ! pgrep -x igneumd >/dev/null; then
say "node"
mkdir -p $F/node
nohup $B/igneumd --devnet --appdir=$F/node --rpclisten=127.0.0.1:26610 --evm-rpclisten=127.0.0.1:26790 --listen=0.0.0.0:26611 \
--addpeer=$SEED --override-params-file=$F/override.json --nodnsseed --disable-upnp --nologfiles --yes > $F/node.log 2>&1 &
sleep 8
fi
echo "RESULT node pid=$(pgrep -x igneumd | head -1) digest=$(grep -o 'digest: [0-9a-f]*' $F/node.log | head -1 | awk '{print $2}') fresh=$(grep -c 'fresh-record rule' $F/node.log)"
# 3. rust and go
export PATH="$HOME/.cargo/bin:/opt/igneum-floor/go/bin:$PATH" RUSTUP_TOOLCHAIN=stable # the SP1 repo's rust-toolchain.toml would pull llvm-tools, rustc-dev and clippy from a CDN some hosts reach badly (the 3060 box, 12:00Z); plain stable builds the server
if [ ! -x "$HOME/.cargo/bin/cargo" ]; then say "rustup"; curl -sSf https://sh.rustup.rs | sh -s -- -y --profile minimal >/dev/null 2>&1 || fail rustup; fi
if [ ! -x /opt/igneum-floor/go/bin/go ]; then
say "go"
curl -sSL -o /opt/igneum-floor/go.tgz https://go.dev/dl/go1.27.1.linux-amd64.tar.gz || fail go_download
echo "63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445 /opt/igneum-floor/go.tgz" | sha256sum -c - >/dev/null || fail go_sha256
tar -xzf /opt/igneum-floor/go.tgz -C /opt/igneum-floor && rm -f /opt/igneum-floor/go.tgz
fi
export GOPATH=/opt/igneum-floor/gopath GOCACHE=/opt/igneum-floor/gocache GOFLAGS=-mod=mod
echo "RESULT toolchain cargo=$(cargo --version) go=$(go version | awk '{print $3}')"
# 4. the patched SP1 GPU server
SRC=/opt/igneum-floor/sp1
echo "RESULT patch_sha256 $(sha256sum $IN/floor.patch | cut -c1-64) expected $PATCH_SHA"
[ "$(sha256sum $IN/floor.patch | cut -c1-64)" = "$PATCH_SHA" ] || fail patch_sha256
if [ ! -x /opt/igneum-floor/bin/sp1-gpu-server ]; then
if [ ! -d "$SRC/.git" ]; then say "clone sp1"; git clone -q --depth 1 --branch v6.8.1 https://github.com/succinctlabs/sp1 "$SRC" || fail clone; fi
cd "$SRC" && git checkout -q -- . && git clean -qfd sp1-gpu/crates >/dev/null 2>&1
echo "RESULT source $(git describe --tags --always) $(git rev-parse HEAD)"
git apply $IN/floor.patch || fail patch_apply
touch sp1-gpu/crates/prover_components/src/builder.rs sp1-gpu/crates/jagged_tracegen/src/lib.rs sp1-gpu/crates/server/src/server.rs sp1-gpu/crates/cuda/src/task.rs
echo "RESULT patched $(git diff --stat | tail -1)"
export CUDA_ARCHS="$ARCHS" CARGO_TARGET_DIR=/opt/igneum-floor/target CUDA_PATH=/usr/local/cuda CUDACXX=/usr/local/cuda/bin/nvcc
J=$(nproc); [ "$J" -gt 16 ] && J=16
say "build server jobs=$J archs=$ARCHS"; t0=$(date +%s)
cargo build --release --bin sp1-gpu-server -j $J > /opt/igneum-floor/logs/build-server.log 2>&1 &
BP=$!; while kill -0 $BP 2>/dev/null; do sleep 60; echo "STAGE server $(stamp) $(( ($(date +%s) - t0) / 60 )) min $(grep -c '^ Compiling' /opt/igneum-floor/logs/build-server.log) crates"; done
wait $BP; rc=$?
echo "RESULT server_build_exit $rc time_s=$(( $(date +%s) - t0 ))"
[ $rc -eq 0 ] || { grep -n -A8 '^error' /opt/igneum-floor/logs/build-server.log | head -60; fail server_build; }
cp /opt/igneum-floor/target/release/sp1-gpu-server /opt/igneum-floor/bin/ && cp /opt/igneum-floor/bin/sp1-gpu-server /opt/igneum-floor/home/.sp1/bin/ && chmod +x /opt/igneum-floor/bin/sp1-gpu-server /opt/igneum-floor/home/.sp1/bin/sp1-gpu-server
fi
echo "RESULT server bytes=$(stat -c %s /opt/igneum-floor/bin/sp1-gpu-server) sha256=$(sha256sum /opt/igneum-floor/bin/sp1-gpu-server | cut -c1-64) version=$(/opt/igneum-floor/bin/sp1-gpu-server --version 2>&1 | head -1) elf=$(cuobjdump --list-elf /opt/igneum-floor/bin/sp1-gpu-server 2>/dev/null | grep -o 'sm_[0-9]*' | sort -u | tr '\n' ' ')"
# 5. the host and the exporter (cuda feature), from the prover-floor bundle
H=/opt/igneum-floor/prove
unset CARGO_TARGET_DIR # the server build's target dir must not catch the host (the 3080, 12:09Z: the host landed under /opt/igneum-floor/target)
if [ ! -x $H/proving/igneum-prove/target/release/igneum-prove-host ] && [ ! -x /opt/igneum-floor/target/release/igneum-prove-host ]; then
say "host"
rm -rf $H && mkdir -p $H && unzip -q -o $IN/igneum-prove-wsl2-floor.zip -d $F/unz && cp -r $F/unz/igneum-prove-wsl2/package/. $H/ && find $H -type f -exec touch {} +
cd $H/proving/igneum-prove || fail host_src
t0=$(date +%s)
cargo build --release -p igneum-prove-export -p igneum-prove-host --features igneum-prove-host/cuda -j $(( $(nproc) > 16 ? 16 : $(nproc) )) > /opt/igneum-floor/logs/build-host.log 2>&1 &
BP=$!; while kill -0 $BP 2>/dev/null; do sleep 60; echo "STAGE host $(stamp) $(( ($(date +%s) - t0) / 60 )) min $(grep -c '^ Compiling' /opt/igneum-floor/logs/build-host.log) crates"; done
wait $BP; rc=$?
echo "RESULT host_build_exit $rc time_s=$(( $(date +%s) - t0 ))"
[ $rc -eq 0 ] || { grep -n -A8 '^error' /opt/igneum-floor/logs/build-host.log | head -60; fail host_build; }
fi
HOST=$H/proving/igneum-prove/target/release/igneum-prove-host; EXPORT=$H/proving/igneum-prove/target/release/igneum-prove-export
[ -x "$HOST" ] || { HOST=/opt/igneum-floor/target/release/igneum-prove-host; EXPORT=/opt/igneum-floor/target/release/igneum-prove-export; }
ln -sfn $HOST /opt/igneum-floor/bin/igneum-prove-host; ln -sfn $EXPORT /opt/igneum-floor/bin/igneum-prove-export
echo "RESULT host sha256=$(sha256sum $HOST | cut -c1-16) export=$(sha256sum $EXPORT | cut -c1-16) ids=$($HOST --mode id 2>&1 | grep -o '0x[0-9a-f]*' | tr '\n' ' ')"
echo "RESULT fixtures $(ls $H/proving/fixtures | wc -l) files, v1=$(sha256sum $H/proving/fixtures/fees-v1-shards2.json | cut -c1-16) empty=$(sha256sum $H/proving/fixtures/block-72854-empty-block-first.json | cut -c1-16)"
[ "${NO_NODE:-0}" = 1 ] || echo "RESULT node_now $($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1)"
echo "RESULT setup_done $(stamp)"

95
tools/fleet/box-standing.sh Executable file
View file

@ -0,0 +1,95 @@
#!/usr/bin/env bash
# The standing-fleet supervisor on a box (the project lead's ruling, 6 October 2026 19:50 UK: rented cards stay up; a standing box is
# never destroyed on a job's end). One loop, started once by lib/standing.py and left alone:
# every 60 s the live-devnet node (ROLE=live: /opt/igneum/pkg/bin/igneumd-<ver> on the live object) or the Devnet 2
# node (ROLE=dn2: --devnet-suffix=2 on dn2-override.json) is restarted if its process is gone; the miner
# loop likewise (the package's igneum-miner on the CUDA worker, --exit-on-seed-change, pack re-exported on
# rc 42); PROVER=1 keeps box-prover.py up (the floor host on /opt/igneum-floor).
# every 10 min the recovery recipe (box-exec-snapshot.sh, the shipper's) runs when the node's exec tip reads 0 "from
# snapshot" while consensus is synced above 1,000 blocks (the dead-exec class of the fleet night, row 1/5).
# every 10 min RESULT standing lines to /root/fleet/out/standing.log: uptime, node pid/version/digest, blocks, daa,
# peers, synced, exec tip, miner MH/s and accepted/rejected, prover segments; lib/standing.py reads them.
# The version follows the live manifest from the Mac (lib/standing.py update step): it stages the new binary as
# /root/fleet/in/igneumd-<sha16> and writes its path into /root/fleet/standing.node; this loop picks it up at the next
# check and restarts the node on it (data dir kept, so seconds). Patterns are anchored on paths (the pgrep rule).
set -uo pipefail
F=/root/fleet; OUT=$F/out; B=/opt/igneum/pkg/bin; mkdir -p $OUT $F/mine/packs; exec >> $OUT/standing-supervisor.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
ROLE="${ROLE:-live}"; LABEL="${LABEL:-standing}"; WALLET="${WALLET:-0x1919191919191919191919191919191919191919}"; PROVER="${PROVER:-0}"; MINE="${MINE:-1}"
HUB_PEER="${HUB_PEER:-}"; SEED="${SEED:-}"; P2P_LISTEN="${P2P_LISTEN:-0.0.0.0:26611}"; STARTED=$(date -u +%s) # P2P_LISTEN: a box with a mapped inbound port listens on it (pool-1: 4463)
if [ "$ROLE" = dn2 ]; then APP=$F/dn2; LOG=$F/dn2-node.log; OVR=$F/dn2-override.json; SUFFIX="--devnet-suffix=2"; PACK=dn2; NET=igneum-devnet-2
else APP=$F/node; LOG=$F/node.log; OVR=$F/override.json; SUFFIX=""; PACK=devnet; NET=igneum-devnet; fi
echo "RESULT standing_start $(stamp) role=$ROLE label=$LABEL prover=$PROVER mine=$MINE"
NODE_RE='^/(opt/igneum/pkg/bin|opt/igneum/pkg\.prev/bin|root/fleet/in)/igneumd(-0313|-[0-9a-f]{16})? --devnet --appdir=' # the LIVE node only, never igneumd-v4 (the rehearsal node, --devnet-suffix=400) nor a Devnet 2 node: at 19:5xZ the hub's live node died and the old pattern took the rehearsal node for it
node_running_bin() { pgrep -af "$NODE_RE" | head -1 | awk '{print $2}'; }
# adopt the node that already runs (a canary or swap may have put a newer binary than the package's in place): never downgrade it
[ -s $F/standing.node ] || { rb=$(node_running_bin); [ -n "$rb" ] && echo "$rb" > $F/standing.node; }
node_bin() { [ -s $F/standing.node ] && cat $F/standing.node || { [ -x $B/igneumd-0313 ] && echo $B/igneumd-0313 || echo $B/igneumd; }; }
node_pid() { pgrep -f "$NODE_RE" | head -1; }
start_node() {
local nb; nb=$(node_bin); local peers=""; [ -n "$HUB_PEER" ] && peers="--addpeer=$HUB_PEER"; [ -n "$SEED" ] && peers="$peers --addpeer=$SEED"
[ "$ROLE" = live ] && peers="$peers --addpeer=188.245.5.161:26611"
local ver; ver="$HOST_VERIFIER"; [ -x /opt/igneum-floor/bin/igneum-prove-host ] && ver=/opt/igneum-floor/bin/igneum-prove-host
pkill -9 -f "$NODE_RE" 2>/dev/null; sleep 2
IGNEUM_PROOF_VERIFIER=${ver:-} setsid nohup $nb --devnet $SUFFIX --appdir=$APP --rpclisten=0.0.0.0:26610 --evm-rpclisten=127.0.0.1:26790 --listen=$P2P_LISTEN $peers --override-params-file=$OVR --nodnsseed --disable-upnp --nologfiles --yes </dev/null >> $LOG 2>&1 &
sleep 8; echo "RESULT standing_node_started $(stamp) bin=$nb pid=$(node_pid) digest=$(grep -o 'digest: [0-9a-f]*' $LOG | tail -1 | awk '{print substr($2,1,16)}')"
# a standalone igneum-miner keeps a dead template subscription after its node restarts (20:24Z: templates frozen,
# fetch_errors climbing, no submits), so the miner and its worker are restarted with the node; the loop brings them back
pkill -9 -f '^/opt/igneum/pkg/bin/igneum-miner mine' 2>/dev/null; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-worker-cuda --serve --device 0 --pack packs' 2>/dev/null; true
}
HOST_VERIFIER=""
miner_loop() {
cd $F/mine; [ -d packs/$PACK ] || $B/igneum-miner export-pack grpc://127.0.0.1:26610 packs/$PACK > $OUT/export-pack.log 2>&1
while :; do
pkill -9 -f '^/opt/igneum/pkg/bin/igneum-miner mine' 2>/dev/null; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-worker-cuda --serve --device 0 --pack packs' 2>/dev/null # one miner per GPU: a restart race left two on three boxes (21:52Z)
$B/igneum-miner mine grpc://127.0.0.1:26610 1 100000000 "$LABEL" --worker $B/igneum-worker-cuda --worker-args "--device 0 --pack packs/$PACK" --prepare-packs packs/$PACK-prepare --exit-on-seed-change --evm-address "$WALLET" --payout-label "$LABEL" --status-secs 60 >> $OUT/miner-0.log 2>&1; rc=$?
echo "RESULT standing_miner_exit $(stamp) rc=$rc"; [ $rc = 42 ] && { rm -rf packs/$PACK; $B/igneum-miner export-pack grpc://127.0.0.1:26610 packs/$PACK >> $OUT/export-pack.log 2>&1; } || sleep 10
done
}
n=0
while :; do
cur=$(node_bin); rb=$(node_running_bin)
if [ -z "$rb" ]; then
# a crash loop on a corrupt datadir (22:10Z: p2-4090-1b, "Corruption: Can't access /029017.sst", 24 restarts): three starts in five
# minutes with a Corruption or "No space left" line in the log's tail move the datadir aside and the node joins fresh
if [ "$(grep -c standing_node_started $OUT/standing-supervisor.log 2>/dev/null)" -ge 3 ] && tail -40 $LOG 2>/dev/null | grep -qE 'Corruption|No such file or directory: while stat' && [ "$(find $OUT/standing-supervisor.log -mmin -5 | wc -l)" = 1 ] && [ "$(grep standing_node_started $OUT/standing-supervisor.log | tail -3 | awk '{print $3}' | sort -u | wc -l)" = 3 ]; then
mv $APP $APP.corrupt-$(date -u +%H%M); mkdir -p $APP; echo "RESULT standing_datadir_reset $(stamp) corrupt datadir moved to $APP.corrupt-$(date -u +%H%M); the node joins fresh"
fi
start_node
elif [ -n "$cur" ] && [ "$rb" != "$cur" ] && [ -s $F/standing.node ]; then echo "RESULT standing_node_update $(stamp) from=$rb to=$cur"; start_node; fi
if [ "$MINE" = 1 ]; then
wl="$($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1)"; wb=$(grep -oE 'blocks=[0-9]+' <<< "$wl" | cut -d= -f2); ws=$(grep -oE 'synced=[a-z]+' <<< "$wl" | cut -d= -f2)
if [ "$ws" = true ] && [ "${wb:-0}" -gt 1000 ]; then
if ! pgrep -f '^/opt/igneum/pkg/bin/igneum-miner mine' >/dev/null && [ -z "${MINER_LOOP_PID:-}" -o ! -e "/proc/${MINER_LOOP_PID:-0}" ]; then miner_loop & MINER_LOOP_PID=$!; echo "RESULT standing_miner_started $(stamp) loop=$MINER_LOOP_PID"; fi
else
# a box never mines from a genesis-only or unsynced view (the block-rate runs' late-pod reorgs; 22:38Z: a miner on an IBD node)
if pgrep -f '^/opt/igneum/pkg/bin/igneum-miner mine' >/dev/null || [ -n "${MINER_LOOP_PID:-}" -a -e "/proc/${MINER_LOOP_PID:-0}" ]; then
[ -n "${MINER_LOOP_PID:-}" ] && kill -9 "$MINER_LOOP_PID" 2>/dev/null; MINER_LOOP_PID=""
pkill -9 -f '^/opt/igneum/pkg/bin/igneum-miner mine' 2>/dev/null; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-worker-cuda --serve --device 0 --pack packs' 2>/dev/null
echo "RESULT standing_miner_stopped $(stamp) node unsynced ($wl): the miner waits for the tip"
fi
fi
fi
if [ "$PROVER" = 1 ] && [ -x /opt/igneum-floor/bin/igneum-prove-host ] && ! pgrep -f '^python3 -u /root/fleet/in/box-prover.py' >/dev/null; then
cd $F && LABEL="$LABEL" WALLET="$WALLET" EXPORT_FROM="${EXPORT_FROM:-0}" CHAIN_NAME=$NET MINER=none RUN_HOURS=240 setsid nohup python3 -u $F/in/box-prover.py </dev/null >> $OUT/prover-launch.log 2>&1 &
echo "RESULT standing_prover_started $(stamp)"
fi
n=$((n+1))
if [ $((n % 10)) = 1 ]; then
# disk: the prover's segment exports (/root/fleet/out/segs, 50 to 500 MB each) filled 60 GB boxes in nine hours (20:42Z: the
# hub's node died on "No space left on device"); keep the last 20 minutes of them and the node log under 300 MB
[ -d $OUT/segs ] && find $OUT/segs -mindepth 1 -maxdepth 1 -mmin +20 -exec rm -rf {} + 2>/dev/null
[ "$(stat -c %s $LOG 2>/dev/null || echo 0)" -gt 300000000 ] && { tail -c 100000000 $LOG > $LOG.tail && mv $LOG.tail $LOG; echo "RESULT standing_log_trimmed $(stamp)"; }
DISK="$(df -h / | tail -1 | awk '{print $4"/"$5}')"
w="$($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1 | sed -E 's/difficulty=[0-9.]* sink=[0-9a-f]* //')"
ex=$(curl -s -m 5 127.0.0.1:26790 -H 'content-type: application/json' -d '{"jsonrpc":"2.0","id":1,"method":"eth_blockNumber","params":[]}' | grep -oE '"result":"0x[0-9a-f]+"' | cut -d'"' -f4)
exn=$((${ex:-0x0})); blocks=$(printf '%s' "$w" | grep -oE 'blocks=[0-9]+' | cut -d= -f2)
m="$(grep STATUS $OUT/miner-0.log 2>/dev/null | tail -1 | grep -oE 'accepted=[0-9]+ rejected=[0-9]+|\([0-9.]+ MH/s inside jobs' | tr '\n' ' ' | tr -d '(')"
echo "RESULT standing $(stamp) up=$(( $(date -u +%s) - STARTED ))s role=$ROLE node=$(node_pid) bin=$cur version=$($cur --version 2>&1 | head -1 | awk '{print $2}') $w exec=$exn miner=\"$m\" prover=$(pgrep -c -f '^python3 -u /root/fleet/in/box-prover.py') disk=$DISK gpu=$(nvidia-smi --query-gpu=utilization.gpu,power.draw --format=csv,noheader,nounits 2>/dev/null | head -1 | tr -d ' ')" >> $OUT/standing.log
# the dead-exec class: synced consensus, exec at 0 from snapshot -> the shipper's recovery recipe (live devnet only)
if [ "$ROLE" = live ] && [ "${blocks:-0}" -gt 1000 ] && [ "$exn" = 0 ] && [ -x $F/in/box-exec-snapshot.sh ] && [ $((n / 10)) -gt 2 ]; then
echo "RESULT standing_recovery $(stamp) exec=0 blocks=$blocks: running box-exec-snapshot.sh"; HUB_PEER="$HUB_PEER" bash $F/in/box-exec-snapshot.sh; sleep 30
fi
fi
sleep 60
done

16
tools/fleet/box-wave-pool.sh Executable file
View file

@ -0,0 +1,16 @@
#!/usr/bin/env bash
# A wave pod moved onto the fleet pool (pool-v0 on the pool pod): the solo miner loop stops, the pool fork's igneum-miner
# (pushed as /root/fleet/in/igneum-miner-pool) mines against the pool with its own node as the verifier (templates and
# seeds checked against grpc://127.0.0.1:26610), one worker, the vote key on. RESULT lines in /root/fleet/out/wave.log.
set -uo pipefail
F=/root/fleet; OUT=$F/out; B=/opt/igneum/pkg/bin; POOL="${POOL:?host:port}"; LABEL="${LABEL:-wave}"; WALLET="${WALLET:?}"
exec >> $OUT/wave.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
pkill -f '^bash in/box-wave.sh'; pkill -x igneum-miner; sleep 2
chmod +x $F/in/igneum-miner-pool
echo "RESULT pool_mode $(stamp) pool=$POOL miner=$(sha256sum $F/in/igneum-miner-pool | cut -c1-16)"
cd $F/mine
while :; do
$F/in/igneum-miner-pool mine grpc://127.0.0.1:26610 1 100000000 "$LABEL" --pool "$POOL" --evm-address "$WALLET" --worker-name "$LABEL" --worker $B/igneum-worker-cuda --worker-args "--device 0 --pack packs/devnet" --prepare-packs packs/prepare --exit-on-seed-change --status-secs 30 >> $OUT/pool-miner.log 2>&1; rc=$?
echo "RESULT pool_miner_exit $(stamp) rc=$rc"; [ $rc = 42 ] && { rm -rf packs/devnet; $B/igneum-miner export-pack grpc://127.0.0.1:26610 packs/devnet >> $OUT/export-pack.log 2>&1; } || sleep 10
done

32
tools/fleet/box-wave.sh Executable file
View file

@ -0,0 +1,32 @@
#!/usr/bin/env bash
# A wave miner (6 October 2026): the 0.3.12 Linux package's workers with the 0.3.13 node (igneumd-0313, pushed) on the
# thirteen-field override, peered with the fleet hub and the seed, one igneum-miner with its vote key on (the default);
# no builds, no prover. Prints RESULT lines to /root/fleet/out/wave.log: node digest, sync, the miner's rate.
set -uo pipefail
export DEBIAN_FRONTEND=noninteractive
F=/root/fleet; OUT=$F/out; mkdir -p $F/in $OUT $F/mine/packs $F/node; exec >> $OUT/wave.log 2>&1
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
LABEL="${LABEL:-wave}"; WALLET="${WALLET:-0x1919191919191919191919191919191919191919}"; HUB_PEER="${HUB_PEER:-213.173.107.74:16516}"
echo "RESULT wave_start $(stamp) label=$LABEL gpu=$(nvidia-smi --query-gpu=name,memory.total,driver_version --format=csv,noheader | head -1)"
command -v curl >/dev/null || { apt-get update -qq >/dev/null 2>&1; apt-get install -y -qq curl ca-certificates python3 >/dev/null 2>&1; }
if [ ! -x /opt/igneum/pkg/bin/igneum-miner ]; then
# the current hive package from the public index (0.3.12's file left the host when 0.3.13 was published, 16:17Z)
read -r PKG_PATH PKG_SHA PKG_VER <<< "$(curl -fsSL -m 30 https://dl.igneum.network/dl/public/igneum-downloads.json | python3 -c 'import sys,json; d=json.load(sys.stdin)["files"]["miner-hive"]; print(d["path"], d["sha256"], d["version"])')"
curl -fsSL -o $F/pkg.tgz "https://dl.igneum.network$PKG_PATH" || { echo "RESULT wave_failed package $PKG_PATH"; exit 2; }
echo "$PKG_SHA $F/pkg.tgz" | sha256sum -c - >/dev/null || { echo "RESULT wave_failed sha256"; exit 2; }
echo "RESULT package $(stamp) $PKG_VER $PKG_PATH ${PKG_SHA:0:16}"
mkdir -p /opt/igneum/pkg && tar -C /opt/igneum/pkg --strip-components=1 -xzf $F/pkg.tgz
fi
B=/opt/igneum/pkg/bin
[ "$(sha256sum $F/in/igneumd-0313 | cut -c1-16)" = d6350586fe837b1f ] || { echo "RESULT wave_failed node binary"; exit 2; }
cp $F/in/igneumd-0313 $B/igneumd-0313; chmod +x $B/igneumd-0313; cp $F/in/ov13.json $F/override.json
ldconfig -p | grep -q libnvrtc.so.12 || echo "RESULT note no libnvrtc.so.12 on the path (the worker dlopens it)"
pgrep -f '^/opt/igneum/pkg/bin/igneumd-0313' >/dev/null || nohup $B/igneumd-0313 --devnet --appdir=$F/node --rpclisten=127.0.0.1:26610 --listen=0.0.0.0:26611 --addpeer=$HUB_PEER --addpeer=188.245.5.161:26611 --override-params-file=$F/override.json --nodnsseed --disable-upnp --nologfiles --yes > $F/node.log 2>&1 &
sleep 10; echo "RESULT node $(stamp) digest=$(grep -o 'digest: [0-9a-f]*' $F/node.log | tail -1 | awk '{print substr($2,1,16)}') pid=$(pgrep -f '^/opt/igneum/pkg/bin/igneumd-0313' | head -1)"
for i in $(seq 1 90); do w="$($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1)"; [[ "$w" == *synced=true* ]] && break; sleep 10; done
echo "RESULT synced $(stamp) $(printf '%s' "$w" | sed -E 's/difficulty=[0-9.]* sink=[0-9a-f]* //') after $((i*10)) s"
cd $F/mine && rm -rf packs/devnet && $B/igneum-miner export-pack grpc://127.0.0.1:26610 packs/devnet > $OUT/export-pack.log 2>&1
while :; do
$B/igneum-miner mine grpc://127.0.0.1:26610 1 100000000 "$LABEL" --worker $B/igneum-worker-cuda --worker-args "--device 0 --pack packs/devnet" --prepare-packs packs/prepare --exit-on-seed-change --evm-address "$WALLET" --payout-label "$LABEL" --status-secs 30 >> $OUT/miner.log 2>&1; rc=$?
echo "RESULT miner_exit $(stamp) rc=$rc"; [ $rc = 42 ] && { rm -rf packs/devnet; $B/igneum-miner export-pack grpc://127.0.0.1:26610 packs/devnet >> $OUT/export-pack.log 2>&1; } || sleep 10
done

View file

@ -0,0 +1,55 @@
#!/usr/bin/env python3
"""The block-rate run's collector (docs/analysis/block-rate-devnet2.md). Every 60 s, from the Devnet 2 seed and every
Devnet 2 miner box (the library's chain-side readers): height, daa, blue score, tips, peers, exec tip; the seed's
counts of accepted blocks and red blocks, reorg depths, finality lock lines (delay from determined to LOCKED, voters
and weight), p2p bytes (the node's /proc/<pid>/net or ss), CPU and RSS of the node. One JSON line per minute to
~/Desktop/fleet/bps/<run>.jsonl; `--report <run>` prints the run's table and the payout intervals per tier.
"""
import sys, os, json, time, datetime, re
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))); from lib import Box, Registry
ROOT = os.path.expanduser("~/Desktop/fleet/bps"); os.makedirs(ROOT, exist_ok=True)
def now(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
SEED_CMD = r"""L=${SEED_LOG:-/root/fleet/dn2-node.log}; P=$(pgrep -f "${SEED_PAT:-^/root/fleet/in/igneumd}" | head -1)
echo accepted=$(grep -c 'PoW accepted' $L) red=$(grep -ciE 'red block|colored red|is red' $L) reorg_max=$(grep -oE 'selected-chain reorg: [0-9]+ chain blocks' $L | grep -oE '[0-9]+ chain' | awk '{if ($1>m) m=$1} END {print m+0}') locks=$(grep -c 'LOCKED' $L)
grep -E 'Finality: (checkpoint [0-9]+ determined|checkpoint [0-9]+ LOCKED|certificate at index [0-9]+ received)' $L | tail -40 | sed -E 's/^([0-9-]+ [0-9:.]+)[^]]*\] //' | cut -c1-170
echo cpu_rss=$(ps -o %cpu=,rss= -p $P 2>/dev/null | tr -s ' ') rx_tx=$(cat /proc/$P/net/dev 2>/dev/null | awk 'NR>2 && $1!="lo:" {rx+=$2; tx+=$10} END {print rx, tx}')
${SEED_MINER:-/opt/igneum/pkg/bin/igneum-miner} watch 1 grpc://127.0.0.1:${SEED_RPC:-26610} 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1
curl -s -m 6 -X POST -H 'Content-Type: application/json' --data '{"jsonrpc":"2.0","id":1,"method":"igneum_getExecStatus","params":[]}' http://127.0.0.1:${SEED_EVM:-26790}/ | python3 -c 'import sys,json; r=sys.stdin.read(); d=json.loads(r).get("result",{}) if r.strip() else {}; print("exec", int(d.get("executedTip","0x0"),16))' 2>/dev/null"""
def sample(run):
reg = Registry.load()
sb = next((b for i, b in reg.items() if b.get("dn2_seed") and b.get("state") != "destroyed"), None)
seed = Box(sb["ssh_host"], sb["ssh_port"], sb["label"], None, sb.get("wallet"), sb.get("provider"), sb.get("ssh_user", "root")) if sb else None; seed_env = (sb or {}).get("seed_env")
row = {"t": now(), "run": run}
if seed:
env = " ".join(f"{k}={v}" for k, v in (seed_env or {}).items()); rc, out, err = seed.run((env + " " if env else "") + SEED_CMD, 60); row["seed_raw"] = out[-4000:]
for l in out.split("\n"):
if l.startswith("accepted="): row.update(dict(kv.split("=") for kv in l.split()))
elif l.startswith("blocks="):
row.update({"w_" + k: v for k, v in (kv.split("=") for kv in l.split() if "=" in kv)})
try: row["red_share"] = round(1 - int(row["w_blue"]) / max(int(row["w_daa"]), 1), 4) # the node logs no red line; red = 1 - blue score / DAA score
except Exception: pass
elif l.startswith("exec "): row["exec_tip"] = int(l.split()[1])
elif l.startswith("cpu_rss="): row["cpu_rss"] = l
fin = [l for l in out.split("\n") if "Finality:" in l]; row["finality_tail"] = fin[-12:]
miners = [b for b in reg.values() if b.get("state") != "destroyed" and b.get("ssh_host") and (b.get("devnet2") or b.get("stage") == "bps")]
row["miner_boxes"] = len(miners)
with open(os.path.join(ROOT, f"{run}.jsonl"), "a") as f: f.write(json.dumps(row) + "\n")
print(row["t"], run, {k: row.get(k) for k in ("accepted", "red", "reorg_max", "locks", "w_blocks", "w_daa", "w_blue", "w_tips", "w_peers", "exec_tip", "miner_boxes")}, flush=True)
def report(run):
rows = [json.loads(l) for l in open(os.path.join(ROOT, f"{run}.jsonl"))]
if len(rows) < 2: print("too few rows"); return
a, z = rows[0], rows[-1]
secs = (datetime.datetime.strptime(z["t"], "%Y-%m-%dT%H:%M:%SZ") - datetime.datetime.strptime(a["t"], "%Y-%m-%dT%H:%M:%SZ")).total_seconds()
blocks = int(z.get("w_blocks", 0)) - int(a.get("w_blocks", 0)); blue = int(z.get("w_blue", 0)) - int(a.get("w_blue", 0))
print(f"run {run}: {secs/60:.1f} min, blocks {blocks} ({blocks/max(secs,1):.2f}/s), blue score +{blue} ({blue/max(secs,1):.2f}/s), red share {100*(1-blue/max(blocks,1)):.1f}%, max reorg {max(int(r.get('reorg_max',0)) for r in rows)}, exec lag at end {int(z.get('w_blocks',0)) - int(z.get('exec_tip',0))} blocks, locks {int(z.get('locks',0)) - int(a.get('locks',0))}")
# payout interval per tier at the measured block rate: blocks/s x (miner hash / network hash)
bps = blocks / max(secs, 1)
for net_name, net in (("tonight (2.0 GH/s)", 2.0e9), ("1 TH/s", 1e12), ("10 TH/s", 1e13), ("100 TH/s", 1e14)):
for card, mhs in (("4070 (28 MH/s)", 28e6), ("5090 (128 MH/s)", 128e6), ("8x 4090 rig (459 MH/s)", 459e6)):
per_s = bps * mhs / net; iv = 1 / per_s if per_s else float("inf")
print(f" {net_name:<20} {card:<24} one block every {iv/60:,.1f} min ({iv/3600:,.2f} h)")
if __name__ == "__main__":
if sys.argv[1] == "--report": report(sys.argv[2])
else:
run = sys.argv[1]
while True: sample(run); time.sleep(60)

76
tools/fleet/canary-next.sh Executable file
View file

@ -0,0 +1,76 @@
#!/bin/bash
# bash 3.2 (the Mac's): plain arrays, no mapfile, no associative arrays (lookup files under $WORK)
# The 0.3.14 canary gate on the LIVE devnet (6 October 2026, 17:00Z, the coordinator's form for a publish that moves no
# digest and no height): the release's Linux igneumd goes onto the six restart-path provers and two miners within two
# minutes of the URL (the override file kept), runs TEN minutes, then one line: PASS only with zero rejected blocks on
# every box, exec state roots equal on every box and on the hub at a common height, at least one segment record paid on
# the new binary (a prover's paidSegments up), every box's node on the new version; FAIL names the box and the line.
# canary.sh --binary <url|file> --sha256 <hex> [--minutes 10]
set -uo pipefail
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"; ROOT="$HOME/Desktop/fleet"
BIN=""; SHA=""; MINUTES=10
while [[ $# -gt 0 ]]; do case "$1" in --binary) BIN="$2"; shift 2 ;; --sha256) SHA="$2"; shift 2 ;; --minutes) MINUTES="$2"; shift 2 ;; *) echo "unknown $1"; exit 2 ;; esac; done
[[ -n "$BIN" && -n "$SHA" ]] || { echo "usage: $0 --binary <url|file> --sha256 <hex> [--minutes 10]"; exit 2; }
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
LOG="$ROOT/canary-$(date -u +%Y%m%dT%H%M%SZ).log"; exec > >(tee -a "$LOG") 2>&1
PROVERS="${PROVERS:-p1-3090 p1-5090 p2-3090-1 p2-3090-2 p2-3090-3 p2-4090-1b}"; MINERS="${MINERS:-p1-3080 p1-a5000}"
fail() { echo "FAIL $(stamp) $*"; exit 1; }
LOCAL="$ROOT/canary-igneumd"
if [[ "$BIN" == http* ]]; then curl -fsSL -o "$LOCAL" "$BIN" || fail "download $BIN"; else cp "$BIN" "$LOCAL"; fi
[[ "$(shasum -a 256 "$LOCAL" | cut -c1-64)" == "$SHA" ]] || fail "sha256 of the binary is $(shasum -a 256 "$LOCAL" | cut -c1-16), not ${SHA:0:16}"
WANT="${SHA:0:16}_$(file "$LOCAL" >/dev/null; echo igneumd_2.1.0)" # the running binary's sha256 (first 16) and its --version word, per box
echo "CANARY start $(stamp) binary=${SHA:0:16} version=$WANT boxes: $PROVERS $MINERS"
ROWS=(); while IFS= read -r line; do ROWS+=("$line"); done < <(python3 - "$PROVERS $MINERS" <<'PY'
import json, os, sys; reg=json.load(open(os.path.expanduser("~/Desktop/fleet/boxes.json"))); want=sys.argv[1].split()
for lab in want:
b=next((v for v in reg.values() if v["label"]==lab and v.get("state")!="destroyed"), None)
if b: print(lab, b["ssh_host"], b["ssh_port"], b["wallet"])
hub=next(v for v in reg.values() if v.get("hub") and v.get("state")!="destroyed" and v.get("hub_peer")); print("hub-1", hub["ssh_host"], hub["ssh_port"], hub["wallet"], hub["hub_peer"])
PY
)
HUB="$(printf '%s\n' "${ROWS[@]}" | awk '$1=="hub-1"')"; HUB_HOST="$(awk '{print $2}' <<< "$HUB")"; HUB_PORT="$(awk '{print $3}' <<< "$HUB")"; HUB_PEER="$(awk '{print $5}' <<< "$HUB")"
SSH() { ssh -i ~/.ssh/igneum-fleet -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=20 -o ServerAliveInterval=10 -o ServerAliveCountMax=3 -o BatchMode=yes -p "$2" "root@$1" "${@:3}"; }
read_box() { SSH "$1" "$2" 'B=/opt/igneum/pkg/bin; w=$($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o "blocks=[0-9]*.*synced=[a-z]*" | tail -1); d=$(grep -o "daa=[0-9]*" <<< "$w" | cut -d= -f2); e=$(curl -s -m 6 -X POST -H "Content-Type: application/json" --data "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"igneum_getProvingStatus\",\"params\":[]}" http://127.0.0.1:26790/ | python3 -c "import sys,json; r=sys.stdin.read(); x=json.loads(r).get(\"result\",{}) if r.strip() else {}; print(int(x.get(\"tipDaa\",\"0x0\"),16), x.get(\"v1\",{}).get(\"paidSegments\",0))" 2>/dev/null); ex=$(curl -s -m 6 -X POST -H "Content-Type: application/json" --data "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"igneum_getExecStatus\",\"params\":[]}" http://127.0.0.1:26790/ | python3 -c "import sys,json; r=sys.stdin.read(); x=json.loads(r).get(\"result\",{}) if r.strip() else {}; print(int(x.get(\"executedTip\",\"0x0\"),16))" 2>/dev/null); v=$(pgrep -fa "^/opt/igneum/pkg/bin/igneumd" | head -1 | awk "{print \$2}"); ver=$(sha256sum "$v" 2>/dev/null | cut -c1-16)_$($v --version 2>&1 | head -1 | tr " " "_"); rej=$(grep -cE "got reject message|PoW rejected|block rejected|invalid block" /root/fleet/node.log 2>/dev/null); echo "${d:-0} ${e:-0 0} ${ex:-0} ${ver:-none} ${rej:-0}"' 2>/dev/null; }
root_at() { SSH "$1" "$2" "curl -s -m 6 -X POST -H 'Content-Type: application/json' --data '{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"eth_getBlockByNumber\",\"params\":[\"$3\",false]}' http://127.0.0.1:26790/ | python3 -c 'import sys,json; b=json.load(sys.stdin).get(\"result\") or {}; print(b.get(\"stateRoot\",\"none\"))'" 2>/dev/null; }
WORK="$(mktemp -d)"; paid0() { cat "$WORK/paid0.$1" 2>/dev/null || echo 0; }; rej0() { cat "$WORK/rej0.$1" 2>/dev/null || echo 0; }
echo "CANARY baseline $(stamp)"
for r in "${ROWS[@]}"; do set -- $r; x="$(read_box "$2" "$3")"; echo " $1: daa=$(awk '{print $1}' <<< "$x") tip=$(awk '{print $2}' <<< "$x") paidSeg=$(awk '{print $3}' <<< "$x") exec=$(awk '{print $4}' <<< "$x") ver=$(awk '{print $5}' <<< "$x") rej=$(awk '{print $6}' <<< "$x")"; awk '{print $3}' <<< "$x" > "$WORK/paid0.$1"; awk '{print $6}' <<< "$x" > "$WORK/rej0.$1"; done
# install: the binary swap with the override file kept (box-node-swap.sh STEP=binary, which refuses on a digest change);
# SKIP_INSTALL=1 when every box already runs the binary (the window alone)
T0=$(date +%s)
if [[ "${SKIP_INSTALL:-0}" == 1 ]]; then
for r in "${ROWS[@]}"; do set -- $r; [[ "$1" == hub-1 ]] && continue; x="$(read_box "$2" "$3")"; v="$(awk '{print $5}' <<< "$x")"; [[ "$v" == "$WANT" ]] || fail "$1: SKIP_INSTALL but the box runs $v, not $WANT"; done
echo "CANARY install skipped $(stamp): every box already runs $WANT"
else
for r in "${ROWS[@]}"; do set -- $r; [[ "$1" == hub-1 ]] && continue
okc=0; for try in 1 2 3; do scp -i ~/.ssh/igneum-fleet -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=20 -o ServerAliveInterval=10 -o ServerAliveCountMax=3 -P "$3" "$LOCAL" "root@$2:/root/fleet/in/igneumd-0313" >/dev/null 2>&1 && { okc=1; break; }; sleep 5; done # Vast's ssh proxy drops a connection now and then (17:14Z); three tries before a FAIL
[[ $okc == 1 ]] || fail "$1: scp failed three times"
oks=0; for try in 1 2 3; do SSH "$2" "$3" "cd /root/fleet && chmod +x in/igneumd-0313 && [ \"\$(sha256sum in/igneumd-0313 | cut -c1-64)\" = $SHA ] && mv out/node-swap.log out/node-swap-canary-prev.log 2>/dev/null; STEP=binary NODE_SHA256=$SHA EXPECT_DIGEST=b18ed271f75dd46406d230f4156c37472127415a4c32c558bac662f6f840e61c HUB_PEER=$HUB_PEER HUB_SSH=$HUB_HOST HUB_PORT=$HUB_PORT setsid nohup in/box-node-swap.sh </dev/null >/dev/null 2>&1 & echo swapped" >/dev/null 2>&1 && { oks=1; break; }; sleep 5; done
[[ $oks == 1 ]] || fail "$1: install failed three times"
done
echo "CANARY installed $(stamp) on $(( ${#ROWS[@]} - 1 )) boxes in $(( $(date +%s) - T0 )) s"
sleep 75
fi
if [[ "${SKIP_INSTALL:-0}" != 1 ]]; then
for r in "${ROWS[@]}"; do set -- $r; [[ "$1" == hub-1 ]] && continue
l="$(SSH "$2" "$3" "grep -E '^RESULT (node_started|swap_failed)' /root/fleet/out/node-swap.log | tail -1")"; echo " $1: ${l:0:160}"; [[ "$l" == *digest_ok* ]] || fail "$1: the swap did not reach digest_ok (${l:0:200})"
done
# the provers back on (box-prover.sh keeps a running igneumd-0313 with the verifier)
for p in $PROVERS; do r="$(printf '%s\n' "${ROWS[@]}" | awk -v l="$p" '$1==l')"; set -- $r; SSH "$2" "$3" "cd /root/fleet && pkill -f '^python3 -u /root/fleet/in/box-prover.py'; pkill -x igneum-miner; sleep 2; mv out/prover.log out/prover-canary-prev.log 2>/dev/null; LABEL=$1 WALLET=$4 HUB_PEER=$HUB_PEER setsid nohup in/box-prover.sh </dev/null >/dev/null 2>&1 & echo prover" >/dev/null || fail "$1: prover relaunch"; done
for p in $MINERS; do r="$(printf '%s\n' "${ROWS[@]}" | awk -v l="$p" '$1==l')"; set -- $r; SSH "$2" "$3" "cd /root/fleet/mine && pkill -x igneum-miner; sleep 2; rm -rf packs/devnet; /opt/igneum/pkg/bin/igneum-miner export-pack grpc://127.0.0.1:26610 packs/devnet >/dev/null 2>&1; setsid nohup /opt/igneum/pkg/bin/igneum-miner mine grpc://127.0.0.1:26610 1 100000000 $1 --worker /opt/igneum/pkg/bin/igneum-worker-cuda --worker-args '--device 0 --pack packs/devnet' --prepare-packs packs/prepare --exit-on-seed-change --evm-address $4 --payout-label $1 --status-secs 30 </dev/null >> /root/fleet/out/mine-only-0.log 2>&1 & echo miner" >/dev/null || fail "$1: miner relaunch"; done
fi
echo "CANARY running $(stamp) $MINUTES min"; sleep $((MINUTES * 60))
echo "CANARY checks $(stamp)"; ok=1; paid_any=0; H=""
for r in "${ROWS[@]}"; do set -- $r; x="$(read_box "$2" "$3")"; paid="$(awk '{print $3}' <<< "$x")"; ex="$(awk '{print $4}' <<< "$x")"; ver="$(awk '{print $5}' <<< "$x")"; rej="$(awk '{print $6}' <<< "$x")"
echo " $1: tip=$(awk '{print $2}' <<< "$x") paidSeg=$paid exec=$ex ver=$ver rej=$rej (was $(rej0 $1))"
[[ "$1" == hub-1 ]] || { [[ "$ver" == "$WANT" ]] || { echo "FAIL $1: version $ver, want $WANT"; ok=0; }; }
[[ "${rej:-0}" -le "$(rej0 $1)" ]] || { echo "FAIL $1: $(( rej - $(rej0 $1) )) new rejected blocks (grep 'reject' /root/fleet/node.log)"; ok=0; }
[[ "$1" == hub-1 ]] || { for p in $PROVERS; do [[ "$1" == "$p" && "${paid:-0}" -gt "$(paid0 $1)" ]] && paid_any=1; done; }
[[ -z "$H" || "${ex:-0}" -lt "$H" ]] && H="${ex:-0}"
done
H=$((H > 20 ? H - 20 : 1)); HX="$(printf '0x%x' "$H")"; first=""
for r in "${ROWS[@]}"; do set -- $r; rt="$(root_at "$2" "$3" "$HX")"; echo " $1 root@$H ${rt:0:18}"; [[ -z "$first" ]] && first="$rt"; [[ "$rt" == "$first" ]] || { echo "FAIL $1: exec root at $H ${rt:0:18} differs from ${first:0:18}"; ok=0; }; done
[[ $paid_any == 1 ]] || { echo "FAIL no segment record paid on any prover during the $MINUTES min (paidSegments did not rise)"; ok=0; }
res=FAIL; [[ $ok == 1 ]] && res=PASS
echo "$res $(stamp) canary 0.3.14 ${SHA:0:16} version $WANT: $(( ${#ROWS[@]} - 1 )) boxes, $MINUTES min, rejects 0, roots equal at $H incl. the hub, segment paid $paid_any; log $LOG"
[[ $ok == 1 ]]

View file

@ -0,0 +1,33 @@
"""The post-deploy confirmation (coordinator + shipper, 21:4xZ): the forced poisoned-peer case against one 0.3.15 standing box,
in the honest layout. Relay R = pool-1 (inbound 213.173.98.36:18451) temporarily on 6615571c (7961c5f1: accepts and relays 1026
blocks) over its own 1026-block datadir (/root/fleet/node.1026-2109, kept from 20:34Z to 21:09Z). Poisoned miner M = p2-3090-3 on
713ef876 (06211d55, stamps 1026) with a wiped datadir, peered with R only, mining five bounded minutes. Target T = p2-4090-3 on
7f0bde70, which holds R as a peer. Decisive reads every minute: T's log naming version 1026 as rejected from R's address and never
'Accepted ... via relay' for M's hashes; T never rejected or disconnected by the hands or the hub; the hub's count by peer (R's own
line may climb: it is poisoned and old-rule-blind, expected). Then everything wiped and back on 7f0bde70 / 0.3.14 as before.
Usage: poison-confirm.py (after the publish-1 table is in)."""
import sys, os, time, subprocess; sys.path.insert(0, "/Users/joshm/Projects/igneum-wt-gpu-fleet/tools/fleet"); from lib import Box, Registry
def st(): return time.strftime("%H:%M:%SZ", time.gmtime())
R=Box.from_registry("pool-1"); M=Box.from_registry("p2-3090-3"); T=Box.from_registry("p2-4090-3"); hub=Box.from_registry("hub-1")
_, rM = Registry.find("p2-3090-3"); _, rR = Registry.find("pool-1")
W="/opt/igneum/pkg/bin/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -oE 'blocks=[0-9]+|peers=[0-9]+|synced=[a-z]+' | tr '\\n' ' '"
KILL="pkill -9 -f '^bash in/box-standing.sh'; pkill -9 -f '^python3 -u /root/fleet/in/box-prover.py'; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-miner mine'; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-worker-cuda --serve --device 0 --pack packs'; pkill -9 -f '^/(opt/igneum/pkg/bin|root/fleet/in)/igneumd(-0313|-[0-9a-f]{16})? --devnet --appdir='; sleep 3"
print(st(), "T before:", T.run("echo vline=$(grep -oE 'igneumd/2\\.1\\.0-[0-9a-f]+' /root/fleet/node.log | tail -1) rej1026=$(grep -c 'wrong block version' /root/fleet/node.log) got=$(grep -c 'got reject message' /root/fleet/node.log); "+W, 30)[1].replace("\n"," "), flush=True)
# R: the relay on 6615571c over its 1026 datadir (kept aside as node.clean while it runs)
rc,out,err=R.run(f"cd /root/fleet && {KILL}; [ -d node.clean ] || mv node node.clean; rm -rf node; cp -a node.1026-2109 node; setsid nohup in/igneumd-6615571ced35f694 --devnet --appdir=/root/fleet/node --rpclisten=0.0.0.0:26610 --evm-rpclisten=127.0.0.1:26790 --listen=0.0.0.0:4463 --addpeer=213.173.107.74:16516 --addpeer=188.245.5.161:26611 --override-params-file=/root/fleet/override.json --nodnsseed --disable-upnp --nologfiles --yes </dev/null >> node.log 2>&1 & for i in $(seq 1 40); do sleep 10; w=$({W}); case \"$w\" in *synced=true*) break;; esac; done; echo R vline=$(grep -oE 'igneumd/2\\.1\\.0-[0-9a-f]+' node.log | tail -1) after=$((i*10))s $w", 500)
print(st(), "relay:", out.replace("\n"," "), err[:80], flush=True)
# M: the 713ef876 miner, wiped datadir, peered with R only; syncs from R then mines
rc,out,err=M.run(f"cd /root/fleet && {KILL}; mv node node.pre-poison-$(date -u +%H%M); mkdir node; setsid nohup in/igneumd-06211d55b8cce1c1 --devnet --appdir=/root/fleet/node --rpclisten=0.0.0.0:26610 --evm-rpclisten=127.0.0.1:26790 --listen=0.0.0.0:26611 --addpeer=213.173.98.36:18451 --override-params-file=/root/fleet/override.json --nodnsseed --disable-upnp --nologfiles --yes </dev/null >> node.log 2>&1 & for i in $(seq 1 90); do sleep 10; w=$({W}); case \"$w\" in *synced=true*) break;; esac; done; echo M vline=$(grep -oE 'igneumd/2\\.1\\.0-[0-9a-f]+' node.log | tail -1) synced_after=$((i*10))s $w; cd mine && rm -rf packs/devnet && /opt/igneum/pkg/bin/igneum-miner export-pack grpc://127.0.0.1:26610 packs/devnet >/dev/null 2>&1; setsid nohup /opt/igneum/pkg/bin/igneum-miner mine grpc://127.0.0.1:26610 1 100000000 p2-3090-3 --worker /opt/igneum/pkg/bin/igneum-worker-cuda --worker-args '--device 0 --pack packs/devnet' --prepare-packs packs/prepare --exit-on-seed-change --evm-address {rM.get('wallet')} --payout-label p2-3090-3 --status-secs 30 </dev/null >> /root/fleet/out/miner-poison.log 2>&1 & echo M mining T0=$(date -u +%T)", 1100)
print(st(), "poisoned miner:", out.replace("\n"," "), err[:80], flush=True)
T0=time.time()
while time.time()-T0 < 300:
time.sleep(60)
m=M.run("grep STATUS /root/fleet/out/miner-poison.log | tail -1 | grep -oE 'accepted=[0-9]+ rejected=[0-9]+'; echo got=$(grep -c 'got reject message' /root/fleet/node.log) "+W, 30)[1].replace("\n"," ")
r=R.run("echo R_wbv=$(grep -c 'wrong block version' /root/fleet/node.log) R_got=$(grep -c 'got reject message' /root/fleet/node.log) R_relayed=$(grep -c 'via relay' /root/fleet/node.log) "+W, 30)[1].replace("\n"," ")
t=T.run("echo T_rej1026=$(grep -c 'wrong block version' /root/fleet/node.log) T_from_R=$(grep 'wrong block version' /root/fleet/node.log | grep -c 213.173.98.36) T_got=$(grep -c 'got reject message' /root/fleet/node.log) T_peers=$(grep -oE 'Connected to (outgoing|incoming) peer [0-9.:]+' /root/fleet/node.log | awk '{print $NF}' | cut -d: -f1 | sort -u | tr '\\n' ','); grep 'wrong block version' /root/fleet/node.log | tail -1 | cut -c12-200", 30)[1].replace("\n"," | ")
h=hub.run("awk '$2 >= \"'$(date -u -d '-6 min' +%H:%M:%S)'\"' /root/fleet/node.log | grep 'wrong block version' | grep -oE 'peer [0-9.]+' | sort | uniq -c | sort -rn | tr '\\n' ';'", 30)[1].strip()
print(st(), f"+{int(time.time()-T0)}s M: {m} | R: {r} | T: {t} | hub by peer: {h}", flush=True)
# restore: M wiped back to 0.3.14 fresh under its supervisor; R wiped back to 7f0bde70 (its standing pointer) under its supervisor
rc,out,err=M.run(f"cd /root/fleet && {KILL}; mv node node.1026-$(date -u +%H%M); mkdir node; echo /opt/igneum/pkg/bin/igneumd-0313 > standing.node; ROLE=live LABEL=p2-3090-3 WALLET={rM.get('wallet')} PROVER=1 MINE=1 setsid nohup bash in/box-standing.sh </dev/null >/dev/null 2>&1 & echo M restored", 60); print(st(), out.strip(), flush=True)
rc,out,err=R.run(f"cd /root/fleet && {KILL}; rm -rf node.1026-replay; mv node node.1026-replay; rm -rf node; mkdir node; cat standing.node; ROLE=live LABEL=pool-1 WALLET={rR.get('wallet')} PROVER=0 MINE=1 P2P_LISTEN=0.0.0.0:4463 setsid nohup bash in/box-standing.sh </dev/null >/dev/null 2>&1 & echo R restored on its pointer with a fresh datadir", 60); print(st(), out.replace("\n"," "), flush=True)
print(st(), "POISON CONFIRM done", flush=True)

View file

@ -0,0 +1,42 @@
"""The 0.3.15 retry in the poisoned-peer form (shipper, 20:4xZ): usage retry-0315-v3.py <igneumd path> <sha256> <igneum-miner path>
New node p2-4090-3 (wiped datadir, mining; peers the hub, the seed box and pool-1), poisoned peer pool-1 as it stands (1026 blocks,
6615571c), restart case p2-3090-2 (wiped datadir, restarted at +5 min, the hub its only re-sync peer). Reads every minute for 12:
p2-4090-3's own log naming version 1026 as rejected and never 'Accepted ... via relay' of a 1026 block, the hands and the hub never
rejecting or disconnecting p2-4090-3, the hub holding a block by p2-4090-3's key mined in the window, p2-3090-2 back at the tip,
the hub's 'wrong block version' count split by peer."""
import sys, os, time, hashlib; sys.path.insert(0, "/Users/joshm/Projects/igneum-wt-gpu-fleet/tools/fleet"); from lib import Box, Registry
def st(): return time.strftime("%H:%M:%SZ", time.gmtime())
NODE, SHA, MINER = sys.argv[1], sys.argv[2], sys.argv[3]
assert hashlib.sha256(open(NODE,"rb").read()).hexdigest()==SHA, "node sha"
WATCH="/opt/igneum/pkg/bin/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -oE 'blocks=[0-9]+|peers=[0-9]+|synced=[a-z]+' | tr '\\n' ' '"
KILL="pkill -9 -f '^bash in/box-standing.sh'; pkill -9 -f '^python3 -u /root/fleet/in/box-prover.py'; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-miner mine'; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-worker-cuda --serve --device 0 --pack packs'; pkill -9 -f '^/(opt/igneum/pkg/bin|root/fleet/in)/igneumd(-0313|-[0-9a-f]{16})? --devnet --appdir='; sleep 3"
hub=Box.from_registry("hub-1"); pool=Box.from_registry("pool-1"); A=Box.from_registry("p2-4090-3"); B=Box.from_registry("p2-3090-2")
# pool-1: the poisoned peer, as it stands (6615571c, its datadir), on its inbound port
rc,out,err=pool.run("cd /root/fleet && pgrep -f '^/root/fleet/in/igneumd-6615571ced35f694 ' >/dev/null || { setsid nohup /root/fleet/in/igneumd-6615571ced35f694 --devnet --appdir=/root/fleet/node --rpclisten=0.0.0.0:26610 --evm-rpclisten=127.0.0.1:26790 --listen=0.0.0.0:4463 --addpeer=213.173.107.74:16516 --addpeer=188.245.5.161:26611 --override-params-file=/root/fleet/override.json --nodnsseed --disable-upnp --nologfiles --yes </dev/null >> node.log 2>&1 & sleep 15; }; echo pool1 $(grep -oE 'igneumd/2\\.1\\.0-[0-9a-f]+' node.log | tail -1) "+WATCH, 60); print(st(), "poisoned peer:", out.replace("\n"," "), flush=True)
def hubsplit(): return hub.run("echo total=$(grep -c 'wrong block version' /root/fleet/node.log); awk '$2 >= \"'$(date -u -d '-12 min' +%H:%M:%S)'\"' /root/fleet/node.log | grep 'wrong block version' | grep -oE 'peer [0-9.]+' | sort | uniq -c | sort -rn | tr '\\n' ';'", 30)[1].replace("\n"," ").strip()
print(st(), "hub before:", hubsplit(), flush=True)
for box, l, peers in ((A, "p2-4090-3", "--addpeer=213.173.107.74:16516 --addpeer=188.245.5.161:26611 --addpeer=213.173.98.36:18451"), (B, "p2-3090-2", "--addpeer=213.173.107.74:16516 --addpeer=188.245.5.161:26611")):
path=box.install_payload(NODE, SHA); box.put([MINER], "/root/fleet/in/igneum-miner-new"); iid,r=Registry.find(l)
already = "ALREADY=yes" in box.run(f"pgrep -f '^{path} --devnet --appdir=' >/dev/null && echo ALREADY=yes || echo ALREADY=no", 20)[1] # a marker: the Vast proxy prints a banner on stdout # the node already runs this binary (a relaunch): keep its datadir and sync
if already: print(st(), l, "already on the binary, keeping its sync", flush=True)
rc,out,err=(0,"kept","") if already else box.run(f"cd /root/fleet && {KILL}; [ -d node ] && mv node node.pre-$(date -u +%H%M); mkdir -p node; mv node.log node.pre-$(date -u +%H%M).log 2>/dev/null; cp in/igneum-miner-new /opt/igneum/pkg/bin/igneum-miner.new && chmod +x /opt/igneum/pkg/bin/igneum-miner.new && mv /opt/igneum/pkg/bin/igneum-miner.new /opt/igneum/pkg/bin/igneum-miner; echo {path} > standing.node; IGNEUM_PROOF_VERIFIER=/opt/igneum-floor/bin/igneum-prove-host setsid nohup {path} --devnet --appdir=/root/fleet/node --rpclisten=0.0.0.0:26610 --evm-rpclisten=127.0.0.1:26790 --listen=0.0.0.0:26611 {peers} --override-params-file=/root/fleet/override.json --nodnsseed --disable-upnp --nologfiles --yes </dev/null >> node.log 2>&1 & echo started T0=$(date -u +%T)", 90)
print(st(), l, out.replace("\n"," "), flush=True)
for i in range(90):
time.sleep(30); rc,out,err=box.run(WATCH, 30)
if "synced=true" in out and "blocks=0 " not in out: break
rc,out,err=box.run(f"cd /root/fleet && echo synced_after={(i+1)*30}s {WATCH} vline=$(grep -oE 'igneumd/2\\.1\\.0-[0-9a-f]+' node.log | tail -1) digest=$(grep -o 'digest: [0-9a-f]*' node.log | tail -1 | awk '{{print substr($2,1,8)}}'); cd mine 2>/dev/null || mkdir -p /root/fleet/mine/packs && cd /root/fleet/mine; rm -rf packs/devnet; /opt/igneum/pkg/bin/igneum-miner export-pack grpc://127.0.0.1:26610 packs/devnet >/dev/null 2>&1; setsid nohup /opt/igneum/pkg/bin/igneum-miner mine grpc://127.0.0.1:26610 1 100000000 {l} --worker /opt/igneum/pkg/bin/igneum-worker-cuda --worker-args '--device 0 --pack packs/devnet' --prepare-packs packs/prepare --exit-on-seed-change --evm-address {r.get('wallet')} --payout-label {l} --status-secs 60 </dev/null >> /root/fleet/out/miner-0.log 2>&1 & echo miner started; /opt/igneum/pkg/bin/igneum-miner key-hash {l} 2>/dev/null | grep -oE '[0-9a-f]{{64}}' | head -1", 120)
print(st(), l, "synced + mining:", out.replace("\n"," "), flush=True)
if l=="p2-4090-3": keyA=out.strip().split()[-1]
T0=time.time(); restarted=False
while time.time()-T0 < 720:
time.sleep(60)
rc,a,err=A.run("echo rej1026=$(grep -c 'wrong block version' /root/fleet/node.log) got_reject=$(grep -c 'got reject message' /root/fleet/node.log) relayed1026=$(grep -ciE 'version 1026.*(accept|relay)|accepted.*1026' /root/fleet/node.log) peers_now=$(grep -oE 'Connected to (outgoing|incoming) peer [0-9.:]+' /root/fleet/node.log | awk '{print $NF}' | cut -d: -f1 | sort -u | tr '\\n' ','); grep -E 'wrong block version|1026' /root/fleet/node.log | tail -1 | cut -c12-200; "+WATCH, 40)
rc,b,err=B.run(WATCH+"; echo got_reject=$(grep -c 'got reject message' /root/fleet/node.log)", 30)
print(st(), f"+{int(time.time()-T0)}s A:", a.replace("\n"," | ").strip(), "| B:", b.replace("\n"," ").strip(), "| hub:", hubsplit(), flush=True)
if not restarted and time.time()-T0 >= 300:
rc,out,err=B.run("cd /root/fleet && pkill -9 -f '^/root/fleet/in/igneumd-[0-9a-f]{16} --devnet --appdir='; sleep 20; BIN=$(cat standing.node); IGNEUM_PROOF_VERIFIER=/opt/igneum-floor/bin/igneum-prove-host setsid nohup $BIN --devnet --appdir=/root/fleet/node --rpclisten=0.0.0.0:26610 --evm-rpclisten=127.0.0.1:26790 --listen=0.0.0.0:26611 --addpeer=213.173.107.74:16516 --override-params-file=/root/fleet/override.json --nodnsseed --disable-upnp --nologfiles --yes </dev/null >> node.log 2>&1 & for i in $(seq 1 24); do sleep 5; w=$("+WATCH+"); case \"$w\" in *synced=true*) break;; esac; done; echo resync_after=$((i*5+20))s $w; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-miner mine'; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-worker-cuda --serve'", 200)
print(st(), "B restart (hub-only re-sync):", out.replace("\n"," "), flush=True); restarted=True
hub_pan=hub.run("echo hub_panics=$(grep -ciE 'panicked at consensus' /root/fleet/node.log)", 20)[1].strip(); print(st(), hub_pan, flush=True)
rc,out,err=hub.run(f"timeout 150 /opt/igneum/pkg/bin/igneum-miner inspect 900 grpc://127.0.0.1:26610 2>&1 | grep -c 'voteKeyHash={keyA}'; awk '$2 >= \"'$(date -u -d '-14 min' +%H:%M:%S)'\"' /root/fleet/node.log | grep -E 'wrong block version|disconnecting' | grep -c '216.232.211.246'", 170)
print(st(), f"hub holds {out.split()[0] if out.split() else '?'} blocks by p2-4090-3's key {keyA[:16]} in its last 900; hub rejects/disconnects naming p2-4090-3 in the window: {out.split()[1] if len(out.split())>1 else '?'}", flush=True)
print(st(), "RETRY V3 done", flush=True)

68
tools/fleet/canary.sh Executable file
View file

@ -0,0 +1,68 @@
#!/bin/bash
# bash 3.2 (the Mac's): plain arrays, no mapfile, no associative arrays (lookup files under $WORK)
# The 0.3.14 canary gate on the LIVE devnet (6 October 2026, 17:00Z, the coordinator's form for a publish that moves no
# digest and no height): the release's Linux igneumd goes onto the six restart-path provers and two miners within two
# minutes of the URL (the override file kept), runs TEN minutes, then one line: PASS only with zero rejected blocks on
# every box, exec state roots equal on every box and on the hub at a common height, at least one segment record paid on
# the new binary (a prover's paidSegments up), every box's node on the new version; FAIL names the box and the line.
# canary.sh --binary <url|file> --sha256 <hex> [--minutes 10]
set -uo pipefail
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"; ROOT="$HOME/Desktop/fleet"
BIN=""; SHA=""; MINUTES=10
while [[ $# -gt 0 ]]; do case "$1" in --binary) BIN="$2"; shift 2 ;; --sha256) SHA="$2"; shift 2 ;; --minutes) MINUTES="$2"; shift 2 ;; *) echo "unknown $1"; exit 2 ;; esac; done
[[ -n "$BIN" && -n "$SHA" ]] || { echo "usage: $0 --binary <url|file> --sha256 <hex> [--minutes 10]"; exit 2; }
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
LOG="$ROOT/canary-$(date -u +%Y%m%dT%H%M%SZ).log"; exec > >(tee -a "$LOG") 2>&1
PROVERS="p1-3090 p1-5090 p2-3090-1 p2-3090-2 p2-3090-3 p2-4090-1b"; MINERS="p1-3080 p1-a5000"
fail() { echo "FAIL $(stamp) $*"; exit 1; }
LOCAL="$ROOT/canary-igneumd"
if [[ "$BIN" == http* ]]; then curl -fsSL -o "$LOCAL" "$BIN" || fail "download $BIN"; else cp "$BIN" "$LOCAL"; fi
[[ "$(shasum -a 256 "$LOCAL" | cut -c1-64)" == "$SHA" ]] || fail "sha256 of the binary is $(shasum -a 256 "$LOCAL" | cut -c1-16), not ${SHA:0:16}"
WANT="${SHA:0:16}_$(file "$LOCAL" >/dev/null; echo igneumd_2.1.0)" # the running binary's sha256 (first 16) and its --version word, per box
echo "CANARY start $(stamp) binary=${SHA:0:16} version=$WANT boxes: $PROVERS $MINERS"
ROWS=(); while IFS= read -r line; do ROWS+=("$line"); done < <(python3 - "$PROVERS $MINERS" <<'PY'
import json, os, sys; reg=json.load(open(os.path.expanduser("~/Desktop/fleet/boxes.json"))); want=sys.argv[1].split()
for lab in want:
b=next((v for v in reg.values() if v["label"]==lab and v.get("state")!="destroyed"), None)
if b: print(lab, b["ssh_host"], b["ssh_port"], b["wallet"])
hub=next(v for v in reg.values() if v.get("hub") and v.get("state")!="destroyed" and v.get("hub_peer")); print("hub-1", hub["ssh_host"], hub["ssh_port"], hub["wallet"], hub["hub_peer"])
PY
)
HUB="$(printf '%s\n' "${ROWS[@]}" | awk '$1=="hub-1"')"; HUB_HOST="$(awk '{print $2}' <<< "$HUB")"; HUB_PORT="$(awk '{print $3}' <<< "$HUB")"; HUB_PEER="$(awk '{print $5}' <<< "$HUB")"
SSH() { ssh -i ~/.ssh/igneum-fleet -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=20 -o ServerAliveInterval=10 -o ServerAliveCountMax=3 -o BatchMode=yes -p "$2" "root@$1" "${@:3}"; }
read_box() { SSH "$1" "$2" 'B=/opt/igneum/pkg/bin; w=$($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o "blocks=[0-9]*.*synced=[a-z]*" | tail -1); d=$(grep -o "daa=[0-9]*" <<< "$w" | cut -d= -f2); e=$(curl -s -m 6 -X POST -H "Content-Type: application/json" --data "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"igneum_getProvingStatus\",\"params\":[]}" http://127.0.0.1:26790/ | python3 -c "import sys,json; r=sys.stdin.read(); x=json.loads(r).get(\"result\",{}) if r.strip() else {}; print(int(x.get(\"tipDaa\",\"0x0\"),16), x.get(\"v1\",{}).get(\"paidSegments\",0))" 2>/dev/null); ex=$(curl -s -m 6 -X POST -H "Content-Type: application/json" --data "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"igneum_getExecStatus\",\"params\":[]}" http://127.0.0.1:26790/ | python3 -c "import sys,json; r=sys.stdin.read(); x=json.loads(r).get(\"result\",{}) if r.strip() else {}; print(int(x.get(\"executedTip\",\"0x0\"),16))" 2>/dev/null); v=$(pgrep -fa "^/opt/igneum/pkg/bin/igneumd" | head -1 | awk "{print \$2}"); ver=$(sha256sum "$v" 2>/dev/null | cut -c1-16)_$($v --version 2>&1 | head -1 | tr " " "_"); rej=$(grep -cE "got reject message|PoW rejected|block rejected|invalid block" /root/fleet/node.log 2>/dev/null); echo "${d:-0} ${e:-0 0} ${ex:-0} ${ver:-none} ${rej:-0}"' 2>/dev/null; }
root_at() { SSH "$1" "$2" "curl -s -m 6 -X POST -H 'Content-Type: application/json' --data '{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"eth_getBlockByNumber\",\"params\":[\"$3\",false]}' http://127.0.0.1:26790/ | python3 -c 'import sys,json; b=json.load(sys.stdin).get(\"result\") or {}; print(b.get(\"stateRoot\",\"none\"))'" 2>/dev/null; }
WORK="$(mktemp -d)"; paid0() { cat "$WORK/paid0.$1" 2>/dev/null || echo 0; }; rej0() { cat "$WORK/rej0.$1" 2>/dev/null || echo 0; }
echo "CANARY baseline $(stamp)"
for r in "${ROWS[@]}"; do set -- $r; x="$(read_box "$2" "$3")"; echo " $1: daa=$(awk '{print $1}' <<< "$x") tip=$(awk '{print $2}' <<< "$x") paidSeg=$(awk '{print $3}' <<< "$x") exec=$(awk '{print $4}' <<< "$x") ver=$(awk '{print $5}' <<< "$x") rej=$(awk '{print $6}' <<< "$x")"; awk '{print $3}' <<< "$x" > "$WORK/paid0.$1"; awk '{print $6}' <<< "$x" > "$WORK/rej0.$1"; done
# install: the binary swap with the override file kept (box-node-swap.sh STEP=binary, which refuses on a digest change)
T0=$(date +%s)
for r in "${ROWS[@]}"; do set -- $r; [[ "$1" == hub-1 ]] && continue
okc=0; for try in 1 2 3; do scp -i ~/.ssh/igneum-fleet -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=20 -o ServerAliveInterval=10 -o ServerAliveCountMax=3 -P "$3" "$LOCAL" "root@$2:/root/fleet/in/igneumd-0313" >/dev/null 2>&1 && { okc=1; break; }; sleep 5; done # Vast's ssh proxy drops a connection now and then (17:14Z); three tries before a FAIL
[[ $okc == 1 ]] || fail "$1: scp failed three times"
oks=0; for try in 1 2 3; do SSH "$2" "$3" "cd /root/fleet && chmod +x in/igneumd-0313 && [ \"\$(sha256sum in/igneumd-0313 | cut -c1-64)\" = $SHA ] && mv out/node-swap.log out/node-swap-canary-prev.log 2>/dev/null; STEP=binary NODE_SHA256=$SHA EXPECT_DIGEST=b18ed271f75dd46406d230f4156c37472127415a4c32c558bac662f6f840e61c HUB_PEER=$HUB_PEER HUB_SSH=$HUB_HOST HUB_PORT=$HUB_PORT setsid nohup in/box-node-swap.sh </dev/null >/dev/null 2>&1 & echo swapped" >/dev/null 2>&1 && { oks=1; break; }; sleep 5; done
[[ $oks == 1 ]] || fail "$1: install failed three times"
done
echo "CANARY installed $(stamp) on $(( ${#ROWS[@]} - 1 )) boxes in $(( $(date +%s) - T0 )) s"
sleep 75
for r in "${ROWS[@]}"; do set -- $r; [[ "$1" == hub-1 ]] && continue
l="$(SSH "$2" "$3" "grep -E '^RESULT (node_started|swap_failed)' /root/fleet/out/node-swap.log | tail -1 | cut -c1-140")"; echo " $1: $l"; [[ "$l" == *digest_ok* ]] || fail "$1: the swap did not reach digest_ok ($l)"
done
# the provers back on (box-prover.sh keeps a running igneumd-0313 with the verifier)
for p in $PROVERS; do r="$(printf '%s\n' "${ROWS[@]}" | awk -v l="$p" '$1==l')"; set -- $r; SSH "$2" "$3" "cd /root/fleet && pkill -f '^python3 -u /root/fleet/in/box-prover.py'; pkill -x igneum-miner; sleep 2; mv out/prover.log out/prover-canary-prev.log 2>/dev/null; LABEL=$1 WALLET=$4 HUB_PEER=$HUB_PEER setsid nohup in/box-prover.sh </dev/null >/dev/null 2>&1 & echo prover" >/dev/null || fail "$1: prover relaunch"; done
for p in $MINERS; do r="$(printf '%s\n' "${ROWS[@]}" | awk -v l="$p" '$1==l')"; set -- $r; SSH "$2" "$3" "cd /root/fleet/mine && pkill -x igneum-miner; sleep 2; rm -rf packs/devnet; /opt/igneum/pkg/bin/igneum-miner export-pack grpc://127.0.0.1:26610 packs/devnet >/dev/null 2>&1; setsid nohup /opt/igneum/pkg/bin/igneum-miner mine grpc://127.0.0.1:26610 1 100000000 $1 --worker /opt/igneum/pkg/bin/igneum-worker-cuda --worker-args '--device 0 --pack packs/devnet' --prepare-packs packs/prepare --exit-on-seed-change --evm-address $4 --payout-label $1 --status-secs 30 </dev/null >> /root/fleet/out/mine-only-0.log 2>&1 & echo miner" >/dev/null || fail "$1: miner relaunch"; done
echo "CANARY running $(stamp) $MINUTES min"; sleep $((MINUTES * 60))
echo "CANARY checks $(stamp)"; ok=1; paid_any=0; H=""
for r in "${ROWS[@]}"; do set -- $r; x="$(read_box "$2" "$3")"; paid="$(awk '{print $3}' <<< "$x")"; ex="$(awk '{print $4}' <<< "$x")"; ver="$(awk '{print $5}' <<< "$x")"; rej="$(awk '{print $6}' <<< "$x")"
echo " $1: tip=$(awk '{print $2}' <<< "$x") paidSeg=$paid exec=$ex ver=$ver rej=$rej (was $(rej0 $1))"
[[ "$1" == hub-1 ]] || { [[ "$ver" == "$WANT" ]] || { echo "FAIL $1: version $ver, want $WANT"; ok=0; }; }
[[ "${rej:-0}" -le "$(rej0 $1)" ]] || { echo "FAIL $1: $(( rej - $(rej0 $1) )) new rejected blocks (grep 'reject' /root/fleet/node.log)"; ok=0; }
[[ "$1" == hub-1 ]] || { for p in $PROVERS; do [[ "$1" == "$p" && "${paid:-0}" -gt "$(paid0 $1)" ]] && paid_any=1; done; }
[[ -z "$H" || "${ex:-0}" -lt "$H" ]] && H="${ex:-0}"
done
H=$((H > 20 ? H - 20 : 1)); HX="$(printf '0x%x' "$H")"; first=""
for r in "${ROWS[@]}"; do set -- $r; rt="$(root_at "$2" "$3" "$HX")"; echo " $1 root@$H ${rt:0:18}"; [[ -z "$first" ]] && first="$rt"; [[ "$rt" == "$first" ]] || { echo "FAIL $1: exec root at $H ${rt:0:18} differs from ${first:0:18}"; ok=0; }; done
[[ $paid_any == 1 ]] || { echo "FAIL no segment record paid on any prover during the $MINUTES min (paidSegments did not rise)"; ok=0; }
res=FAIL; [[ $ok == 1 ]] && res=PASS
echo "$res $(stamp) canary 0.3.14 ${SHA:0:16} version $WANT: $(( ${#ROWS[@]} - 1 )) boxes, $MINUTES min, rejects 0, roots equal at $H incl. the hub, segment paid $paid_any; log $LOG"
[[ $ok == 1 ]]

68
tools/fleet/collect.py Executable file
View file

@ -0,0 +1,68 @@
#!/usr/bin/env python3
"""Phase 1 collector: reads ~/Desktop/fleet/<instance>/matrix.json (and ember.json when present) for every phase-1 box,
writes ~/Desktop/fleet/results.json (the fleet page's rows, keyed by card) and prints the markdown table for
docs/analysis/prover-tiers-real-cards.md. The verdict rule, per card (the 9.0 GB mine-and-prove line and the tier
gates of docs/analysis/prover-floor.md), on the card's OWN numbers:
proves alone = a compressed proof of the v1 shard at 2^26 or 2^27 verified with the server's own working set
plus the idle under the card's memory, but mine-and-prove over the line
mines and proves = the compressed v1 shard beside the miner verified (peak under the card's memory)
mines and proves core-only = only the core-only point beside the miner fits
mining only = no proof point verified on the card (every point refused or out of memory)
"""
import json, os, sys, glob
ROOT = os.path.expanduser("~/Desktop/fleet")
reg = json.load(open(f"{ROOT}/boxes.json"))
def gb(mib): return round(mib / 1024, 1)
rows = {}; table = []
for iid, b in reg.items():
mp = f"{ROOT}/{iid}/matrix.json" # every box with a matrix (a box moves to phase 2 when its Ember ladder ends)
rj = f"{ROOT}/{iid}/rows.jsonl"
if os.path.exists(mp): m = json.load(open(mp))
elif os.path.exists(rj): # the box's final step failed (the 4060 reported power_w=[N/A]): rebuild from the raw rows and the start line
ml = f"{ROOT}/{iid}/matrix.log"; start = [l for l in open(ml, errors="replace") if l.startswith("RESULT start")][-1] if os.path.exists(ml) else ""
g = lambda k, d="": (start.split(k + "=")[1].split()[0] if k + "=" in start else d)
idl = [l for l in open(ml, errors="replace") if l.startswith("RESULT idle")][-1] if os.path.exists(ml) else ""
m = {"label": b["label"], "card": g("card", b["card"]), "total_mib": int(g("total_mib", b.get("vram_mb") or 0)), "driver": g("driver"), "idle_mib": int(idl.split("mem_mib=")[1].split()[0]) if "mem_mib=" in idl else 0, "idle_w": 0, "rows": [json.loads(l) for l in open(rj) if l.strip()]}
else: continue
pts = {r["name"]: r for r in m["rows"] if r.get("row") == "point"}; miner = next((r for r in m["rows"] if r.get("row") == "miner"), {})
total = m["total_mib"]; idle = m["idle_mib"]
# the miner's rate from the raw STATUS lines (the first matrices parsed a second now= field): the mean of the last
# 12 STATUS lines before the miner was stopped for the stock point, i.e. the sampled window
ml = f"{ROOT}/{iid}/miner.log"
if os.path.exists(ml):
vals = [float(l.split(" now=")[1].split()[0]) for l in open(ml, errors="replace") if "STATUS" in l and " now=" in l]
# only when the script's own row is broken (the first matrices parsed a second now= field): the miner-alone
# window is the first 21 STATUS lines (60 s warm + 150 s sample at 10 s), minus the warm-up's first 6
if vals and (not miner.get("mhs") or miner.get("mhs", 0) > 10000): w = vals[6:21]; miner = dict(miner, mhs=round(sum(w) / len(w), 2), mhs_recomputed=True)
def ok(n): p = pts.get(n); return p and p.get("verified") == "yes"
def own(n): p = pts.get(n); return p.get("own_mib") if p else None
def secs(n): p = pts.get(n); return p.get("prove_s") if p else None
alone = next((n for n in ("alone-comp-26-v1", "alone-comp-27-v1", "alone-comp-25-v1") if ok(n)), None)
beside = next((n for n in ("miner-comp-26-v1", "miner-comp-27-v1") if ok(n)), None)
core_b = next((n for n in ("miner-core-25-v1", "miner-core-26-v1", "miner-core-24-v1") if ok(n)), None)
stock = pts.get("stock-comp-v1", {})
if beside: verdict = "mines and proves"
elif alone and core_b: verdict = "mines and proves core-only"
elif alone: verdict = "proves alone"
else: verdict = "mining only"
card = b.get("card") or m["card"]
card = card if "4060 Ti" not in card else f"RTX 4060 Ti {round(total / 1024)} GB"
r = {"card": card, "vram_gb": round(total / 1024), "mhs": miner.get("mhs"), "watts": miner.get("watts"), "miner_gb": gb(miner.get("own_mib", 0)),
"stock": ("not measured: the SDK's server download stalled (killed at 553 s)" if stock.get("rc") == 137 else "refused: " + (stock.get("err") or "")[:60]) if stock.get("verified") != "yes" else f"proved {stock.get('prove_s')} s at {gb(stock.get('peak_mib', 0))} GB",
"prove_alone_gb": gb(own(alone)) if alone else None, "shard_s": secs(alone), "prove_alone_point": alone,
"mine_prove_gb": gb(pts[beside]["peak_mib"]) if beside else None, "mine_prove_s": secs(beside), "core_only_gb": gb(own(core_b)) if core_b else None, "core_only_s": secs(core_b),
"verdict": verdict, "points": {n: {"v": p.get("verified"), "own": p.get("own_mib"), "peak": p.get("peak_mib"), "s": p.get("prove_s"), "err": (p.get("err") or "")[:80]} for n, p in pts.items()}}
r["tier_line"] = {"mines and proves": f"{card} ({r['vram_gb']} GB): mines and proves; the v1 shard beside the miner {r['mine_prove_s']} s at {r['mine_prove_gb']} GB peak",
"mines and proves core-only": f"{card} ({r['vram_gb']} GB): proves alone ({r['shard_s']} s) and mines beside a core-only prover ({r['core_only_gb']} GB own)",
"proves alone": f"{card} ({r['vram_gb']} GB): proves alone, {r['shard_s']} s a v1 shard at {r['prove_alone_gb']} GB own; not beside the miner",
"mining only": f"{card} ({r['vram_gb']} GB): mines only ({r['mhs']} MH/s); no proof point fits"}[verdict]
ep = f"{ROOT}/{iid}/ember.json"
if os.path.exists(ep):
e = json.load(open(ep)); c = e["chosen"]
r.update({"tune_w": c["watts"], "tune_clock_mhz": c["clock_mhz"], "tune_mhs": c["mhs"], "tune_mhw": c["eff"], "ladder": [{"w": s["watts"], "mhs": s["mhs"], "mhw": s["eff"], "step": s["label"]} for s in e["steps"]], "tune_plan": e["plan"]})
rows[f"{card} {r['vram_gb']} GB"] = r
table.append(f"| {card} | {r['vram_gb']} | {idle} | {r['mhs']} MH/s at {r['watts']} W, {r['miner_gb']} GB | {r['stock']} | " +
f"{r['prove_alone_gb'] or 'no'} GB, {r['shard_s'] or ''} s ({alone or 'none'}) | {r['mine_prove_gb'] or 'no'} GB peak, {r['mine_prove_s'] or ''} s | {r['core_only_gb'] or 'no'} GB, {r['core_only_s'] or ''} s | {verdict} |")
json.dump(rows, open(f"{ROOT}/results.json", "w"), indent=1)
print("| Card | VRAM GB | Idle MiB | Miner | Stock SP1 6.8.1 | Patched, proves alone (own) | Beside the miner (peak) | Core-only beside the miner (own) | Verdict |\n|---|---|---|---|---|---|---|---|---|")
print("\n".join(sorted(table)))

16
tools/fleet/deploy-gate.sh Executable file
View file

@ -0,0 +1,16 @@
#!/usr/bin/env bash
# The fleet's deploy gate (coordinator, 6 October 2026 23:0xZ, after three comment-on-a-line faults): every script pushed to a box
# passes master's defaults-line check and a syntax check first; lib/box.py's put() calls this for any .sh/.py source under
# tools/fleet and refuses the push on red. Usage: deploy-gate.sh <file...> (exit 0 = clean)
set -u; rc=0; HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"; MASTER_CI="${IGNEUM_MASTER_CI:-/Users/joshm/Projects/igneum/tools/ci}"
for f in "$@"; do
case "$f" in
*.sh) bash -n "$f" 2>&1 | head -2 | sed "s|^|$f: |" ; [ "${PIPESTATUS[0]}" = 0 ] || rc=1 ;;
*.py) python3 -m py_compile "$f" 2>&1 | head -2 | sed "s|^|$f: |"; [ "${PIPESTATUS[0]}" = 0 ] || rc=1 ;;
esac
# a comment inserted mid-line swallows the rest of the line: no code may follow a comment that follows code
if [ -f "$MASTER_CI/defaults-line-check.sh" ]; then bash "$MASTER_CI/defaults-line-check.sh" "$f" 2>&1 | grep -v "no assignment hides" | sed "s|^|$f: |" | head -3; [ "${PIPESTATUS[0]}" = 0 ] || rc=1; fi
if grep -nE '^[^#"'"'"']*[^ #]\s+#\s[^"]*\S+.*(\$\{|>>|&$|;\s*[a-zA-Z_]+=)' "$f" 2>/dev/null | grep -vE '^\s*[0-9]+:\s*#' | grep -qE '#.*(>> |&$|; [A-Z_]+=)'; then echo "$f: a comment with code after it (the 17:18Z/22:10Z/22:52Z class)"; rc=1; fi
done
[ $rc = 0 ] && echo "deploy-gate: clean ($# files)" || echo "deploy-gate: RED, the push is refused"
exit $rc

107
tools/fleet/devnet2-gate.sh Normal file
View file

@ -0,0 +1,107 @@
#!/bin/bash
# The Devnet 2 release gate (6 October 2026). Given a release's Linux igneumd (a download URL or a local file) and its
# manifest (for the activation field and value, or --activation name=value), it: (1) installs the binary on every
# Devnet 2 box and restarts each node on the CURRENT override, checking the version; (2) sets the activation at tip + MARGIN
# in a new override, restarts every box on it (the seed first) and waits for the crossing; (3) runs N minutes past it;
# (4) prints PASS only when: zero rejected blocks on every box, no selected-chain reorg over depth 3, exec state roots
# equal on every box at a common height, at least one segment record paid since the start, every node's version equal
# to the release's. FAIL names the box and the line. The one-line status goes to ~/Desktop/fleet/devnet2-status.json.
#
# devnet2-gate.sh --binary <url|file> --sha256 <hex> [--activation proving_v1_fresh_rule_daa=NNN | --activation-field F]
# [--margin 600] [--minutes 20] [--dry-run] (dry-run: steps 1 and 4 against the current object, no activation)
set -uo pipefail
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"; ROOT="$HOME/Desktop/fleet"
BIN=""; SHA=""; ACT=""; ACT_FIELD=""; MARGIN=600; MINUTES=20; DRY=0
while [[ $# -gt 0 ]]; do case "$1" in
--binary) BIN="$2"; shift 2 ;; --sha256) SHA="$2"; shift 2 ;; --activation) ACT="$2"; shift 2 ;; --activation-field) ACT_FIELD="$2"; shift 2 ;;
--margin) MARGIN="$2"; shift 2 ;; --minutes) MINUTES="$2"; shift 2 ;; --dry-run) DRY=1; shift ;; *) echo "unknown $1"; exit 2 ;; esac; done
[[ -n "$BIN" && -n "$SHA" ]] || { echo "usage: $0 --binary <url|file> --sha256 <hex> [--activation name=value] [--minutes N] [--dry-run]"; exit 2; }
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
LOG="$ROOT/devnet2-gate-$(date -u +%Y%m%dT%H%M%SZ).log"; exec > >(tee -a "$LOG") 2>&1
echo "GATE start $(stamp) binary=$BIN sha256=${SHA:0:16} activation=${ACT:-none} minutes=$MINUTES dry_run=$DRY"
fail() { echo "FAIL $(stamp) $*"; python3 - "$ROOT" "FAIL: $*" <<'PY'
import json, sys, os, datetime; p=os.path.join(sys.argv[1], "devnet2-status.json"); d=json.load(open(p)) if os.path.exists(p) else {}
d.update({"last_gate": sys.argv[2][:200], "last_gate_at": datetime.datetime.utcnow().strftime("%Y-%m-%dT%H:%M:%SZ")}); json.dump(d, open(p, "w"))
PY
exit 1; }
# the binary, local
LOCAL="$ROOT/dn2-gate-igneumd"
if [[ "$BIN" == http* ]]; then curl -fsSL -o "$LOCAL" "$BIN" || fail "download $BIN"; else cp "$BIN" "$LOCAL"; fi
[[ "$(shasum -a 256 "$LOCAL" | cut -c1-64)" == "$SHA" ]] || fail "sha256 of the binary is $(shasum -a 256 "$LOCAL" | cut -c1-16), not ${SHA:0:16}"
VERSION_WANT="${SHA:0:16}_igneumd_2.1.0" # the running binary's sha256 (first 16) and its --version word
# the boxes
BOXES=(); while IFS= read -r line; do BOXES+=("$line"); done < <(python3 - <<'PY'
import json, os; reg=json.load(open(os.path.expanduser("~/Desktop/fleet/boxes.json")))
for iid,b in reg.items():
if b.get("devnet2") and b.get("state")!="destroyed" and b.get("ssh_host"): print(f"{b['label']} {b['ssh_host']} {b['ssh_port']} {int(bool(b.get('dn2_seed')))} {b.get('wallet')} {int(bool(b.get('dn2_prover')))}")
PY
)
[[ ${#BOXES[@]} -gt 0 ]] || fail "no Devnet 2 boxes in the registry"
SSH() { ssh -i ~/.ssh/igneum-fleet -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=20 -o ServerAliveInterval=10 -o ServerAliveCountMax=3 -o BatchMode=yes -p "$2" "root@$1" "${@:3}"; }
SCP() { scp -i ~/.ssh/igneum-fleet -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=20 -o ServerAliveInterval=10 -o ServerAliveCountMax=3 -P "$2" "${@:3}" "root@$1:/root/fleet/in/" >/dev/null; }
seed_line="$(printf '%s\n' "${BOXES[@]}" | awk '$4==1 {print; exit}')"; SEED_HOST="$(awk '{print $2}' <<< "$seed_line")"; SEED_PORT="$(awk '{print $3}' <<< "$seed_line")"
SEED_PEER="$(python3 - <<'PY'
import json, os; reg=json.load(open(os.path.expanduser("~/Desktop/fleet/boxes.json")))
print(next((b.get("dn2_peer","") for b in reg.values() if b.get("dn2_seed") and b.get("state")!="destroyed"), ""))
PY
)"
echo "GATE boxes=${#BOXES[@]} seed=$SEED_HOST:$SEED_PORT peer=$SEED_PEER"
# a reading of every box: height, exec tip and root at a height, rejects, max reorg, version, paid segments
SINCE="${SINCE:-1970-01-01 00:00:00}" # log lines at or after this UTC stamp count (set at step 1)
read_box() { # host port -> "height daa exec_tip paid_seg version rejects max_reorg", rejects and reorgs since $SINCE
SSH "$1" "$2" 'SINCE="'"$SINCE"'"; B=/opt/igneum/pkg/bin; w=$($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o "blocks=[0-9]*.*synced=[a-z]*" | tail -1); h=$(grep -o "blocks=[0-9]*" <<< "$w" | cut -d= -f2); d=$(grep -o "daa=[0-9]*" <<< "$w" | cut -d= -f2); e=$(curl -s -m 6 -X POST -H "Content-Type: application/json" --data "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"igneum_getProvingStatus\",\"params\":[]}" http://127.0.0.1:26790/ | python3 -c "import sys,json; r=sys.stdin.read(); x=json.loads(r).get(\"result\",{}) if r.strip() else {}; print(int(x.get(\"tipDaa\",\"0x0\"),16), x.get(\"v1\",{}).get(\"paidSegments\",0))" 2>/dev/null); v=$(pgrep -fa "^/root/fleet/in/igneumd" | head -1 | awk "{print \$2}"); ver=$(sha256sum "$v" 2>/dev/null | cut -c1-16)_$($v --version 2>&1 | head -1 | tr " " "_"); rej=$(awk -v s="$SINCE" "substr(\$0,1,19) >= s" /root/fleet/dn2-node.log 2>/dev/null | grep -cE "reject message|PoW rejected|block rejected|invalid block"); mr=$(awk -v s="$SINCE" "substr(\$0,1,19) >= s" /root/fleet/dn2-node.log 2>/dev/null | grep -oE "selected-chain reorg: [0-9]+ chain blocks" | grep -oE "[0-9]+ chain" | awk "{if (\$1>m) m=\$1} END {print m+0}"); echo "${h:-0} ${d:-0} ${e:-0 0} ${ver:-none} ${rej:-0} ${mr:-0}"' 2>/dev/null
}
root_at() { SSH "$1" "$2" "curl -s -m 6 -X POST -H 'Content-Type: application/json' --data '{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"eth_getBlockByNumber\",\"params\":[\"$3\",false]}' http://127.0.0.1:26790/ | python3 -c 'import sys,json; b=json.load(sys.stdin).get(\"result\") or {}; print(b.get(\"stateRoot\",\"none\"))'" 2>/dev/null; }
WORK="$(mktemp -d)"; paid0() { cat "$WORK/paid0.$1" 2>/dev/null || echo 0; }
echo "GATE baseline $(stamp)"
for l in "${BOXES[@]}"; do set -- $l; r="$(read_box "$2" "$3")"; echo " $1: $r"; awk '{print $4}' <<< "$r" > "$WORK/paid0.$1"; done
# 1. install the binary on every box, restart on the current override
for l in "${BOXES[@]}"; do set -- $l
SCP "$2" "$3" "$LOCAL" || fail "$1: scp of the binary"
SSH "$2" "$3" "mv /root/fleet/in/dn2-gate-igneumd /root/fleet/in/igneumd-gate && chmod +x /root/fleet/in/igneumd-gate && [ \"\$(sha256sum /root/fleet/in/igneumd-gate | cut -c1-64)\" = $SHA ]" || fail "$1: the binary's sha256 on the box"
done
restart_all() { # with the override file already on each box as /root/fleet/in/dn2-override.json
for l in "${BOXES[@]}"; do set -- $l; [[ "$4" == 1 ]] || continue
SSH "$2" "$3" "cd /root/fleet && pkill -f '^bash in/box-dn2.sh'; pkill -f '^/root/fleet/in/igneumd'; pkill -x igneum-miner; pkill -f '^python3 -u /root/fleet/in/box-prover.py'; sleep 3; LABEL=$1 WALLET=$5 NODE_BIN=/root/fleet/in/igneumd-gate PROVER=$6 setsid nohup in/box-dn2.sh </dev/null >/dev/null 2>&1 & echo restarted" || fail "$1: restart"
sleep 15; done
for l in "${BOXES[@]}"; do set -- $l; [[ "$4" == 1 ]] && continue
SSH "$2" "$3" "cd /root/fleet && pkill -f '^bash in/box-dn2.sh'; pkill -f '^/root/fleet/in/igneumd'; pkill -x igneum-miner; pkill -f '^python3 -u /root/fleet/in/box-prover.py'; sleep 3; LABEL=$1 WALLET=$5 SEED=$SEED_PEER NODE_BIN=/root/fleet/in/igneumd-gate PROVER=$6 setsid nohup in/box-dn2.sh </dev/null >/dev/null 2>&1 & echo restarted" || fail "$1: restart"
done
}
SINCE="$(date -u +'%Y-%m-%d %H:%M:%S')"; echo "GATE step1 $(stamp) restart on the release binary, the current override (rejects and reorgs counted from $SINCE)"; restart_all; sleep 60
for l in "${BOXES[@]}"; do set -- $l; r="$(read_box "$2" "$3")"; echo " $1: $r"; v="$(awk '{print $5}' <<< "$r")"; [[ -z "$VERSION_WANT" || "$v" == "$VERSION_WANT" ]] || fail "$1: version $v, want $VERSION_WANT"; done
# 2. the activation
if [[ $DRY == 0 && -n "$ACT" ]]; then
tip="$(read_box "$SEED_HOST" "$SEED_PORT" | awk '{print $2}')"; name="${ACT%%=*}"; val="${ACT#*=}"; [[ "$val" == auto ]] && val=$((tip + MARGIN))
python3 - "$HERE/devnet2-override.json" "$name" "$val" "$ROOT/dn2-override-next.json" <<'PY'
import json, sys; d=json.load(open(sys.argv[1])); d[sys.argv[2]]=int(sys.argv[3]); json.dump(d, open(sys.argv[4], "w"))
PY
for l in "${BOXES[@]}"; do set -- $l; SCP "$2" "$3" "$ROOT/dn2-override-next.json" && SSH "$2" "$3" "mv /root/fleet/in/dn2-override-next.json /root/fleet/in/dn2-override.json" || fail "$1: override push"; done
echo "GATE step2 $(stamp) activation $name=$val at tip $tip (margin $MARGIN); restart on the new object"; restart_all
for i in $(seq 1 120); do
sleep 15; d="$(read_box "$SEED_HOST" "$SEED_PORT" | awk '{print $2}')"
if [[ "${d:-0}" -ge "$val" ]]; then echo "GATE crossed $name=$val at $(stamp), daa $d"; break; fi
done
fi
# 3. run
echo "GATE step3 $(stamp) running $MINUTES min"; sleep $((MINUTES * 60))
# 4. the checks
echo "GATE step4 $(stamp) checks"; H=""; ok=1; paid_any=0
for l in "${BOXES[@]}"; do set -- $l; r="$(read_box "$2" "$3")"; echo " $1: $r"
rej="$(awk '{print $6}' <<< "$r")"; mr="$(awk '{print $7}' <<< "$r")"; v="$(awk '{print $5}' <<< "$r")"; paid="$(awk '{print $4}' <<< "$r")"; et="$(awk '{print $3}' <<< "$r")"
[[ "$rej" == 0 ]] || { echo "FAIL $1: $rej rejected blocks (grep 'reject' /root/fleet/dn2-node.log)"; ok=0; }
[[ "${mr:-0}" -le 3 ]] || { echo "FAIL $1: a selected-chain reorg of depth $mr"; ok=0; }
[[ -z "$VERSION_WANT" || "$v" == "$VERSION_WANT" ]] || { echo "FAIL $1: version $v"; ok=0; }
[[ "${paid:-0}" -gt "$(paid0 $1)" ]] && paid_any=1
[[ -z "$H" || "$et" -lt "$H" ]] && H="$et"
done
H=$((H > 20 ? H - 20 : 1)); HX="$(printf '0x%x' "$H")"
first=""; for l in "${BOXES[@]}"; do set -- $l; rt="$(root_at "$2" "$3" "$HX")"; echo " $1 root@$H ${rt:0:18}"; [[ -z "$first" ]] && first="$rt"; [[ "$rt" == "$first" ]] || { echo "FAIL $1: exec root at height $H ${rt:0:18} differs from ${first:0:18}"; ok=0; }; done
[[ $paid_any == 1 ]] || { echo "FAIL no segment record paid on any box during the run"; ok=0; }
res="FAIL"; [[ $ok == 1 ]] && res="PASS"
echo "$res $(stamp) release ${SHA:0:16} activation ${ACT:-none} boxes ${#BOXES[@]} minutes $MINUTES log $LOG"
python3 - "$ROOT" "$res: ${SHA:0:16} ${ACT:-current object}, ${#BOXES[@]} boxes, $MINUTES min" <<'PY'
import json, sys, os, datetime; p=os.path.join(sys.argv[1], "devnet2-status.json"); d=json.load(open(p)) if os.path.exists(p) else {}
d.update({"last_gate": sys.argv[2], "last_gate_at": datetime.datetime.utcnow().strftime("%Y-%m-%dT%H:%M:%SZ")}); json.dump(d, open(p, "w"))
PY
[[ $ok == 1 ]]

View file

@ -0,0 +1 @@
{"genesis_bits":505413632,"difficulty_v2_activation_daa":600,"proving_v0_activation_daa":900,"fees_v1_activation_daa":1200,"finality_v3_activation_daa":1500,"program_class_v3_activation_daa":1800,"proving_v1_activation_daa":1800,"proving_v1_segment_blocks":8,"proving_v1_unproven_daa":600,"proving_v1_aggregator_share_bps":1000,"proving_v1_fresh_rule_daa":2400}

36
tools/fleet/disk-sweep.py Normal file
View file

@ -0,0 +1,36 @@
#!/usr/bin/env python3
"""Disk on every live fleet box (6 October 2026, after the hub died on a full disk): one df per box in parallel, written to
~/Desktop/fleet/disk.json ({label: {pct, avail_gb, total_gb, t}}) for the fleet page's disk column; any box at or over 85
percent posts one line to Discord #incidents through tools/community/discord-hooks.mjs (incident open, the fleet's own
sentence), at most one line per box per hour (state in ~/Desktop/fleet/disk-alerts.json). Run from page.py's cadence or
by hand: disk-sweep.py [--no-alert]."""
import sys, os, json, time, subprocess
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))); from lib import Box, Registry
from concurrent.futures import ThreadPoolExecutor
ROOT = os.path.expanduser("~/Desktop/fleet"); OUT = os.path.join(ROOT, "disk.json"); ALERTS = os.path.join(ROOT, "disk-alerts.json")
HOOKS = "/Users/joshm/Projects/igneum/tools/community/discord-hooks.mjs"
def now(): return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
def one(item):
iid, b = item
try:
box = Box(b["ssh_host"], b["ssh_port"], b["label"], iid, None, b.get("provider"), b.get("ssh_user", "root"))
rc, out, err = box.run("df -B1 / | tail -1 | awk '{print $2, $4, $5}'", 25)
tot, avail, pct = out.split()
return b["label"], {"pct": int(pct.rstrip("%")), "avail_gb": round(int(avail) / 1e9, 1), "total_gb": round(int(tot) / 1e9), "t": now()}
except Exception as e: return b["label"], {"pct": None, "error": str(e)[:60], "t": now()}
def sweep(alert=True):
reg = Registry.load(); live = [(i, b) for i, b in reg.items() if b.get("state") != "destroyed" and b.get("ssh_host")]
with ThreadPoolExecutor(20) as ex: res = dict(ex.map(one, live))
json.dump(res, open(OUT, "w"), indent=1)
if alert:
st = json.load(open(ALERTS)) if os.path.exists(ALERTS) else {}
for label, d in res.items():
if d.get("pct") is not None and d["pct"] >= 85 and time.time() - st.get(label, 0) > 3600:
what = f"Fleet box {label} disk at {d['pct']} percent ({d['avail_gb']} GB free of {d['total_gb']})"
r = subprocess.run(["node", HOOKS, "incident", "open", "--what", what, "--affected", "one rented fleet box; no user-facing effect unless it is the hub (then a slower first join for home miners, no loss)", "--doing", "the standing supervisor prunes the prover's segment exports and trims the node log every ten minutes; the fleet agent reads the box", "--id", f"disk-{label}-{int(time.time())}", "--live"], capture_output=True, text=True, timeout=60)
st[label] = time.time(); print("alert", label, d["pct"], "rc", r.returncode, (r.stdout + r.stderr)[-120:].replace("\n", " "))
json.dump(st, open(ALERTS, "w"))
return res
if __name__ == "__main__":
res = sweep(alert="--no-alert" not in sys.argv)
for l, d in sorted(res.items(), key=lambda kv: -(kv[1].get("pct") or 0)): print(f"{l:<12} {str(d.get('pct')):>4}% {d.get('avail_gb', '?')} GB free")

51
tools/fleet/dn2-check.py Normal file
View file

@ -0,0 +1,51 @@
#!/usr/bin/env python3
"""The Devnet 2 standing gate's reads without the install step (6 October 2026, 21:4xZ; the seed moved to igneum-build-1,
user build, rpc 27610, so devnet2-gate.sh's root@ ssh and standard ports no longer fit the seed). Reads every Devnet 2 box
(registry rows with devnet2 or dn2_seed; the seed through its seed_env ports) and prints one line: PASS only with zero
rejected blocks on every box since SINCE, no selected-chain reorg over depth 3, equal exec state roots at a common height,
at least one paid segment record, every node on one version. Usage: dn2-check.py [--since 2026-10-06T21:00:00Z]"""
import sys, os, json, time, datetime
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))); from lib import Box, Registry
from concurrent.futures import ThreadPoolExecutor
def now(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
since = sys.argv[sys.argv.index("--since") + 1] if "--since" in sys.argv else (datetime.datetime.now(datetime.timezone.utc) - datetime.timedelta(minutes=30)).strftime("%Y-%m-%dT%H:%M:%SZ")
since_log = since.replace("T", " ").rstrip("Z")
# a box's log may stamp local time (igneum-build-1 logs +02:00): shift the since stamp by the offset the log's lines carry
SINCE_SHIFT = r"""off=$(grep -oE '^[0-9-]+ [0-9:.]+[+-][0-9]{2}:[0-9]{2}' LOGFILE | tail -1 | grep -oE '[+-][0-9]{2}:[0-9]{2}$'); s='SINCE_UTC'; if [ -n "$off" ] && [ "$off" != "+00:00" ]; then sign=${off:0:1}; h=$((10#${off:1:2})); m=$((10#${off:4:2})); secs=$(( (h*3600+m*60) )); [ "$sign" = "-" ] && secs=$((-secs)); s=$(date -u -d "$s UTC $secs seconds" +'%Y-%m-%d %H:%M:%S' 2>/dev/null || echo "$s"); fi; echo "$s" """
reg = Registry.load()
rows = [(i, b) for i, b in reg.items() if b.get("state") != "destroyed" and b.get("ssh_host") and (b.get("devnet2") or b.get("dn2_seed"))]
def read(item):
i, b = item; env = b.get("seed_env") or {}
log = env.get("SEED_LOG", "/root/fleet/dn2-node.log"); rpc = env.get("SEED_RPC", "27610" if b["label"].startswith("wave") else "26610"); evm = env.get("SEED_EVM", "27790" if b["label"].startswith("wave") else "26790"); miner = env.get("SEED_MINER", "/opt/igneum/pkg/bin/igneum-miner")
if b["label"] == "dn2-seed": rpc, evm = "27610", "27790" # its 26610 is the rehearsal box's live node; the Devnet 2 node sits on the alternate ports
box = Box(b["ssh_host"], b["ssh_port"], b["label"], i, b.get("wallet"), b.get("provider"), b.get("ssh_user", "root"))
shift = SINCE_SHIFT.replace("LOGFILE", log).replace("SINCE_UTC", since_log)
cmd = f"""w=$({miner} watch 1 grpc://127.0.0.1:{rpc} 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1); echo W $w
SL="$({shift})"; echo SINCE_LOCAL $SL
TIP=$(curl -s -m 6 -X POST -H 'Content-Type: application/json' --data '{{"jsonrpc":"2.0","id":1,"method":"igneum_getExecStatus","params":[]}}' http://127.0.0.1:{evm}/ | grep -oE '"executedTip":"0x[0-9a-f]+"' | grep -oE '0x[0-9a-f]+'); TIPN=$((${{TIP:-0x0}}))
echo REJ $(awk -v s="$SL" '($1" "$2) >= s' {log} | grep -c 'PoW rejected') REORG $(awk -v s="$SL" '($1" "$2) >= s' {log} | grep -oE 'selected-chain reorg: [0-9]+ chain blocks removed, unwinding to height [0-9]+' | awk -v t=$TIPN 'BEGIN{{m=0}} {{d=$3; h=$NF; if (t-h < 300 && d>m) m=d}} END{{print m}}')
# a reorg line more than 300 blocks below the exec tip is the follower re-walking the chain's old history after a restart (22:17Z: "unwinding to height 33" on a chain at 11,500), not a live reorg
echo EXEC $(curl -s -m 6 -X POST -H 'Content-Type: application/json' --data '{{"jsonrpc":"2.0","id":1,"method":"igneum_getExecStatus","params":[]}}' http://127.0.0.1:{evm}/ | grep -oE '"executedTip":"0x[0-9a-f]+"' | grep -oE '0x[0-9a-f]+')
echo PAID $(curl -s -m 6 -X POST -H 'Content-Type: application/json' --data '{{"jsonrpc":"2.0","id":1,"method":"igneum_getProvingStatus","params":[]}}' http://127.0.0.1:{evm}/ | grep -oE '"paidSegments":[0-9]+' | grep -oE '[0-9]+')
echo VER $(grep -oE 'igneumd/2\\.1\\.0-[0-9a-f]+' {log} | tail -1) DIGEST $(grep -o 'digest: [0-9a-f]*' {log} | tail -1 | awk '{{print substr($2,1,16)}}')"""
rc, out, err = box.run(cmd, 60); d = {"label": b["label"], "seed": bool(b.get("dn2_seed")), "evm": evm, "box": box}
for ln in out.splitlines():
p = ln.split()
if not p: continue
if p[0] == "W": d.update({k: v for k, v in (kv.split("=", 1) for kv in p[1:] if "=" in kv)})
elif p[0] == "REJ" and len(p) >= 4: d["rej"] = int(p[1] or 0); d["reorg"] = int(p[3] or 0)
elif p[0] == "EXEC" and len(p) > 1: d["exec"] = int(p[1], 16)
elif p[0] == "PAID" and len(p) > 1: d["paid"] = int(p[1])
elif p[0] == "VER": d["ver"] = p[1] if len(p) > 1 and p[1] != "DIGEST" else ""; d["digest"] = p[-1] if "DIGEST" in p else ""
return d
with ThreadPoolExecutor(8) as ex: res = list(ex.map(read, rows))
# a common height for the state root: the lowest exec tip minus a little
execs = [d.get("exec", 0) for d in res if d.get("exec")]; h = max(0, min(execs) - 2) if execs else 0
def root(d):
rc, out, err = d["box"].run(f"curl -s -m 6 -X POST -H 'Content-Type: application/json' --data '{{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"eth_getBlockByNumber\",\"params\":[\"{hex(h)}\",false]}}' http://127.0.0.1:{d['evm']}/ | grep -oE '\"stateRoot\":\"0x[0-9a-f]+\"' | cut -d'\"' -f4", 30); return out.strip()[:18]
with ThreadPoolExecutor(8) as ex:
for d, r in zip(res, ex.map(root, res)): d["root"] = r
for d in res: print(f" {d['label']:<10} {'seed ' if d['seed'] else ' '} blocks={d.get('blocks','?'):<7} daa={d.get('daa','?'):<7} synced={d.get('synced','?'):<5} exec={d.get('exec','?'):<7} root@{h}={d.get('root','?'):<18} paid={d.get('paid','?'):<4} rej={d.get('rej','?')} reorg={d.get('reorg','?')} {d.get('ver','')} {d.get('digest','')}")
rej = sum(d.get("rej", 0) for d in res); reorg = max((d.get("reorg", 0) for d in res), default=0); roots = {d.get("root") for d in res if d.get("root")}; vers = {d.get("ver") for d in res if d.get("ver")}; paid = max((d.get("paid", 0) for d in res), default=0)
ok = rej == 0 and reorg <= 3 and len(roots) == 1 and paid >= 1 and len(vers) == 1
print(f"{'PASS' if ok else 'FAIL'} {now()} since {since}: boxes {len(res)}, rejected {rej}, max reorg {reorg}, state roots at {h}: {len(roots)} distinct, paid segments {paid}, versions {sorted(vers)}")

4
tools/fleet/dn2-kill.sh Executable file
View file

@ -0,0 +1,4 @@
#!/usr/bin/env bash
# Stops everything of a Devnet 2 box but sshd: run as a FILE (an inline pkill with "igneumd" in it kills the ssh shell that carries the word).
pkill -9 -f '^bash in/box-dn2.sh'; pkill -9 -f '^/root/fleet/in/igneumd'; pkill -9 -f '^/root/fleet/in/igneumd-gate'; pkill -9 -x igneum-miner; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-worker-cuda'; pkill -9 -f '^python3 -u /root/fleet/in/box-prover.py'; pkill -9 -x sp1-gpu-server
sleep 2; echo "left=$(pgrep -c -f '^/root/fleet/in/igneumd')+$(pgrep -c -x igneum-miner)+$(pgrep -c -f '^bash in/box-dn2.sh')"

View file

@ -0,0 +1,9 @@
#!/usr/bin/env bash
# Restarts the Devnet 2 prover on a box as ONE instance: run as a FILE (an inline pkill carrying these words kills the ssh shell
# that carries them). Env: LABEL WALLET THRESHOLD (optional).
F=/root/fleet; cd $F || exit 2
pkill -9 -f '^python3 -u (/root/fleet/)?in/box-prover.py'; pkill -9 -x sp1-gpu-server; pkill -9 -f '^/opt/igneum-floor/bin/igneum-prove-host'; pkill -9 -f '^/opt/igneum-segal/proving/igneum-prove/target/release/igneum-prove-host'
sleep 3; rm -f /tmp/sp1-cuda-*.sock $F/out/prover.pid $F/out/prover-state.json.*.tmp
LABEL="${LABEL:?}" WALLET="${WALLET:?}" THRESHOLD="${THRESHOLD:-}" EXPORT_FROM=0 CHAIN_NAME=igneum-devnet-2 MINER=none RUN_HOURS=240 setsid nohup python3 -u /root/fleet/in/box-prover.py </dev/null >> $F/out/dn2-prover-launch.log 2>&1 &
sleep 25
echo "RESULT dn2_prover_restart $(date -u +%FT%TZ) instances=$(pgrep -f '^python3 -u (/root/fleet/)?in/box-prover.py' | wc -l) pid=$(cat $F/out/prover.pid 2>/dev/null) server=$(pgrep -c -x sp1-gpu-server) last=\"$(tail -1 $F/out/prover.log | cut -c1-140)\""

9
tools/fleet/dn400-worker.sh Executable file
View file

@ -0,0 +1,9 @@
#!/usr/bin/env bash
# The CUDA worker for the class v4 rehearsal miner (igneum-miner-v4 --worker <this file>): the package's igneum-worker-cuda
# needs --pack <dir> (igneum-miner export-pack writes the first one) and a device; the v4 miner passes only its own flags.
# IGNEUM_CUDA_DEVICE picks the GPU (default 0). The miner's arguments are logged once and passed through.
D="${IGNEUM_CUDA_DEVICE:-0}"; P=/root/fleet/dn400/packs/live
echo "$(date -u +%FT%TZ) worker args from the miner: $*" >> /root/fleet/out/dn400-worker-args.log
case " $* " in *" --serve "*) S="";; *) S="--serve";; esac
W=/opt/igneum/pkg/bin/igneum-worker-cuda; [ -x /root/fleet/in/igneum-worker-cuda-v4 ] && W=/root/fleet/in/igneum-worker-cuda-v4 # the generator-4 worker (release tree 563485b, zig glibc 2.36) once staged; the 0.3.14 package's refuses generator 4
exec $W $S --device "$D" --pack "$P" "$@"

16
tools/fleet/fleet-bg.sh Executable file
View file

@ -0,0 +1,16 @@
#!/usr/bin/env bash
# Mac-side background jobs by NAME with a pid file, so a job is stopped by its pid and never by a pattern (21:09Z: a
# `pkill -f <log name>` matched nothing, the roll-everything script lived on and wiped a box it had been told to hold).
# tools/fleet/fleet-bg.sh start <name> <command...> runs the command detached, stdout+stderr to ~/Desktop/fleet/<name>.log, pid to ~/Desktop/fleet/pids/<name>.pid
# tools/fleet/fleet-bg.sh stop <name> kills that pid (and its process group) if it is alive, removes the pid file
# tools/fleet/fleet-bg.sh list every job and whether its pid is alive
# macOS has no setsid, so the job runs under nohup alone there (a comment at the end of a case line swallows its ;; : none here).
set -u; P="$HOME/Desktop/fleet/pids"; L="$HOME/Desktop/fleet"; mkdir -p "$P"
case "${1:-}" in
start) n="$2"; shift 2; [ -f "$P/$n.pid" ] && kill -0 "$(cat "$P/$n.pid")" 2>/dev/null && { echo "$n already running (pid $(cat "$P/$n.pid"))"; exit 1; }
if command -v setsid >/dev/null; then nohup setsid "$@" > "$L/$n.log" 2>&1 < /dev/null & else nohup "$@" > "$L/$n.log" 2>&1 < /dev/null & fi; echo $! > "$P/$n.pid"; echo "$n started pid $(cat "$P/$n.pid") log $L/$n.log" ;;
stop) n="$2"; [ -f "$P/$n.pid" ] || { echo "$n: no pid file"; exit 1; }; pid=$(cat "$P/$n.pid")
if kill -0 "$pid" 2>/dev/null; then kill -TERM -- "-$pid" 2>/dev/null || kill -TERM "$pid"; sleep 2; kill -0 "$pid" 2>/dev/null && kill -KILL -- "-$pid" 2>/dev/null; echo "$n stopped (pid $pid)"; else echo "$n not running (pid $pid gone)"; fi; rm -f "$P/$n.pid" ;;
list) for f in "$P"/*.pid; do [ -f "$f" ] || continue; n=$(basename "$f" .pid); pid=$(cat "$f"); kill -0 "$pid" 2>/dev/null && echo "$n pid $pid alive" || echo "$n pid $pid gone"; done ;;
*) echo "usage: fleet-bg.sh start <name> <cmd...> | stop <name> | list"; exit 2 ;;
esac

187
tools/fleet/fleet.py Executable file
View file

@ -0,0 +1,187 @@
#!/usr/bin/env python3
"""The fleet orchestrator (Mac side). Registry: ~/Desktop/fleet/boxes.json ({instance: {...}}); per-box raw logs in
~/Desktop/fleet/<instance>/. Uses vast.py for the provider and ssh with ~/.ssh/igneum-fleet for the boxes.
fleet.py rent-card "<gpu name>" <label> <archs> [--min-ram MB] [--max-ram MB] [--disk 60] [--phase 1]
fleet.py wait [labels...] until ssh answers on every (named) box
fleet.py setup [labels...] push the inputs and start box-setup.sh (nohup) on every (named) box
fleet.py status [labels...] the last RESULT or STAGE line of setup.log and the node's sync line
fleet.py run <script> [labels...] push tools/fleet/<script> and start it under nohup (box-matrix.sh, box-ember.sh, ...)
fleet.py tail <label> [file] the last 30 lines of a box's log
fleet.py sh <label> "<cmd>" run one command on a box
fleet.py pull [labels...] rsync /root/fleet/out and the logs into ~/Desktop/fleet/<instance>/
fleet.py destroy <labels...> pull, then destroy, then mark the registry
"""
import json, os, sys, subprocess, time, secrets, datetime
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import vast
HERE = os.path.dirname(os.path.abspath(__file__))
ROOT = os.path.expanduser("~/Desktop/fleet"); REG = os.path.join(ROOT, "boxes.json")
SSH_OPTS = ["-i", os.path.expanduser("~/.ssh/igneum-fleet"), "-o", "StrictHostKeyChecking=no", "-o", "UserKnownHostsFile=/dev/null",
"-o", "LogLevel=ERROR", "-o", "ConnectTimeout=20", "-o", "ServerAliveInterval=15"]
INPUTS = [os.path.expanduser("~/Desktop/igneum-prove-wsl2-floor.zip"), os.path.join(HERE, "floor.patch"), os.path.join(HERE, "override.json")]
def now(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
import fcntl
def load():
if not os.path.exists(REG): return {}
for attempt in range(6): # a reader can meet a half-written file only if a writer bypasses save(); retry anyway
try: return json.load(open(REG))
except json.JSONDecodeError:
time.sleep(0.2 * (attempt + 1))
raise
def save(reg):
os.makedirs(ROOT, exist_ok=True); tmp = REG + f".tmp.{os.getpid()}"
json.dump(reg, open(tmp, "w"), indent=1); os.replace(tmp, REG)
def patch(iid, **fields):
with open(REG + ".lock", "w") as lk:
fcntl.flock(lk, fcntl.LOCK_EX)
reg = load(); reg.setdefault(str(iid), {}).update(fields); save(reg); return reg[str(iid)]
def boxes(labels, reg=None):
reg = reg or load()
sel = {k: v for k, v in reg.items() if v.get("state") != "destroyed" and (not labels or v["label"] in labels)}
if labels and len(sel) != len(set(labels)): print("unknown or destroyed labels:", set(labels) - {v["label"] for v in sel.values()}, file=sys.stderr)
return sel
def refresh_ssh(reg):
live = {str(i.get("id")): i for i in vast.instances()}
for iid, b in reg.items():
if b.get("provider", "vast") == "vast" and iid in live:
i = live[iid]; patch(iid, ssh_host=i.get("ssh_host"), ssh_port=i.get("ssh_port"), actual_status=i.get("actual_status"), status_msg=(i.get("status_msg") or "")[:120])
if any(b.get("provider") == "runpod" and b.get("state") != "destroyed" for b in reg.values()):
import runpod
for p in runpod.pods():
pid = str(p.get("id"))
if pid in reg and reg[pid].get("state") != "destroyed":
pm = p.get("portMappings") or {}
patch(pid, ssh_host=p.get("publicIp") or None, ssh_port=pm.get("22") if p.get("publicIp") else None, actual_status=p.get("desiredStatus"), port_map=pm, status_msg="")
return load()
def ssh(b, cmd, timeout=120, capture=True):
try:
r = subprocess.run(["ssh"] + SSH_OPTS + ["-p", str(b["ssh_port"]), f"root@{b['ssh_host']}", cmd], capture_output=capture, text=True, timeout=timeout, stdin=subprocess.DEVNULL)
except subprocess.TimeoutExpired: return 124, "", "timeout"
return r.returncode, (r.stdout or ""), (r.stderr or "")
def scp(b, files, dest):
return subprocess.run(["scp"] + SSH_OPTS + ["-P", str(b["ssh_port"])] + files + [f"root@{b['ssh_host']}:{dest}"], capture_output=True, text=True, timeout=900).returncode
def rent_card(gpu, label, archs, min_ram=None, max_ram=None, disk=60, phase="1", min_cores=4, gpus=1):
offers = vast.search(gpu, n=12, min_cores=min_cores, gpus=gpus)
if min_ram: offers = [o for o in offers if (o.get("gpu_ram") or 0) >= min_ram]
if max_ram: offers = [o for o in offers if (o.get("gpu_ram") or 0) <= max_ram]
if not offers: print(f"{label}: no offer for {gpu}"); return None
# prefer a few more cores for the build when the price is close: score = price + 0.01 per missing core under 8
offers.sort(key=lambda o: o["dph_total"] + 0.01 * max(0, 8 - (o.get("cpu_cores_effective") or 0)))
o = offers[0]; print(label, "->", vast.fmt_offer(o))
iid = vast.rent(o["id"], label, disk, offer=o)
reg = load()
reg[str(iid)] = {"label": label, "card": gpu, "vram_mb": o.get("gpu_ram"), "archs": archs, "offer": o["id"], "dph": o["dph_total"], "num_gpus": o.get("num_gpus"),
"cores": o.get("cpu_cores_effective"), "ram_gb": round((o.get("cpu_ram") or 0) / 1024), "geo": o.get("geolocation"), "driver": o.get("driver_version"),
"provider": "vast", "phase": phase, "state": "renting", "rented_at": now(), "wallet": "0x" + secrets.token_hex(20)}
save(reg); os.makedirs(os.path.join(ROOT, str(iid)), exist_ok=True); return iid
def wait(labels, limit=1200):
t0 = time.time()
while True:
reg = refresh_ssh(load()); pending = []
for iid, b in boxes(labels, reg).items():
if b.get("state") in ("destroyed",): continue
if b.get("ssh_ok"): continue
if not b.get("ssh_host") or not b.get("ssh_port"): pending.append((b["label"], b.get("actual_status"), b.get("status_msg"))); continue
try: rc, out, err = ssh(b, "nvidia-smi --query-gpu=name,memory.total --format=csv,noheader", timeout=40)
except subprocess.TimeoutExpired: rc, out, err = 1, "", "timeout"
if rc == 0 and out.strip(): patch(iid, ssh_ok=True, state="installing", gpu_seen=out.strip().replace("\n", ";")); print(f"{b['label']}: ssh ok, {out.strip()[:60]}")
else: pending.append((b["label"], b.get("actual_status"), (err or out).strip()[:80]))
if not pending: print("all up"); return
if time.time() - t0 > limit: print("still pending:", pending); return
print(f"waiting ({int(time.time()-t0)} s): " + "; ".join(f"{l} {s} {m}" for l, s, m in pending)); time.sleep(20)
def setup(labels):
for iid, b in boxes(labels).items():
if not b.get("ssh_ok"): print(b["label"], "no ssh yet"); continue
ssh(b, "mkdir -p /root/fleet/in /root/fleet/out")
if scp(b, INPUTS + [os.path.join(HERE, "box-setup.sh")], "/root/fleet/in/") != 0: print(b["label"], "scp failed"); continue
rc, out, err = ssh(b, f"cd /root/fleet && chmod +x in/box-setup.sh && if pgrep -f '^bash in/box-setup.sh' >/dev/null; then echo already; else RUSTUP_TOOLCHAIN=stable ARCHS={b['archs']} LABEL={b['label']} WALLET={b['wallet']} setsid nohup in/box-setup.sh </dev/null >/dev/null 2>&1 & echo started; fi", timeout=60)
print(b["label"], out.strip()[:40])
patch(iid, setup_started=now())
def status(labels):
reg = load()
from concurrent.futures import ThreadPoolExecutor
def one(item):
iid, b = item
if not b.get("ssh_ok"): return f"{b['label']:<14} {iid} no ssh", None
rc, out, err = ssh(b, "grep -E '^(RESULT|STAGE)' /root/fleet/setup.log 2>/dev/null | tail -1; grep -c '^RESULT setup_done' /root/fleet/setup.log 2>/dev/null; /opt/igneum/pkg/bin/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1; grep -E '^(RESULT|STAGE)' /root/fleet/out/matrix.log 2>/dev/null | tail -1; grep -E '^RESULT' /root/fleet/out/ember.log 2>/dev/null | tail -1; grep -E '^RESULT' /root/fleet/out/prover.log 2>/dev/null | tail -1", timeout=45)
if rc == 124: return f"{b['label']:<14} {iid} ssh timeout", None
lines = out.strip().split("\n")
done = len(lines) > 1 and lines[1].strip() == "1"
return f"{b['label']:<14} {iid} | " + " | ".join(l[:120] for l in lines if l.strip() and l.strip() not in ("0", "1")), done
items = list(boxes(labels, reg).items())
with ThreadPoolExecutor(max_workers=16) as ex: results = list(ex.map(one, items))
for (iid, b), (line, done) in zip(items, results):
print(line)
if done and b.get("state") == "installing": patch(iid, state="running")
def run(script, labels, env=""):
hub = next((v for v in load().values() if v.get("hub") and v.get("hub_peer")), None)
if hub and "HUB_PEER" not in env: env = f"HUB_PEER={hub['hub_peer']} " + env
for iid, b in boxes(labels).items():
files = [os.path.join(HERE, script), os.path.join(HERE, "box-kill.sh")] + ([os.path.join(HERE, "box-prover.py")] if "prover" in script else [])
if scp(b, files, "/root/fleet/in/") != 0: print(b["label"], "scp failed"); continue
name = os.path.basename(script)
logname = {"box-matrix.sh": "matrix", "box-ember.sh": "ember", "box-prover.sh": "prover"}.get(name, name)
# rotate the stage's log before the start: an old run's end line must never be read as this run's (the 5070, 12:33Z)
ssh(b, f"cd /root/fleet/out 2>/dev/null && for f in {logname}.log {logname}-launch.log rows.jsonl; do [ -f $f ] && mv $f $f.$(date +%s).old; done; true", timeout=30)
ssh(b, f"cd /root/fleet && chmod +x in/{name} && LABEL={b['label']} WALLET={b['wallet']} ARCHS={b['archs']} {env} setsid nohup in/{name} </dev/null >/dev/null 2>&1 & echo ok", timeout=60)
print(b["label"], name, "started")
def pull(labels):
for iid, b in boxes(labels).items():
d = os.path.join(ROOT, iid); os.makedirs(d, exist_ok=True)
r = subprocess.run(["rsync", "-az", "-e", "ssh " + " ".join(SSH_OPTS) + f" -p {b['ssh_port']}", f"root@{b['ssh_host']}:/root/fleet/out/", f"root@{b['ssh_host']}:/root/fleet/setup.log", f"root@{b['ssh_host']}:/root/fleet/node.log", d + "/"], capture_output=True, text=True, timeout=600)
print(b["label"], "pulled" if r.returncode == 0 else f"pull failed: {r.stderr[:200]}")
def destroy(labels):
reg = load()
for iid, b in boxes(labels, reg).items():
if b.get("ssh_ok"):
try: pull([b["label"]])
except Exception as e: print("pull failed", e)
vast.destroy(iid); t1 = now()
t0 = datetime.datetime.strptime(b["rented_at"], "%Y-%m-%dT%H:%M:%SZ"); h = (datetime.datetime.strptime(t1, "%Y-%m-%dT%H:%M:%SZ") - t0).total_seconds() / 3600
patch(iid, state="destroyed", destroyed_at=t1, hours=round(h, 2), cost_usd=round(h * b["dph"], 3)); print(b["label"], f"destroyed after {h:.2f} h, USD {h * b['dph']:.2f}")
if __name__ == "__main__":
a = sys.argv[1:]
if not a: print(__doc__); sys.exit(1)
c = a[0]
if c == "rent-card":
kw = {}; pos = []
i = 1
while i < len(a):
if a[i] == "--min-ram": kw["min_ram"] = int(a[i+1]); i += 2
elif a[i] == "--max-ram": kw["max_ram"] = int(a[i+1]); i += 2
elif a[i] == "--disk": kw["disk"] = int(a[i+1]); i += 2
elif a[i] == "--phase": kw["phase"] = a[i+1]; i += 2
elif a[i] == "--min-cores": kw["min_cores"] = int(a[i+1]); i += 2
elif a[i] == "--gpus": kw["gpus"] = int(a[i+1]); i += 2
else: pos.append(a[i]); i += 1
rent_card(pos[0], pos[1], pos[2], **kw)
elif c == "wait": wait(a[1:])
elif c == "setup": setup(a[1:])
elif c == "status": status(a[1:])
elif c == "run": run(a[1], a[2:])
elif c == "pull": pull(a[1:])
elif c == "destroy": destroy(a[1:])
elif c == "tail":
b = [v for v in boxes([a[1]]).values()][0]; f = a[2] if len(a) > 2 else "/root/fleet/setup.log"; print(ssh(b, f"tail -n 30 {f}")[1])
elif c == "sh":
b = [v for v in boxes([a[1]]).values()][0]; rc, out, err = ssh(b, a[2], timeout=600); print(out, err)
elif c == "dn2-version": # the version string a Linux igneumd prints, read from its strings (the gate compares every box to it)
import re as _re
data = open(a[1], "rb").read(); m = _re.search(rb"igneumd/2\.1\.0-[0-9a-f]{8}", data); print(m.group(0).decode().replace("igneumd/", "igneumd_") if m else "")
elif c == "list":
for iid, b in load().items(): print(iid, b["label"], b["card"], b.get("state"), f"${b['dph']:.3f}/h", b.get("ssh_host"), b.get("ssh_port"))

425
tools/fleet/floor-v5.patch Normal file
View file

@ -0,0 +1,425 @@
diff --git a/sp1-gpu/crates/cuda/src/task.rs b/sp1-gpu/crates/cuda/src/task.rs
index a503a86..813016b 100644
--- a/sp1-gpu/crates/cuda/src/task.rs
+++ b/sp1-gpu/crates/cuda/src/task.rs
@@ -149,7 +149,15 @@ pub enum GlobalTaskPoolBuildError {
impl TaskPoolBuilder {
pub fn new() -> Self {
- Self { capacity: None, device: CudaDevice(0), mem_release_threshold: u64::MAX }
+ // Igneum prover-floor patch: upstream keeps every freed device allocation in the pool for the process's
+ // life (threshold u64::MAX), so the prover holds its high-water mark between shards on a card it shares
+ // with a miner. `SP1_GPU_MEM_RELEASE_THRESHOLD=<bytes>` sets the pool's release threshold (0 returns
+ // freed memory to the driver at once); unset, upstream's behaviour.
+ let mem_release_threshold = std::env::var("SP1_GPU_MEM_RELEASE_THRESHOLD")
+ .ok()
+ .and_then(|s| s.parse::<u64>().ok())
+ .unwrap_or(u64::MAX);
+ Self { capacity: None, device: CudaDevice(0), mem_release_threshold }
}
pub fn num_tasks(mut self, num_tasks: usize) -> Self {
diff --git a/sp1-gpu/crates/jagged_tracegen/src/lib.rs b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
index 579f70a..2264044 100644
--- a/sp1-gpu/crates/jagged_tracegen/src/lib.rs
+++ b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
@@ -481,6 +481,33 @@ async fn device_preprocessed_tracegen<A: CudaTracegenAir<Felt>>(
named_traces
}
+/// Igneum prover-floor patch: the dense elements a set of traces will occupy once `generate_jagged_traces`
+/// has laid them out, that is the sum of their buffers padded to the next multiple of 2^log_stacking_height
+/// (the "final padding" step below). Each phase (preprocessed, then main) is padded on its own.
+pub fn padded_trace_elements(
+ traces: &BTreeMap<String, Trace<TaskScope>>,
+ log_stacking_height: u32,
+) -> usize {
+ let total: usize = traces
+ .values()
+ .map(|t| match t {
+ Trace::Real(trace) => trace.guts().as_buffer().len(),
+ Trace::Padding(_) => 0,
+ })
+ .sum();
+ total.next_multiple_of(1 << log_stacking_height)
+}
+
+/// Igneum prover-floor patch: the capacity to allocate for a trace set: the exact padded size plus one
+/// stacking height of slack, never more than the prover's `max_trace_size`. `SP1_GPU_FLOOR_EXACT=0` restores
+/// upstream's full-capacity allocation.
+fn floor_capacity(max_trace_size: usize, needed: usize, log_stacking_height: u32) -> usize {
+ if std::env::var("SP1_GPU_FLOOR_EXACT").map(|v| v == "0").unwrap_or(false) {
+ return max_trace_size;
+ }
+ max_trace_size.min(needed + (1 << log_stacking_height))
+}
+
async fn allocate_and_initialize_traces(
preprocessed_traces: BTreeMap<String, Trace<TaskScope>>,
max_trace_size: usize,
@@ -494,6 +521,11 @@ async fn allocate_and_initialize_traces(
let total_gb = total_bytes as f64 / (1 << 30) as f64;
tracing::debug!("Allocating {:?} GB of traces", total_gb);
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
+ eprintln!(
+ "FLOOR tracegen alloc capacity_elements={max_trace_size} bytes={total_bytes} ({total_gb:.3} GB)"
+ );
+ }
let mut dense_data: Buffer<Felt, TaskScope> =
Buffer::with_capacity_in(max_trace_size, backend.clone());
let mut col_index: Buffer<u32, TaskScope> =
@@ -677,9 +709,14 @@ pub async fn setup_tracegen<A: CudaTracegenAir<Felt>>(
let preprocessed_traces =
device_preprocessed_tracegen(program, host_phase_tracegen, backend).await;
+ let capacity = floor_capacity(
+ max_trace_size,
+ padded_trace_elements(&preprocessed_traces, log_stacking_height),
+ log_stacking_height,
+ );
let jagged_traces = allocate_and_initialize_traces(
preprocessed_traces,
- max_trace_size,
+ capacity,
log_stacking_height,
max_log_row_count,
backend,
@@ -906,6 +943,11 @@ pub async fn main_tracegen<GC: IopCtx<F = Felt>, A: CudaTracegenAir<Felt>>(
log_chip_stats(machine, &chip_set, &traces);
+ // Igneum prover-floor patch: the key's buffer is sized to its preprocessed traces at setup (upstream sized it
+ // for a whole shard), so grow it here to what this shard needs before the main traces are appended: a bigger
+ // dense buffer and column index, the preprocessed region copied device to device, swapped into the key.
+ grow_for_main(&mut jagged_traces.preprocessed_traces, &traces, log_stacking_height, backend);
+
copy_main_jagged_traces(
traces,
&mut jagged_traces.preprocessed_traces,
@@ -918,6 +960,61 @@ pub async fn main_tracegen<GC: IopCtx<F = Felt>, A: CudaTracegenAir<Felt>>(
(public_values, chip_set, permit)
}
+/// Igneum prover-floor patch: see `main_tracegen`. The need is the preprocessed phase as laid out (its padded
+/// end, `preprocessed_offset`) plus the main traces padded to the stacking height plus one stacking height of
+/// slack; a buffer at least that big is left alone. The process aborts, loudly, if the copy cannot be made,
+/// because a panic inside a prover task is what left sweep 2 hanging on the client's socket.
+fn grow_for_main(
+ jagged: &mut JaggedTraceMle<Felt, TaskScope>,
+ main_traces: &BTreeMap<String, Trace<TaskScope>>,
+ log_stacking_height: u32,
+ backend: &TaskScope,
+) {
+ let pre_end = jagged.dense().preprocessed_offset;
+ let needed = pre_end
+ + padded_trace_elements(main_traces, log_stacking_height)
+ + (1 << log_stacking_height);
+ let have = jagged.dense().dense.capacity();
+ if have >= needed {
+ return;
+ }
+ let mut new_dense: Buffer<Felt, TaskScope> = Buffer::with_capacity_in(needed, backend.clone());
+ let mut new_col_index: Buffer<u32, TaskScope> =
+ Buffer::with_capacity_in(needed >> 1, backend.clone());
+ unsafe {
+ new_dense.assume_init();
+ new_col_index.assume_init();
+ }
+ {
+ let JaggedMle { dense_data, col_index, .. } = &mut **jagged;
+ let src_dense: &Slice<_, _> = &dense_data.dense[..pre_end];
+ let dst_dense: &mut Slice<_, _> = &mut new_dense[..pre_end];
+ let src_col: &Slice<_, _> = &col_index[..pre_end >> 1];
+ let dst_col: &mut Slice<_, _> = &mut new_col_index[..pre_end >> 1];
+ unsafe {
+ if dst_dense.copy_from_slice(src_dense, backend).is_err()
+ || dst_col.copy_from_slice(src_col, backend).is_err()
+ {
+ eprintln!("FLOOR grow FAILED: could not copy the preprocessed region ({pre_end} elements) into the grown buffer ({needed} elements); aborting instead of hanging");
+ std::process::abort();
+ }
+ }
+ }
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
+ eprintln!(
+ "FLOOR grow key buffer {have} -> {needed} elements (preprocessed {pre_end}, {} bytes)",
+ needed * 6
+ );
+ }
+ let JaggedMle { dense_data, col_index, .. } = &mut **jagged;
+ dense_data.dense = new_dense;
+ *col_index = new_col_index;
+ unsafe {
+ dense_data.dense.set_len(pre_end);
+ col_index.set_len(pre_end >> 1);
+ }
+}
+
#[allow(clippy::too_many_arguments)]
pub async fn main_tracegen_permit<GC: IopCtx<F = Felt>, A: CudaTracegenAir<Felt>>(
machine: &Machine<Felt, A>,
@@ -984,9 +1081,15 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
log_chip_stats(machine, &chip_set, &main_traces);
+ let capacity = floor_capacity(
+ max_trace_size,
+ padded_trace_elements(&preprocessed_traces, log_stacking_height)
+ + padded_trace_elements(&main_traces, log_stacking_height),
+ log_stacking_height,
+ );
let mut jagged_mle = allocate_and_initialize_traces(
preprocessed_traces,
- max_trace_size,
+ capacity,
log_stacking_height,
max_log_row_count,
backend,
@@ -1002,6 +1105,18 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
)
.await;
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
+ let dense = jagged_mle.dense();
+ let (free, total) = sp1_gpu_cudart::cuda_memory_info().unwrap_or((0, 0));
+ eprintln!(
+ "FLOOR tracegen used preprocessed_elements={} main_elements={} dense_len={} capacity_elements={capacity} max_trace_size={max_trace_size} device_used_mib={}",
+ dense.preprocessed_offset,
+ dense.main_size(),
+ dense.dense.len(),
+ (total - free) >> 20
+ );
+ }
+
(public_values, jagged_mle, chip_set, permit)
}
diff --git a/sp1-gpu/crates/prover_components/src/builder.rs b/sp1-gpu/crates/prover_components/src/builder.rs
index 5dccd9d..574d4fa 100644
--- a/sp1-gpu/crates/prover_components/src/builder.rs
+++ b/sp1-gpu/crates/prover_components/src/builder.rs
@@ -23,28 +23,75 @@ use crate::{
SP1CudaProverComponents,
};
+/// Igneum prover-floor patch (5 October 2026). Upstream sizes every device buffer for a 24 GB card or larger
+/// and panics below that, whatever the shard. Here the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks a
+/// tier, and `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly.
+/// The proof format, the verifier and the program ids do not change: the element threshold only decides where
+/// the executor splits shards, as upstream's own 24 GB tier already does.
+fn env_usize(name: &str) -> Option<usize> {
+ std::env::var(name).ok().and_then(|s| s.parse::<usize>().ok())
+}
+
+fn env_f64(name: &str) -> Option<f64> {
+ std::env::var(name).ok().and_then(|s| s.parse::<f64>().ok())
+}
+
+/// The core element threshold for a memory budget in GB (upstream's own figure for the budget, +4, as it
+/// computed it: a 32 GB card is 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16).
+pub fn element_threshold_for_budget(gpu_memory_gb: usize, full_size_shards: bool) -> u64 {
+ if gpu_memory_gb > 30 || (full_size_shards && gpu_memory_gb >= 24) {
+ ELEMENT_THRESHOLD
+ } else if gpu_memory_gb >= 24 {
+ ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24)
+ } else if gpu_memory_gb >= 18 {
+ (1 << 27) + (1 << 26)
+ } else {
+ 1 << 27
+ }
+}
+
+/// The recursion trace allocation (elements) for a memory budget.
+pub fn recursion_trace_allocation_for_budget(gpu_memory_gb: usize) -> usize {
+ if gpu_memory_gb >= 24 {
+ RECURSION_TRACE_ALLOCATION
+ } else {
+ RECURSION_TRACE_ALLOCATION
+ }
+}
+
+pub fn gpu_memory_gb() -> usize {
+ let gb = 1024.0 * 1024.0 * 1024.0;
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
+ Some(b) => (b.ceil() as usize) + 4,
+ None => (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4,
+ }
+}
+
+pub fn recursion_trace_allocation() -> usize {
+ env_usize("SP1_GPU_RECURSION_TRACE_ALLOCATION")
+ .unwrap_or_else(|| recursion_trace_allocation_for_budget(gpu_memory_gb()))
+}
+
pub fn local_gpu_opts() -> SP1CoreOpts {
let mut opts = SP1CoreOpts::default();
let log2_shard_size = 24;
opts.shard_size = 1 << log2_shard_size;
- let gb = 1024.0 * 1024.0 * 1024.0;
-
- // Get the amount of memory on the GPU.
- let gpu_memory_gb: usize = (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4;
-
- if gpu_memory_gb < 24 {
- panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB");
- }
+ // The card's memory plus 4, as upstream computed it (a 32 GB card reads 36), or the budget given.
+ let gpu_memory_gb = gpu_memory_gb();
- let shard_threshold = if !opts.full_size_shards && gpu_memory_gb <= 30 {
- ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24)
- } else {
- ELEMENT_THRESHOLD
+ let shard_threshold = match env_usize("SP1_GPU_ELEMENT_THRESHOLD") {
+ Some(t) => t as u64,
+ None => element_threshold_for_budget(gpu_memory_gb, opts.full_size_shards),
};
+ let height_threshold = opts.sharding_threshold.height_threshold;
- tracing::debug!("Shard threshold: {shard_threshold}");
+ eprintln!(
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
+ recursion_trace_allocation(),
+ opts.full_size_shards
+ );
opts.sharding_threshold.element_threshold = shard_threshold;
opts.global_dependencies_opt = true;
@@ -92,7 +139,7 @@ pub async fn recursion_prover_and_verifier(
) {
let recursion_verifier = SP1CudaProverComponents::compress_verifier();
(
- new_cuda_prover(&recursion_verifier, RECURSION_TRACE_ALLOCATION, 4, false, false, scope)
+ new_cuda_prover(&recursion_verifier, recursion_trace_allocation(), 4, false, false, scope)
.await,
recursion_verifier,
)
diff --git a/sp1-gpu/crates/server/src/main.rs b/sp1-gpu/crates/server/src/main.rs
index 65e94f5..498357c 100644
--- a/sp1-gpu/crates/server/src/main.rs
+++ b/sp1-gpu/crates/server/src/main.rs
@@ -17,9 +17,55 @@ struct Args {
version: bool,
}
+/// Igneum prover-floor patch (6 October 2026, the GPU fleet's finding): a panic inside a prover task (an allocation
+/// the card cannot meet, `cudaMallocAsync` failing and `Buffer::with_capacity_in` panicking in a tokio worker) left
+/// the request's future waiting for ever, the client on its socket, the card at 0% for the 15 minutes until someone
+/// killed it (the 8 GB and 10 GB cards at threshold 2^27). The server must fail the shard instead: this hook names
+/// the stage from the panic's location and exits, so the client's proof fails at once with the server's last line.
+fn stage_of(location: &str) -> &'static str {
+ let l = location.to_ascii_lowercase();
+ if l.contains("jagged_tracegen") || l.contains("/tracegen") {
+ "trace generation"
+ } else if l.contains("commit") || l.contains("basefold") || l.contains("merkle") {
+ "the commit (codewords and Merkle trees)"
+ } else if l.contains("logup_gkr") {
+ "LogUp GKR"
+ } else if l.contains("zerocheck") {
+ "the zerocheck"
+ } else if l.contains("jagged") {
+ "the jagged sumcheck"
+ } else if l.contains("prover_components") || l.contains("recursion") || l.contains("sp1-prover") || l.contains("sp1_prover") {
+ "the recursion (compression)"
+ } else if l.contains("cuda") || l.contains("slop") {
+ "a device allocation"
+ } else {
+ "the prover"
+ }
+}
+
+fn install_abort_on_panic() {
+ std::panic::set_hook(Box::new(|info| {
+ let location = info.location().map(|l| format!("{}:{}", l.file(), l.line())).unwrap_or_else(|| "unknown".into());
+ let message = info
+ .payload()
+ .downcast_ref::<&str>()
+ .map(|s| s.to_string())
+ .or_else(|| info.payload().downcast_ref::<String>().cloned())
+ .unwrap_or_default();
+ let oom = message.to_ascii_lowercase().contains("alloc") || message.contains("MemoryAllocation") || message.contains("OUT_OF_MEMORY");
+ eprintln!(
+ "FLOOR abort: {} failed at {location}: {message}{}; the server exits so the client's proof fails instead of waiting",
+ stage_of(&location),
+ if oom { " (the card's memory could not meet an allocation: lower SP1_GPU_ELEMENT_THRESHOLD one notch)" } else { "" }
+ );
+ std::process::exit(70);
+ }));
+}
+
#[tokio::main]
#[allow(clippy::print_stdout)]
async fn main() {
+ install_abort_on_panic();
tracing_subscriber::fmt::init();
let args = Args::parse();
@@ -40,3 +86,19 @@ async fn main() {
eprintln!("Error running server: {e}");
}
}
+
+#[cfg(test)]
+mod floor_tests {
+ use super::stage_of;
+
+ #[test]
+ fn the_stage_is_named_from_the_panic_location() {
+ assert_eq!(stage_of("sp1-gpu/crates/jagged_tracegen/src/lib.rs:240"), "trace generation");
+ assert_eq!(stage_of("sp1-gpu/crates/basefold/src/fri.rs:97"), "the commit (codewords and Merkle trees)");
+ assert_eq!(stage_of("sp1-gpu/crates/logup_gkr/src/tracegen.rs:72"), "LogUp GKR");
+ assert_eq!(stage_of("sp1-gpu/crates/zerocheck/src/prover.rs:1163"), "the zerocheck");
+ assert_eq!(stage_of("sp1-gpu/crates/prover_components/src/builder.rs:70"), "the recursion (compression)");
+ assert_eq!(stage_of("sp1-gpu/crates/cuda/src/stream.rs:330"), "a device allocation");
+ assert_eq!(stage_of("somewhere/else.rs:1"), "the prover");
+ }
+}
diff --git a/sp1-gpu/crates/server/src/server.rs b/sp1-gpu/crates/server/src/server.rs
index 4035f1f..0d0d907 100644
--- a/sp1-gpu/crates/server/src/server.rs
+++ b/sp1-gpu/crates/server/src/server.rs
@@ -157,6 +157,7 @@ impl Server {
};
let pk = CachedProgram { elf: Arc::new(Elf::Dynamic(elf.into())), vk: vk.clone() };
ctx.pk_cache.insert(elf_hash, pk);
+ floor_memory_line("after setup");
Response::Setup { id: elf_hash, vk }
}
Request::Destroy { key } => {
@@ -177,15 +178,31 @@ impl Server {
);
};
let context = SP1Context::builder().proof_nonce(proof_nonce).build();
- match prover.prove_with_mode(&cached.elf, stdin, context, mode).await {
+ let started = std::time::Instant::now();
+ let response = match prover.prove_with_mode(&cached.elf, stdin, context, mode).await {
Ok(proof) => Response::Proof { proof },
Err(e) => Response::ProverError(e.to_string()),
- }
+ };
+ floor_memory_line(&format!("after prove {:?} in {:.1} s", mode, started.elapsed().as_secs_f64()));
+ response
}
}
}
}
+/// Igneum prover-floor patch: the device memory in use (total minus free, as the driver reports it) at the
+/// points that bound a proof, so a run's log carries the terms of the peak without a sampler.
+fn floor_memory_line(what: &str) {
+ if let Ok((free, total)) = sp1_gpu_cudart::cuda_memory_info() {
+ eprintln!(
+ "FLOOR memory {what}: device_used_mib={} free_mib={} total_mib={}",
+ (total - free) >> 20,
+ free >> 20,
+ total >> 20
+ );
+ }
+}
+
fn sha256(data: &[u8]) -> [u8; 32] {
use sha2::{Digest, Sha256};
let mut hasher = Sha256::new();

345
tools/fleet/floor.patch Normal file
View file

@ -0,0 +1,345 @@
diff --git a/sp1-gpu/crates/cuda/src/task.rs b/sp1-gpu/crates/cuda/src/task.rs
index a503a86..813016b 100644
--- a/sp1-gpu/crates/cuda/src/task.rs
+++ b/sp1-gpu/crates/cuda/src/task.rs
@@ -149,7 +149,15 @@ pub enum GlobalTaskPoolBuildError {
impl TaskPoolBuilder {
pub fn new() -> Self {
- Self { capacity: None, device: CudaDevice(0), mem_release_threshold: u64::MAX }
+ // Igneum prover-floor patch: upstream keeps every freed device allocation in the pool for the process's
+ // life (threshold u64::MAX), so the prover holds its high-water mark between shards on a card it shares
+ // with a miner. `SP1_GPU_MEM_RELEASE_THRESHOLD=<bytes>` sets the pool's release threshold (0 returns
+ // freed memory to the driver at once); unset, upstream's behaviour.
+ let mem_release_threshold = std::env::var("SP1_GPU_MEM_RELEASE_THRESHOLD")
+ .ok()
+ .and_then(|s| s.parse::<u64>().ok())
+ .unwrap_or(u64::MAX);
+ Self { capacity: None, device: CudaDevice(0), mem_release_threshold }
}
pub fn num_tasks(mut self, num_tasks: usize) -> Self {
diff --git a/sp1-gpu/crates/jagged_tracegen/src/lib.rs b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
index 579f70a..2264044 100644
--- a/sp1-gpu/crates/jagged_tracegen/src/lib.rs
+++ b/sp1-gpu/crates/jagged_tracegen/src/lib.rs
@@ -481,6 +481,33 @@ async fn device_preprocessed_tracegen<A: CudaTracegenAir<Felt>>(
named_traces
}
+/// Igneum prover-floor patch: the dense elements a set of traces will occupy once `generate_jagged_traces`
+/// has laid them out, that is the sum of their buffers padded to the next multiple of 2^log_stacking_height
+/// (the "final padding" step below). Each phase (preprocessed, then main) is padded on its own.
+pub fn padded_trace_elements(
+ traces: &BTreeMap<String, Trace<TaskScope>>,
+ log_stacking_height: u32,
+) -> usize {
+ let total: usize = traces
+ .values()
+ .map(|t| match t {
+ Trace::Real(trace) => trace.guts().as_buffer().len(),
+ Trace::Padding(_) => 0,
+ })
+ .sum();
+ total.next_multiple_of(1 << log_stacking_height)
+}
+
+/// Igneum prover-floor patch: the capacity to allocate for a trace set: the exact padded size plus one
+/// stacking height of slack, never more than the prover's `max_trace_size`. `SP1_GPU_FLOOR_EXACT=0` restores
+/// upstream's full-capacity allocation.
+fn floor_capacity(max_trace_size: usize, needed: usize, log_stacking_height: u32) -> usize {
+ if std::env::var("SP1_GPU_FLOOR_EXACT").map(|v| v == "0").unwrap_or(false) {
+ return max_trace_size;
+ }
+ max_trace_size.min(needed + (1 << log_stacking_height))
+}
+
async fn allocate_and_initialize_traces(
preprocessed_traces: BTreeMap<String, Trace<TaskScope>>,
max_trace_size: usize,
@@ -494,6 +521,11 @@ async fn allocate_and_initialize_traces(
let total_gb = total_bytes as f64 / (1 << 30) as f64;
tracing::debug!("Allocating {:?} GB of traces", total_gb);
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
+ eprintln!(
+ "FLOOR tracegen alloc capacity_elements={max_trace_size} bytes={total_bytes} ({total_gb:.3} GB)"
+ );
+ }
let mut dense_data: Buffer<Felt, TaskScope> =
Buffer::with_capacity_in(max_trace_size, backend.clone());
let mut col_index: Buffer<u32, TaskScope> =
@@ -677,9 +709,14 @@ pub async fn setup_tracegen<A: CudaTracegenAir<Felt>>(
let preprocessed_traces =
device_preprocessed_tracegen(program, host_phase_tracegen, backend).await;
+ let capacity = floor_capacity(
+ max_trace_size,
+ padded_trace_elements(&preprocessed_traces, log_stacking_height),
+ log_stacking_height,
+ );
let jagged_traces = allocate_and_initialize_traces(
preprocessed_traces,
- max_trace_size,
+ capacity,
log_stacking_height,
max_log_row_count,
backend,
@@ -906,6 +943,11 @@ pub async fn main_tracegen<GC: IopCtx<F = Felt>, A: CudaTracegenAir<Felt>>(
log_chip_stats(machine, &chip_set, &traces);
+ // Igneum prover-floor patch: the key's buffer is sized to its preprocessed traces at setup (upstream sized it
+ // for a whole shard), so grow it here to what this shard needs before the main traces are appended: a bigger
+ // dense buffer and column index, the preprocessed region copied device to device, swapped into the key.
+ grow_for_main(&mut jagged_traces.preprocessed_traces, &traces, log_stacking_height, backend);
+
copy_main_jagged_traces(
traces,
&mut jagged_traces.preprocessed_traces,
@@ -918,6 +960,61 @@ pub async fn main_tracegen<GC: IopCtx<F = Felt>, A: CudaTracegenAir<Felt>>(
(public_values, chip_set, permit)
}
+/// Igneum prover-floor patch: see `main_tracegen`. The need is the preprocessed phase as laid out (its padded
+/// end, `preprocessed_offset`) plus the main traces padded to the stacking height plus one stacking height of
+/// slack; a buffer at least that big is left alone. The process aborts, loudly, if the copy cannot be made,
+/// because a panic inside a prover task is what left sweep 2 hanging on the client's socket.
+fn grow_for_main(
+ jagged: &mut JaggedTraceMle<Felt, TaskScope>,
+ main_traces: &BTreeMap<String, Trace<TaskScope>>,
+ log_stacking_height: u32,
+ backend: &TaskScope,
+) {
+ let pre_end = jagged.dense().preprocessed_offset;
+ let needed = pre_end
+ + padded_trace_elements(main_traces, log_stacking_height)
+ + (1 << log_stacking_height);
+ let have = jagged.dense().dense.capacity();
+ if have >= needed {
+ return;
+ }
+ let mut new_dense: Buffer<Felt, TaskScope> = Buffer::with_capacity_in(needed, backend.clone());
+ let mut new_col_index: Buffer<u32, TaskScope> =
+ Buffer::with_capacity_in(needed >> 1, backend.clone());
+ unsafe {
+ new_dense.assume_init();
+ new_col_index.assume_init();
+ }
+ {
+ let JaggedMle { dense_data, col_index, .. } = &mut **jagged;
+ let src_dense: &Slice<_, _> = &dense_data.dense[..pre_end];
+ let dst_dense: &mut Slice<_, _> = &mut new_dense[..pre_end];
+ let src_col: &Slice<_, _> = &col_index[..pre_end >> 1];
+ let dst_col: &mut Slice<_, _> = &mut new_col_index[..pre_end >> 1];
+ unsafe {
+ if dst_dense.copy_from_slice(src_dense, backend).is_err()
+ || dst_col.copy_from_slice(src_col, backend).is_err()
+ {
+ eprintln!("FLOOR grow FAILED: could not copy the preprocessed region ({pre_end} elements) into the grown buffer ({needed} elements); aborting instead of hanging");
+ std::process::abort();
+ }
+ }
+ }
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
+ eprintln!(
+ "FLOOR grow key buffer {have} -> {needed} elements (preprocessed {pre_end}, {} bytes)",
+ needed * 6
+ );
+ }
+ let JaggedMle { dense_data, col_index, .. } = &mut **jagged;
+ dense_data.dense = new_dense;
+ *col_index = new_col_index;
+ unsafe {
+ dense_data.dense.set_len(pre_end);
+ col_index.set_len(pre_end >> 1);
+ }
+}
+
#[allow(clippy::too_many_arguments)]
pub async fn main_tracegen_permit<GC: IopCtx<F = Felt>, A: CudaTracegenAir<Felt>>(
machine: &Machine<Felt, A>,
@@ -984,9 +1081,15 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
log_chip_stats(machine, &chip_set, &main_traces);
+ let capacity = floor_capacity(
+ max_trace_size,
+ padded_trace_elements(&preprocessed_traces, log_stacking_height)
+ + padded_trace_elements(&main_traces, log_stacking_height),
+ log_stacking_height,
+ );
let mut jagged_mle = allocate_and_initialize_traces(
preprocessed_traces,
- max_trace_size,
+ capacity,
log_stacking_height,
max_log_row_count,
backend,
@@ -1002,6 +1105,18 @@ pub async fn full_tracegen<A: CudaTracegenAir<Felt>>(
)
.await;
+ if std::env::var("SP1_GPU_FLOOR_LOG").is_ok() {
+ let dense = jagged_mle.dense();
+ let (free, total) = sp1_gpu_cudart::cuda_memory_info().unwrap_or((0, 0));
+ eprintln!(
+ "FLOOR tracegen used preprocessed_elements={} main_elements={} dense_len={} capacity_elements={capacity} max_trace_size={max_trace_size} device_used_mib={}",
+ dense.preprocessed_offset,
+ dense.main_size(),
+ dense.dense.len(),
+ (total - free) >> 20
+ );
+ }
+
(public_values, jagged_mle, chip_set, permit)
}
diff --git a/sp1-gpu/crates/prover_components/src/builder.rs b/sp1-gpu/crates/prover_components/src/builder.rs
index 5dccd9d..574d4fa 100644
--- a/sp1-gpu/crates/prover_components/src/builder.rs
+++ b/sp1-gpu/crates/prover_components/src/builder.rs
@@ -23,28 +23,75 @@ use crate::{
SP1CudaProverComponents,
};
+/// Igneum prover-floor patch (5 October 2026). Upstream sizes every device buffer for a 24 GB card or larger
+/// and panics below that, whatever the shard. Here the card's memory (or `SP1_GPU_MEMORY_BUDGET_GB`) picks a
+/// tier, and `SP1_GPU_ELEMENT_THRESHOLD` / `SP1_GPU_RECURSION_TRACE_ALLOCATION` set the two buffers directly.
+/// The proof format, the verifier and the program ids do not change: the element threshold only decides where
+/// the executor splits shards, as upstream's own 24 GB tier already does.
+fn env_usize(name: &str) -> Option<usize> {
+ std::env::var(name).ok().and_then(|s| s.parse::<usize>().ok())
+}
+
+fn env_f64(name: &str) -> Option<f64> {
+ std::env::var(name).ok().and_then(|s| s.parse::<f64>().ok())
+}
+
+/// The core element threshold for a memory budget in GB (upstream's own figure for the budget, +4, as it
+/// computed it: a 32 GB card is 36, a 24 GB card 28, a 16 GB card 20, a 12 GB card 16).
+pub fn element_threshold_for_budget(gpu_memory_gb: usize, full_size_shards: bool) -> u64 {
+ if gpu_memory_gb > 30 || (full_size_shards && gpu_memory_gb >= 24) {
+ ELEMENT_THRESHOLD
+ } else if gpu_memory_gb >= 24 {
+ ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24)
+ } else if gpu_memory_gb >= 18 {
+ (1 << 27) + (1 << 26)
+ } else {
+ 1 << 27
+ }
+}
+
+/// The recursion trace allocation (elements) for a memory budget.
+pub fn recursion_trace_allocation_for_budget(gpu_memory_gb: usize) -> usize {
+ if gpu_memory_gb >= 24 {
+ RECURSION_TRACE_ALLOCATION
+ } else {
+ RECURSION_TRACE_ALLOCATION
+ }
+}
+
+pub fn gpu_memory_gb() -> usize {
+ let gb = 1024.0 * 1024.0 * 1024.0;
+ match env_f64("SP1_GPU_MEMORY_BUDGET_GB") {
+ Some(b) => (b.ceil() as usize) + 4,
+ None => (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4,
+ }
+}
+
+pub fn recursion_trace_allocation() -> usize {
+ env_usize("SP1_GPU_RECURSION_TRACE_ALLOCATION")
+ .unwrap_or_else(|| recursion_trace_allocation_for_budget(gpu_memory_gb()))
+}
+
pub fn local_gpu_opts() -> SP1CoreOpts {
let mut opts = SP1CoreOpts::default();
let log2_shard_size = 24;
opts.shard_size = 1 << log2_shard_size;
- let gb = 1024.0 * 1024.0 * 1024.0;
-
- // Get the amount of memory on the GPU.
- let gpu_memory_gb: usize = (((cuda_memory_info().unwrap().1 as f64) / gb).ceil() as usize) + 4;
-
- if gpu_memory_gb < 24 {
- panic!("Unsupported GPU memory: {gpu_memory_gb}, must be at least 24GB");
- }
+ // The card's memory plus 4, as upstream computed it (a 32 GB card reads 36), or the budget given.
+ let gpu_memory_gb = gpu_memory_gb();
- let shard_threshold = if !opts.full_size_shards && gpu_memory_gb <= 30 {
- ELEMENT_THRESHOLD - (1 << 26) - (1 << 25) - (1 << 24)
- } else {
- ELEMENT_THRESHOLD
+ let shard_threshold = match env_usize("SP1_GPU_ELEMENT_THRESHOLD") {
+ Some(t) => t as u64,
+ None => element_threshold_for_budget(gpu_memory_gb, opts.full_size_shards),
};
+ let height_threshold = opts.sharding_threshold.height_threshold;
- tracing::debug!("Shard threshold: {shard_threshold}");
+ eprintln!(
+ "FLOOR opts gpu_memory_gb={gpu_memory_gb} element_threshold={shard_threshold} height_threshold={height_threshold} recursion_trace_allocation={} full_size_shards={}",
+ recursion_trace_allocation(),
+ opts.full_size_shards
+ );
opts.sharding_threshold.element_threshold = shard_threshold;
opts.global_dependencies_opt = true;
@@ -92,7 +139,7 @@ pub async fn recursion_prover_and_verifier(
) {
let recursion_verifier = SP1CudaProverComponents::compress_verifier();
(
- new_cuda_prover(&recursion_verifier, RECURSION_TRACE_ALLOCATION, 4, false, false, scope)
+ new_cuda_prover(&recursion_verifier, recursion_trace_allocation(), 4, false, false, scope)
.await,
recursion_verifier,
)
diff --git a/sp1-gpu/crates/server/src/server.rs b/sp1-gpu/crates/server/src/server.rs
index 4035f1f..0d0d907 100644
--- a/sp1-gpu/crates/server/src/server.rs
+++ b/sp1-gpu/crates/server/src/server.rs
@@ -157,6 +157,7 @@ impl Server {
};
let pk = CachedProgram { elf: Arc::new(Elf::Dynamic(elf.into())), vk: vk.clone() };
ctx.pk_cache.insert(elf_hash, pk);
+ floor_memory_line("after setup");
Response::Setup { id: elf_hash, vk }
}
Request::Destroy { key } => {
@@ -177,15 +178,31 @@ impl Server {
);
};
let context = SP1Context::builder().proof_nonce(proof_nonce).build();
- match prover.prove_with_mode(&cached.elf, stdin, context, mode).await {
+ let started = std::time::Instant::now();
+ let response = match prover.prove_with_mode(&cached.elf, stdin, context, mode).await {
Ok(proof) => Response::Proof { proof },
Err(e) => Response::ProverError(e.to_string()),
- }
+ };
+ floor_memory_line(&format!("after prove {:?} in {:.1} s", mode, started.elapsed().as_secs_f64()));
+ response
}
}
}
}
+/// Igneum prover-floor patch: the device memory in use (total minus free, as the driver reports it) at the
+/// points that bound a proof, so a run's log carries the terms of the peak without a sampler.
+fn floor_memory_line(what: &str) {
+ if let Ok((free, total)) = sp1_gpu_cudart::cuda_memory_info() {
+ eprintln!(
+ "FLOOR memory {what}: device_used_mib={} free_mib={} total_mib={}",
+ (total - free) >> 20,
+ free >> 20,
+ total >> 20
+ );
+ }
+}
+
fn sha256(data: &[u8]) -> [u8; 32] {
use sha2::{Digest, Sha256};
let mut hasher = Sha256::new();

View file

@ -0,0 +1,20 @@
"""tools/fleet/lib: the ONE tested library for operations on a rented box (the project lead's standing rule, CLAUDE.md e1d0b23,
6 October 2026). Every fleet playbook (canary, wave, matrix, recovery, devnet2-gate) calls these; a watcher reads the
chain-side fact (height, peers, exec tip, paid records), never a reported MH/s or a process name.
from lib import Box
b = Box.from_registry("dn2-1") # or Box(host, port, label)
b.run("hostname") # ssh, with ServerAlive and a timeout; (rc, out, err)
b.put([local, ...], "/root/fleet/in/") # scp with three tries
b.install_payload(url_or_path, sha256) # the node binary into /root/fleet/in/igneumd-<sha16>, checked on the box
b.start_node(binary, override, peers, devnet_suffix=None, verifier=None, extra=[]) # anchored kill of the old one, then start
b.wait_synced(limit_s) # the watch line's synced=true (chain-side)
b.height(), b.peers(), b.exec_tip(), b.exec_status(), b.proving_status(), b.state_root(height), b.paid_segments()
b.start_miner(label, wallet, device=0, pack_dir=...) # one igneum-miner with the CUDA worker; returns the pid
b.start_prover(label, wallet, hub_peer, threshold="", miner="keep") # box-prover.sh
b.stop_all() # every stage process (anchored), never sshd
b.destroy() # the provider call plus the registry row
Process patterns are anchored on the executable's path or use -x (tools/ci/pgrep-self-match-check.sh).
"""
from .box import Box, Registry, SshError # noqa: F401

150
tools/fleet/lib/box.py Normal file
View file

@ -0,0 +1,150 @@
import json, os, subprocess, time, datetime, hashlib, shlex
ROOT = os.path.expanduser("~/Desktop/fleet"); REG = os.path.join(ROOT, "boxes.json")
KEY = os.path.expanduser("~/.ssh/igneum-fleet")
SSH_OPTS = ["-i", KEY, "-o", "StrictHostKeyChecking=no", "-o", "UserKnownHostsFile=/dev/null", "-o", "LogLevel=ERROR",
"-o", "ConnectTimeout=20", "-o", "ServerAliveInterval=10", "-o", "ServerAliveCountMax=3", "-o", "BatchMode=yes"]
B = "/opt/igneum/pkg/bin"; F = "/root/fleet"
NODE_PATH_PATTERN = "^/root/fleet/in/igneumd|^/opt/igneum/pkg/bin/igneumd"
def now(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
class SshError(RuntimeError): pass
def hexi(v):
if v is None: return 0
return int(v, 16) if isinstance(v, str) and v.startswith("0x") else int(v or 0)
class Registry:
"""~/Desktop/fleet/boxes.json: one row per rented instance; writes are atomic and locked."""
@staticmethod
def load():
if not os.path.exists(REG): return {}
for i in range(6):
try: return json.load(open(REG))
except json.JSONDecodeError: time.sleep(0.2 * (i + 1))
raise
@staticmethod
def patch(iid, **fields):
import fcntl
with open(REG + ".lock", "w") as lk:
fcntl.flock(lk, fcntl.LOCK_EX)
reg = Registry.load(); reg.setdefault(str(iid), {}).update(fields)
tmp = REG + f".tmp.{os.getpid()}"; json.dump(reg, open(tmp, "w"), indent=1); os.replace(tmp, REG)
return reg[str(iid)]
@staticmethod
def find(label):
for iid, b in Registry.load().items():
if b.get("label") == label and b.get("state") != "destroyed": return iid, b
raise KeyError(label)
class Box:
def __init__(self, host, port, label="box", iid=None, wallet=None, provider="vast", user="root"):
self.host, self.port, self.label, self.iid, self.wallet, self.provider, self.user = host, int(port), label, iid, wallet, provider, user
@classmethod
def from_registry(cls, label):
iid, b = Registry.find(label)
if not b.get("ssh_host") or not b.get("ssh_port"): raise SshError(f"{label}: no ssh endpoint in the registry")
return cls(b["ssh_host"], b["ssh_port"], label, iid, b.get("wallet"), b.get("provider", "vast"), b.get("ssh_user", "root"))
# ---- transport ----
def run(self, cmd, timeout=60):
try:
r = subprocess.run(["ssh"] + SSH_OPTS + ["-p", str(self.port), f"{self.user}@{self.host}", cmd], capture_output=True, text=True, timeout=timeout, stdin=subprocess.DEVNULL)
except subprocess.TimeoutExpired: return 124, "", "timeout"
return r.returncode, r.stdout or "", r.stderr or ""
def check(self, cmd, timeout=60):
rc, out, err = self.run(cmd, timeout)
if rc != 0: raise SshError(f"{self.label}: rc {rc}: {err.strip()[:200]}")
return out
def put(self, files, dest, tries=3, timeout=900):
# the deploy gate: a fleet script (.sh/.py under tools/fleet) goes to a box only when deploy-gate.sh reads clean
srcs = [f for f in files if (f.endswith(".sh") or f.endswith(".py")) and "/tools/fleet/" in os.path.abspath(f)]
if srcs:
g = subprocess.run(["bash", os.path.join(os.path.dirname(os.path.dirname(os.path.abspath(__file__))), "deploy-gate.sh")] + srcs, capture_output=True, text=True, timeout=120)
if g.returncode != 0: raise SshError(f"{self.label}: deploy gate RED, push refused: {(g.stdout + g.stderr).strip()[:300]}")
for i in range(tries):
r = subprocess.run(["scp"] + SSH_OPTS + ["-P", str(self.port)] + list(files) + [f"{self.user}@{self.host}:{dest}"], capture_output=True, text=True, timeout=timeout)
if r.returncode == 0: return True
time.sleep(5)
raise SshError(f"{self.label}: scp failed {tries} times: {r.stderr.strip()[:120]}")
def alive(self): return self.run("true", 30)[0] == 0
# ---- payload ----
def install_payload(self, src, sha256):
"""The node binary onto the box as /root/fleet/in/igneumd-<sha16>, verified there; src is a local file or an http(s) URL."""
local = src
if src.startswith("http"):
local = os.path.join(ROOT, f"payload-{sha256[:16]}"); subprocess.run(["curl", "-fsSL", "-o", local, src], check=True, timeout=600)
got = hashlib.sha256(open(local, "rb").read()).hexdigest()
if got != sha256: raise SshError(f"payload sha256 is {got[:16]}, not {sha256[:16]}")
name = f"igneumd-{sha256[:16]}"
self.check(f"mkdir -p {F}/in {F}/out"); self.put([local], f"{F}/in/{name}.new")
self.check(f"cd {F}/in && [ \"$(sha256sum {name}.new | cut -c1-64)\" = {sha256} ] && mv {name}.new {name} && chmod +x {name}")
return f"{F}/in/{name}"
# ---- node ----
def stop_node(self):
self.run(f"pkill -f '{NODE_PATH_PATTERN}'; sleep 3; pkill -9 -f '{NODE_PATH_PATTERN}' 2>/dev/null; true", 40)
def start_node(self, binary, override_local, peers, devnet_suffix=None, verifier=None, extra=(), appdir=None, log=None, unsynced_mining=False):
"""Writes the override file, kills the old node (anchored on its path), starts the new one; returns the digest it prints."""
appdir = appdir or (f"{F}/dn2" if devnet_suffix else f"{F}/node"); log = log or (f"{F}/dn2-node.log" if devnet_suffix else f"{F}/node.log")
self.put([override_local], f"{F}/override-live.json")
self.stop_node()
flags = ["--devnet"] + ([f"--devnet-suffix={devnet_suffix}"] if devnet_suffix else []) + [f"--appdir={appdir}", "--rpclisten=0.0.0.0:26610", "--evm-rpclisten=127.0.0.1:26790", "--listen=0.0.0.0:26611"]
flags += [f"--addpeer={p}" for p in peers] + [f"--override-params-file={F}/override-live.json", "--nodnsseed", "--disable-upnp", "--nologfiles", "--yes"] + (["--enable-unsynced-mining"] if unsynced_mining else []) + list(extra)
env = f"IGNEUM_PROOF_VERIFIER={shlex.quote(verifier)} " if verifier else ""
self.check(f"cd {F} && {env}setsid nohup {shlex.quote(binary)} {' '.join(shlex.quote(x) for x in flags)} </dev/null >> {log} 2>&1 & sleep 8; echo started", 60)
out = self.run(f"grep -o 'digest: [0-9a-f]*' {log} | tail -1 | awk '{{print $2}}'", 30)[1].strip()
return out
def watch(self):
out = self.run(f"{B}/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1", 40)[1].strip()
return dict(x.split("=", 1) for x in out.split() if "=" in x) if out else {}
def height(self): return int(self.watch().get("blocks", 0))
def daa(self): return int(self.watch().get("daa", 0))
def peers(self): return int(self.watch().get("peers", 0))
def synced(self): return self.watch().get("synced") == "true"
def wait_synced(self, limit_s=1200, poll=15):
t0 = time.time()
while time.time() - t0 < limit_s:
if self.synced(): return True
time.sleep(poll)
return False
# ---- exec and proving (chain-side facts) ----
def rpc(self, method, params=None, timeout=20):
body = json.dumps({"jsonrpc": "2.0", "id": 1, "method": method, "params": params or []})
out = self.run(f"curl -s -m {timeout} -X POST -H 'Content-Type: application/json' --data {shlex.quote(body)} http://127.0.0.1:26790/", timeout + 20)[1]
try: return json.loads(out).get("result")
except json.JSONDecodeError: return None
def exec_status(self): return self.rpc("igneum_getExecStatus") or {}
def exec_tip(self): return hexi(self.exec_status().get("executedTip"))
def proving_status(self): return self.rpc("igneum_getProvingStatus") or {}
def paid_segments(self): return int((self.proving_status().get("v1") or {}).get("paidSegments") or 0)
def paid_shards(self): return int(self.proving_status().get("paidShards") or 0)
def state_root(self, height):
b = self.rpc("eth_getBlockByNumber", [hex(height), False]) or {}
return b.get("stateRoot")
def node_version(self):
return self.run(f"v=$(pgrep -fa '{NODE_PATH_PATTERN}' | head -1 | awk '{{print $2}}'); [ -n \"$v\" ] && echo \"$(sha256sum \"$v\" | cut -c1-16) $($v --version 2>&1 | head -1)\"", 30)[1].strip()
def rejects_since(self, since_utc, log=None):
log = log or f"{F}/node.log"
out = self.run(f"awk -v s='{since_utc}' 'substr($0,1,19) >= s' {log} 2>/dev/null | grep -cE 'got reject message|PoW rejected|block rejected|invalid block'", 30)[1].strip()
return int(out or 0)
def max_reorg_since(self, since_utc, log=None):
log = log or f"{F}/node.log"
out = self.run(f"awk -v s='{since_utc}' 'substr($0,1,19) >= s' {log} 2>/dev/null | grep -oE 'selected-chain reorg: [0-9]+ chain blocks' | grep -oE '[0-9]+ chain' | awk '{{if ($1>m) m=$1}} END {{print m+0}}'", 30)[1].strip()
return int(out or 0)
# ---- miner and prover ----
def start_miner(self, label, wallet, device=0, grpc="grpc://127.0.0.1:26610", pack="devnet", log=None):
log = log or f"{F}/out/miner-{device}.log"
self.check(f"cd {F}/mine 2>/dev/null || mkdir -p {F}/mine/packs && cd {F}/mine; [ -d packs/{pack} ] || {B}/igneum-miner export-pack {grpc} packs/{pack} >/dev/null 2>&1; setsid nohup {B}/igneum-miner mine {grpc} 1 100000000 {shlex.quote(label)} --worker {B}/igneum-worker-cuda --worker-args '--device {device} --pack packs/{pack}' --prepare-packs packs/prepare-{device} --exit-on-seed-change --evm-address {wallet} --payout-label {shlex.quote(label)} --status-secs 30 </dev/null >> {log} 2>&1 & echo $!", 90)
return int(self.run("pgrep -x igneum-miner | tail -1", 20)[1].strip() or 0)
def stop_miners(self): self.run("pkill -x igneum-miner; pkill -f '^/opt/igneum/pkg/bin/igneum-worker-cuda'; true", 30)
def start_prover(self, label, wallet, hub_peer="", threshold="", miner="keep"):
self.check(f"cd {F} && chmod +x in/box-prover.sh && LABEL={shlex.quote(label)} WALLET={wallet} HUB_PEER={hub_peer} THRESHOLD={threshold} MINER={miner} setsid nohup in/box-prover.sh </dev/null >/dev/null 2>&1 & echo started", 60)
def stop_all(self):
self.run(f"bash {F}/in/box-kill.sh >/dev/null 2>&1; pkill -f '^python3 -u /root/fleet/in/box-prover.py'; pkill -x igneum-miner; pkill -f '^/opt/igneum/pkg/bin/igneum-worker-cuda'; pkill -x sp1-gpu-server; true", 60)
# ---- provider ----
def destroy(self):
import sys; sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
if self.provider == "runpod":
import runpod; runpod.destroy(self.iid)
else:
import vast; vast.destroy(self.iid)
b = Registry.load().get(str(self.iid), {}); t1 = now()
h = (datetime.datetime.strptime(t1, "%Y-%m-%dT%H:%M:%SZ") - datetime.datetime.strptime(b.get("rented_at", t1), "%Y-%m-%dT%H:%M:%SZ")).total_seconds() / 3600 if b.get("rented_at") else 0
Registry.patch(self.iid, state="destroyed", destroyed_at=t1, hours=round(h, 2), cost_usd=round(h * (b.get("dph") or 0), 2))
return round(h * (b.get("dph") or 0), 2)

180
tools/fleet/lib/standing.py Normal file
View file

@ -0,0 +1,180 @@
"""The standing fleet (the project lead's ruling, 6 October 2026, 19:50 UK): rented cards that stay up and are never destroyed on a
job's end. A box is standing when its registry row carries standing=true and a role: live (the live devnet: node, miner,
prover, voter) or dn2 (Devnet 2: seed, miner or prover). This module is the Mac side of the rule; box-standing.sh is the
box side (the supervisor loop). Every call goes through lib.box.
roster() the standing rows with role, card, price and uptime
install(label) puts box-standing.sh (and the recovery recipe) on the box and starts the supervisor once
check() one pass: ssh alive? supervisor up? node synced? exec moving? miner and prover present?
version against the live manifest (dl.igneum.network/dl/public/igneum-downloads.json,
files.miner-hive.version); returns one dict per box and writes ~/Desktop/fleet/standing.jsonl
update(label, pkg) the version follow: the box downloads the manifest's hive package (sha256 checked), unpacks it
to /opt/igneum/pkg.new, swaps it in and writes /root/fleet/standing.node; the supervisor
restarts the node on it (data dir kept, seconds)
rerent(label) the host died (ssh dead for two checks, or the provider says the instance is gone): rent the
same shape (card, VRAM, provider) with the label suffixed "-r<n>", set it up, install the
supervisor, mark the old row destroyed; the Devnet 2 roles take their seed from the registry
loop(every=600) check, then rerent what died and update what is behind, every ten minutes
weight_check(labels) the 10 percent rule (20:00Z): the labels' share of the live devnet's weight (blue blocks per
key over the window, from the hub) plus what was removed in the last hour must stay under
10 percent, or the job is refused; every job that stops or shares a standing miner calls it
Nothing here destroys a standing box; destroy() in lib.box stays for the one-shot boxes.
"""
import os, sys, json, time, datetime, subprocess
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from lib.box import Box, Registry, F, SshError
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
ROOT = os.path.expanduser("~/Desktop/fleet"); MANIFEST = "https://dl.igneum.network/dl/public/igneum-downloads.json"
def now(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def roster():
reg = Registry.load(); rows = []
for iid, b in reg.items():
if not b.get("standing") or b.get("state") == "destroyed": continue
since = b.get("standing_since") or b.get("rented_at") or now()
up_h = (datetime.datetime.now(datetime.timezone.utc) - datetime.datetime.fromisoformat(since.replace("Z", "+00:00"))).total_seconds() / 3600
rows.append({"iid": iid, "label": b["label"], "role": b.get("role", "live"), "card": b.get("card"), "provider": b.get("provider"), "dph": round(float(b.get("dph") or 0), 3), "uptime_h": round(up_h, 1), "since": since})
return sorted(rows, key=lambda r: (r["role"], r["label"]))
def manifest():
try: return json.loads(subprocess.run(["curl", "-fsSL", "-m", "20", MANIFEST], capture_output=True, text=True, timeout=30).stdout)["files"]["miner-hive"]
except Exception: return {}
def install(label, role=None, prover=None, mine=None):
"""Mark the row standing and start the supervisor on the box; idempotent (a running supervisor is left alone)."""
iid, b = Registry.find(label); role = role or b.get("role", "live"); b_ = Box.from_registry(label)
prover = "1" if (prover if prover is not None else role == "live" or b.get("dn2_prover")) else "0"
mine = "1" if (mine if mine is not None else not b.get("dn2_seed")) else "0"
b_.put([os.path.join(HERE, "box-standing.sh"), os.path.join(HERE, "box-exec-snapshot.sh")], f"{F}/in/")
hub = next((v for v in Registry.load().values() if v.get("hub") and v.get("state") != "destroyed"), {})
seed = b.get("dn2_seed_addr", "")
rc, out, err = b_.run(f"cd {F} && chmod +x in/box-standing.sh in/box-exec-snapshot.sh; pgrep -f '^bash in/box-standing.sh' >/dev/null && echo supervisor-already || {{ ROLE={role} LABEL={label} WALLET={b.get('wallet')} PROVER={prover} MINE={mine} HUB_PEER={hub.get('hub_peer','')} SEED={seed} setsid nohup bash in/box-standing.sh </dev/null >/dev/null 2>&1 & sleep 2; echo supervisor-started; }}", 60)
Registry.patch(iid, standing=True, role=role, standing_since=b.get("standing_since") or now(), doing=f"standing {role}: node, {'miner, ' if mine == '1' else ''}{'prover' if prover == '1' else 'no prover'} (supervised)")
return out.strip()
DIGEST_ALERTS = os.path.join(ROOT, "digest-alerts.json")
def digest_alert(label, got, expected):
"""One #incidents line per box per hour when a standing box's handshake digest is not the live one (the seed's 7bd98cc4 at 15:3xZ froze a home miner's tip for 53 minutes)."""
st = json.load(open(DIGEST_ALERTS)) if os.path.exists(DIGEST_ALERTS) else {}
if time.time() - st.get(label, 0) < 3600: return
what = f"Fleet box {label} answers handshakes with consensus digest {got[:8]}, the live object is {expected[:8]}: peers on the live object refuse it and a node whose only peer it is sits on a frozen tip"
subprocess.run(["node", "/Users/joshm/Projects/igneum/tools/community/discord-hooks.mjs", "incident", "open", "--what", what, "--affected", "nodes that peer only with that box (a home miner reads synced on a frozen tip)", "--doing", "the fleet agent restarts the box's node on the live object; the supervisor's five-minute check found it", "--id", f"digest-{label}-{int(time.time())}", "--live"], capture_output=True, text=True, timeout=60)
st[label] = time.time(); json.dump(st, open(DIGEST_ALERTS, "w")); print(now(), "digest alert", label, got[:8], "expected", expected[:8], flush=True)
def check_one(r):
b = Box.from_registry(r["label"]); row = dict(r); row["t"] = now()
rc, out, err = b.run("echo alive; pgrep -fc '^bash in/box-standing.sh'; tail -1 /root/fleet/out/standing.log 2>/dev/null | cut -c1-400; echo DIGEST=$(grep -o 'digest: [0-9a-f]*' /root/fleet/node.log | tail -1 | awk '{print substr($2,1,16)}')", 45)
if rc != 0 or "alive" not in out: row.update(alive=False); return row
lines = [l for l in out.strip().split("\n") if not l.startswith("DIGEST=")]; row["digest"] = next((l.split("=", 1)[1] for l in out.strip().split("\n") if l.startswith("DIGEST=")), "")
row.update(alive=True, supervisor=int(lines[1] or 0) if len(lines) > 1 else 0, last=lines[2] if len(lines) > 2 else "")
kv = dict(x.split("=", 1) for x in row["last"].split() if "=" in x and not x.startswith("miner="))
row.update(blocks=int(kv.get("blocks", 0) or 0), synced=kv.get("synced") == "true", exec_tip=int(kv.get("exec", 0) or 0), version=kv.get("version", ""), node_pid=kv.get("node", ""), prover=int(kv.get("prover", 0) or 0), bin=kv.get("bin", ""))
row["bin_sha16"] = os.path.basename(row["bin"]).replace("igneumd-", "") if "igneumd-" in row["bin"] and len(os.path.basename(row["bin"])) == 24 else ""
return row
def check(write=True):
from concurrent.futures import ThreadPoolExecutor
rows = roster(); m = manifest(); want = m.get("version", "")
with ThreadPoolExecutor(16) as ex: res = list(ex.map(check_one, rows))
reg = Registry.load(); expected = open(os.path.join(ROOT, "expected-digest")).read().strip()[:16] if os.path.exists(os.path.join(ROOT, "expected-digest")) else ""
for r in res:
r["expected_digest"] = expected; r["digest_ok"] = (not expected) or (not r.get("digest")) or r["digest"][:16] == expected
if expected and r.get("digest") and not r["digest_ok"]: digest_alert(r["label"], r["digest"], expected)
r["manifest_version"] = want; wanted = (reg.get(r["iid"]) or {}).get("node_sha16_wanted")
r["behind"] = bool(wanted and r.get("bin_sha16") and r["bin_sha16"] != wanted) # only an explicit publish sets the want; the package version string is not a node version
if write:
with open(os.path.join(ROOT, "standing.jsonl"), "a") as f:
for r in res: f.write(json.dumps(r) + "\n")
return res
def update(label):
"""The version follow: the manifest's hive package onto the box, the node pointer rewritten, the supervisor restarts it."""
m = manifest(); b = Box.from_registry(label)
if not m: raise SshError("no manifest")
rc, out, err = b.run(f"set -e; cd {F}; curl -fsSL -m 600 -o pkg-{m['version']}.tgz https://dl.igneum.network{m['path']}; echo '{m['sha256']} pkg-{m['version']}.tgz' | sha256sum -c - >/dev/null; rm -rf /opt/igneum/pkg.new; mkdir -p /opt/igneum/pkg.new; tar -C /opt/igneum/pkg.new --strip-components=1 -xzf pkg-{m['version']}.tgz; rm -rf /opt/igneum/pkg.prev; mv /opt/igneum/pkg /opt/igneum/pkg.prev; mv /opt/igneum/pkg.new /opt/igneum/pkg; echo /opt/igneum/pkg/bin/igneumd > {F}/standing.node; pkill -9 -f '^/opt/igneum/pkg/bin/igneumd' 2>/dev/null; pkill -9 -f '^/opt/igneum/pkg.prev/bin/igneumd' 2>/dev/null; echo updated-to-{m['version']}", 900)
iid, _ = Registry.find(label); Registry.patch(iid, version_wanted=m["version"], updated_at=now()); return out.strip()
def rerent(label):
"""Same shape again on the same provider; the old row is marked destroyed. Returns the new label or None."""
import fleet
iid, b = Registry.find(label); n = int(b.get("rerents", 0)) + 1; new = f"{label.split('-r')[0]}-r{n}"
if b.get("provider") == "runpod":
import runpod; d = runpod.rent(b.get("gpu_type_id") or f"NVIDIA GeForce {b.get('card')}", new, count=1, disk=40)
if not d.get("id"): return None
reg = fleet.load(); reg[str(d["id"])] = {**{k: b[k] for k in ("card", "num_gpus", "provider", "phase", "wallet", "role", "standing") if k in b}, "label": new, "dph": d.get("costPerHr") or b.get("dph"), "state": "renting", "rented_at": now(), "standing_since": now(), "rerents": n, "rerent_of": label}; fleet.save(reg)
else:
nid = fleet.rent_card(b.get("card"), new, b.get("archs") or [], min_ram=b.get("vram_mb"), phase=b.get("phase", "2"))
if not nid: return None
Registry.patch(str(nid), standing=True, role=b.get("role", "live"), standing_since=now(), rerents=n, rerent_of=label)
Registry.patch(iid, state="destroyed", destroyed_at=now(), destroyed_reason="host died, re-rented as " + new)
return new
# ---- the 10 percent rule (20:00Z): never remove more than 10 percent of the live devnet's 30-day weight in any hour ----
WEIGHT_LOG = os.path.join(ROOT, "weight-removals.jsonl")
def weight_shares(blocks=2000):
"""Vote-key-hash shares over the last `blocks` blocks of the live devnet, from the hub's node (igneum-miner inspect:
one line per block with voteKeyHash=<64 hex>); the hub is pruned, so the 30-day window is approximated by the largest
read the RPC serves (2,000 blocks is about 33 minutes at 1 block/s; pass more when the RPC allows). Returns
({key_hash: share}, blocks_read)."""
hub = next((v for v in Registry.load().values() if v.get("hub") and v.get("state") != "destroyed"), None)
if not hub: raise SshError("no hub box")
b = Box(hub["ssh_host"], hub["ssh_port"], "hub-1", None, None, hub.get("provider"))
rc, out, err = b.run(f"timeout 170 /opt/igneum/pkg/bin/igneum-miner inspect {blocks} grpc://127.0.0.1:26610 2>&1 | grep -oE 'voteKeyHash=[0-9a-f]{{64}}'", 190)
keys = [ln.split("=", 1)[1] for ln in out.split() if ln.startswith("voteKeyHash=")]
n = len(keys) or 1; rows = {}
for k in keys: rows[k] = rows.get(k, 0) + 1
return {k: v / n for k, v in rows.items()}, len(keys)
def vote_key_hash(label):
"""The box's vote key hash: its miner votes with the key its payout label derives (igneum-miner key-hash <label>)."""
b = Box.from_registry(label); rc, out, err = b.run(f"/opt/igneum/pkg/bin/igneum-miner key-hash {label} 2>/dev/null | grep -oE '[0-9a-f]{{64}}' | head -1", 30)
return out.strip()
def weight_check(labels, limit=0.10, hours=1.0):
"""Refuses when the labels' weight plus what was removed in the last `hours` exceeds `limit`. Records an allowed removal."""
shares, nblocks = weight_shares()
keys = {l: vote_key_hash(l) for l in labels}
want = sum(shares.get(k, 0) for k in keys.values() if k)
since = time.time() - hours * 3600; recent = 0.0
if os.path.exists(WEIGHT_LOG):
for ln in open(WEIGHT_LOG):
try: j = json.loads(ln)
except Exception: continue
if j.get("ts", 0) >= since: recent += float(j.get("share", 0))
ok = (want + recent) <= limit
verdict = {"labels": labels, "keys": {l: k[:16] for l, k in keys.items()}, "share": round(want, 4), "removed_last_hour": round(recent, 4), "limit": limit, "ok": ok, "blocks_read": nblocks, "t": now()}
if ok:
with open(WEIGHT_LOG, "a") as f: f.write(json.dumps({"ts": time.time(), "labels": labels, "share": want}) + "\n")
return verdict
def table_signed_pct():
"""The signed share of the frozen voter table at the hub's last lock (the "% of total" of the LOCKED line), or None."""
hub = next((v for v in Registry.load().values() if v.get("hub") and v.get("state") != "destroyed"), None)
if not hub: return None
out = Box(hub["ssh_host"], hub["ssh_port"], "hub-1", None, None, hub.get("provider")).run("grep -E 'Finality: checkpoint [0-9]+ LOCKED:' /root/fleet/node.log | tail -1 | grep -oE '[0-9.]+% of total' | head -1", 30)[1].strip()
try: return float(out.split("%")[0])
except Exception: return None
def table_gate(min_pct=75.0):
"""The coordinator's rule (6 October 2026, 23:0xZ): nothing of the fleet's moves on a standing box, pool slices included, until the
table reads over min_pct signed. Returns (ok, pct)."""
pct = table_signed_pct(); return (pct is not None and pct > min_pct), pct
def loop(every=600):
dead = {}
while True:
try:
res = check()
for r in res:
if not r["alive"]:
dead[r["label"]] = dead.get(r["label"], 0) + 1
if dead[r["label"]] >= 2: print(now(), "re-renting", r["label"], "->", rerent(r["label"]), flush=True); dead.pop(r["label"], None)
else:
dead.pop(r["label"], None)
if r["behind"]: print(now(), "behind:", r["label"], r.get("bin_sha16"), "wanted", (Registry.load().get(r["iid"]) or {}).get("node_sha16_wanted"), "(a publish script moves it; the loop only reports)", flush=True)
print(now(), "standing check:", len(res), "boxes,", sum(1 for r in res if r["alive"]), "alive,", sum(1 for r in res if r.get("synced")), "synced,", sum(1 for r in res if r["behind"]), "behind,", sum(1 for r in res if not r.get("digest_ok", True)), "off the live digest", flush=True)
except Exception as e: print(now(), "loop error", str(e)[:200], flush=True)
time.sleep(every)
if __name__ == "__main__":
a = sys.argv[1:]
if not a or a[0] == "roster":
for r in roster(): print(f"{r['label']:<12} {r['role']:<5} {str(r['card']):<22} {r['provider']:<7} USD {r['dph']:.3f}/h up {r['uptime_h']:.1f} h")
rs = roster(); print(f"{len(rs)} standing boxes, USD {sum(r['dph'] for r in rs):.2f}/h, USD {24*sum(r['dph'] for r in rs):.0f}/day")
elif a[0] == "install": print(install(a[1], *(a[2:3] or [None])))
elif a[0] == "check":
for r in check(): print({k: r.get(k) for k in ("label", "role", "alive", "supervisor", "blocks", "synced", "exec_tip", "version", "behind", "prover")})
elif a[0] == "update": print(update(a[1]))
elif a[0] == "rerent": print(rerent(a[1]))
elif a[0] == "loop": loop(int(a[1]) if len(a) > 1 else 600)
elif a[0] == "table_gate":
ok, pct = table_gate(float(a[1]) if len(a) > 1 else 75.0); print(json.dumps({"ok": ok, "signed_pct_of_table": pct, "min": float(a[1]) if len(a) > 1 else 75.0})); sys.exit(0 if ok else 1)
elif a[0] == "weight_check":
v = weight_check(a[1:]); print(json.dumps(v)); sys.exit(0 if v["ok"] else 1)
elif a[0] == "weights":
sh, n = weight_shares(); print("blocks read", n, "voters", len(sh)); print(json.dumps({k[:16]: round(v, 4) for k, v in sorted(sh.items(), key=lambda kv: -kv[1])[:40]}, indent=1))

View file

@ -0,0 +1,39 @@
"""Exercises the library against a live Devnet 2 box (the chain the fleet may touch freely): transport, the chain-side
readers, a miner start and stop, a node restart on the current binary and object with the digest read back.
python3 -m lib.test_box dn2-3 (run from tools/fleet)
The known-finished case is a PASS line; the known-failed case is a label that does not exist (a SshError).
"""
import sys, os, time
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from lib import Box, Registry, SshError
def main(label):
try: Box.from_registry("no-such-box-" + label)
except (KeyError, SshError) as e: print("ok known-failed: a missing label raises", type(e).__name__)
else: print("FAIL known-failed case did not raise"); return 1
b = Box.from_registry(label); fails = 0
def chk(name, cond, detail=""):
nonlocal fails
print(("ok " if cond else "FAIL") + f" {name} {detail}"); fails += 0 if cond else 1
chk("alive", b.alive())
w = b.watch(); chk("watch has blocks and daa", "blocks" in w and "daa" in w, str(w)[:100])
h0 = b.height(); chk("height > 0", h0 > 0, h0)
chk("peers >= 1", b.peers() >= 1, b.peers())
chk("exec tip within 60 of the height", abs(b.exec_tip() - h0) <= 60, f"exec {b.exec_tip()} height {h0}")
chk("state root at height-20 reads", bool(b.state_root(max(1, h0 - 20))))
chk("proving status answers", "tipDaa" in b.proving_status())
v = b.node_version(); chk("node version reads", "igneumd" in v, v)
chk("rejects since 1970 is an int", isinstance(b.rejects_since("1970-01-01 00:00:00", f"/root/fleet/dn2-node.log"), int))
b.stop_miners(); pid = b.start_miner(label + "-libtest", b.wallet, pack="dn2", log="/root/fleet/out/libtest-miner.log"); time.sleep(20)
chk("miner started (pid)", pid > 0, pid)
rc, out, _ = b.run("pgrep -c -x igneum-miner", 20); chk("one miner process", out.strip() == "1", out.strip())
b.stop_miners(); rc, out, _ = b.run("pgrep -c -x igneum-miner", 20); chk("miner stopped", out.strip() == "0", out.strip())
seed_peer = next((x.get("dn2_peer") for x in Registry.load().values() if x.get("dn2_seed") and x.get("state") != "destroyed"), "")
binary = b.run("ls /root/fleet/in/igneumd-0313 /root/fleet/in/igneumd-gate 2>/dev/null | head -1", 20)[1].strip()
override = os.path.join(os.path.dirname(os.path.dirname(os.path.abspath(__file__))), "devnet2-override.json")
d = b.start_node(binary, override, [seed_peer] if seed_peer else [], devnet_suffix=2, verifier="/opt/igneum-floor/bin/igneum-prove-host", unsynced_mining=True)
chk("node restarted, digest read", d.startswith("4a0b8726"), d[:16])
chk("height resumes within 60 s", (time.sleep(45) or b.height()) >= h0, b.height())
b.start_miner(label, b.wallet, pack="dn2", log="/root/fleet/out/dn2-miner.log") # the box back as it was
print(("PASS" if fails == 0 else "FAIL") + f" lib.test_box on {label}: {fails} failures")
return 1 if fails else 0
if __name__ == "__main__": sys.exit(main(sys.argv[1] if len(sys.argv) > 1 else "dn2-3"))

67
tools/fleet/night.py Normal file
View file

@ -0,0 +1,67 @@
#!/usr/bin/env python3
"""The fleet night loop (phase 2): every 60 s reads each box's swap log and prover state; launches the prover on a box
whose replay reached the tip (once); re-runs the exec restart on a box whose executed tip fell back to 0 for more than
two minutes (the deep-reorg class of 15:43Z, 6 October 2026); every 15 minutes writes a night row (provers up, claims,
submitted, paid segments, shards paid, the chain's segment window from the hub) to ~/Desktop/fleet/night.json and the
page, and appends the hourly row to the plan's fleet night section when the hour turns. Log: ~/Desktop/fleet/night.log."""
import sys, os, time, json, datetime, subprocess
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))); import fleet
from concurrent.futures import ThreadPoolExecutor
SHA = "d6350586fe837b1f71696546628ec071b6accce2531aef776cf2b1e5487a8cdc"; DIG = "b18ed271f75dd46406d230f4156c37472127415a4c32c558bac662f6f840e61c"
LOG = os.path.join(fleet.ROOT, "night.log"); NIGHT = os.path.join(fleet.ROOT, "night.json")
def now(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def log(s): open(LOG, "a").write(f"{now()} {s}\n"); print(now(), s, flush=True)
PROBE = r"""grep -E '^RESULT (swap_done|swap_failed)' /root/fleet/out/node-swap.log 2>/dev/null | grep -v step=binary | tail -1 | cut -c1-60
curl -s -m 6 -X POST -H 'Content-Type: application/json' --data '{"jsonrpc":"2.0","id":1,"method":"igneum_getExecStatus","params":[]}' http://127.0.0.1:26790/ | python3 -c 'import sys,json; r=sys.stdin.read(); d=json.loads(r).get("result",{}) if r.strip() else {}; print("exec", int(d.get("executedTip","0x0"),16), "blocked", "yes" if d.get("blocked") else "no", "from", str(d.get("startedFrom",""))[:30])' 2>/dev/null
curl -s -m 6 -X POST -H 'Content-Type: application/json' --data '{"jsonrpc":"2.0","id":1,"method":"igneum_getProvingStatus","params":[]}' http://127.0.0.1:26790/ | python3 -c 'import sys,json; r=sys.stdin.read(); d=json.loads(r).get("result",{}) if r.strip() else {}; v=d.get("v1",{}); w=v.get("segmentsInWindow",{}); print("tipDaa", int(d.get("tipDaa","0x0"),16), "fresh", v.get("freshRuleActive"), "paidShards", d.get("paidShards"), "pending", w.get("pending"), "proven", w.get("proven"), "unproven", w.get("unproven"), "paidSeg", v.get("paidSegments"))' 2>/dev/null
/opt/igneum/pkg/bin/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'daa=[0-9]*' | tail -1
pgrep -c -f '^python3 -u /root/fleet/in/box-prover.py'
for f in /root/fleet/out/prover-state.json /root/fleet/card*/out/prover-state.json; do [ -f $f ] && python3 -c 'import json,sys; s=json.load(open(sys.argv[1])); print("ps", s.get("claimed",0), s.get("submitted",0), s.get("paid",0), s.get("shards_accepted",0), s.get("shards_refused",0), s.get("segment_refused",0), s.get("held",0), s.get("miner_mhs",0))' $f; done; true"""
def probe(b):
r, out, err = fleet.ssh(b, PROBE, timeout=45)
if r != 0 or not out.strip(): return None
l = out.strip().split("\n"); d = {"swap": "", "exec": None, "blocked": "", "from": "", "tipDaa": None, "fresh": "", "paidShards": None, "pending": None, "proven": None, "unproven": None, "paidSeg": None, "daa": None, "provers": 0, "ps": []}
for x in l:
if x.startswith("RESULT swap"): d["swap"] = x
elif x.startswith("exec "): p = x.split(); d["exec"] = int(p[1]); d["blocked"] = p[3]; d["from"] = " ".join(p[5:])
elif x.startswith("tipDaa"): p = x.split(); d.update(tipDaa=int(p[1]), fresh=p[3], paidShards=p[5], pending=p[7], proven=p[9], unproven=p[11], paidSeg=p[13])
elif x.startswith("daa="): d["daa"] = int(x[4:])
elif x.isdigit(): d["provers"] = int(x)
elif x.startswith("ps "): d["ps"].append([float(v) for v in x.split()[1:]])
return d
zero_since = {}; launched = set(); last_row = 0; last_hour = None
hub = next(b for b in fleet.load().values() if b.get("hub") and b.get("state") != "destroyed" and b.get("hub_peer"))
while True:
reg = fleet.load()
live = [(iid, b) for iid, b in reg.items() if b.get("state") != "destroyed" and b.get("ssh_host") and (b.get("phase") in ("1", "2") or b.get("hub"))]
with ThreadPoolExecutor(14) as ex: res = dict(zip([i for i, _ in live], ex.map(lambda x: probe(x[1]), live)))
tot = {"provers": 0, "claimed": 0, "submitted": 0, "paid": 0, "shards": 0, "refused": 0, "segref": 0, "held": 0, "exec_ok": 0, "boxes": 0}
hubd = None
for iid, b in live:
d = res.get(iid)
if not d: continue
tot["boxes"] += 1
if b.get("hub"): hubd = d
at_tip = d["exec"] is not None and d["tipDaa"] and d["daa"] and d["tipDaa"] >= d["daa"] - 30 and d["exec"] > 0
if at_tip: tot["exec_ok"] += 1; zero_since.pop(iid, None)
elif d["exec"] == 0 and "swap_done" in d["swap"]:
zero_since.setdefault(iid, time.time())
if time.time() - zero_since[iid] > 150:
fleet.ssh(b, f"cd /root/fleet && bash in/box-kill.sh >/dev/null 2>&1; mv out/node-swap.log out/node-swap-reset-$(date +%s).log 2>/dev/null; STEP=file NODE_SHA256={SHA} EXPECT_DIGEST={DIG} HUB_PEER={hub['hub_peer']} HUB_SSH={hub['ssh_host']} HUB_PORT={hub['ssh_port']} setsid nohup in/box-node-swap.sh </dev/null >/dev/null 2>&1 &", timeout=60)
log(f"{b['label']}: exec at 0 for 150 s after a completed replay; exec restart re-run"); zero_since.pop(iid, None); launched.discard(iid)
if at_tip and iid not in launched and d["provers"] == 0 and b.get("stage") != "mine-only":
if (b.get("num_gpus") or 1) > 1: fleet.ssh(b, f"cd /root/fleet && LABEL={b['label']} WALLET={b['wallet']} setsid nohup in/box-rig-prover.sh </dev/null > out/rig-provers.log 2>&1 &", timeout=60)
else: fleet.ssh(b, f"cd /root/fleet && mv out/prover.log out/prover-prev-$(date +%s).log 2>/dev/null; LABEL={b['label']} WALLET={b['wallet']} HUB_PEER={hub['hub_peer']} setsid nohup in/box-prover.sh </dev/null >/dev/null 2>&1 &", timeout=60)
launched.add(iid); fleet.patch(iid, stage="prover", state="running", doing=f"phase 2 prover on the executed tip ({d['exec']})"); log(f"{b['label']}: replay at the tip (exec {d['exec']}), prover started")
tot["provers"] += d["provers"]
for p in d["ps"]: tot["claimed"] += p[0]; tot["submitted"] += p[1]; tot["paid"] += p[2]; tot["shards"] += p[3]; tot["refused"] += p[4]; tot["segref"] += p[5]; tot["held"] += p[6]
fleet.patch(iid, last_line=f"exec {d['exec']} tip {d['tipDaa']} provers {d['provers']} " + (" ".join(f"{int(p[0])}/{int(p[1])}/{int(p[2])}" for p in d["ps"])))
if time.time() - last_row > 900:
row = {"hour": now()[:13], "at": now(), "provers": tot["provers"], "boxes_executing": tot["exec_ok"], "boxes": tot["boxes"], "claimed": tot["claimed"], "segments_submitted": tot["submitted"], "segments_done": tot["paid"], "shards_paid": tot["shards"], "shards_refused": tot["refused"], "segment_records_refused": tot["segref"], "held": tot["held"],
"chain": ({"tipDaa": hubd["tipDaa"], "fresh": hubd["fresh"], "pending": hubd["pending"], "proven": hubd["proven"], "unproven": hubd["unproven"], "paidSeg": hubd["paidSeg"], "paidShards": hubd["paidShards"]} if hubd else {}),
"coverage_pct": (round(100 * int(hubd["proven"]) / max(1, int(hubd["proven"]) + int(hubd["pending"]) + int(hubd["unproven"])), 1) if hubd and hubd["proven"] not in (None, "None") else None),
"note": f"{tot['exec_ok']} of {tot['boxes']} boxes executing"}
rows = json.load(open(NIGHT)) if os.path.exists(NIGHT) else []; rows.append(row); json.dump(rows, open(NIGHT, "w"), indent=1)
log("row " + json.dumps(row)); last_row = time.time()
subprocess.run([sys.executable, os.path.join(fleet.HERE, "page.py"), "publish"], capture_output=True)
time.sleep(60)

View file

@ -0,0 +1 @@
{"difficulty_v2_activation_daa":33000,"proving_v0_activation_daa":84100,"fees_v1_activation_daa":210000,"finality_v3_activation_daa":135200,"program_class_v3_activation_daa":154800,"proving_v1_activation_daa":154800,"proving_v1_segment_blocks":8,"proving_v1_unproven_daa":600,"proving_v1_aggregator_share_bps":1000,"proving_v1_fresh_rule_daa":198000}

74
tools/fleet/page.py Executable file
View file

@ -0,0 +1,74 @@
#!/usr/bin/env python3
"""Builds ~/Desktop/fleet/fleet.json (the fleet page's single source) from the registry (boxes.json), the ledger, the
result rows (results.json: {card_key: {...}}), the night rows (night.json: [..]), the phases (phases.json) and the log
(log.jsonl), then publishes it with the master checkout's tools/fleet/publish-fleet.sh (at most once a minute).
page.py log "<text>" append a log line (and rebuild, no publish)
page.py phase p1 state=running done=3 note="..."
page.py build rebuild only
page.py publish rebuild and publish (skipped when the last publish was under 60 s ago unless --force)
"""
import time, json, os, sys, time, datetime, subprocess
ROOT = os.path.expanduser("~/Desktop/fleet"); OUT = os.path.join(ROOT, "fleet.json")
PUBLISH = "/Users/joshm/Projects/igneum/tools/fleet/publish-fleet.sh"
def now(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def jl(name, default):
p = os.path.join(ROOT, name)
if not os.path.exists(p): return default
if name.endswith(".jsonl"): return [json.loads(l) for l in open(p) if l.strip()]
return json.load(open(p))
def hours(a, b=None):
t0 = datetime.datetime.strptime(a, "%Y-%m-%dT%H:%M:%SZ")
t1 = datetime.datetime.strptime(b, "%Y-%m-%dT%H:%M:%SZ") if b else datetime.datetime.utcnow()
return max((t1 - t0).total_seconds(), 0) / 3600
def build():
reg = jl("boxes.json", {}); res = jl("results.json", {}); night = jl("night.json", []); log = jl("log.jsonl", [])
phases = jl("phases.json", [
{"order": 1, "name": "Memory matrix on real cards", "state": "planned", "planned": 11, "done": 0, "note": ""},
{"order": 2, "name": "Prover fleet on the devnet", "state": "planned", "planned": 25, "done": 0, "note": ""},
{"order": 3, "name": "The 8x rigs", "state": "planned", "planned": 2, "done": 0, "note": ""},
{"order": 4, "name": "AMD mining", "state": "planned", "planned": 3, "done": 0, "note": ""}])
dj = os.path.join(ROOT, "disk.json")
if not os.path.exists(dj) or time.time() - os.path.getmtime(dj) > 300:
try:
import importlib.util; spec = importlib.util.spec_from_file_location("disk_sweep", os.path.join(os.path.dirname(os.path.abspath(__file__)), "disk-sweep.py")); m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m); m.sweep(alert=True)
except Exception as e: print("disk sweep failed:", str(e)[:120])
disk = jl("disk.json", {})
boxes = []; vast = 0.0; runpod = 0.0
for iid, b in reg.items():
h = b.get("hours") if b.get("state") == "destroyed" else hours(b["rented_at"])
cost = round(h * b["dph"], 3)
if b.get("provider", "vast") == "vast": vast += cost
else: runpod += cost
boxes.append({"id": iid, "label": b["label"], "card": b["card"], "vram_gb": round((b.get("vram_mb") or 0) / 1024), "provider": b.get("provider", "vast"),
"rate_usd_h": b["dph"], "phase": b.get("phase", "1"), "state": "done" if b.get("state") == "destroyed" else b.get("state", "renting"),
"disk_pct": (disk.get(b["label"]) or {}).get("pct"), "disk_free_gb": (disk.get(b["label"]) or {}).get("avail_gb"), "disk_red": bool(((disk.get(b["label"]) or {}).get("pct") or 0) >= 85),
"standing": bool(b.get("standing")), "role": b.get("role", ""), "standing_since": b.get("standing_since"), "uptime_h": round(hours(b.get("standing_since") or b["rented_at"]), 1) if b.get("standing") and b.get("state") != "destroyed" else None,
"started_at": b["rented_at"], "ended_at": b.get("destroyed_at"), "hours": round(h, 2), "cost_usd": cost, "doing": (("done, destroyed: " + (b.get("doing") or "its measurement is in")) if b.get("state") == "destroyed" else b.get("doing", ""))})
running = [b for b in boxes if b["state"] not in ("done", "failed")]
stand = [b for b in running if b["standing"]]
standing_block = {"count": len(stand), "usd_h": round(sum(b["rate_usd_h"] for b in stand), 2), "usd_day": round(24 * sum(b["rate_usd_h"] for b in stand)), "roles": {r: sum(1 for b in stand if b["role"] == r) for r in sorted(set(b["role"] for b in stand))},
"rule": "Standing boxes (the project lead, 6 October 2026) stay up and are never destroyed on a job's end: 16 on the live devnet (mining, proving, voting) and 6 on Devnet 2, about USD 8/h; a dead host is re-rented in the same shape, the node runs under a supervisor with the recovery recipe, the version follows the live manifest by the update job, and no standing box joins an experiment (experiments use wave boxes); each should carry a mapped p2p port where the provider allows, so the live devnet stops being a star."}
page = {"spend": {"vast_usd": round(vast, 2), "runpod_usd": round(runpod, 2), "cap_vast": 1000, "cap_runpod": 500, "updated_at": now(), "running": len(running), "running_note": open(os.path.join(ROOT, "why.txt")).read().strip() if os.path.exists(os.path.join(ROOT, "why.txt")) else ""},
"phases": phases, "standing": standing_block, "boxes": boxes, "results": list(res.values()), "night": night, "log": log[-200:],
"devnet2": jl("devnet2-status.json", {"state": "building", "boxes": 0, "height": None, "last_gate": "none yet", "last_gate_at": None})}
json.dump(page, open(OUT, "w"), indent=1); return page
def publish(force=False):
stamp = os.path.join(ROOT, ".last-publish")
if not force and os.path.exists(stamp) and time.time() - os.path.getmtime(stamp) < 60: print("published under a minute ago; skipped"); return
build(); open(stamp, "w").write(now())
r = subprocess.run([PUBLISH, OUT], capture_output=True, text=True, timeout=300); print((r.stdout + r.stderr).strip()[-300:])
if __name__ == "__main__":
c = sys.argv[1] if len(sys.argv) > 1 else "build"
if c == "log":
with open(os.path.join(ROOT, "log.jsonl"), "a") as f: f.write(json.dumps({"at": now(), "text": sys.argv[2]}) + "\n")
build()
elif c == "phase":
ph = jl("phases.json", None) or build()["phases"]
for p in ph:
if f"p{p['order']}" == sys.argv[2]:
for kv in sys.argv[3:]:
k, v = kv.split("=", 1); p[k] = int(v) if v.isdigit() else v
json.dump(ph, open(os.path.join(ROOT, "phases.json"), "w")); build()
elif c == "publish": publish("--force" in sys.argv)
else: build(); print(OUT)

View file

@ -0,0 +1,39 @@
#!/usr/bin/env python3
"""Publish 1 of 0.3.15 on the standing live-devnet boxes (the shipper's set, 6 October 2026, 21:4xZ): the Linux igneumd 7f0bde70
(release-0.3.15-node f1ea7a38; the first two cuts 713ef876 and 7961c5f1 failed the canary), igneum-miner e827d9db, the generator-4 workers (cuda 97e036e2, opencl ef9fbd03); the override
file stays the live thirteen-field object (digest b18ed271 on the new binary). Per box, through lib.box: the node binary to
/root/fleet/in/igneumd-<sha16> (sha-checked), the miner and workers over /opt/igneum/pkg/bin (the old ones kept as *.0314),
/root/fleet/standing.node rewritten so the supervisor restarts the node on it (data dir kept), the miner killed once so
the supervisor's loop restarts it on the new miner and worker. Then the read-back table: version line, digest, worker sha.
publish-0315.py [labels...] default: every standing box with role live
publish-0315.py --readback [labels...]
"""
import sys, os, hashlib
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))); from lib import Box, Registry, standing
from concurrent.futures import ThreadPoolExecutor
SHIP = "/Users/joshm/Projects/igneum-wt-ship0315"
NODE = (SHIP + "/vendor/igneum-node-0315/target-remote/release/igneumd", "7f0bde70cf2e72a3a9479ba6421f2b37162c3dd4391bf413717c99e6f168668d") # release-0.3.15-node f1ea7a38 (the stamp gate, the receive-side version rule, the pruned-node sync fix)
MINER = (SHIP + "/vendor/igneum-node-0315/target-remote/release/igneum-miner", "3976c1b270761218")
CUDA = (SHIP + "/infra/cross/out-workers/igneum-worker-cuda", "97e036e23f4ace66")
OCL = (SHIP + "/infra/cross/out-workers/igneum-worker-opencl", "ef9fbd03e7ce14d1")
def sha16(p): return hashlib.sha256(open(p, "rb").read()).hexdigest()[:16]
def publish(label):
b = Box.from_registry(label)
for p, s in (MINER, CUDA, OCL):
if sha16(p) != s: return label, f"local {os.path.basename(p)} sha {sha16(p)} is not {s}"
path = b.install_payload(NODE[0], NODE[1])
b.put([MINER[0], CUDA[0], OCL[0]], "/root/fleet/in/")
rc, out, err = b.run(f"""set -e; cd /opt/igneum/pkg/bin; for f in igneum-miner igneum-worker-cuda igneum-worker-opencl; do [ -e $f.0314 ] || cp $f $f.0314; cp /root/fleet/in/$f $f.new && chmod +x $f.new && mv $f.new $f; done
echo {path} > /root/fleet/standing.node
pkill -9 -f '^/(opt/igneum/pkg/bin|root/fleet/in)/igneumd(-0313|-[0-9a-f]{16})? --devnet --appdir=' 2>/dev/null; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-miner mine' 2>/dev/null; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-worker-cuda --serve' 2>/dev/null; true # the supervisor restarts the node on the new pointer within a minute, and the miner with it
echo swapped""", 120)
return label, out.strip() or err[:120]
def readback(label):
b = Box.from_registry(label)
rc, out, err = b.run(r"""P=$(pgrep -af '^/(opt/igneum/pkg/bin|root/fleet/in)/igneumd[^ ]* ' | head -1 | awk '{print $2}'); echo bin=$P sha=$(sha256sum $P 2>/dev/null | cut -c1-16) version="$($P --version 2>&1 | head -1)" vline=$(grep -oE 'igneumd/2\.1\.0-[0-9a-f]+' /root/fleet/node.log | tail -1) digest=$(grep -o 'digest: [0-9a-f]*' /root/fleet/node.log | tail -1 | awk '{print substr($2,1,16)}') miner=$(sha256sum /opt/igneum/pkg/bin/igneum-miner | cut -c1-16) cuda=$(sha256sum /opt/igneum/pkg/bin/igneum-worker-cuda | cut -c1-16) ocl=$(sha256sum /opt/igneum/pkg/bin/igneum-worker-opencl | cut -c1-16) $(/opt/igneum/pkg/bin/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -oE 'blocks=[0-9]+|synced=[a-z]+' | tr '\n' ' ') miner_up=$(pgrep -fc '^/opt/igneum/pkg/bin/igneum-miner mine') prover_up=$(pgrep -fc '^python3 -u /root/fleet/in/box-prover.py') rej=$(grep -c 'PoW rejected' /root/fleet/node.log)""", 60)
return label, out.strip() or err[:120]
if __name__ == "__main__":
a = sys.argv[1:]; rb = a and a[0] == "--readback"; a = a[1:] if rb else a
labels = a or [r["label"] for r in standing.roster() if r["role"] == "live"]
with ThreadPoolExecutor(13) as ex:
for l, o in ex.map(readback if rb else publish, labels): print(l, "::", o, flush=True)

View file

@ -0,0 +1,27 @@
"""Publish 2 on the standing live boxes (shipper, 22:43Z): the sixteen-field object replaces /root/fleet/override.json on every
live standing box, the node is restarted on it (the supervisor's start_node, data dir kept; the miner restarts with the node and
is held until synced), then the read-back: digest line (must be eada4bda...), synced, miner up. Devnet 2 boxes untouched."""
import sys, os, time; sys.path.insert(0, "/Users/joshm/Projects/igneum-wt-gpu-fleet/tools/fleet"); from lib import Box, Registry, standing
from concurrent.futures import ThreadPoolExecutor
def st(): return time.strftime("%H:%M:%SZ", time.gmtime())
F16=os.path.expanduser("~/Desktop/fleet/override-16.json"); WANT="eada4bda8aa8368c"
live=[r["label"] for r in standing.roster() if r["role"]=="live"]
def move(l):
try:
b=Box.from_registry(l); b.put([F16], "/root/fleet/override-16.json")
rc,out,err=b.run("cd /root/fleet && cp override.json override-13.json && cp override-16.json override.json && python3 -c \"import json; d=json.load(open('/root/fleet/override.json')); print('fields', len(d))\" && pkill -9 -f '^/(opt/igneum/pkg/bin|root/fleet/in)/igneumd(-0313|-[0-9a-f]{16})? --devnet --appdir=' ; echo killed-for-restart $(date -u +%T)", 60)
return l, out.replace("\n"," ").strip()+(" ERR "+err[:60] if err.strip() else "")
except Exception as e: return l, "EXC "+str(e)[:60]
print(st(), "moving", len(live), "boxes", flush=True)
with ThreadPoolExecutor(14) as ex:
for l,o in ex.map(move, live): print(st(), l, o, flush=True)
time.sleep(150)
def read(l):
try:
rc,out,err=Box.from_registry(l).run("echo digest=$(grep -o 'digest: [0-9a-f]*' /root/fleet/node.log | tail -1 | awk '{print substr($2,1,16)}') vline=$(grep -oE 'igneumd/2\\.1\\.0-[0-9a-f]+' /root/fleet/node.log | tail -1) fields=$(python3 -c \"import json; print(len(json.load(open('/root/fleet/override.json'))))\") $(/opt/igneum/pkg/bin/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -oE 'blocks=[0-9]+|peers=[0-9]+|synced=[a-z]+' | tr '\\n' ' ') miner=$(pgrep -f '^/opt/igneum/pkg/bin/igneum-miner mine' | wc -l) rej_since=$(awk '$2 >= \"22:47:00\"' /root/fleet/node.log | grep -c 'got reject message') v4=\"$(grep -E 'Program class v4 (signal window|from the override)' /root/fleet/node.log | tail -2 | sed -E 's/^[0-9: .+-]*//' | cut -c1-70 | tr '\\n' ';')\"", 60)
return l, out.strip().split("\n")[-1]
except Exception as e: return l, "EXC "+str(e)[:60]
with ThreadPoolExecutor(14) as ex: res=list(ex.map(read, live))
ok=0
for l,o in res: print(st(), l, "::", o, flush=True); ok += ("digest="+WANT) in o
print(st(), f"TABLE {ok} of {len(live)} on {WANT}", flush=True)

View file

@ -0,0 +1 @@
{"difficulty_v2_activation_daa":0,"proving_v0_activation_daa":0,"fees_v1_activation_daa":0,"finality_v3_activation_daa":0,"program_class_v3_activation_daa":0,"proving_v1_activation_daa":0,"proving_v1_segment_blocks":8,"proving_v1_unproven_daa":600,"proving_v1_aggregator_share_bps":1000,"proving_v1_fresh_rule_daa":0,"exec_restart_number":18446744073709551615,"exec_restart_hash":"","exec_restart_trust_daa":18446744073709551615,"pow_epoch_blocks":600,"pow_epoch_lead":100,"program_class_v4_activation_daa":2400,"program_class_v4_signal_window_daa":600}

View file

@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Stops a rehearsal box's miner loops, miner, worker and sampler but NOT its dn400 node: run as a FILE (an inline pkill
# carrying these words kills the ssh shell that carries them). Then MINER_ONLY=1 bash in/box-rehearsal.sh restarts the miner.
pkill -9 -f '^bash in/box-rehearsal.sh'; pkill -9 -f '^/root/fleet/in/igneum-miner-v4 '; pkill -9 -f '^/opt/igneum/pkg/bin/igneum-worker-cuda --serve .*dn400'
sleep 1; echo "left=$(pgrep -c -f '^bash in/box-rehearsal.sh')+$(pgrep -c -f '^/root/fleet/in/igneum-miner-v4 ') node=$(pgrep -c -f '^/root/fleet/in/igneumd')"

18
tools/fleet/restage.py Normal file
View file

@ -0,0 +1,18 @@
#!/usr/bin/env python3
"""Replaces /root/fleet/in/igneumd-0313 on every live phase-1/2 box and the hub with the given local file, checked by
sha256 on the box before the move (a wrong binary is never left under the swap's name)."""
import sys, os, subprocess, hashlib
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))); import fleet
from concurrent.futures import ThreadPoolExecutor
src, want = sys.argv[1], sys.argv[2]
assert hashlib.sha256(open(src, "rb").read()).hexdigest() == want, "local file does not match"
reg = fleet.load()
targets = [(iid, b) for iid, b in reg.items() if b.get("state") != "destroyed" and b.get("ssh_ok") and (b.get("phase") in ("1", "2") or b.get("hub"))]
def one(x):
iid, b = x
p = subprocess.run(["scp"] + fleet.SSH_OPTS + ["-P", str(b["ssh_port"]), src, f"root@{b['ssh_host']}:/root/fleet/in/igneumd-0313.new"], capture_output=True, text=True, timeout=900)
if p.returncode: return b["label"], "scp failed " + p.stderr.strip()[:60]
r, out, _ = fleet.ssh(b, f"cd /root/fleet/in && s=$(sha256sum igneumd-0313.new | cut -c1-64) && [ \"$s\" = {want} ] && mv igneumd-0313.new igneumd-0313 && chmod +x igneumd-0313 && echo replaced ${{s:0:16}} || echo BAD $s; rm -f /opt/igneum/pkg/bin/igneumd-0313", timeout=60)
return b["label"], out.strip()
with ThreadPoolExecutor(10) as ex:
for lab, out in sorted(ex.map(one, targets)): print(f"{lab:<12} {out}", flush=True)

38
tools/fleet/rig-retry.py Normal file
View file

@ -0,0 +1,38 @@
#!/usr/bin/env python3
"""Retries the two 8-card rigs on Vast and RunPod every ten minutes until both are held or RIG_RETRY_UNTIL passes."""
import time, sys, os, secrets, datetime
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import fleet, vast, runpod
START = ["bash", "-c", "apt-get update -qq >/dev/null 2>&1; DEBIAN_FRONTEND=noninteractive apt-get install -y -qq openssh-server >/dev/null 2>&1; mkdir -p /run/sshd /root/.ssh; echo \"$PUBLIC_KEY\" > /root/.ssh/authorized_keys; chmod 700 /root/.ssh; chmod 600 /root/.ssh/authorized_keys; sed -i 's/^#\\?PermitRootLogin.*/PermitRootLogin prohibit-password/' /etc/ssh/sshd_config; exec /usr/sbin/sshd -D"]
def log(s): print(datetime.datetime.utcnow().strftime("%H:%M:%SZ"), s, flush=True)
def have(label): return any(b["label"] == label and b.get("state") != "destroyed" for b in fleet.load().values())
def try_vast(gpu, label, archs):
offers = [o for o in vast.search(gpu, n=12, min_cores=16, gpus=8, min_disk=150) if (o.get("cpu_ram") or 0)/1024 >= 96]
offers.sort(key=lambda o: (o.get("reliability2", 0) < 0.98, o["dph_total"]))
for o in offers[:3]:
try: d = vast.call("PUT", f"/v0/asks/{o['id']}/", {"client_id": "me", "image": vast.IMAGE, "disk": 150, "label": label, "runtype": "ssh", "cancel_unavail": True})
except SystemExit as e: log(f"{label} vast: {str(e)[-60:]}"); return False
if not d.get("success"): continue
iid = d["new_contract"]; time.sleep(40)
if not any(str(i.get("id")) == str(iid) for i in vast.instances()): log(f"{label} vast {iid} cancelled"); continue
vast.ledger({"t": fleet.now(), "event": "rent", "instance": iid, "offer": o["id"], "label": label, "gpu": gpu, "num_gpus": 8, "dph": o["dph_total"], "disk": 150, "geo": o.get("geolocation")})
fleet.patch(iid, label=label, card=gpu, vram_mb=o.get("gpu_ram"), archs=archs, offer=o["id"], dph=o["dph_total"], num_gpus=8, provider="vast", phase="3", state="renting", rented_at=fleet.now(), wallet="0x"+secrets.token_hex(20), geo=o.get("geolocation"), doing=f"phase 3: the 8x {gpu.split()[-1]} rig")
log(f"{label} -> vast {iid} ${o['dph_total']:.2f}/h {o.get('geolocation')}"); return True
return False
def try_runpod(gpu, label, archs):
for cloud in ("SECURE", "COMMUNITY"):
body = {"name": label, "imageName": "nvidia/cuda:12.8.1-devel-ubuntu24.04", "gpuTypeIds": [gpu], "gpuCount": 8, "containerDiskInGb": 120, "volumeInGb": 0, "cloudType": cloud, "ports": ["22/tcp"], "env": {"PUBLIC_KEY": runpod.PUB}, "supportPublicIp": True, "computeType": "GPU", "dockerStartCmd": START}
try: d = runpod.call("POST", "/pods", body)
except SystemExit as e: continue
iid = d.get("id")
runpod.ledger({"t": fleet.now(), "event": "rent", "provider": "runpod", "instance": iid, "label": label, "gpu": gpu, "num_gpus": 8, "dph": d.get("costPerHr"), "disk": 120})
fleet.patch(iid, label=label, card=gpu.replace("NVIDIA GeForce ", ""), vram_mb=32607 if "5090" in gpu else 24564, archs=archs, dph=d.get("costPerHr") or 0, num_gpus=8, provider="runpod", phase="3", state="renting", rented_at=fleet.now(), wallet="0x"+secrets.token_hex(20), cloud=cloud, doing=f"phase 3: the 8x {gpu.split()[-1]} rig")
log(f"{label} -> runpod {iid} {cloud} ${d.get('costPerHr')}/h"); return True
return False
until = float(os.environ.get("RIG_RETRY_UNTIL", time.time() + 6 * 3600))
while time.time() < until:
for gpu, label, archs, rpg in (("RTX 4090", "rig-4090x8", "89", "NVIDIA GeForce RTX 4090"), ("RTX 5090", "rig-5090x8", "120", "NVIDIA GeForce RTX 5090")):
if have(label): continue
if not try_vast(gpu, label, archs): try_runpod(rpg, label, archs)
if have("rig-4090x8") and have("rig-5090x8"): log("both rigs held"); break
time.sleep(600)

55
tools/fleet/runpod.py Executable file
View file

@ -0,0 +1,55 @@
#!/usr/bin/env python3
"""RunPod REST client for the fleet (https://rest.runpod.io/v1, Bearer from ~/.config/runpod/credentials; the GraphQL
endpoint refuses the Bearer header, so nothing here touches account settings: the fleet key goes in per pod through
the PUBLIC_KEY env, which RunPod's images write to authorized_keys). Ledger lines as vast.py.
runpod.py gpus the GPU types with availability and prices
runpod.py rent <gpuTypeId> <label> [--count 8] [--disk 80] [--image ...] [--secure]
runpod.py list pods: id, name, status, ssh endpoint
runpod.py destroy <pod_id>...
"""
import json, os, sys, time, urllib.request, urllib.error, datetime, argparse
KEY = open(os.path.expanduser("~/.config/runpod/credentials")).read().strip()
BASE = "https://rest.runpod.io/v1"
LEDGER = os.path.expanduser("~/Desktop/fleet/ledger.jsonl")
IMAGE = "runpod/pytorch:2.8.0-py3.11-cuda12.8.1-cudnn-devel-ubuntu22.04"
PUB = open(os.path.expanduser("~/.ssh/igneum-fleet.pub")).read().strip()
def call(method, path, body=None):
req = urllib.request.Request(BASE + path, data=(json.dumps(body).encode() if body is not None else None),
headers={"Authorization": "Bearer " + KEY, "Content-Type": "application/json", "User-Agent": "curl/8.7.1", "Accept": "*/*"}, method=method) # Cloudflare 1010 refuses urllib's default agent
try:
with urllib.request.urlopen(req, timeout=90) as r:
t = r.read(); return json.loads(t) if t else {}
except urllib.error.HTTPError as e: raise SystemExit(f"{method} {path}: HTTP {e.code}: {e.read().decode(errors='replace')[:400]}")
def now(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def ledger(row):
with open(LEDGER, "a") as f: f.write(json.dumps(row) + "\n")
def gpus():
for g in call("GET", "/gputypes"):
print(f"{g.get('id'):<34} {g.get('displayName','')!s:<22} {g.get('memoryInGb')}GB secure ${g.get('securePrice')}/h community ${g.get('communityPrice')}/h max {g.get('maxGpuCount')} avail {g.get('lowestPrice', {}) if isinstance(g.get('lowestPrice'), dict) else ''}")
def rent(gpu, label, count=8, disk=80, image=IMAGE, secure=False, ports=("22/tcp",)):
body = {"name": label, "imageName": image, "gpuTypeIds": [gpu], "gpuCount": count, "containerDiskInGb": disk, "volumeInGb": 0,
"cloudType": "SECURE" if secure else "COMMUNITY", "ports": list(ports), "env": {"PUBLIC_KEY": PUB}, "supportPublicIp": True, "computeType": "GPU"}
d = call("POST", "/pods", body)
row = {"t": now(), "event": "rent", "provider": "runpod", "instance": d.get("id"), "label": label, "gpu": gpu, "num_gpus": count, "dph": d.get("costPerHr"), "disk": disk}
ledger(row); print(json.dumps(row)); print(json.dumps(d)[:600]); return d
def pods():
out = call("GET", "/pods")
return out if isinstance(out, list) else out.get("pods", [])
def ssh_of(p):
for m in (p.get("portMappings") or {}).items() if isinstance(p.get("portMappings"), dict) else []:
if str(m[0]) == "22": return p.get("publicIp"), m[1]
return p.get("publicIp"), None
def destroy(pid):
d = call("DELETE", f"/pods/{pid}"); ledger({"t": now(), "event": "destroy", "provider": "runpod", "instance": pid}); print("destroyed", pid, d)
if __name__ == "__main__":
a = sys.argv[1:]
if a[0] == "gpus": gpus()
elif a[0] == "rent":
ap = argparse.ArgumentParser(); ap.add_argument("gpu"); ap.add_argument("label"); ap.add_argument("--count", type=int, default=8); ap.add_argument("--disk", type=int, default=80); ap.add_argument("--image", default=IMAGE); ap.add_argument("--secure", action="store_true")
n = ap.parse_args(a[1:]); rent(n.gpu, n.label, n.count, n.disk, n.image, n.secure)
elif a[0] == "list":
for p in pods(): print(p.get("id"), p.get("name"), p.get("desiredStatus"), p.get("gpuCount"), p.get("machine", {}).get("gpuTypeId") if isinstance(p.get("machine"), dict) else "", f"${p.get('costPerHr')}/h", ssh_of(p), (p.get("portMappings")))
elif a[0] == "destroy":
for i in a[1:]: destroy(i)
elif a[0] == "raw": print(json.dumps(pods(), indent=1)[:3000])

40
tools/fleet/swap.py Normal file
View file

@ -0,0 +1,40 @@
#!/usr/bin/env python3
"""The phase-2 node swap over the fleet, one call per shipper word:
swap.py binary on "hands on 0.3.13": every phase-2 box (the hub first) restarts on igneumd-0313 with the ten-field file
swap.py file on "hands on b18ed271": the thirteen-field file swapped in and the node restarted; the replay timed
swap.py status the swap log's last line and the exec status per box
The expected sha256 and digests are pinned here; a box whose binary or digest differs refuses and is reported."""
import sys, os, time, json
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))); import fleet
from concurrent.futures import ThreadPoolExecutor
SHA = "d6350586fe837b1f71696546628ec071b6accce2531aef776cf2b1e5487a8cdc"
DIG = {"binary": "7bd98cc4118616455709d5e32a30b799e6e67caa42d2b5d09875cd49848a7ed7", "file": "b18ed271f75dd46406d230f4156c37472127415a4c32c558bac662f6f840e61c"}
def targets():
reg = fleet.load()
t = [(iid, b) for iid, b in reg.items() if b.get("state") != "destroyed" and b.get("ssh_ok") and (b.get("phase") in ("1", "2") or b.get("hub"))]
return sorted(t, key=lambda x: (not x[1].get("hub"), x[1]["label"]))
def start(step):
hub = next((b for _, b in targets() if b.get("hub")), None)
env = f"STEP={step} NODE_SHA256={SHA} EXPECT_DIGEST={DIG[step]} HUB_PEER={hub['hub_peer']} HUB_SSH={hub['ssh_host']} HUB_PORT={hub['ssh_port']}" if hub else f"STEP={step} NODE_SHA256={SHA} EXPECT_DIGEST={DIG[step]}"
def one(x):
iid, b = x
fleet.scp(b, [os.path.join(fleet.HERE, "box-node-swap.sh"), os.path.join(fleet.HERE, "box-kill.sh")], "/root/fleet/in/")
r, out, err = fleet.ssh(b, f"cd /root/fleet && chmod +x in/box-node-swap.sh && mv out/node-swap.log out/node-swap-{step}-prev.log 2>/dev/null; {env} setsid nohup in/box-node-swap.sh </dev/null >/dev/null 2>&1 & sleep 1; echo started", timeout=60)
fleet.patch(iid, stage="swap-" + step, doing=f"0.3.13 swap, step {step}: restart, digest check, the exec replay timed")
return b["label"], out.strip() or err[:60]
ts = targets()
if hub: # the hub first, then the rest in parallel 20 s later (it is every box's second peer)
print(one(next(x for x in ts if x[1].get("hub")))); time.sleep(20)
with ThreadPoolExecutor(12) as ex:
for lab, out in ex.map(one, [x for x in ts if not x[1].get("hub")]): print(lab, out)
def status():
def one(x):
iid, b = x
r, out, err = fleet.ssh(b, "grep -E '^RESULT (node_started|swap_done|swap_failed|datadir_failed)' /root/fleet/out/node-swap.log 2>/dev/null | tail -2 | cut -c1-200; grep -E '^RESULT replay' /root/fleet/out/node-swap.log 2>/dev/null | tail -1 | sed -E 's/difficulty=[0-9.]* sink=[0-9a-f]* //' | cut -c1-120; grep -E '^RESULT exec_status' /root/fleet/out/node-swap.log 2>/dev/null | tail -1 | cut -c1-200", timeout=40)
return b["label"], out.strip().replace("\n", " | ")
with ThreadPoolExecutor(12) as ex:
for lab, out in ex.map(one, targets()): print(f"{lab:<12} {out}")
if __name__ == "__main__":
c = sys.argv[1]
if c in ("binary", "file"): start(c)
elif c == "status": status()

126
tools/fleet/vast.py Executable file
View file

@ -0,0 +1,126 @@
#!/usr/bin/env python3
"""Vast.ai client for the Igneum GPU fleet (6 October 2026). The key is read from ~/.config/vast/credentials and never
printed. Every rent and destroy is appended to ~/Desktop/fleet/ledger.jsonl with the offer's hourly price, so the spend
can be summed without the provider's statement.
vast.py search "RTX 3090" [--n 8] [--min-cores 6] [--max-dph 0.5] [--gpus 1]
vast.py rent <offer_id> --label <label> [--disk 60] [--image nvidia/cuda:12.8.1-devel-ubuntu24.04] [--onstart FILE]
vast.py list every instance: id, label, gpu, state, ssh host and port, dph
vast.py ssh <instance_id> prints the ssh command line (host, port) for the fleet key
vast.py destroy <instance_id>... destroys and writes the ledger line
vast.py spend the ledger's running total and the account's credit
"""
import json, os, sys, time, urllib.request, urllib.parse, urllib.error, argparse, datetime
KEY = open(os.path.expanduser("~/.config/vast/credentials")).read().strip()
BASE = "https://console.vast.ai/api"
LEDGER = os.path.expanduser("~/Desktop/fleet/ledger.jsonl")
IMAGE = "nvidia/cuda:12.8.1-devel-ubuntu24.04"
def call(method, path, body=None, timeout=90):
req = urllib.request.Request(BASE + path, data=(json.dumps(body).encode() if body is not None else None),
headers={"Authorization": "Bearer " + KEY, "Content-Type": "application/json", "Accept": "application/json"},
method=method)
for attempt in range(4):
try:
with urllib.request.urlopen(req, timeout=timeout) as r:
raw = r.read() or b"{}"
try: return json.loads(raw)
except json.JSONDecodeError:
if attempt < 3: time.sleep(3 * (attempt + 1)); continue
raise SystemExit(f"{method} {path}: non-JSON reply: {raw[:120]!r}")
except urllib.error.HTTPError as e:
txt = e.read().decode(errors="replace")
if e.code in (429, 502, 503, 504) and attempt < 3:
time.sleep(3 * (attempt + 1)); continue
raise SystemExit(f"{method} {path}: HTTP {e.code}: {txt[:400]}")
except (urllib.error.URLError, TimeoutError) as e:
if attempt < 3: time.sleep(3 * (attempt + 1)); continue
raise
def now(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def ledger(row):
os.makedirs(os.path.dirname(LEDGER), exist_ok=True)
with open(LEDGER, "a") as f: f.write(json.dumps(row) + "\n")
def search(gpu, n=8, min_cores=4, max_dph=None, gpus=1, min_disk=70, min_rel=0.95, min_down=200, cuda=12.8):
q = {"gpu_name": {"eq": gpu}, "rentable": {"eq": True}, "verified": {"eq": True}, "num_gpus": {"eq": gpus},
"cuda_max_good": {"gte": cuda}, "disk_space": {"gte": min_disk}, "reliability2": {"gte": min_rel},
"inet_down": {"gte": min_down}, "cpu_cores_effective": {"gte": min_cores},
"order": [["dph_total", "asc"]], "type": "on-demand", "limit": n}
if max_dph: q["dph_total"] = {"lte": max_dph}
d = call("GET", "/v0/bundles/?q=" + urllib.parse.quote(json.dumps(q)))
return d.get("offers", [])
def fmt_offer(o):
return (f"id {o['id']} ${o['dph_total']:.3f}/h {o.get('num_gpus')}x {o.get('gpu_name')} {o.get('gpu_ram')}MB cuda{o.get('cuda_max_good')} "
f"cores{o.get('cpu_cores_effective',0):.0f} ram{o.get('cpu_ram',0)/1024:.0f}G disk{o.get('disk_space',0):.0f} dl{o.get('inet_down',0):.0f} "
f"ul{o.get('inet_up',0):.0f} rel{o.get('reliability2',0):.3f} {o.get('geolocation')} drv{o.get('driver_version')}")
def rent(offer_id, label, disk=60, image=IMAGE, onstart=None, offer=None):
body = {"client_id": "me", "image": image, "disk": disk, "label": label, "runtype": "ssh", "image_login": None,
"python_utf8": False, "lang_utf8": False, "use_jupyter_lab": False, "cancel_unavail": True}
if onstart: body["onstart"] = open(onstart).read()
# the offer's price for the ledger
o = offer or {}
d = call("PUT", f"/v0/asks/{offer_id}/", body)
if not d.get("success"): raise SystemExit(f"rent failed: {d}")
iid = d.get("new_contract")
row = {"t": now(), "event": "rent", "instance": iid, "offer": int(offer_id), "label": label, "gpu": o.get("gpu_name"),
"num_gpus": o.get("num_gpus"), "dph": o.get("dph_total"), "disk": disk, "geo": o.get("geolocation")}
ledger(row); print(json.dumps(row)); return iid
def instances():
try:
d = call("GET", "/v0/instances/?owner=me")
return d.get("instances", [])
except SystemExit:
d = call("GET", "/v1/instances/")
return d.get("instances", d if isinstance(d, list) else [])
def destroy(iid):
try: d = call("DELETE", f"/v0/instances/{iid}/", {})
except SystemExit: d = call("DELETE", f"/v1/instances/{iid}/", {})
row = {"t": now(), "event": "destroy", "instance": int(iid), "ok": bool(d.get("success", True))}
ledger(row); print(json.dumps(row))
def spend():
rows = [json.loads(l) for l in open(LEDGER)] if os.path.exists(LEDGER) else []
starts = {r["instance"]: r for r in rows if r["event"] == "rent"}
stops = {r["instance"]: r for r in rows if r["event"] == "destroy"}
total = 0.0; lines = []
tnow = datetime.datetime.now(datetime.timezone.utc)
for iid, r in starts.items():
t0 = datetime.datetime.strptime(r["t"], "%Y-%m-%dT%H:%M:%SZ").replace(tzinfo=datetime.timezone.utc)
t1 = datetime.datetime.strptime(stops[iid]["t"], "%Y-%m-%dT%H:%M:%SZ").replace(tzinfo=datetime.timezone.utc) if iid in stops else tnow
h = max((t1 - t0).total_seconds(), 60) / 3600
cost = h * float(r.get("dph") or 0) + (r.get("disk") or 0) * 0.0002 * h # storage about USD 0.2 per TB-hour, approximate
total += cost
lines.append(f"{iid} {r['label']:<22} {r.get('gpu')!s:<16} {h:5.2f} h x ${r.get('dph') or 0:.3f} = ${cost:6.2f} {'running' if iid not in stops else 'destroyed'}")
me = call("GET", "/v0/users/current/")
print("\n".join(lines)); print(f"ledger total USD {total:.2f}; account credit USD {me.get('credit')}; balance {me.get('balance')}")
if __name__ == "__main__":
ap = argparse.ArgumentParser(); sub = ap.add_subparsers(dest="cmd", required=True)
s = sub.add_parser("search"); s.add_argument("gpu"); s.add_argument("--n", type=int, default=8); s.add_argument("--min-cores", type=int, default=4)
s.add_argument("--max-dph", type=float); s.add_argument("--gpus", type=int, default=1); s.add_argument("--min-disk", type=int, default=70); s.add_argument("--json", action="store_true")
r = sub.add_parser("rent"); r.add_argument("offer"); r.add_argument("--label", required=True); r.add_argument("--disk", type=int, default=60); r.add_argument("--image", default=IMAGE); r.add_argument("--onstart")
sub.add_parser("list"); sub.add_parser("spend")
d = sub.add_parser("destroy"); d.add_argument("ids", nargs="+")
h = sub.add_parser("ssh"); h.add_argument("id")
a = ap.parse_args()
if a.cmd == "search":
offs = search(a.gpu, a.n, a.min_cores, a.max_dph, a.gpus, a.min_disk)
print(json.dumps(offs) if a.json else "\n".join(fmt_offer(o) for o in offs) or "no offers")
elif a.cmd == "rent": rent(a.offer, a.label, a.disk, a.image, a.onstart)
elif a.cmd == "list":
for i in instances():
print(f"{i.get('id')} {str(i.get('label')):<22} {i.get('num_gpus')}x {i.get('gpu_name')} {i.get('actual_status')}/{i.get('cur_state')} "
f"ssh {i.get('ssh_host')}:{i.get('ssh_port')} ${i.get('dph_total',0):.3f}/h {i.get('geolocation')} {i.get('status_msg') or ''}"[:200])
elif a.cmd == "ssh":
for i in instances():
if str(i.get("id")) == a.id: print(f"ssh -i ~/.ssh/igneum-fleet -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -p {i.get('ssh_port')} root@{i.get('ssh_host')}")
elif a.cmd == "destroy":
for i in a.ids: destroy(i)
elif a.cmd == "spend": spend()

49
tools/fleet/winddown.py Normal file
View file

@ -0,0 +1,49 @@
#!/usr/bin/env python3
"""The wave's wind-down by the standing rule (6 October 2026, 21:5xZ, the coordinator's order): never remove more than 10 percent of
the live devnet's voter weight in any hour, so finality never sees tonight's departure again. Every hour: read each candidate's
share of the last 2,000 blocks (lib.standing.weight_shares, the conservative proxy for the table), pick the largest slice under
LIMIT (default 9.5 percent) that weight_check allows (it adds what the last hour removed), hold the slice if the hub's last lock
signed under HOLD_PCT of the table after the removal, destroy the slice's pods (RunPod, registry rows marked), then read the lock
line ten minutes later. Candidates: the wave pods, then the 8x rig (one voter), never a standing box or a Devnet 2 box; rz-4090 waits
for the proving agent's word. Log: ~/Desktop/fleet/winddown.log; schedule and finality reads in ~/Desktop/fleet/winddown.jsonl."""
import sys, os, time, json, datetime
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))); from lib import Box, Registry, standing
LIMIT = float(os.environ.get("LIMIT", "0.095")); HOLD_PCT = float(os.environ.get("HOLD_PCT", "75.0")) # the coordinator's rule of 23:0xZ: nothing moves on a standing box, pool slices included, until the table reads over 75 percent signed; EVERY = int(os.environ.get("EVERY", "3600"))
ROOT = os.path.expanduser("~/Desktop/fleet"); J = os.path.join(ROOT, "winddown.jsonl")
def st(): return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
def log(**row): row["t"] = st(); open(J, "a").write(json.dumps(row) + "\n"); print(row["t"], {k: v for k, v in row.items() if k != "t"}, flush=True)
hub = Box.from_registry("hub-1")
def finality():
rc, out, err = hub.run("grep -E 'Finality: checkpoint [0-9]+ LOCKED:' /root/fleet/node.log | tail -1 | grep -oE 'checkpoint [0-9]+|signed [0-9]+ = [0-9.]+% of active, [0-9.]+% of total' | tr '\\n' ' '; grep -E 'certificate at index [0-9]+ received' /root/fleet/node.log | tail -1 | grep -oE '[0-9]+ of [0-9]+ voters, weight [0-9]+ of [0-9]+'; grep -E 'Finality: checkpoint [0-9]+ LOCKED:' /root/fleet/node.log | tail -1 | cut -c1-19", 40)
s = out.replace("\n", " | ").strip(); import re; m = re.search(r"([0-9.]+)% of total", s); return s, float(m.group(1)) if m else None
def candidates():
reg = Registry.load(); pods = sorted((i, b) for i, b in reg.items() if b.get("state") != "destroyed" and b.get("ssh_host") and b["label"].startswith("wave"))
rig = [(i, b) for i, b in reg.items() if b.get("state") != "destroyed" and b.get("label") == "rig-4090x8"]
return pods + rig
def keys_of(labels):
rc, out, err = hub.run("for l in " + " ".join(labels) + "; do echo $l $(/opt/igneum/pkg/bin/igneum-miner key-hash $l 2>/dev/null | grep -oE '[0-9a-f]{64}' | head -1); done", 120)
return dict(ln.split() for ln in out.splitlines() if len(ln.split()) == 2)
while True:
cands = candidates()
if not cands: log(event="done", note="no candidates left: the wave and the rig are down; the standing 14 and the Devnet 2 four stay"); break
labels = [b["label"] for _, b in cands]; sh, n = standing.weight_shares(2000); keys = keys_of(labels); share = {l: sh.get(keys.get(l, ""), 0.0) for l in labels}
fin, signed_total = finality()
recent = 0.0
if os.path.exists(standing.WEIGHT_LOG):
for ln in open(standing.WEIGHT_LOG):
try: j = json.loads(ln)
except Exception: continue
if j.get("ts", 0) >= time.time() - 3600: recent += float(j.get("share", 0))
budget = LIMIT - recent; slice_, acc = [], 0.0
for l in sorted(labels, key=lambda x: share[x]): # smallest shares first: more pods per slice, the same weight
if acc + share[l] <= budget: slice_.append(l); acc += share[l]
if signed_total is not None and signed_total - 100 * acc < HOLD_PCT: log(event="hold", reason=f"the slice would take the signed share from {signed_total}% to {signed_total - 100*acc:.1f}%, under {HOLD_PCT}%", finality=fin); time.sleep(EVERY); continue
if not slice_: log(event="hold", reason=f"no pod fits the budget ({100*budget:.1f}% left this hour)", finality=fin); time.sleep(EVERY); continue
v = standing.weight_check(slice_, limit=LIMIT)
if not v["ok"]: log(event="hold", reason="weight_check refused", verdict=v, finality=fin); time.sleep(EVERY); continue
log(event="slice", pods=slice_, share_pct=round(100 * acc, 2), remaining=len(labels) - len(slice_), finality_before=fin, blocks_read=n)
for l in slice_:
try: Box.from_registry(l).destroy(); log(event="destroyed", pod=l, share_pct=round(100 * share[l], 2))
except Exception as e: log(event="destroy_failed", pod=l, error=str(e)[:120])
time.sleep(600); fin2, _ = finality(); log(event="finality_after", slice=slice_, finality=fin2)
time.sleep(EVERY - 600)