diff --git a/docs/bench-log.md b/docs/bench-log.md index 47e317d82..75cc20fac 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -2597,3 +2597,39 @@ USD 20 an hour on community pods, against a devnet of 1.16 GH/s. Consequence: the devnet's hash is rentable for the price of a dinner, so nothing on it is a security result; the counter-ASIC and finality work is tested there for correctness, not for cost. The cost argument only starts at the TH/s scale, where the rental market's supply (not its price) is the limit, and that number belongs in the litepaper with this caveat. + +## Block rate on Devnet 2, 6 October 2026 (branch gpu-fleet): 10 blocks per second against 1 on 42 rented cards + +Run A (10 blocks/s profile, star topology, 65 min): 4.87 DAG blocks/s, 1.09 blue blocks/s, 77.6 percent red, tips 250 to 660, +max reorg 55, difficulty easing all hour (6,719 to 1,307), the exec follower at 0.05 blocks/s (lag 18,901 at the end). Run B +(1 block/s, same boxes, 30 min): 1.41 blocks/s over the window with the join burst, 1.0 blocks/s and under 2 percent red from +minute six, tips 1 to 3, difficulty settled in six minutes (453 to 482 M), the exec follower at 0.46 blocks/s. The network +lane's read: run A's reds came from node throughput (61 to 345 ms CPU per accepted block at mergeset 8 to 200), not from the +star. Per tier the blue rate decides the payout interval (1.09 against 1.19 blue/s: a 4070 at 10 TH/s waits about three days +for a paying block either way), so the higher rate buys the solo miner nothing until the node processes a block in under 50 ms +at mergeset 248. Recommendation (the lane's): 1 block/s for the testnet and the launch, 10 behind three measured gates. +Full tables and sources: `docs/analysis/block-rate-devnet2.md`, rows in `~/Desktop/fleet/bps/{A,B}.jsonl`. + +## The rented fleet is the devnet's finality, 6 October 2026 (branch gpu-fleet) + +Measured at 21:57Z from the hub's last 2,000 blocks: the 38 wave pods held 77.8 percent of the voter weight (mean 2.05 percent +a pod), the 14 standing boxes most of the rest, the hands and the hub the remainder; the last lock signed 93.5 percent of the +active voters and 89.9 percent of the frozen table (53 of 84 voters on the first certificate). Earlier the same evening the +fleet removed 13 miners' GPUs inside three minutes and finality paused for two hours five minutes (18:39:36Z to 20:44:44Z, +42.7 percent of the table gone with earlier leavers; rule v3 holds a full window). From that came the 10 percent rule (never +remove more than 10 percent of the live devnet's weight in an hour, `tools/fleet/lib/standing.py weight_check`) and the wave's +wind-down by hourly slices (`tools/fleet/winddown.py`: slice 1 at 21:58Z took 12 pods and the 8x rig at 8.6 percent of weight). + +When the wave is gone the 14 standing boxes hold about 95 percent of the weight, so from then until public hash arrives the +fleet alone is the devnet's finality: a home miner's lock lands only while the fleet is up. What holds it up: every standing +box runs under `box-standing.sh`, which restarts a dead node within one of its 60-second passes (the hub's three deaths +tonight: 63 s, 41 s and 56 s to the restart line), restarts the miner with the node, prunes the prover's exports and trims +the node log, and runs the exec recovery recipe when the state layer reads zero; `lib/standing.py loop` re-rents a dead host +in the same shape and reports a box behind its wanted binary. + +| Tier | What it means | +|---|---| +| Home miner | your lock depends on 14 rented cards staying up and mining; a finality pause is not your node's fault and nothing you can fix; the rule above is what keeps it from recurring on the fleet's side | +| Rig | the same, and a rig that leaves is itself a weight removal: at 459 MH/s on tonight's devnet it is about 20 percent of the weight, over the hour's budget by itself | +| Pool | a pool node is one voter carrying its members' whole weight; a pool restart is the largest single removal on the network and must be sliced like the fleet's | +| The network | finality by miner weight is only as steady as the miners' uptime; until public hash dwarfs the fleet, the fleet's supervisor is a consensus component | diff --git a/docs/plans/ember-tune.md b/docs/plans/ember-tune.md index 6e437cd0e..dffe4561f 100644 --- a/docs/plans/ember-tune.md +++ b/docs/plans/ember-tune.md @@ -326,3 +326,42 @@ Earlier (shipped in 0.3.12 and 0.3.13 unless marked): - The confirm plan's neighbour is one step; a second neighbour (the other knob) would cost 75 s more and catch a prior that is wrong on both knobs. - Intel: no knob yet; the row says measure only. + +## 7a. One administrator approval, ever (0.3.13; Josh, 6 October 2026, 11:50 UTC) + +What 0.3.12 does: Power control on raises one prompt and sets every cap in that step; every later cap (an app start, a +reboot, a slider move) and every tune's helper is another elevated launch, so another prompt. Not "once, ever". + +What `src/powertask.rs` does: the first approval's elevated step also registers a per-user Windows scheduled task, +`Igneum Power Helper` (principal = the signed-in user, interactive logon, RunLevel Highest, no trigger, hidden, one +hour limit, new starts ignored while one runs), whose action is the app's own exe in the install folder with +`--power-helper`. A task the user owns is started by the user's unelevated engine with `Start-ScheduledTask`, no +prompt, and runs elevated. Every later cap and every tune's helper starts the task and writes the command file +`/app/sweep/cmd.txt` (` dev `, ` pl `, ` lgc `, ` rgc`, `quit`). The task +survives app restarts, updates (the per-user installer replaces the exe in place; the task's action path is the +install folder) and reboots. Power control off starts the task once and sends `remove`: the helper unregisters the +task (elevated) and exits; nothing is left behind. Linux keeps pkexec per step; macOS has no cap. + +Threat note: the helper runs only fixed verbs with digit-only arguments through `Command::new(nvidia-smi).args` +(the driver's own path, never PATH, never a shell); a line that is anything else is ignored; the sequence must rise +(a stale file runs nothing); an attacker running as the user gains the power limit and clock cap of the user's own +NVIDIA cards inside the driver's ranges, which the same user could set with one approved prompt anyway; no file, +process, registry key or other binary is reachable through it. Tests: `powertask::tests` (the parser refuses every +non-digit or extra argument, the arguments reach nvidia-smi as a list, the registration is per-user, highest, +trigger-less and quote-safe, a stale command file runs nothing). + +## Fleet priors from rented cards (6 October 2026, branch gpu-fleet) + +Measured by `tools/fleet/box-ember.sh` on Vast.ai containers, the 0.3.12 CUDA worker against the box's devnet node, 15 s settle and 60 s hold per step. `nvidia-smi -pl` and `-lgc` are refused inside the containers (the host's driver holds the knobs), so every ladder is its baseline step only: the untuned point per model, as `plan: baseline` records in `relay/lib/ember.mjs` terms (they summarise beside a prior, never set one). A tuned prior per model needs bare metal or a VM with the driver inside. + +| Card | Driver | MH/s | W | MH/W | Power default W | Clock max MHz | Plan | +|---|---|---|---|---|---|---|---| +| RTX 3060 | 580.126.20 | 24.59 | 106.1 | 0.2318 | 170.0 | 2100.0 | baseline | +| RTX 4060 Ti | | 17.64 | 73.5 | 0.24 | | | baseline | +| RTX 4070 | 580.159.04 | 24.77 | 91.3 | 0.2713 | 200.0 | 3120.0 | baseline | +| RTX 4090 | 595.71.05 | 52.24 | 179.9 | 0.2904 | 450.0 | 3105.0 | baseline | +| RTX A5000 | 580.82.09 | 47.6 | 222.3 | 0.2141 | 230.0 | 2100.0 | baseline | +| RTX 5090 | 580.173.02 | 98.12 | 258.2 | 0.38 | 575.0 | 3090.0 | baseline | +| RTX 5070 | 595.84 | 56.83 | 136.4 | 0.4166 | 250.0 | 3135.0 | baseline | + +The TUNE records themselves (`ember.json` per instance under `~/Desktop/fleet//`) carry the step, the mean core and memory clock and the hottest reading; `relay/lib/ember.mjs parseRecords` reads the `TUNE {...}` line each ladder prints. diff --git a/docs/plans/gpu-fleet.md b/docs/plans/gpu-fleet.md new file mode 100644 index 000000000..9f364b9a2 --- /dev/null +++ b/docs/plans/gpu-fleet.md @@ -0,0 +1,65 @@ +# The rented GPU fleet: the standing rule and the one-shot rule + +Josh's ruling, 6 October 2026, 19:50 UK ("keep rented cards up"), written as the fleet agent runs it. Tooling in +`tools/fleet/` on branch `gpu-fleet`; every box operation goes through `tools/fleet/lib/` (the box library) and the +box-side scripts it ships. + +## Two kinds of box + +| Kind | Rule | Examples | +|---|---|---| +| Standing | stays up, never destroyed on a job's end; replaced in the same shape when its host dies; node under a supervisor; version follows the live manifest; never joins an experiment | the 13 live-devnet boxes of 6 October (USD 3.50/h), the Devnet 2 set (seed, 3 miners, 2 provers) | +| One-shot | rented for one measurement, destroyed the minute its measurement is in (a running meter is a bug) | the memory-matrix cards, the 8x rigs, the wave pods, the RISC Zero box | + +The standing set's size and cost: 16 on the live devnet (a mix: 5090, 4090s, 3090s, 4070, 3080, A5000, an L4, at least +two AMD when Vast reopens) plus 6 on Devnet 2, about USD 8 an hour, USD 200 a day, inside the daily budget (USD 250 to +1,000). Tonight's 13 cost USD 3.50 an hour (USD 84 a day); the L4, the AMD pair and the Devnet 2 six are owed to the +roster as providers free cards (RunPod gave no pod of 8 card types from 18:59Z, Vast refuses every new rent). + +## What a standing box carries + +| Piece | Where | What it does | +|---|---|---| +| `box-standing.sh` (the supervisor) | `/root/fleet/in/`, started once by `lib/standing.py install` | every 60 s restarts the node, the miner loop (CUDA worker, pack re-exported on seed change) and the prover when gone; every 10 min writes a `RESULT standing` line (uptime, node pid and binary, blocks, DAA, peers, synced, exec tip, miner rate, prover) to `/root/fleet/out/standing.log`; runs the shipper's recovery recipe (`box-exec-snapshot.sh`) when consensus is synced above 1,000 blocks and the exec tip reads 0 | +| `/root/fleet/standing.node` | the box | the node binary the supervisor runs; the supervisor adopts whatever node already runs when the file is absent (a canary's newer binary is never downgraded) and restarts the node only when a publish rewrites the pointer | +| the registry row | `~/Desktop/fleet/boxes.json` | `standing: true`, `role: live|dn2`, `standing_since`, `node_sha16_wanted` after a publish | +| `lib/standing.py` | the Mac | `roster` (role, card, price, uptime), `install`, `check` (ssh alive, supervisor up, synced, exec moving, prover, node binary against the wanted sha), `update` (the manifest's hive package onto the box, pointer rewritten), `rerent` (same card and provider, label suffixed `-r`, old row marked destroyed), `loop` (check every 10 min, re-rent after two dead checks, report what is behind) | +| the fleet page | `dl.igneum.network/fleet-22adafa34bc2/` | a `standing` block (count, USD/h, USD/day, roles, the rule) and `standing`, `role`, `uptime_h` on every box row | + +A publish on the standing set is a script, not the loop: `tools/fleet/publish-0315.py` is the first (the Linux igneumd, +igneum-miner and the generator-4 workers over the package's, the pointer rewritten, the read-back table of version line, +digest and worker sha per box). The loop reports a box behind its wanted sha; it does not move binaries by itself. + +## Ports + +A standing box should carry a mapped p2p port whenever the provider allows it (RunPod: request `26611/tcp` at rent; +Vast: none of tonight's offers mapped one), so that the live devnet stops being a star in which every rented node dials +only the two hand nodes and the hub: the 18:39Z to 19:4xZ finality pause coincided with the hands being down, and run A +(10 blocks/s on a star, 77 to 84 percent red blocks) showed the same weakness on Devnet 2. Tonight's inbound-capable +nodes were igneum-build-1 (26611) and dn2-seed's one mapped port; a true mesh needs pods rented with a p2p port. + +## Clock + +- 19:41Z: the 13 live boxes converted (supervisor started, every node synced, every prover up). +- 20:00Z: the 10 percent rule, after the finality pause the fleet caused (`lib/standing.py weight_check`). +- 20:42Z to 20:48Z: the disk class (the prover's exports) found on the hub's third death; the supervisor prunes, `disk-sweep.py` watches, box-prover.py caps its own directory. +- 20:58Z: Devnet 2 standing at 1 block/s: the seed on igneum-build-1 (188.40.146.49:26611), miners dn2-seed, dn2-1, dn2-2, dn2-3, provers on dn2-1 and dn2-2 (the floor host copied from p1-4090); the gate owed until the rehearsal chain ends and dn2-seed's mapped port is free. +- 21:44Z to 21:53Z: publish 1 of 0.3.15 on the 14 live standing boxes (`publish-0315.py`: node 7f0bde70 = f1ea7a38, miner, generator-4 workers; the thirteen-field object, digest b18ed271). +- 22:18Z to 22:24Z: the forced poisoned-peer confirmation on the live network (relay on 6615571c over a 1026 datadir, a 713ef876 miner into it, the f1ea7a38 target refusing six relayed 1026 blocks with the old rule's line and never itself rejected). +- 22:49Z: publish 2 (the sixteen-field object: exec_restart_state_root, class v4 floor 831,600, window 86,400) on the 14 live standing boxes by `~/Desktop/fleet/move-16.py` (file to /root/fleet/override.json, the previous kept as override-13.json, node restarted by the supervisor, miner with it); expected digest eada4bda; `~/Desktop/fleet/expected-digest` rewritten so the five-minute check posts on any box left on b18ed271. The Devnet 2 five stay on their own object (4a0b8726). +- Owed: the L4 and the AMD pair when a provider reopens; the ports on re-rent. + +## The 10 percent rule (6 October 2026, 20:00Z, from the finality pause) + +What happened: the class v4 rehearsal job took the GPUs of 13 live-devnet miners between 18:27Z and 18:30Z (their +live nodes stayed up and synced, but a voter's weight is its blue blocks, and a miner that stops mining stops being +a voter), and with seven earlier leavers that was 42.7 percent of the frozen voter table; finality rule v3 then holds +the pause for one full window, and the first lock after 18:39:36Z is expected at about 20:40Z. The fleet's read that +"the boxes never left the devnet" was true of the nodes and wrong about the weight. + +The rule: never remove more than 10 percent of the live devnet's 30-day weight in any hour. Weight is counted by blue +blocks per key over the window, read from the hub; any experiment that borrows live miners does it in slices with an +hour between slices; a standing box's miner is never stopped (or its GPU shared with a second worker) by a job +without that check. `lib/standing.py weight_check ` is the gate: it reads the hub's blue blocks per payout key +over the window, sums the share of the labels asked for, and refuses the job when the share since the last hour's +removals exceeds 10 percent; every fleet job that touches a standing box's miner calls it first. diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md index 0c03025ad..bfbe00059 100644 --- a/docs/plans/proving-v1.md +++ b/docs/plans/proving-v1.md @@ -221,3 +221,34 @@ Josh, 5 October 2026: "fix everything else in the numbers tonight". Branch `agg- | The chosen combination | the defaults | the defaults stay: batch-log2 22 and SP1's default knobs. The one knob that moves the prover (2^16) costs a fifth of the hash rate all the time for a prover that is busy a few seconds a minute on the devnet; it is Josh's trade, not a default (below) | Reading. The per-block aggregation is 2.1 s and a block 4.1 s on a 5090 that only proves, 9.7 and 17.5 s on one that also mines; no knob, fold or stream on tonight's SP1 changes the first pair, and only the miner's kernel length changes the second, at 1 MH/s per 0.37 s of block time. So "under 3 s a block" and "under 1.5x" are met on a card that is not mining and are not reachable on one that is. What that means per tier: a 5090 that mines and proves delivers a proven empty block every 17.5 s (6 cards for 1 block/s), the same card proving only every 4.1 s (2 cards, plus the shard work of full blocks: the fleet table above), and a batch fold of the aggregator (a new pinned guest) would bring the proving-only card to about 2.7 s a block and the mining one to about 12 s. What is being done: the app and host defaults are left as measured; the plan's open decision for Josh is whether a card that holds a shard assignment should drop to 2^16 for the proof's minute (1.6x faster proof, 19% of its hash rate for that minute) or whether proving-only cards carry the aggregation (the clean 2.1 s), and the batch fold goes on the next pin's list. The state class found on the way (`/api/state` answering `{}` once `paid_wei` passes u64::MAX, fixed on the app branch at 6714a45) is in the bench log with the rest. + +## The fleet night (6 October 2026, from 11:50 UTC) + +The rented fleet (branch `gpu-fleet`, `tools/fleet/`, raw logs under `~/Desktop/fleet//`): 11 cards for the +memory matrix (`docs/analysis/prover-tiers-real-cards.md`), 15 extras for the prover night, a p2p hub on RunPod with its +port public, the proving agent's three RISC Zero boxes. Every box: the 0.3.12 Linux node 83089544 on the ten-field +override (digest 7bd98cc4...), the 0.3.12 workers, the patched SP1 server, the cuda host, a throwaway payout address +made on the box, a per-card key label. The prover is `tools/fleet/box-prover.py`, a port of `app/igneum-app/src/prover.rs` +(whole-segment claiming, FNV spread over the fleet, `--mode chain --save-shards [--prev]`, sign and submit per shard and +per segment, held fresh records offered again every pass). + +| Row | What was found | Per tier | What is being done | +|---|---|---|---| +| 1, 13:03Z: a node that joined the devnet today never executes the chain | On every fleet node `eth_blockNumber` reads 0x0 and `igneum_getProvingStatus` tipDaa 0, v1 active=false, an hour after consensus synced (blocks = headers, synced=true, blocks flowing). At exec debug the follower says `cannot find header edc4fa84... (genesis)` on a proof-synced node and `the queried hash does not have retention root on its chain` on a node started over a copy of the observer's full datadir, archival or not: the exec state is memory only (`igneum/exec/src/service.rs`: `_db_dir` unused, `IgneumDb::genesis()` at start, the follower walks the virtual chain from genesis), and once the devnet's pruning point left genesis (today) no node can start that walk. The observer's own exec read tipDaa 0 at 12:05Z (before the fleet touched it) and node 1's reads tipDaa 0, both restarted for publish 2 at about 11:37Z: the hand nodes' exec layer has been dead since, and with it every work list the provers read | Home miner (any card, any OS): a 0.3.12 install today mines but never proves, its wallet and the explorer against its own node read zero. Rig: the same. Pool user: nothing visible until the pool's own node restarts. The devnet: proving v1 produced nothing after the publish-2 restarts; the fleet night's "before the switch" window cannot exist on 0.3.12 | The coordinator's decision (13:15Z): no genesis restart; the proving agent builds the follower rebuild (walk the stored blocks where the full history is on disk, persist the exec state), tested on a copy of node 1's datadir tonight, then the hand nodes swap binaries and the PCs take a node-only 0.3.13. The fleet keeps the hub and the observer's full-history tarball (`/root/fleet/share/observer-datadir.tgz` on the hub, sha256 e67cc649...) for the phase-2 nodes, keeps the fifteen extras building until 14:15Z, and runs phase 1 and phase 3 meanwhile | +| 2, 12:24Z: the seed drops every new node every 30 s | `P2P, route error: incoming route capacity for message type IgneumFinality has been reached (peer: 188.245.5.161:26611)` then `P2P Connected to outgoing peer 188.245.5.161:26611` at :06, :36, :06 on every fleet node; IBD through the headers proof restarts at each drop, so 4 of 11 nodes had 0 blocks after 35 minutes while the seed served 26 nodes at once; `IBD with peer 188.245.5.161:26611 completed with error: peer connection is closed` | Every joiner with the seed as its only peer syncs in pieces; a rig the same once; pools unaffected | The hub (a RunPod 4090 with 26611 public) is every fleet node's second peer; the hub itself drew 23 fleet peers within minutes through the seed's address exchange. For the node: the IgneumFinality route's capacity against the per-checkpoint burst | +| 3, 12:38Z: the patched server hangs instead of failing when a profile does not fit | The 3080 (10 GB) and the 4060 Ti 8 GB at 2^27 (a 10.3 GB allocation): the server holds the card's limit at 0% for 568 and 904 s until killed; patch v5 (prover-floor) turns it into `FLOOR abort: a device allocation failed at slop/crates/tensor/src/inner.rs:51 ... AllocError { size: 486586112 }` and exit 70 in 13 s on the 3080 (the known-failed case of its gate) | Every prover run needs a wall-clock timeout (the rig unit has one; the app's prover and this fleet's loop have one) | v5 is the server the fleet ships from here | +| 4, 15:37Z to 15:41Z: the 0.3.13 swap on the fleet | The hands switched to the 0.3.13 node (bb43e9a8) with the thirteen-field override (`exec_restart_number` 27276, `exec_restart_hash` bb45cf0d..., `exec_restart_trust_daa` 200000; digest b18ed271...) at 15:29Z; the fleet's boxes followed in two steps (`tools/fleet/box-node-swap.sh`, `swap.py`): the binary with the ten-field file (digest 7bd98cc4 kept, the exec layer blocked by design), then the file. Replay from the restart to an executed tip at the chain tip, per box: hub (RunPod 4090) 88 s, 3080 87 s, 4070-1 113 s, 4090-3 138 s, the 8x 4090 rig 138 s, A5000 163 s, 3090-4 188 s; a second set 15:45Z to 15:57Z (3090, 5090, 3090-1, 3090-2, 3090-3, 4090-1b) about 10 to 12 min each including a 438 MB datadir pull. Seven boxes failed the first file pass on a partial pre-pull of the full-history tarball (the swap now checks its sha256) | A joiner on 0.3.13 with a full-history datadir executes within 1.5 to 3 minutes; a joiner without one still cannot (the fresh-join path, a snapshot from a peer, is the next item); every tier | The swap and the replay times are the fleet's measurement for the 0.3.13 release note | +| 5, 15:42:57Z: DAA 198,000, the fresh-record rule armed; 15:43:42Z: a 229-block reorg reset every executing node | The hub's first `PoW accepted ... daa 198000` at 15:42:57Z. One-block selected-chain reorgs at 15:42:32, :43, :53 and 15:43:05Z (heights 135,065 to 135,088), then at 15:43:42Z `selected-chain reorg: 229 chain blocks removed, unwinding to height 134884` (our tip blue score 194,395, the last removed 194,117), `reorg deeper than the snapshot ring; replaying from genesis`, genesis executed, then `exec not synced: the executor is at chain block 0 and the bodies below this node's retention root are gone`: executedTip 0, persistedTip 135,028, blocked null. The same on every fleet box that was executing and on the observer and the seed (the shipper's reading). The chain ran two-sided for about a minute after the switch | Every node operator whose node was executing at 15:43Z (home miner, rig, pool) read a zero wallet and an empty work list until a restart through the exec-restart path; the miners of the 229 losing blocks lost those rewards; a deep reorg after a snapshot ring on a pruned node is the class: the fallback must be the exec-restart point, not genesis (the proving agent's item) | The fleet restarted its boxes through the exec-restart path (90 to 190 s each, the night loop `tools/fleet/night.py` re-runs it on any box whose executed tip falls to 0 for 150 s) and the provers started on the executed tip from 15:56Z | +| 6, 16:38:45Z: the first segment records after the fresh-rule switch, accepted and paid | The fleet's RTX 5090 (Vast, a restart-path 0.3.13 node, the segment host from the proving-v1 bundle, the prover loop `tools/fleet/box-prover.py` exporting from the exec restart block) claimed segment 137142..137149 at 16:35:30Z, cut 8 fixtures, ran `--mode chain --save-shards` (8 shards and 8 aggregations on a card that also mines), had 8 of 8 shard records accepted and the fresh segment record accepted at 16:38:45Z: 195.8 s claim to acceptance, aggregator share 1.5076 IGN; a 4090 (137102..137109) followed at 16:39:51Z in 211.2 s, 1.6150 IGN. The hub, a node on the hands' real state, read paidSegments 2, paidSegmentWei 3.12 IGN, pool 2 entries 2 verified, paidShards 1,482 to 1,516 at 16:42Z: carried and paid | A 24 or 32 GB card that mines earns a segment's aggregator share about every 200 s on top of its 8 shards' 90%; six such cards cover about a quarter of the chain's segments (1.8 of 7.5 a minute); 47 mining 24 GB cards or 6 proving-only ones cover it (the fleet table's arithmetic, now with a measured 200 s) | The six restart-path boxes prove through the night; the hourly rows carry segments per hour and the chain's coverage | +| 7, 16:50Z: one exec state, and the export gap on a snapshot-recovered node | The state roots at 27,276 (0xed27bb2d...), 130,272 (0xf0a762da...) and 130,273 (0x2e22e029...) are identical on a restart-path node and on a snapshot-recovered node (export segment root and eth_getBlockByNumber agree), so the two recovery paths of 16:00Z to 16:25Z are one chain state and records from either side are valid on the other (row 6 confirms it). But on a snapshot-recovered node every `igneum_exportSegments` (from 0, 27,276 or 130,272) carries 41 to 42 accounts, the exporter's port state root differs from the node's at the first segment, and from 0 the 27,276 segments below the restart carry zero roots: the prover kit cannot cut against the hands' kind of node. The fleet's first report of this (16:52Z) called it two states; the roots corrected it 10 minutes later | An operator whose node recovered through the snapshot cannot run a prover until the export carries the state (the 0.3.14 account-dump export, the proving agent's item); a node that joined through the exec-restart path proves today | The nine snapshot boxes mine without provers tonight (each prover loop had pulled a 100 MB export every 6 s for nothing); the roots above are the 0.3.14 pin's numbers | +| 8, 17:45Z to 18:45Z: finality at 78 to 93 voters (the 38-pod wave, 1,748 MH/s for USD 20.44/h) | 116 checkpoints determined, 110 locked in the hour; lock delay p50 1.30 s, p90 1.54 s, max 16.6 s (n 107); the first certificate carries p50 59 voters (max 80) and the fold lifts it to 86 of 93; lock share p50 90.1% of active voters (min 76.9%), 77.2% of all (min 67.0%); 1.65 certificate replacements per checkpoint (max 5). The detector's dry run on the window: no alert, no event, the 5090 band 108/115/128 MH/s (n 1,555). The first hour of the wave (16:50Z to 17:45Z) is void: `pgrep -f igneumd-0313` in box-wave.sh matched its own launching shell, so no wave node started for 55 minutes (about USD 20 of pods); the fix is the anchored pattern and the CI check `tools/ci/pgrep-self-match-check.sh` | Home miner: a lock lands about every second block at 1 block/s, 0.2 s later at 93 voters than at 56; what to watch past 100 voters is the certificate's size (the fold to 86 signatures), not the vote count. Rig: the same. Pool: a pool's one node votes once for all its members; its weight is the day's blue blocks, so a pool of 38 pods is one voter with 40% of the weight | the rows are the fleet night's finality baseline for the 100-voter question; the certificate-size series continues in the hour marks | +| 9, 18:14Z: pool-v0 under load, 10 members, 0 shares | igneum-pool (PPLNS, port 4463) with 10 wave pods as members: every member's GPU shares were refused `WORKER MISMATCH` because the pool's job carried no program class and the member's worker hashed class v2 while the devnet's templates are class v3; 0 shares accepted in 25 minutes, the pool's stats page alive. The pool agent rebased the pool-mode miner onto 0.3.14 (pool-v0-rebase c4c92e8: the job carries program_class and era_seed, the member re-checks shares with the template's class); the rerun with both new binaries is scheduled for tonight after the block-rate runs | Pool user: on the old pool binary every share is wasted work; on the rebased pair a member without a node takes the day and the dataset size from the pool's seeds line. Home miner and rig: unaffected (solo mining never touched the pool) | the rerun's counts (shares accepted per member inside the first vardiff interval, mismatches 0, rejected 0) go to main | +| 10, 17:25Z to 17:38Z: the 0.3.14 canary on the live devnet | the release's Linux igneumd (4c6b129d) on five restart-path boxes for ten minutes: 0 new rejected blocks on every box, exec state roots equal to the hub's at a common height, a segment record paid on the new binary, every node on the new version: PASS on all four criteria; the first run's FAIL was the gate counting the hub (whose trait is about 3 rejects a minute) and a `cut -c1-140` that truncated the digest line before the match | Operator on the restart path: the swap is a binary change with the data dir kept, seconds of downtime, no replay. Snapshot-path operator: the same binary, the snapshot recipe unchanged | the canary form (canary-next.sh) is the fleet's release gate; the 0.3.15 run is in flight on four boxes at 19:41Z | +| 11, 18:31Z to 19:35Z: the class v4 rehearsal on igneum-devnet-400 (16 nodes, 38 pods joining, one stale 0.3.13 box) | The flip by miner signal landed at DAA 1,200 (19:00:5xZ) with the identical line on every node ("epoch 2 ... seed block f0606e20..., threshold 9500 bps, 599 of 599 blue blocks"), one program id per epoch across boxes (epoch 2 0x24304f0788ea9408, epoch 3 0xcc266b4f5dbc3447, epoch 4 0x634018bab5e5f283; the Mac's igneum-pow show agrees and the v3 id differs), 0 PoW rejected all run, the stale box refused by digest at every connect (12 by 19:46Z) and refusing the sixteen-field file at parse ("unknown field program_class_v4_activation_daa"), the floor at 2,400 crossed with v4 in force and nothing to print. Two findings outside the protocol: the plan's CPU engine (5 kH/s a box) cannot move a chain whose genesis difficulty is 2^27 (the CUDA worker fixed it: 14 to 98 MH/s a box), and the 0.3.14 package's igneum-worker-cuda refuses a generator-4 pack ("not a generator version this worker runs (2 or 3)"), so the chain stood at DAA 1,202 from 19:00:5xZ until the release tree's generator-4 Linux worker (97e036e2) went on at 19:15Z; a worker restarted on its old --pack after --exit-on-seed-change also answers every job "epoch seed mismatch" until the pack is re-exported | Home miner, any card: at a class flip the old worker stops dead; the flip is safe only when the generator-4 worker ships in the hive, Mac and Windows packages before the signal window closes, and the app re-exports the pack at every seed change. Rig: the same, times eight. Pool: the pool's node decides the class; a member on an old worker mismatches every share | the cut's rule (publish 2 gated on every worker, not every node) is with the shipper; P1 PASSED is on master | +| 12, 18:39:40Z onward: the live devnet's finality paused at 93 voters | On every node read the last lock is checkpoint 6842 at 18:39:36Z; checkpoints keep being determined and no certificate is received or built anywhere after 18:39:40Z. The fleet's first reading (the hands down 18:2xZ to 19:42Z, the star topology) was wrong: the observer rows show 13 fleet keys stopped mining the live devnet between 18:27Z and 18:30Z when the rehearsal job took their GPUs (their live nodes stayed up and synced, but a voter's weight is its blue blocks), and with seven earlier leavers that was 42.7 percent of the frozen voter table; rule v3 then holds the pause for one full window, the first lock expected about 20:40Z | Home miner: a voter that stops mining stops counting within the window, and 10 percent of the weight leaving in an hour is the most the table absorbs without a pause. Rig and pool: a pool is one voter with its members' whole weight; its restart is the biggest single removal on the network | the standing-fleet rule from it (6 October 2026, 20:00Z): never remove more than 10 percent of the live devnet's 30-day weight in any hour; `lib/standing.py weight_check` gates every job that stops or shares a standing miner | +| 13, 19:14Z to 20:30Z: the 0.3.15 canary, FAIL on the first binary, the retry confounded | 713ef876 on four live boxes: every block a 0.3.15 node mined or relayed carried version 1026 (the class v4 signal bit stamped from its first block, no window set) and every 0.3.14 node answered "wrong block version: got 1026 but expected 2" and disconnected it (the hub: 45 such rejects in the first 14 minutes, 468 by 20:29Z), so a 0.3.15 node that fell behind could not re-sync (p2-3090-1: connected and dropped every 30 s for 32 minutes) and every 0.3.15 miner lost every block it found; FAIL, publish 1 held. Three more findings on the way back: (a) a 0.3.14 node whose datadir holds 1026 blocks keeps relaying them and stays a disconnected peer after the rollback, so the four canary datadirs are poisoned until wiped; (b) a pruned 0.3.14 node dies ("consensus/src/processes/sync/mod.rs:87 KeyNotFound(GhostdagCompact/0/)") when a peer syncing a gap asks below its retention (the hub twice, 19:58Z and 20:01Z, restarted by the standing supervisor); (c) a standalone igneum-miner keeps a dead template subscription after its node restarts (templates frozen, fetch_errors climbing, no submits). The shipper's 7961c5f1 gates the stamp on publish 2's object and fixes the serving side of (b); its retry on the poisoned boxes was confounded by (a) and (c), the clean retry runs on two untouched boxes | Home miner: an update that stamps a new block version before the network accepts it is a silent death (the app shows hashing, nothing is paid); the fix is the version gate on the object plus a datadir that never held a bad block. Rig: the same, times eight. Pool: a pool node on the bad version drops every member's share from the network's view | the canary form stays the release gate; the fresh-join line from a wiped datadir is the last read | +| 14, 20:00Z: the standing fleet | Josh's ruling (19:50 UK): rented cards stay up and are never destroyed on a job's end. 13 live-devnet boxes converted at 19:41Z (USD 3.50/h, USD 84/day), each under box-standing.sh (node, miner and prover restarted when gone, the recovery recipe on a dead exec, a status line every 10 min), lib/standing.py on the Mac (roster, check, update, re-rent in the same shape, the 10 percent weight gate), a standing block on the fleet page; the Devnet 2 six and the L4 and AMD cards owed as providers free them. The supervisor's own two faults tonight (it matched any igneumd, so it mistook the rehearsal node for the live one and restarted a canary box's dead node on the 0.3.15 file) are fixed in db58804 and 9a294b4 | Operator of a standing box: the node comes back within a minute of dying, the miner with it, and no job takes its GPU without the weight gate | docs/plans/gpu-fleet.md carries the rule | +| 15, 19:58Z, 20:01Z, 20:42Z: the hub's live node died three times | Twice on a peer's sync request below its retention ("consensus/src/processes/sync/mod.rs:87 KeyNotFound(GhostdagCompact/0/)", while the rolled-back p2-3090-1 synced a 32-minute gap against it; the serving-side fix is in the 0.3.15 node), once on a full disk ("header_processor/processor.rs:534 IO error: No space left on device"): the prover's segment exports under /root/fleet/out/segs (50 to 500 MB a segment, never pruned by box-prover.py) had filled the hub's 60 GB and 3 to 46 GB on every standing box since 11:50Z; two boxes stood at 100 percent. The standing supervisor restarted the node each time (19:59:11Z, 20:02:36Z, 20:44:38Z; the third after 38 GB were freed by hand) and now prunes exports older than 20 minutes and trims the node log every ten minutes (353cc5f); a disk sweep every five minutes writes disk per box to the fleet page and posts one #incidents line per box per hour at 85 percent (disk-sweep.py); the exporter-side cap is the proving lane's | Home miner: the hub is one of three public peers a fresh node dials; a dead hub means a slower first join and nothing lost; a home node's own disk is not at risk (the app's prover does not write exports). Rig: the same. Pool: a pool node that exports segments for its provers has the same disk clock | the three deaths' lines are in the hub's node.log (preserved on the box) and the finality pulls under the scratchpad | +| 16, 19:17Z to 20:58Z: the block rate on Devnet 2, 10 blocks a second against 1 | Run A (the fork's 10 blocks/s profile, fresh genesis, 42 cards, the seed on igneum-build-1, every miner dialling the seed only): 4.87 DAG blocks/s but 1.09 blue blocks/s, 77.6 percent red, tips 250 to 660, difficulty easing all hour, the exec follower at 0.05 blocks/s. Run B (1 block/s, same boxes): 1.0 blocks/s and under 2 percent red from minute six, tips 1 to 3, difficulty settled in six minutes, the follower at 0.46 blocks/s. The network lane's read: the reds came from node throughput (61 to 345 ms of CPU per accepted block), not the star | Home miner: the payout interval follows the blue rate, which the profile did not move (1.09 against 1.19 blue/s), so a 4070 at a 10 TH/s network waits about three days for a paying block at either rate; the pool, not the block rate, is the small card's shorter wait. Rig: 4 to 5 hours at 10 TH/s either way. Everyone: a wallet or prover on a 10 blocks/s chain would read state hours behind within the first hour at tonight's follower rate | `docs/analysis/block-rate-devnet2.md`: 1 block/s for the testnet and the launch, 10 behind three measured gates | +| 17, 21:52Z: two miners on one GPU after a restart | The read-back after publish 1 showed miner_up=2 on p1-4090, p1-a5000 and p2-4090-3: the publish killed the miner once, the supervisor's miner loop restarted it within ten seconds, and the supervisor's main loop, reading "no miner" in the same gap, started a second loop; two igneum-miner processes then shared one GPU at half rate each. The supervisor's miner loop now kills any other miner and worker before it starts its own (409bc3d); redeployed on the 14 standing boxes at 21:55Z | Home miner: the app owns one miner per GPU and never sees this; a hand-run box with two miner loops halves its rate silently and the only sign is the STATUS line's MH/s. Rig: times eight. Pool: a member with two miners doubles its share submissions at half the rate each, the pool sees one member with a jittery rate | one miner per GPU is now the supervisor's invariant, not the operator's care | +| 18, 21:43Z to 21:56Z: two provers on one Devnet 2 box killed each other's GPU server | Repeated by-hand relaunches of box-prover.py on dn2-1 and dn2-2 left two instances on a box: each relaunch's "kill the server, remove the socket" took the other instance's SP1 GPU server away mid-proof ("CudaClientError: early eof" on every chain step from 21:43Z) and both wrote the same prover-state.json.tmp, so one lost the file (FileNotFoundError at the os.replace). box-prover.py now holds a pid file under its out directory and a second instance exits at once, and each process writes its own tmp (c29c6b9); the Devnet 2 provers relaunched clean at 21:54Z | Home miner: the app's prover is one process by construction. Operator of a hand-run prover box: start the prover once; a second start now refuses with the first's pid instead of taking its GPU server down | the Devnet 2 gate's "paid segments" read waits on the first record from these provers | diff --git a/tools/ci/README.md b/tools/ci/README.md new file mode 100644 index 000000000..a2bf50e4f --- /dev/null +++ b/tools/ci/README.md @@ -0,0 +1,7 @@ +# CI checks + +| Check | What it fails | Since | +|---|---|---| + +| `pgrep-self-match-check.sh` | a `pgrep -f` / `pkill -f` with a bare literal pattern, or `ps \| grep ` without a bracket or `grep -v grep`, in tools/, relay/playbooks/, infra/ or packaging/: the pattern matches the shell that runs it (the wave script of 6 October 2026 never started a node on 38 cards because `pgrep -f igneumd-0313` saw the launching shell; a kill file killed its caller the same day). Anchor to the executable's path, bracket the first letter, or use `-x`. Owed (allow-listed, finished measurements): `tools/prover-floor/pc2-*.ps1`, `tools/proving-v1/pc2-*.ps1` (their `pkill -f sp1-gpu-server` becomes `pkill -x`), `tools/repo/fresh-repo.sh:223`, `tools/observer/autosync.sh:22`, and the live devnet's operational scripts the fleet agent does not own: `infra/devnet/restart-hand-nodes.sh`, `infra/seed-nodes/addpeer-from-mac.sh`, `relay/playbooks/shard-test.ps1`, `tools/ci/fixtures/bash-body-ok.ps1`, `tools/ci/pgrep-self-match-check.sh`, `tools/ci/prover-socket-check.sh`, `tools/exec-attacks/net.sh` (their owners anchor the pattern when next touched; the hand-node and seed scripts run tonight and were not edited blind) | 6 October 2026, branch gpu-fleet | +| kill by exact command or pid file (owed as a check) | 6 October 2026, 21:09Z: a Mac-side `pkill -f ` matched nothing (the log name was a redirect, not part of the command line), the roll-everything script lived on and wiped a box it had been told to hold. Rule: a job is stopped by its pid file (`tools/fleet/fleet-bg.sh start|stop `) or by a pattern anchored on its exact command line (`^python3 -u /root/fleet/in/box-prover.py`), never by a word that may or may not appear in it. The check that flags a `pkill -f`/`pgrep -f` whose literal is a path or a name that never starts a command line is owed to the CI lane | diff --git a/tools/ci/pgrep-self-match-check.sh b/tools/ci/pgrep-self-match-check.sh new file mode 100755 index 000000000..3a081910d --- /dev/null +++ b/tools/ci/pgrep-self-match-check.sh @@ -0,0 +1,53 @@ +#!/usr/bin/env bash +# The self-matching process-pattern class (6 October 2026). Three times in one day a script matched its own shell: +# the shipper's recovery at 16:1xZ, the fleet's wave script at 16:25Z (`pgrep -f igneumd-0313 || start the node` matched +# the launching shell's command line, which carried the file name, so no wave pod ever started its node and 38 cards +# hashed against nothing for an hour), and the fleet's Devnet 2 kill step at 17:18Z (`pkill -f '^bash in/box-dn2.sh'` +# inside a file the same script called killed the caller). Rule: a `pgrep -f`, `pkill -f` or `ps ... | grep` whose +# pattern is a literal word matches every process whose command line carries that word, including the shell that +# runs the pattern and any ssh command that carries the script's text; the pattern must therefore exclude itself: +# anchored to the executable's path (`'^/opt/igneum/pkg/bin/igneumd'`), the bracket form (`'[i]gneumd'`), or +# `pgrep -x ` / `pkill -x ` on the binary name (15 characters at most). This check fails CI when a script +# under tools/, relay/playbooks/, infra/ or packaging/ runs pgrep -f / pkill -f with a bare literal pattern (no `^`, +# no bracket, no `$`), or pipes `ps` into `grep ` without a bracket or a `grep -v grep`. +# With file arguments it checks those files only; --self-test runs the two fixtures. +set -euo pipefail +cd "$(dirname "$0")/../.." +fail=0 +bad_pattern() { # the pattern text between the quotes after -f; prints 1 when it is a bare literal + local p="$1" + [[ "$p" == ^* || "$p" == *'['* || "$p" == *'$' || "$p" == '$'* ]] && return 1 + return 0 +} +check_file() { + local f="$1" n=0 + while IFS= read -r line; do + n=$((n + 1)) + [[ "$line" =~ ^[[:space:]]*# ]] && continue + # pgrep -f / pkill -f with a quoted or bare pattern + while read -r pat; do + [ -z "$pat" ] && continue + if bad_pattern "$pat"; then echo "pgrep-self-match: $f:$n: p(grep|kill) -f with the bare pattern '$pat' matches the shell that runs it; anchor it (^/path), bracket it ([x]rest) or use -x"; fail=1; fi + done < <(printf '%s\n' "$line" | grep -oE "p(grep|kill)( -[0-9A-Za-z]+)* -f(a|c|l)? +(\"[^\"]*\"|'[^']*'|[^ |;)]+)" | sed -E "s/^p(grep|kill)( -[0-9A-Za-z]+)* -f[acl]* +//; s/^[\"']//; s/[\"']$//") + # ps | grep word + if printf '%s\n' "$line" | grep -qE 'ps [^|]*\| *grep ' && ! printf '%s\n' "$line" | grep -qE "grep +(-[a-zA-Z]+ +)*['\"]?\[" && ! printf '%s\n' "$line" | grep -q 'grep -v grep'; then + echo "pgrep-self-match: $f:$n: ps | grep without a bracket pattern or 'grep -v grep' matches the grep itself"; fail=1 + fi + done < "$f" +} +if [ "${1:-}" = "--self-test" ]; then + t="$(mktemp -d)" + printf 'pgrep -f igneumd-0313 >/dev/null || start\npkill -f "bash in/box-x.sh"\nps aux | grep igneumd\n' > "$t/bad.sh" + printf "pgrep -f '^/opt/igneum/pkg/bin/igneumd' || start\npkill -x igneum-miner\npkill -f '[i]gneumd-0313'\nps aux | grep '[i]gneumd'\nps -eo cmd | grep igneumd | grep -v grep\n" > "$t/good.sh" + fail=0; check_file "$t/bad.sh"; [ "$fail" = 1 ] || { echo "pgrep-self-match: self-test FAILED: the bad fixture passed"; exit 1; } + fail=0; check_file "$t/good.sh"; [ "$fail" = 0 ] || { echo "pgrep-self-match: self-test FAILED: the good fixture was flagged"; exit 1; } + echo "pgrep-self-match: self-test ok (the bad fixture fails, the good one passes)"; exit 0 +fi +# Owed, not exempt: the PC 2 playbooks of 5 and 6 October use `pkill -f sp1-gpu-server` (the prover-socket check's own +# required line) inside a WSL `bash -c` whose command line carries the word, so the pkill kills that shell too when it +# runs first; they are finished measurements and get `pkill -x sp1-gpu-server` when next touched (tools/ci/README.md). +ALLOW='^(tools/prover-floor/pc2-.*\.ps1|tools/proving-v1/pc2-.*\.ps1|tools/repo/fresh-repo\.sh|tools/observer/autosync\.sh|infra/devnet/restart\-hand\-nodes\.sh|infra/seed\-nodes/addpeer\-from\-mac\.sh|relay/playbooks/shard\-test\.ps1|tools/ci/fixtures/bash\-body\-ok\.ps1|tools/ci/pgrep\-self\-match\-check\.sh|tools/ci/prover\-socket\-check\.sh|tools/exec\-attacks/net\.sh)$' +list_files() { if [ $# -gt 0 ]; then printf '%s\n' "$@"; else git ls-files 'tools/**' 'relay/playbooks/**' 'infra/**' 'packaging/**' | grep -E '\.(sh|bash|ps1|mjs|py)$'; fi; } +while IFS= read -r f; do [ -f "$f" ] || continue; [[ "$f" =~ $ALLOW ]] && continue; check_file "$f"; done < <(list_files "$@") +[ "$fail" = 0 ] && echo "pgrep-self-match: no script matches its own shell" +exit $fail diff --git a/tools/fleet/autorun.py b/tools/fleet/autorun.py new file mode 100644 index 000000000..fdbbc0ad2 --- /dev/null +++ b/tools/fleet/autorun.py @@ -0,0 +1,61 @@ +#!/usr/bin/env python3 +"""The fleet's orchestrator loop (runs on the Mac in the background, one pass a minute): advances every live box +through its stages without a human, pulls the results of each finished stage into ~/Desktop/fleet//, runs the +collector after a matrix or an Ember ladder lands, and refreshes the fleet page every 5 minutes. + + phase 1: setup_done -> box-matrix.sh -> matrix_done -> box-ember.sh -> ember_done -> box-prover.sh (joins phase 2) + phase 2: setup_done -> box-prover.sh + setup_failed: retried once (the toolchain CDN class), then marked failed and left for a human +Writes ~/Desktop/fleet/autorun.log; every change is a line there and in the page log. +""" +import json, os, sys, time, subprocess, datetime +sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))) +import fleet +ROOT = fleet.ROOT; LOG = os.path.join(ROOT, "autorun.log") +def now(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") +def log(s, page=False): + line = f"{now()} {s}"; open(LOG, "a").write(line + "\n"); print(line, flush=True) + if page: subprocess.run([sys.executable, os.path.join(fleet.HERE, "page.py"), "log", s]) +def probe(b): + rc, out, err = fleet.ssh(b, "tail -1 /root/fleet/setup.log 2>/dev/null | grep -o '^RESULT setup_[a-z]*'; grep -c . /root/fleet/out/matrix.log 2>/dev/null; grep -o '^RESULT matrix_[a-z]*' /root/fleet/out/matrix.log 2>/dev/null | tail -1; grep -c . /root/fleet/out/ember.log 2>/dev/null; grep -o '^RESULT ember_done' /root/fleet/out/ember.log 2>/dev/null | tail -1; pgrep -c -f '^(bash in/box-|python3 -u /root/fleet/in/box-prover)'; grep -E '^RESULT (point|miner|step|choice|claim|submitted|paid|segment_refused|seg [0-9]+ (chain|shards))' /root/fleet/out/matrix.log /root/fleet/out/ember.log /root/fleet/out/prover.log 2>/dev/null | tail -1 | cut -c1-200", timeout=40) + if rc != 0 and not out.strip(): return None + l = out.split("\n") + g = lambda i: l[i].strip() if i < len(l) else "" + return {"setup": g(0), "matrix_lines": g(1), "matrix": g(2), "ember_lines": g(3), "ember": g(4), "running": g(5), "last": g(6)} +def start(b, script, env=""): + fleet.run(script, [b["label"]]) if not env else fleet.run_env(script, [b["label"]], env) +last_publish = 0; retried = set() +while True: + reg = fleet.load(); changed = False + live = [(iid, b) for iid, b in reg.items() if b.get("state") not in ("destroyed", "failed") and b.get("ssh_ok") and b.get("phase") != "5"] + from concurrent.futures import ThreadPoolExecutor + with ThreadPoolExecutor(max_workers=16) as ex: probes = dict(zip([i for i, _ in live], ex.map(lambda x: probe(x[1]), live))) + for iid, b in live: + p = probes.get(iid) + if p is None: continue + stage = b.get("stage", "setup") + if p["last"]: fleet.patch(iid, last_line=p["last"]) + if p["running"] not in ("0", ""): # a stage script is running on the box: never start another (the 5070 ran two matrices at once, 12:23Z) + if stage == "setup" and p["setup"] == "RESULT setup_done": fleet.patch(iid, stage="matrix" if b["phase"] == "1" else "prover", state="running") + continue + if stage == "setup": + if p["setup"] == "RESULT setup_done": + nxt = "matrix" if b["phase"] == "1" else "prover" + fleet.patch(iid, stage=nxt, state="running", doing=("phase 1 matrix: idle, miner, stock server, patched server alone and beside the miner" if nxt == "matrix" else "phase 2: node + miner + segment prover on the devnet")) + fleet.run("box-matrix.sh" if nxt == "matrix" else "box-prover.sh", [b["label"]]); log(f"{b['label']}: setup done, {nxt} started", page=True); changed = True + elif p["setup"] == "RESULT setup_failed": + if iid not in retried: retried.add(iid); fleet.setup([b["label"]]); log(f"{b['label']}: setup failed once, retried") + else: fleet.patch(iid, state="failed", doing="setup failed twice; see setup.log"); log(f"{b['label']}: setup failed twice", page=True) + elif stage == "matrix": + if p["matrix"] == "RESULT matrix_done": + fleet.pull([b["label"]]); subprocess.run([sys.executable, os.path.join(fleet.HERE, "collect.py")], capture_output=True) + fleet.patch(iid, stage="ember", doing="Ember two-knob ladder: 6 power steps, 4 clock caps, 75 s each"); fleet.run("box-ember.sh", [b["label"]]); log(f"{b['label']}: matrix done ({p['last'][:80]}), Ember ladder started", page=True); changed = True + elif p["matrix"] == "RESULT matrix_failed": + fleet.pull([b["label"]]); fleet.patch(iid, stage="ember", doing="matrix failed (see matrix.log); Ember ladder started"); fleet.run("box-ember.sh", [b["label"]]); log(f"{b['label']}: matrix FAILED ({p['last'][:80]}), Ember started", page=True) + elif stage == "ember": + if p["ember"] == "RESULT ember_done" or (p["ember_lines"].isdigit() and int(p["ember_lines"]) > 0 and "ember_failed" in p["last"]): + fleet.pull([b["label"]]); subprocess.run([sys.executable, os.path.join(fleet.HERE, "collect.py")], capture_output=True) + fleet.patch(iid, stage="prover", phase="2", doing="phase 2: node + miner + segment prover on the devnet"); fleet.run("box-prover.sh", [b["label"]]); log(f"{b['label']}: Ember done ({p['last'][:80]}), prover started", page=True); changed = True + if time.time() - last_publish > 300 or changed: + subprocess.run([sys.executable, os.path.join(fleet.HERE, "page.py"), "publish"], capture_output=True); last_publish = time.time() + time.sleep(60) diff --git a/tools/fleet/box-datadir.sh b/tools/fleet/box-datadir.sh new file mode 100755 index 000000000..ee6dab95f --- /dev/null +++ b/tools/fleet/box-datadir.sh @@ -0,0 +1,38 @@ +#!/usr/bin/env bash +# The joiner stop-gap of 6 October 2026 (docs/analysis/prover-tiers-real-cards.md, the fleet night): a node that synced +# through the headers proof holds no genesis header, so the in-memory exec state never rebuilds and eth_blockNumber +# stays 0. This swaps the box's consensus datadir for a copy of the observer's full-history datadir (pulled from the +# hub over the throwaway fleet-internal key), restarts the node with the hub and the seed as peers and the proof +# verifier, and reports the follower's replay: eth_blockNumber every 15 s until it reaches the node's DAA tip. +set -uo pipefail +F=/root/fleet; OUT=$F/out; B=/opt/igneum/pkg/bin; FLOOR=/opt/igneum-floor; HOST=$FLOOR/bin/igneum-prove-host +HUB_SSH="${HUB_SSH:-213.173.107.74}"; HUB_PORT="${HUB_PORT:-16515}"; HUB_PEER="${HUB_PEER:-213.173.107.74:16516}" +mkdir -p $OUT; exec >> $OUT/datadir.log 2>&1 +stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; } +echo "RESULT datadir_start $(stamp)" +t0=$(date +%s) +if [ ! -s $F/observer-datadir.tgz ]; then + scp -i $F/in/fleet-internal -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -P "$HUB_PORT" "root@$HUB_SSH:/root/fleet/share/observer-datadir.tgz" $F/observer-datadir.tgz || { echo "RESULT datadir_failed pull"; exit 2; } +fi +echo "RESULT datadir_pulled $(stamp) bytes=$(stat -c %s $F/observer-datadir.tgz) s=$(( $(date +%s) - t0 ))" +bash $F/in/box-kill.sh >/dev/null 2>&1 # every stage process; never the node (next line) +pkill -x igneumd; sleep 4; pkill -9 -x igneumd 2>/dev/null; sleep 1 +rm -rf $F/node.proof && mv $F/node $F/node.proof && mkdir -p $F/node && tar -C $F/node -xzf $F/observer-datadir.tgz || { echo "RESULT datadir_failed untar"; exit 2; } +echo "RESULT datadir_swapped $(stamp) $(du -sh $F/node | cut -f1)" +IGNEUM_PROOF_VERIFIER=$HOST nohup $B/igneumd --devnet --appdir=$F/node --rpclisten=127.0.0.1:26610 --evm-rpclisten=127.0.0.1:26790 --listen=0.0.0.0:26611 \ + --addpeer=$HUB_PEER --addpeer=188.245.5.161:26611 --override-params-file=$F/override.json --nodnsseed --disable-upnp --nologfiles --yes >> $F/node.log 2>&1 & +t1=$(date +%s) +bn() { curl -s -m 8 -X POST -H 'Content-Type: application/json' --data '{"jsonrpc":"2.0","id":1,"method":"eth_blockNumber","params":[]}' http://127.0.0.1:26790/ | grep -o '"result":"[^"]*"' | cut -d'"' -f4; } +for i in $(seq 1 240); do + sleep 15 + h="$(bn)"; w="$($B/igneum-miner watch 1 grpc://127.0.0.1:26610 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1)" + n=$(( ${h:-0x0} )) + echo "RESULT replay $(stamp) s=$(( $(date +%s) - t1 )) evm_block=$n $w" + if [ "$n" -gt 0 ] && [[ "$w" == *synced=true* ]]; then + st="$(curl -s -m 8 -X POST -H 'Content-Type: application/json' --data '{"jsonrpc":"2.0","id":1,"method":"igneum_getProvingStatus","params":[]}' http://127.0.0.1:26790/ | python3 -c 'import sys,json; d=json.load(sys.stdin).get("result",{}); print("tipDaa", int(d.get("tipDaa","0x0"),16), "active", d.get("v1",{}).get("active"), "fresh", d.get("v1",{}).get("freshRuleActive"))' 2>/dev/null)" + daa=$(printf '%s' "$w" | grep -o 'daa=[0-9]*' | cut -d= -f2) + tip=$(printf '%s' "$st" | awk '{print $2}') + if [ -n "$tip" ] && [ "$tip" -ge $(( ${daa:-0} - 20 )) ]; then echo "RESULT datadir_done $(stamp) replay_s=$(( $(date +%s) - t1 )) total_s=$(( $(date +%s) - t0 )) $st"; exit 0; fi + fi +done +echo "RESULT datadir_failed replay did not reach the tip in 60 min" diff --git a/tools/fleet/box-dn2.sh b/tools/fleet/box-dn2.sh new file mode 100755 index 000000000..c0a79f6f3 --- /dev/null +++ b/tools/fleet/box-dn2.sh @@ -0,0 +1,36 @@ +#!/usr/bin/env bash +# A Devnet 2 box (6 October 2026, Josh's standing structure: the rented fleet is the staging chain every release and +# activation crosses before the live devnet). The node runs igneum-devnet-2 (--devnet --devnet-suffix=2: own handshake +# magic, own data directory, a live-devnet peer refuses it at the handshake) with /root/fleet/dn2-override.json (its own +# genesis bits, every activation at a low DAA, NO exec-restart fields: a fresh chain executes from genesis), peered with +# the Devnet 2 seed; one miner on card 0 with its vote key; PROVER=1 adds the segment prover loop (box-prover.py with +# EXPORT_FROM=0 and the Devnet 2 chain name). NODE_BIN names the igneumd to run (the gate swaps it). +set -uo pipefail +F=/root/fleet; OUT=$F/out; mkdir -p $F/in $OUT $F/dn2 $F/mine/packs; exec >> $OUT/dn2.log 2>&1 +stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; } +LABEL="${LABEL:-dn2}"; WALLET="${WALLET:-0x1919191919191919191919191919191919191919}"; SEED="${SEED:-}"; NODE_BIN="${NODE_BIN:-$F/in/igneumd-0313}"; # UNSYNCED=1 only for a fresh genesis (a node that mines while unsynced forks at every restart: Devnet 2 reorgs after the 22:1xZ restarts) +PROVER="${PROVER:-0}"; BPS="${BPS:-}"; UNSYNCED="${UNSYNCED:-}"; RPC_PORT="${RPC_PORT:-26610}"; P2P_PORT="${P2P_PORT:-26611}"; EVM_PORT="${EVM_PORT:-26790}" # other ports on a box whose live node holds 26610/26611 (the wave pods in run A) # BPS=10|5 -> the devnet2-bps fork's IGNEUMD_DEVNET_BPS profile (block-rate run A); unset = 1 block/s +echo "RESULT dn2_start $(stamp) label=$LABEL node=$(sha256sum $NODE_BIN | cut -c1-16) seed=${SEED:-none} prover=$PROVER" +command -v curl >/dev/null || { apt-get update -qq >/dev/null 2>&1; apt-get install -y -qq curl ca-certificates python3 >/dev/null 2>&1; } +if [ ! -x /opt/igneum/pkg/bin/igneum-miner ]; then + read -r PKG_PATH PKG_SHA PKG_VER <<< "$(curl -fsSL -m 30 https://dl.igneum.network/dl/public/igneum-downloads.json | python3 -c 'import sys,json; d=json.load(sys.stdin)["files"]["miner-hive"]; print(d["path"], d["sha256"], d["version"])')" + curl -fsSL -o $F/pkg.tgz "https://dl.igneum.network$PKG_PATH" && echo "$PKG_SHA $F/pkg.tgz" | sha256sum -c - >/dev/null && mkdir -p /opt/igneum/pkg && tar -C /opt/igneum/pkg --strip-components=1 -xzf $F/pkg.tgz || { echo "RESULT dn2_failed package"; exit 2; } +fi +B=/opt/igneum/pkg/bin; cp $F/in/dn2-override.json $F/dn2-override.json +pkill -9 -f '^/root/fleet/in/igneumd-(0313|[0-9a-f]{16}) ' 2>/dev/null; pkill -9 -f "^/opt/igneum/pkg/bin/igneum-miner mine grpc://127.0.0.1:$RPC_PORT " 2>/dev/null; sleep 2 # the Devnet 2 node only, never igneumd-v4 (the rehearsal node beside it, 19:00Z) +[ "${FRESH:-0}" = 1 ] && { rm -rf $F/dn2; mkdir -p $F/dn2; [ -s $F/dn2-node.log ] && mv $F/dn2-node.log $F/dn2-node.prev.log; echo "RESULT dn2_fresh $(stamp) appdir wiped for a new genesis"; } # never the script's own pattern (17:18Z: dn2-kill.sh killed its caller) +PEER=""; for sd in ${SEED//,/ }; do PEER="$PEER --addpeer=$sd"; done # SEED may be a comma list (run A2: two relays per box) +# the statement's program ids (proving agent, 22:1xZ): on 4c6b129d they come from the environment or the verifier host, else zero, and a zero id refuses every segment record +env ${BPS:+IGNEUMD_DEVNET_BPS=$BPS} IGNEUM_PROOF_PROGRAM_IDS="${PROGRAM_IDS:-0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a,0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896}" IGNEUM_PROOF_VERIFIER=/opt/igneum-floor/bin/igneum-prove-host nohup $NODE_BIN --devnet --devnet-suffix=2 --appdir=$F/dn2 --rpclisten=0.0.0.0:$RPC_PORT --evm-rpclisten=127.0.0.1:$EVM_PORT --listen=0.0.0.0:$P2P_PORT $PEER --override-params-file=$F/dn2-override.json --nodnsseed --disable-upnp --nologfiles --yes ${UNSYNCED:+--enable-unsynced-mining} ${NODE_EXTRA:-} >> $F/dn2-node.log 2>&1 & +# (--enable-unsynced-mining: a fresh chain's nodes start unsynced and must mine anyway, Reject(IsInIBD) on the seed at 16:52Z; a comment put inside this line at 17:00Z swallowed the redirect and the ampersand, so the node ran in the foreground and the script never reached the miner) +sleep 10 +echo "RESULT dn2_node $(stamp) pid=$(pgrep -f '^/root/fleet/in/igneumd-(0313|[0-9a-f]{16}) ' | head -1) version=$($NODE_BIN --version 2>&1 | head -1) digest=$(grep -o 'digest: [0-9a-f]*' $F/dn2-node.log | tail -1 | awk '{print substr($2,1,16)}') network=$(grep -oiE 'igneum-devnet-2[^ ,]*' $F/dn2-node.log | head -1) genesis=$(grep -oiE 'genesis [0-9a-f]{16}' $F/dn2-node.log | head -1)" +for i in $(seq 1 30); do w="$($B/igneum-miner watch 1 grpc://127.0.0.1:$RPC_PORT 2>/dev/null | grep -o 'blocks=[0-9]*.*synced=[a-z]*' | tail -1)"; [ -n "$w" ] && break; sleep 5; done +echo "RESULT dn2_watch $(stamp) $(printf '%s' "$w" | sed -E 's/difficulty=[0-9.]* sink=[0-9a-f]* //')" +cd $F/mine && rm -rf packs/dn2 && $B/igneum-miner export-pack grpc://127.0.0.1:$RPC_PORT packs/dn2 > $OUT/dn2-export-pack.log 2>&1 +( while :; do $B/igneum-miner mine grpc://127.0.0.1:$RPC_PORT 1 100000000 "$LABEL" --worker $B/igneum-worker-cuda --worker-args "--device 0 --pack packs/dn2" --prepare-packs packs/dn2-prepare --exit-on-seed-change --evm-address "$WALLET" --payout-label "$LABEL" --status-secs 30 >> $OUT/dn2-miner.log 2>&1; rc=$?; echo "RESULT dn2_miner_exit $(stamp) rc=$rc" >> $OUT/dn2.log; [ $rc = 42 ] && { rm -rf packs/dn2; $B/igneum-miner export-pack grpc://127.0.0.1:$RPC_PORT packs/dn2 >> $OUT/dn2-export-pack.log 2>&1; } || sleep 10; done ) & +echo "RESULT dn2_miner_started $(stamp)" +if [ "$PROVER" = 1 ] && [ -x /opt/igneum-floor/bin/igneum-prove-host ]; then + cd $F && LABEL="$LABEL" WALLET="$WALLET" EXPORT_FROM=0 CHAIN_NAME=igneum-devnet-2 MINER=none RUN_HOURS=48 setsid nohup python3 -u $F/in/box-prover.py > $OUT/dn2-prover-launch.log 2>&1 & + echo "RESULT dn2_prover_started $(stamp)" +fi diff --git a/tools/fleet/box-ember.sh b/tools/fleet/box-ember.sh new file mode 100755 index 000000000..97a281b01 --- /dev/null +++ b/tools/fleet/box-ember.sh @@ -0,0 +1,86 @@ +#!/usr/bin/env bash +# Ember Tune's two-knob ladder on a rented NVIDIA card (docs/plans/ember-tune.md, branch ember-tune): the miner runs +# throughout; the power ladder 100, 90, 80, 70, 60, 50% of the default limit at the unlocked clock (clamped at the +# card's reported minimum), then the clock ladder 90, 80, 70, 60% of the maximum graphics clock at the power the +# first ladder chose; 15 s settle and 60 s hold per step; the choice is the best MH/W among the steps whose rate is +# within 1% of the fastest. Linux root: nvidia-smi -pl and -lgc, no prompt. Every step prints a RESULT line; the +# ladder and the choice go to /root/fleet/out/ember.json as a TUNE-shaped record (relay/lib/ember.mjs parseRecords). +set -uo pipefail +F=/root/fleet; OUT=$F/out; LOG=$OUT/ember.log; B=/opt/igneum/pkg/bin +LABEL="${LABEL:-box}"; WALLET="${WALLET:-0x1919191919191919191919191919191919191919}" +SETTLE="${SETTLE:-15}"; HOLD="${HOLD:-60}" +mkdir -p $OUT $F/mine/packs +exec > >(tee -a $LOG) 2>&1 +stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; } +say() { echo "$(stamp) $*"; } +q() { nvidia-smi --query-gpu="$1" --format=csv,noheader,nounits -i 0 2>/dev/null | head -1 | tr -d ' '; } +NAME="$(q name)"; DRV="$(q driver_version)"; PDEF="$(q power.default_limit)"; PMIN="$(q power.min_limit)"; PMAX="$(q power.max_limit)"; CMAX="$(q clocks.max.graphics)" +pkill -x sp1-gpu-server 2>/dev/null; rm -f /tmp/sp1-cuda-*.sock +echo "RESULT start $(stamp) card=$NAME driver=$DRV power_default_w=$PDEF min_w=$PMIN max_w=$PMAX clock_max_mhz=$CMAX" +# can we set anything? +nvidia-smi -i 0 -pl "$PDEF" >/dev/null 2>&1 && PL_OK=1 || PL_OK=0 +nvidia-smi -i 0 -lgc 0,"$CMAX" >/dev/null 2>&1 && LGC_OK=1 || LGC_OK=0 +nvidia-smi -i 0 -rgc >/dev/null 2>&1 +echo "RESULT knobs power_limit_settable=$PL_OK clock_cap_settable=$LGC_OK" +cd $F/mine +rm -rf packs/devnet; $B/igneum-miner export-pack grpc://127.0.0.1:26610 packs/devnet > $OUT/ember-export.log 2>&1 +nohup $B/igneum-miner mine grpc://127.0.0.1:26610 1 100000000 "$LABEL" --worker $B/igneum-worker-cuda --worker-args "--device 0 --pack packs/devnet" \ + --prepare-packs packs/prepare --exit-on-seed-change --evm-address "$WALLET" --payout-label "$LABEL" --status-secs 10 > $OUT/ember-miner.log 2>&1 & +MPID=$!; cd $F +cleanup() { nvidia-smi -i 0 -rgc >/dev/null 2>&1; [ "$PL_OK" = 1 ] && nvidia-smi -i 0 -pl "$PDEF" >/dev/null 2>&1; kill $MPID 2>/dev/null; pkill -f '^/opt/igneum/pkg/bin/igneum-worker-cuda' 2>/dev/null; } +trap cleanup EXIT +say "miner warming 90 s"; sleep 90 +STEPS=$OUT/ember-steps.jsonl; : > $STEPS +step() { #