igneum/docs/plans/gpu-fleet.md

5.3 KiB

The rented GPU fleet: the standing rule and the one-shot rule

the project lead's ruling, 6 October 2026, 19:50 UK ("keep rented cards up"), written as the fleet agent runs it. Tooling in tools/fleet/ on branch gpu-fleet; every box operation goes through tools/fleet/lib/ (the box library) and the box-side scripts it ships.

Two kinds of box

Kind Rule Examples
Standing stays up, never destroyed on a job's end; replaced in the same shape when its host dies; node under a supervisor; version follows the live manifest; never joins an experiment the 13 live-devnet boxes of 6 October (USD 3.50/h), the Devnet 2 set (seed, 3 miners, 2 provers)
One-shot rented for one measurement, destroyed the minute its measurement is in (a running meter is a bug) the memory-matrix cards, the 8x rigs, the wave pods, the RISC Zero box

The standing set's size and cost: 16 on the live devnet (a mix: 5090, 4090s, 3090s, 4070, 3080, A5000, an L4, at least two AMD when Vast reopens) plus 6 on Devnet 2, about USD 8 an hour, USD 200 a day, inside the daily budget (USD 250 to 1,000). Tonight's 13 cost USD 3.50 an hour (USD 84 a day); the L4, the AMD pair and the Devnet 2 six are owed to the roster as providers free cards (RunPod gave no pod of 8 card types from 18:59Z, Vast refuses every new rent).

What a standing box carries

Piece Where What it does
box-standing.sh (the supervisor) /root/fleet/in/, started once by lib/standing.py install every 60 s restarts the node, the miner loop (CUDA worker, pack re-exported on seed change) and the prover when gone; every 10 min writes a RESULT standing line (uptime, node pid and binary, blocks, DAA, peers, synced, exec tip, miner rate, prover) to /root/fleet/out/standing.log; runs the shipper's recovery recipe (box-exec-snapshot.sh) when consensus is synced above 1,000 blocks and the exec tip reads 0
/root/fleet/standing.node the box the node binary the supervisor runs; the supervisor adopts whatever node already runs when the file is absent (a canary's newer binary is never downgraded) and restarts the node only when a publish rewrites the pointer
the registry row ~/Desktop/fleet/boxes.json standing: true, `role: live
lib/standing.py the Mac roster (role, card, price, uptime), install, check (ssh alive, supervisor up, synced, exec moving, prover, node binary against the wanted sha), update (the manifest's hive package onto the box, pointer rewritten), rerent (same card and provider, label suffixed -r<n>, old row marked destroyed), loop (check every 10 min, re-rent after two dead checks, report what is behind)
the fleet page dl.igneum.network/fleet-22adafa34bc2/ a standing block (count, USD/h, USD/day, roles, the rule) and standing, role, uptime_h on every box row

A publish on the standing set is a script, not the loop: tools/fleet/publish-0315.py is the first (the Linux igneumd, igneum-miner and the generator-4 workers over the package's, the pointer rewritten, the read-back table of version line, digest and worker sha per box). The loop reports a box behind its wanted sha; it does not move binaries by itself.

Ports

A standing box should carry a mapped p2p port whenever the provider allows it (RunPod: request 26611/tcp at rent; Vast: none of tonight's offers mapped one), so that the live devnet stops being a star in which every rented node dials only the two hand nodes and the hub: the 18:39Z to 19:4xZ finality pause coincided with the hands being down, and run A (10 blocks/s on a star, 77 to 84 percent red blocks) showed the same weakness on Devnet 2. Tonight's inbound-capable nodes were igneum-build-1 (26611) and dn2-seed's one mapped port; a true mesh needs pods rented with a p2p port.

Clock

  • 19:41Z: the 13 live boxes converted (supervisor started, every node synced, every prover up).
  • Owed: the Devnet 2 six after the block-rate runs; the L4 and the AMD pair when a provider reopens; the ports on re-rent.

The 10 percent rule (6 October 2026, 20:00Z, from the finality pause)

What happened: the class v4 rehearsal job took the GPUs of 13 live-devnet miners between 18:27Z and 18:30Z (their live nodes stayed up and synced, but a voter's weight is its blue blocks, and a miner that stops mining stops being a voter), and with seven earlier leavers that was 42.7 percent of the frozen voter table; finality rule v3 then holds the pause for one full window, and the first lock after 18:39:36Z is expected at about 20:40Z. The fleet's read that "the boxes never left the devnet" was true of the nodes and wrong about the weight.

The rule: never remove more than 10 percent of the live devnet's 30-day weight in any hour. Weight is counted by blue blocks per key over the window, read from the hub; any experiment that borrows live miners does it in slices with an hour between slices; a standing box's miner is never stopped (or its GPU shared with a second worker) by a job without that check. lib/standing.py weight_check <labels> is the gate: it reads the hub's blue blocks per payout key over the window, sums the share of the labels asked for, and refuses the job when the share since the last hour's removals exceeds 10 percent; every fleet job that touches a standing box's miner calls it first.