Fleet night rows 17 (two miners on one GPU after a restart) and 18 (two provers on one Devnet 2 box)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
76723b3264
commit
2b724114dd
1 changed files with 2 additions and 0 deletions
|
|
@ -170,3 +170,5 @@ per segment, held fresh records offered again every pass).
|
|||
| 14, 20:00Z: the standing fleet | the project lead's ruling (19:50 UK): rented cards stay up and are never destroyed on a job's end. 13 live-devnet boxes converted at 19:41Z (USD 3.50/h, USD 84/day), each under box-standing.sh (node, miner and prover restarted when gone, the recovery recipe on a dead exec, a status line every 10 min), lib/standing.py on the Mac (roster, check, update, re-rent in the same shape, the 10 percent weight gate), a standing block on the fleet page; the Devnet 2 six and the L4 and AMD cards owed as providers free them. The supervisor's own two faults tonight (it matched any igneumd, so it mistook the rehearsal node for the live one and restarted a canary box's dead node on the 0.3.15 file) are fixed in db58804 and 9a294b4 | Operator of a standing box: the node comes back within a minute of dying, the miner with it, and no job takes its GPU without the weight gate | docs/plans/gpu-fleet.md carries the rule |
|
||||
| 15, 19:58Z, 20:01Z, 20:42Z: the hub's live node died three times | Twice on a peer's sync request below its retention ("consensus/src/processes/sync/mod.rs:87 KeyNotFound(GhostdagCompact/0/<genesis>)", while the rolled-back p2-3090-1 synced a 32-minute gap against it; the serving-side fix is in the 0.3.15 node), once on a full disk ("header_processor/processor.rs:534 IO error: No space left on device"): the prover's segment exports under /root/fleet/out/segs (50 to 500 MB a segment, never pruned by box-prover.py) had filled the hub's 60 GB and 3 to 46 GB on every standing box since 11:50Z; two boxes stood at 100 percent. The standing supervisor restarted the node each time (19:59:11Z, 20:02:36Z, 20:44:38Z; the third after 38 GB were freed by hand) and now prunes exports older than 20 minutes and trims the node log every ten minutes (353cc5f); a disk sweep every five minutes writes disk per box to the fleet page and posts one #incidents line per box per hour at 85 percent (disk-sweep.py); the exporter-side cap is the proving lane's | Home miner: the hub is one of three public peers a fresh node dials; a dead hub means a slower first join and nothing lost; a home node's own disk is not at risk (the app's prover does not write exports). Rig: the same. Pool: a pool node that exports segments for its provers has the same disk clock | the three deaths' lines are in the hub's node.log (preserved on the box) and the finality pulls under the scratchpad |
|
||||
| 16, 19:17Z to 20:58Z: the block rate on Devnet 2, 10 blocks a second against 1 | Run A (the fork's 10 blocks/s profile, fresh genesis, 42 cards, the seed on igneum-build-1, every miner dialling the seed only): 4.87 DAG blocks/s but 1.09 blue blocks/s, 77.6 percent red, tips 250 to 660, difficulty easing all hour, the exec follower at 0.05 blocks/s. Run B (1 block/s, same boxes): 1.0 blocks/s and under 2 percent red from minute six, tips 1 to 3, difficulty settled in six minutes, the follower at 0.46 blocks/s. The network lane's read: the reds came from node throughput (61 to 345 ms of CPU per accepted block), not the star | Home miner: the payout interval follows the blue rate, which the profile did not move (1.09 against 1.19 blue/s), so a 4070 at a 10 TH/s network waits about three days for a paying block at either rate; the pool, not the block rate, is the small card's shorter wait. Rig: 4 to 5 hours at 10 TH/s either way. Everyone: a wallet or prover on a 10 blocks/s chain would read state hours behind within the first hour at tonight's follower rate | `docs/analysis/block-rate-devnet2.md`: 1 block/s for the testnet and the launch, 10 behind three measured gates |
|
||||
| 17, 21:52Z: two miners on one GPU after a restart | The read-back after publish 1 showed miner_up=2 on p1-4090, p1-a5000 and p2-4090-3: the publish killed the miner once, the supervisor's miner loop restarted it within ten seconds, and the supervisor's main loop, reading "no miner" in the same gap, started a second loop; two igneum-miner processes then shared one GPU at half rate each. The supervisor's miner loop now kills any other miner and worker before it starts its own (409bc3d); redeployed on the 14 standing boxes at 21:55Z | Home miner: the app owns one miner per GPU and never sees this; a hand-run box with two miner loops halves its rate silently and the only sign is the STATUS line's MH/s. Rig: times eight. Pool: a member with two miners doubles its share submissions at half the rate each, the pool sees one member with a jittery rate | one miner per GPU is now the supervisor's invariant, not the operator's care |
|
||||
| 18, 21:43Z to 21:56Z: two provers on one Devnet 2 box killed each other's GPU server | Repeated by-hand relaunches of box-prover.py on dn2-1 and dn2-2 left two instances on a box: each relaunch's "kill the server, remove the socket" took the other instance's SP1 GPU server away mid-proof ("CudaClientError: early eof" on every chain step from 21:43Z) and both wrote the same prover-state.json.tmp, so one lost the file (FileNotFoundError at the os.replace). box-prover.py now holds a pid file under its out directory and a second instance exits at once, and each process writes its own tmp (c29c6b9); the Devnet 2 provers relaunched clean at 21:54Z | Home miner: the app's prover is one process by construction. Operator of a hand-run prover box: start the prover once; a second start now refuses with the first's pid instead of taking its GPU server down | the Devnet 2 gate's "paid segments" read waits on the first record from these provers |
|
||||
|
|
|
|||
Loading…
Reference in a new issue