Fleet docs: the gpu-fleet clock (10 percent rule, the disk class, Devnet 2 standing); fleet night row 16 (the block rate)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
b7144f5dc6
commit
f7389dd962
2 changed files with 5 additions and 1 deletions
|
|
@ -41,7 +41,10 @@ nodes were igneum-build-1 (26611) and dn2-seed's one mapped port; a true mesh ne
|
|||
## Clock
|
||||
|
||||
- 19:41Z: the 13 live boxes converted (supervisor started, every node synced, every prover up).
|
||||
- Owed: the Devnet 2 six after the block-rate runs; the L4 and the AMD pair when a provider reopens; the ports on re-rent.
|
||||
- 20:00Z: the 10 percent rule, after the finality pause the fleet caused (`lib/standing.py weight_check`).
|
||||
- 20:42Z to 20:48Z: the disk class (the prover's exports) found on the hub's third death; the supervisor prunes, `disk-sweep.py` watches, box-prover.py caps its own directory.
|
||||
- 20:58Z: Devnet 2 standing at 1 block/s: the seed on igneum-build-1 (188.40.146.49:26611), miners dn2-seed, dn2-1, dn2-2, dn2-3, provers on dn2-1 and dn2-2 (the floor host copied from p1-4090); the gate owed until the rehearsal chain ends and dn2-seed's mapped port is free.
|
||||
- Owed: the L4 and the AMD pair when a provider reopens; the ports on re-rent; the standing flag on the Devnet 2 rows after the gate.
|
||||
|
||||
## The 10 percent rule (6 October 2026, 20:00Z, from the finality pause)
|
||||
|
||||
|
|
|
|||
|
|
@ -169,3 +169,4 @@ per segment, held fresh records offered again every pass).
|
|||
| 13, 19:14Z to 20:30Z: the 0.3.15 canary, FAIL on the first binary, the retry confounded | 713ef876 on four live boxes: every block a 0.3.15 node mined or relayed carried version 1026 (the class v4 signal bit stamped from its first block, no window set) and every 0.3.14 node answered "wrong block version: got 1026 but expected 2" and disconnected it (the hub: 45 such rejects in the first 14 minutes, 468 by 20:29Z), so a 0.3.15 node that fell behind could not re-sync (p2-3090-1: connected and dropped every 30 s for 32 minutes) and every 0.3.15 miner lost every block it found; FAIL, publish 1 held. Three more findings on the way back: (a) a 0.3.14 node whose datadir holds 1026 blocks keeps relaying them and stays a disconnected peer after the rollback, so the four canary datadirs are poisoned until wiped; (b) a pruned 0.3.14 node dies ("consensus/src/processes/sync/mod.rs:87 KeyNotFound(GhostdagCompact/0/<genesis>)") when a peer syncing a gap asks below its retention (the hub twice, 19:58Z and 20:01Z, restarted by the standing supervisor); (c) a standalone igneum-miner keeps a dead template subscription after its node restarts (templates frozen, fetch_errors climbing, no submits). The shipper's 7961c5f1 gates the stamp on publish 2's object and fixes the serving side of (b); its retry on the poisoned boxes was confounded by (a) and (c), the clean retry runs on two untouched boxes | Home miner: an update that stamps a new block version before the network accepts it is a silent death (the app shows hashing, nothing is paid); the fix is the version gate on the object plus a datadir that never held a bad block. Rig: the same, times eight. Pool: a pool node on the bad version drops every member's share from the network's view | the canary form stays the release gate; the fresh-join line from a wiped datadir is the last read |
|
||||
| 14, 20:00Z: the standing fleet | the project lead's ruling (19:50 UK): rented cards stay up and are never destroyed on a job's end. 13 live-devnet boxes converted at 19:41Z (USD 3.50/h, USD 84/day), each under box-standing.sh (node, miner and prover restarted when gone, the recovery recipe on a dead exec, a status line every 10 min), lib/standing.py on the Mac (roster, check, update, re-rent in the same shape, the 10 percent weight gate), a standing block on the fleet page; the Devnet 2 six and the L4 and AMD cards owed as providers free them. The supervisor's own two faults tonight (it matched any igneumd, so it mistook the rehearsal node for the live one and restarted a canary box's dead node on the 0.3.15 file) are fixed in db58804 and 9a294b4 | Operator of a standing box: the node comes back within a minute of dying, the miner with it, and no job takes its GPU without the weight gate | docs/plans/gpu-fleet.md carries the rule |
|
||||
| 15, 19:58Z, 20:01Z, 20:42Z: the hub's live node died three times | Twice on a peer's sync request below its retention ("consensus/src/processes/sync/mod.rs:87 KeyNotFound(GhostdagCompact/0/<genesis>)", while the rolled-back p2-3090-1 synced a 32-minute gap against it; the serving-side fix is in the 0.3.15 node), once on a full disk ("header_processor/processor.rs:534 IO error: No space left on device"): the prover's segment exports under /root/fleet/out/segs (50 to 500 MB a segment, never pruned by box-prover.py) had filled the hub's 60 GB and 3 to 46 GB on every standing box since 11:50Z; two boxes stood at 100 percent. The standing supervisor restarted the node each time (19:59:11Z, 20:02:36Z, 20:44:38Z; the third after 38 GB were freed by hand) and now prunes exports older than 20 minutes and trims the node log every ten minutes (353cc5f); a disk sweep every five minutes writes disk per box to the fleet page and posts one #incidents line per box per hour at 85 percent (disk-sweep.py); the exporter-side cap is the proving lane's | Home miner: the hub is one of three public peers a fresh node dials; a dead hub means a slower first join and nothing lost; a home node's own disk is not at risk (the app's prover does not write exports). Rig: the same. Pool: a pool node that exports segments for its provers has the same disk clock | the three deaths' lines are in the hub's node.log (preserved on the box) and the finality pulls under the scratchpad |
|
||||
| 16, 19:17Z to 20:58Z: the block rate on Devnet 2, 10 blocks a second against 1 | Run A (the fork's 10 blocks/s profile, fresh genesis, 42 cards, the seed on igneum-build-1, every miner dialling the seed only): 4.87 DAG blocks/s but 1.09 blue blocks/s, 77.6 percent red, tips 250 to 660, difficulty easing all hour, the exec follower at 0.05 blocks/s. Run B (1 block/s, same boxes): 1.0 blocks/s and under 2 percent red from minute six, tips 1 to 3, difficulty settled in six minutes, the follower at 0.46 blocks/s. The network lane's read: the reds came from node throughput (61 to 345 ms of CPU per accepted block), not the star | Home miner: the payout interval follows the blue rate, which the profile did not move (1.09 against 1.19 blue/s), so a 4070 at a 10 TH/s network waits about three days for a paying block at either rate; the pool, not the block rate, is the small card's shorter wait. Rig: 4 to 5 hours at 10 TH/s either way. Everyone: a wallet or prover on a 10 blocks/s chain would read state hours behind within the first hour at tonight's follower rate | `docs/analysis/block-rate-devnet2.md`: 1 block/s for the testnet and the launch, 10 behind three measured gates |
|
||||
|
|
|
|||
Loading…
Reference in a new issue