Fleet: disk sweep every five minutes (disk per box on the page, one #incidents line per box per hour at 85 percent); fleet night row 15 (the hub's three deaths)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
fe5e481528
commit
853c0217d9
3 changed files with 45 additions and 1 deletions
|
|
@ -168,3 +168,4 @@ per segment, held fresh records offered again every pass).
|
|||
| 12, 18:39:40Z onward: the live devnet's finality paused at 93 voters | On every node read the last lock is checkpoint 6842 at 18:39:36Z; checkpoints keep being determined and no certificate is received or built anywhere after 18:39:40Z. The fleet's first reading (the hands down 18:2xZ to 19:42Z, the star topology) was wrong: the observer rows show 13 fleet keys stopped mining the live devnet between 18:27Z and 18:30Z when the rehearsal job took their GPUs (their live nodes stayed up and synced, but a voter's weight is its blue blocks), and with seven earlier leavers that was 42.7 percent of the frozen voter table; rule v3 then holds the pause for one full window, the first lock expected about 20:40Z | Home miner: a voter that stops mining stops counting within the window, and 10 percent of the weight leaving in an hour is the most the table absorbs without a pause. Rig and pool: a pool is one voter with its members' whole weight; its restart is the biggest single removal on the network | the standing-fleet rule from it (6 October 2026, 20:00Z): never remove more than 10 percent of the live devnet's 30-day weight in any hour; `lib/standing.py weight_check` gates every job that stops or shares a standing miner |
|
||||
| 13, 19:14Z to 20:30Z: the 0.3.15 canary, FAIL on the first binary, the retry confounded | 713ef876 on four live boxes: every block a 0.3.15 node mined or relayed carried version 1026 (the class v4 signal bit stamped from its first block, no window set) and every 0.3.14 node answered "wrong block version: got 1026 but expected 2" and disconnected it (the hub: 45 such rejects in the first 14 minutes, 468 by 20:29Z), so a 0.3.15 node that fell behind could not re-sync (p2-3090-1: connected and dropped every 30 s for 32 minutes) and every 0.3.15 miner lost every block it found; FAIL, publish 1 held. Three more findings on the way back: (a) a 0.3.14 node whose datadir holds 1026 blocks keeps relaying them and stays a disconnected peer after the rollback, so the four canary datadirs are poisoned until wiped; (b) a pruned 0.3.14 node dies ("consensus/src/processes/sync/mod.rs:87 KeyNotFound(GhostdagCompact/0/<genesis>)") when a peer syncing a gap asks below its retention (the hub twice, 19:58Z and 20:01Z, restarted by the standing supervisor); (c) a standalone igneum-miner keeps a dead template subscription after its node restarts (templates frozen, fetch_errors climbing, no submits). The shipper's 7961c5f1 gates the stamp on publish 2's object and fixes the serving side of (b); its retry on the poisoned boxes was confounded by (a) and (c), the clean retry runs on two untouched boxes | Home miner: an update that stamps a new block version before the network accepts it is a silent death (the app shows hashing, nothing is paid); the fix is the version gate on the object plus a datadir that never held a bad block. Rig: the same, times eight. Pool: a pool node on the bad version drops every member's share from the network's view | the canary form stays the release gate; the fresh-join line from a wiped datadir is the last read |
|
||||
| 14, 20:00Z: the standing fleet | the project lead's ruling (19:50 UK): rented cards stay up and are never destroyed on a job's end. 13 live-devnet boxes converted at 19:41Z (USD 3.50/h, USD 84/day), each under box-standing.sh (node, miner and prover restarted when gone, the recovery recipe on a dead exec, a status line every 10 min), lib/standing.py on the Mac (roster, check, update, re-rent in the same shape, the 10 percent weight gate), a standing block on the fleet page; the Devnet 2 six and the L4 and AMD cards owed as providers free them. The supervisor's own two faults tonight (it matched any igneumd, so it mistook the rehearsal node for the live one and restarted a canary box's dead node on the 0.3.15 file) are fixed in db58804 and 9a294b4 | Operator of a standing box: the node comes back within a minute of dying, the miner with it, and no job takes its GPU without the weight gate | docs/plans/gpu-fleet.md carries the rule |
|
||||
| 15, 19:58Z, 20:01Z, 20:42Z: the hub's live node died three times | Twice on a peer's sync request below its retention ("consensus/src/processes/sync/mod.rs:87 KeyNotFound(GhostdagCompact/0/<genesis>)", while the rolled-back p2-3090-1 synced a 32-minute gap against it; the serving-side fix is in the 0.3.15 node), once on a full disk ("header_processor/processor.rs:534 IO error: No space left on device"): the prover's segment exports under /root/fleet/out/segs (50 to 500 MB a segment, never pruned by box-prover.py) had filled the hub's 60 GB and 3 to 46 GB on every standing box since 11:50Z; two boxes stood at 100 percent. The standing supervisor restarted the node each time (19:59:11Z, 20:02:36Z, 20:44:38Z; the third after 38 GB were freed by hand) and now prunes exports older than 20 minutes and trims the node log every ten minutes (353cc5f); a disk sweep every five minutes writes disk per box to the fleet page and posts one #incidents line per box per hour at 85 percent (disk-sweep.py); the exporter-side cap is the proving lane's | Home miner: the hub is one of three public peers a fresh node dials; a dead hub means a slower first join and nothing lost; a home node's own disk is not at risk (the app's prover does not write exports). Rig: the same. Pool: a pool node that exports segments for its provers has the same disk clock | the three deaths' lines are in the hub's node.log (preserved on the box) and the finality pulls under the scratchpad |
|
||||
|
|
|
|||
36
tools/fleet/disk-sweep.py
Normal file
36
tools/fleet/disk-sweep.py
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Disk on every live fleet box (6 October 2026, after the hub died on a full disk): one df per box in parallel, written to
|
||||
~/Desktop/fleet/disk.json ({label: {pct, avail_gb, total_gb, t}}) for the fleet page's disk column; any box at or over 85
|
||||
percent posts one line to Discord #incidents through tools/community/discord-hooks.mjs (incident open, the fleet's own
|
||||
sentence), at most one line per box per hour (state in ~/Desktop/fleet/disk-alerts.json). Run from page.py's cadence or
|
||||
by hand: disk-sweep.py [--no-alert]."""
|
||||
import sys, os, json, time, subprocess
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))); from lib import Box, Registry
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
ROOT = os.path.expanduser("~/Desktop/fleet"); OUT = os.path.join(ROOT, "disk.json"); ALERTS = os.path.join(ROOT, "disk-alerts.json")
|
||||
HOOKS = "/Users/joshm/Projects/igneum/tools/community/discord-hooks.mjs"
|
||||
def now(): return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
|
||||
def one(item):
|
||||
iid, b = item
|
||||
try:
|
||||
box = Box(b["ssh_host"], b["ssh_port"], b["label"], iid, None, b.get("provider"), b.get("ssh_user", "root"))
|
||||
rc, out, err = box.run("df -B1 / | tail -1 | awk '{print $2, $4, $5}'", 25)
|
||||
tot, avail, pct = out.split()
|
||||
return b["label"], {"pct": int(pct.rstrip("%")), "avail_gb": round(int(avail) / 1e9, 1), "total_gb": round(int(tot) / 1e9), "t": now()}
|
||||
except Exception as e: return b["label"], {"pct": None, "error": str(e)[:60], "t": now()}
|
||||
def sweep(alert=True):
|
||||
reg = Registry.load(); live = [(i, b) for i, b in reg.items() if b.get("state") != "destroyed" and b.get("ssh_host")]
|
||||
with ThreadPoolExecutor(20) as ex: res = dict(ex.map(one, live))
|
||||
json.dump(res, open(OUT, "w"), indent=1)
|
||||
if alert:
|
||||
st = json.load(open(ALERTS)) if os.path.exists(ALERTS) else {}
|
||||
for label, d in res.items():
|
||||
if d.get("pct") is not None and d["pct"] >= 85 and time.time() - st.get(label, 0) > 3600:
|
||||
what = f"Fleet box {label} disk at {d['pct']} percent ({d['avail_gb']} GB free of {d['total_gb']})"
|
||||
r = subprocess.run(["node", HOOKS, "incident", "open", "--what", what, "--affected", "one rented fleet box; no user-facing effect unless it is the hub (then a slower first join for home miners, no loss)", "--doing", "the standing supervisor prunes the prover's segment exports and trims the node log every ten minutes; the fleet agent reads the box", "--id", f"disk-{label}-{int(time.time())}", "--live"], capture_output=True, text=True, timeout=60)
|
||||
st[label] = time.time(); print("alert", label, d["pct"], "rc", r.returncode, (r.stdout + r.stderr)[-120:].replace("\n", " "))
|
||||
json.dump(st, open(ALERTS, "w"))
|
||||
return res
|
||||
if __name__ == "__main__":
|
||||
res = sweep(alert="--no-alert" not in sys.argv)
|
||||
for l, d in sorted(res.items(), key=lambda kv: -(kv[1].get("pct") or 0)): print(f"{l:<12} {str(d.get('pct')):>4}% {d.get('avail_gb', '?')} GB free")
|
||||
|
|
@ -8,7 +8,7 @@ result rows (results.json: {card_key: {...}}), the night rows (night.json: [..])
|
|||
page.py build rebuild only
|
||||
page.py publish rebuild and publish (skipped when the last publish was under 60 s ago unless --force)
|
||||
"""
|
||||
import json, os, sys, time, datetime, subprocess
|
||||
import time, json, os, sys, time, datetime, subprocess
|
||||
ROOT = os.path.expanduser("~/Desktop/fleet"); OUT = os.path.join(ROOT, "fleet.json")
|
||||
PUBLISH = "/Users/joshm/Projects/igneum/tools/fleet/publish-fleet.sh"
|
||||
def now(): return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
|
||||
|
|
@ -28,6 +28,12 @@ def build():
|
|||
{"order": 2, "name": "Prover fleet on the devnet", "state": "planned", "planned": 25, "done": 0, "note": ""},
|
||||
{"order": 3, "name": "The 8x rigs", "state": "planned", "planned": 2, "done": 0, "note": ""},
|
||||
{"order": 4, "name": "AMD mining", "state": "planned", "planned": 3, "done": 0, "note": ""}])
|
||||
dj = os.path.join(ROOT, "disk.json")
|
||||
if not os.path.exists(dj) or time.time() - os.path.getmtime(dj) > 300:
|
||||
try:
|
||||
import importlib.util; spec = importlib.util.spec_from_file_location("disk_sweep", os.path.join(os.path.dirname(os.path.abspath(__file__)), "disk-sweep.py")); m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m); m.sweep(alert=True)
|
||||
except Exception as e: print("disk sweep failed:", str(e)[:120])
|
||||
disk = jl("disk.json", {})
|
||||
boxes = []; vast = 0.0; runpod = 0.0
|
||||
for iid, b in reg.items():
|
||||
h = b.get("hours") if b.get("state") == "destroyed" else hours(b["rented_at"])
|
||||
|
|
@ -36,6 +42,7 @@ def build():
|
|||
else: runpod += cost
|
||||
boxes.append({"id": iid, "label": b["label"], "card": b["card"], "vram_gb": round((b.get("vram_mb") or 0) / 1024), "provider": b.get("provider", "vast"),
|
||||
"rate_usd_h": b["dph"], "phase": b.get("phase", "1"), "state": "done" if b.get("state") == "destroyed" else b.get("state", "renting"),
|
||||
"disk_pct": (disk.get(b["label"]) or {}).get("pct"), "disk_free_gb": (disk.get(b["label"]) or {}).get("avail_gb"), "disk_red": bool(((disk.get(b["label"]) or {}).get("pct") or 0) >= 85),
|
||||
"standing": bool(b.get("standing")), "role": b.get("role", ""), "standing_since": b.get("standing_since"), "uptime_h": round(hours(b.get("standing_since") or b["rented_at"]), 1) if b.get("standing") and b.get("state") != "destroyed" else None,
|
||||
"started_at": b["rented_at"], "ended_at": b.get("destroyed_at"), "hours": round(h, 2), "cost_usd": cost, "doing": (("done, destroyed: " + (b.get("doing") or "its measurement is in")) if b.get("state") == "destroyed" else b.get("doing", ""))})
|
||||
running = [b for b in boxes if b["state"] not in ("done", "failed")]
|
||||
|
|
|
|||
Loading…
Reference in a new issue