diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md index 0000ef8c..6cf04b67 100644 --- a/docs/plans/proving-v1.md +++ b/docs/plans/proving-v1.md @@ -165,4 +165,6 @@ per segment, held fresh records offered again every pass). | 9, 18:14Z: pool-v0 under load, 10 members, 0 shares | igneum-pool (PPLNS, port 4463) with 10 wave pods as members: every member's GPU shares were refused `WORKER MISMATCH` because the pool's job carried no program class and the member's worker hashed class v2 while the devnet's templates are class v3; 0 shares accepted in 25 minutes, the pool's stats page alive. The pool agent rebased the pool-mode miner onto 0.3.14 (pool-v0-rebase c4c92e8: the job carries program_class and era_seed, the member re-checks shares with the template's class); the rerun with both new binaries is scheduled for tonight after the block-rate runs | Pool user: on the old pool binary every share is wasted work; on the rebased pair a member without a node takes the day and the dataset size from the pool's seeds line. Home miner and rig: unaffected (solo mining never touched the pool) | the rerun's counts (shares accepted per member inside the first vardiff interval, mismatches 0, rejected 0) go to main | | 10, 17:25Z to 17:38Z: the 0.3.14 canary on the live devnet | the release's Linux igneumd (4c6b129d) on five restart-path boxes for ten minutes: 0 new rejected blocks on every box, exec state roots equal to the hub's at a common height, a segment record paid on the new binary, every node on the new version: PASS on all four criteria; the first run's FAIL was the gate counting the hub (whose trait is about 3 rejects a minute) and a `cut -c1-140` that truncated the digest line before the match | Operator on the restart path: the swap is a binary change with the data dir kept, seconds of downtime, no replay. Snapshot-path operator: the same binary, the snapshot recipe unchanged | the canary form (canary-next.sh) is the fleet's release gate; the 0.3.15 run is in flight on four boxes at 19:41Z | | 11, 18:31Z to 19:35Z: the class v4 rehearsal on igneum-devnet-400 (16 nodes, 38 pods joining, one stale 0.3.13 box) | The flip by miner signal landed at DAA 1,200 (19:00:5xZ) with the identical line on every node ("epoch 2 ... seed block f0606e20..., threshold 9500 bps, 599 of 599 blue blocks"), one program id per epoch across boxes (epoch 2 0x24304f0788ea9408, epoch 3 0xcc266b4f5dbc3447, epoch 4 0x634018bab5e5f283; the Mac's igneum-pow show agrees and the v3 id differs), 0 PoW rejected all run, the stale box refused by digest at every connect (12 by 19:46Z) and refusing the sixteen-field file at parse ("unknown field program_class_v4_activation_daa"), the floor at 2,400 crossed with v4 in force and nothing to print. Two findings outside the protocol: the plan's CPU engine (5 kH/s a box) cannot move a chain whose genesis difficulty is 2^27 (the CUDA worker fixed it: 14 to 98 MH/s a box), and the 0.3.14 package's igneum-worker-cuda refuses a generator-4 pack ("not a generator version this worker runs (2 or 3)"), so the chain stood at DAA 1,202 from 19:00:5xZ until the release tree's generator-4 Linux worker (97e036e2) went on at 19:15Z; a worker restarted on its old --pack after --exit-on-seed-change also answers every job "epoch seed mismatch" until the pack is re-exported | Home miner, any card: at a class flip the old worker stops dead; the flip is safe only when the generator-4 worker ships in the hive, Mac and Windows packages before the signal window closes, and the app re-exports the pack at every seed change. Rig: the same, times eight. Pool: the pool's node decides the class; a member on an old worker mismatches every share | the cut's rule (publish 2 gated on every worker, not every node) is with the shipper; P1 PASSED is on master | -| 12, 18:39:40Z onward: the live devnet's finality paused at 93 voters | On every node read (hub, p1-3080, wave-05) the last lock is checkpoint 6842 at 18:39:36Z; checkpoints keep being determined (6967 at 19:43Z) and no certificate is received or built anywhere after 18:39:40Z. The fleet's boxes never left the live devnet (every live node synced throughout). The hand nodes (node 1, the observer) were down from about 18:2xZ to 19:42Z with the desktop app; every rented node dials only them and the hub (no rented box has an inbound port), so the votes had no path to the aggregators: the star topology that run A (10 blocks/s, 77 to 84% red blocks) shows on Devnet 2 | Home miner: with the hands down, nothing a home node does restores locks; the fix is more public peers, which is the standing fleet's ports rule (a mapped p2p port on every standing box the provider allows) and the hands' move to a box. Rig and pool: the same | the re-read at 19:52Z decides whether the owner is routed (no certificate ten minutes after the hands' return) | +| 12, 18:39:40Z onward: the live devnet's finality paused at 93 voters | On every node read the last lock is checkpoint 6842 at 18:39:36Z; checkpoints keep being determined and no certificate is received or built anywhere after 18:39:40Z. The fleet's first reading (the hands down 18:2xZ to 19:42Z, the star topology) was wrong: the observer rows show 13 fleet keys stopped mining the live devnet between 18:27Z and 18:30Z when the rehearsal job took their GPUs (their live nodes stayed up and synced, but a voter's weight is its blue blocks), and with seven earlier leavers that was 42.7 percent of the frozen voter table; rule v3 then holds the pause for one full window, the first lock expected about 20:40Z | Home miner: a voter that stops mining stops counting within the window, and 10 percent of the weight leaving in an hour is the most the table absorbs without a pause. Rig and pool: a pool is one voter with its members' whole weight; its restart is the biggest single removal on the network | the standing-fleet rule from it (6 October 2026, 20:00Z): never remove more than 10 percent of the live devnet's 30-day weight in any hour; `lib/standing.py weight_check` gates every job that stops or shares a standing miner | +| 13, 19:14Z to 20:30Z: the 0.3.15 canary, FAIL on the first binary, the retry confounded | 713ef876 on four live boxes: every block a 0.3.15 node mined or relayed carried version 1026 (the class v4 signal bit stamped from its first block, no window set) and every 0.3.14 node answered "wrong block version: got 1026 but expected 2" and disconnected it (the hub: 45 such rejects in the first 14 minutes, 468 by 20:29Z), so a 0.3.15 node that fell behind could not re-sync (p2-3090-1: connected and dropped every 30 s for 32 minutes) and every 0.3.15 miner lost every block it found; FAIL, publish 1 held. Three more findings on the way back: (a) a 0.3.14 node whose datadir holds 1026 blocks keeps relaying them and stays a disconnected peer after the rollback, so the four canary datadirs are poisoned until wiped; (b) a pruned 0.3.14 node dies ("consensus/src/processes/sync/mod.rs:87 KeyNotFound(GhostdagCompact/0/)") when a peer syncing a gap asks below its retention (the hub twice, 19:58Z and 20:01Z, restarted by the standing supervisor); (c) a standalone igneum-miner keeps a dead template subscription after its node restarts (templates frozen, fetch_errors climbing, no submits). The shipper's 7961c5f1 gates the stamp on publish 2's object and fixes the serving side of (b); its retry on the poisoned boxes was confounded by (a) and (c), the clean retry runs on two untouched boxes | Home miner: an update that stamps a new block version before the network accepts it is a silent death (the app shows hashing, nothing is paid); the fix is the version gate on the object plus a datadir that never held a bad block. Rig: the same, times eight. Pool: a pool node on the bad version drops every member's share from the network's view | the canary form stays the release gate; the fresh-join line from a wiped datadir is the last read | +| 14, 20:00Z: the standing fleet | the project lead's ruling (19:50 UK): rented cards stay up and are never destroyed on a job's end. 13 live-devnet boxes converted at 19:41Z (USD 3.50/h, USD 84/day), each under box-standing.sh (node, miner and prover restarted when gone, the recovery recipe on a dead exec, a status line every 10 min), lib/standing.py on the Mac (roster, check, update, re-rent in the same shape, the 10 percent weight gate), a standing block on the fleet page; the Devnet 2 six and the L4 and AMD cards owed as providers free them. The supervisor's own two faults tonight (it matched any igneumd, so it mistook the rehearsal node for the live one and restarted a canary box's dead node on the 0.3.15 file) are fixed in db58804 and 9a294b4 | Operator of a standing box: the node comes back within a minute of dying, the miner with it, and no job takes its GPU without the weight gate | docs/plans/gpu-fleet.md carries the rule |