Proving v1 plan: fleet night rows 8 to 12 (finality at 93 voters, pool-v0 under load, the 0.3.14 canary, the class v4 rehearsal, the finality pause)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
a249904d89
commit
4fbaabbde9
1 changed files with 5 additions and 0 deletions
|
|
@ -161,3 +161,8 @@ per segment, held fresh records offered again every pass).
|
|||
| 5, 15:42:57Z: DAA 198,000, the fresh-record rule armed; 15:43:42Z: a 229-block reorg reset every executing node | The hub's first `PoW accepted ... daa 198000` at 15:42:57Z. One-block selected-chain reorgs at 15:42:32, :43, :53 and 15:43:05Z (heights 135,065 to 135,088), then at 15:43:42Z `selected-chain reorg: 229 chain blocks removed, unwinding to height 134884` (our tip blue score 194,395, the last removed 194,117), `reorg deeper than the snapshot ring; replaying from genesis`, genesis executed, then `exec not synced: the executor is at chain block 0 and the bodies below this node's retention root are gone`: executedTip 0, persistedTip 135,028, blocked null. The same on every fleet box that was executing and on the observer and the seed (the shipper's reading). The chain ran two-sided for about a minute after the switch | Every node operator whose node was executing at 15:43Z (home miner, rig, pool) read a zero wallet and an empty work list until a restart through the exec-restart path; the miners of the 229 losing blocks lost those rewards; a deep reorg after a snapshot ring on a pruned node is the class: the fallback must be the exec-restart point, not genesis (the proving agent's item) | The fleet restarted its boxes through the exec-restart path (90 to 190 s each, the night loop `tools/fleet/night.py` re-runs it on any box whose executed tip falls to 0 for 150 s) and the provers started on the executed tip from 15:56Z |
|
||||
| 6, 16:38:45Z: the first segment records after the fresh-rule switch, accepted and paid | The fleet's RTX 5090 (Vast, a restart-path 0.3.13 node, the segment host from the proving-v1 bundle, the prover loop `tools/fleet/box-prover.py` exporting from the exec restart block) claimed segment 137142..137149 at 16:35:30Z, cut 8 fixtures, ran `--mode chain --save-shards` (8 shards and 8 aggregations on a card that also mines), had 8 of 8 shard records accepted and the fresh segment record accepted at 16:38:45Z: 195.8 s claim to acceptance, aggregator share 1.5076 IGN; a 4090 (137102..137109) followed at 16:39:51Z in 211.2 s, 1.6150 IGN. The hub, a node on the hands' real state, read paidSegments 2, paidSegmentWei 3.12 IGN, pool 2 entries 2 verified, paidShards 1,482 to 1,516 at 16:42Z: carried and paid | A 24 or 32 GB card that mines earns a segment's aggregator share about every 200 s on top of its 8 shards' 90%; six such cards cover about a quarter of the chain's segments (1.8 of 7.5 a minute); 47 mining 24 GB cards or 6 proving-only ones cover it (the fleet table's arithmetic, now with a measured 200 s) | The six restart-path boxes prove through the night; the hourly rows carry segments per hour and the chain's coverage |
|
||||
| 7, 16:50Z: one exec state, and the export gap on a snapshot-recovered node | The state roots at 27,276 (0xed27bb2d...), 130,272 (0xf0a762da...) and 130,273 (0x2e22e029...) are identical on a restart-path node and on a snapshot-recovered node (export segment root and eth_getBlockByNumber agree), so the two recovery paths of 16:00Z to 16:25Z are one chain state and records from either side are valid on the other (row 6 confirms it). But on a snapshot-recovered node every `igneum_exportSegments` (from 0, 27,276 or 130,272) carries 41 to 42 accounts, the exporter's port state root differs from the node's at the first segment, and from 0 the 27,276 segments below the restart carry zero roots: the prover kit cannot cut against the hands' kind of node. The fleet's first report of this (16:52Z) called it two states; the roots corrected it 10 minutes later | An operator whose node recovered through the snapshot cannot run a prover until the export carries the state (the 0.3.14 account-dump export, the proving agent's item); a node that joined through the exec-restart path proves today | The nine snapshot boxes mine without provers tonight (each prover loop had pulled a 100 MB export every 6 s for nothing); the roots above are the 0.3.14 pin's numbers |
|
||||
| 8, 17:45Z to 18:45Z: finality at 78 to 93 voters (the 38-pod wave, 1,748 MH/s for USD 20.44/h) | 116 checkpoints determined, 110 locked in the hour; lock delay p50 1.30 s, p90 1.54 s, max 16.6 s (n 107); the first certificate carries p50 59 voters (max 80) and the fold lifts it to 86 of 93; lock share p50 90.1% of active voters (min 76.9%), 77.2% of all (min 67.0%); 1.65 certificate replacements per checkpoint (max 5). The detector's dry run on the window: no alert, no event, the 5090 band 108/115/128 MH/s (n 1,555). The first hour of the wave (16:50Z to 17:45Z) is void: `pgrep -f igneumd-0313` in box-wave.sh matched its own launching shell, so no wave node started for 55 minutes (about USD 20 of pods); the fix is the anchored pattern and the CI check `tools/ci/pgrep-self-match-check.sh` | Home miner: a lock lands about every second block at 1 block/s, 0.2 s later at 93 voters than at 56; what to watch past 100 voters is the certificate's size (the fold to 86 signatures), not the vote count. Rig: the same. Pool: a pool's one node votes once for all its members; its weight is the day's blue blocks, so a pool of 38 pods is one voter with 40% of the weight | the rows are the fleet night's finality baseline for the 100-voter question; the certificate-size series continues in the hour marks |
|
||||
| 9, 18:14Z: pool-v0 under load, 10 members, 0 shares | igneum-pool (PPLNS, port 4463) with 10 wave pods as members: every member's GPU shares were refused `WORKER MISMATCH` because the pool's job carried no program class and the member's worker hashed class v2 while the devnet's templates are class v3; 0 shares accepted in 25 minutes, the pool's stats page alive. The pool agent rebased the pool-mode miner onto 0.3.14 (pool-v0-rebase c4c92e8: the job carries program_class and era_seed, the member re-checks shares with the template's class); the rerun with both new binaries is scheduled for tonight after the block-rate runs | Pool user: on the old pool binary every share is wasted work; on the rebased pair a member without a node takes the day and the dataset size from the pool's seeds line. Home miner and rig: unaffected (solo mining never touched the pool) | the rerun's counts (shares accepted per member inside the first vardiff interval, mismatches 0, rejected 0) go to main |
|
||||
| 10, 17:25Z to 17:38Z: the 0.3.14 canary on the live devnet | the release's Linux igneumd (4c6b129d) on five restart-path boxes for ten minutes: 0 new rejected blocks on every box, exec state roots equal to the hub's at a common height, a segment record paid on the new binary, every node on the new version: PASS on all four criteria; the first run's FAIL was the gate counting the hub (whose trait is about 3 rejects a minute) and a `cut -c1-140` that truncated the digest line before the match | Operator on the restart path: the swap is a binary change with the data dir kept, seconds of downtime, no replay. Snapshot-path operator: the same binary, the snapshot recipe unchanged | the canary form (canary-next.sh) is the fleet's release gate; the 0.3.15 run is in flight on four boxes at 19:41Z |
|
||||
| 11, 18:31Z to 19:35Z: the class v4 rehearsal on igneum-devnet-400 (16 nodes, 38 pods joining, one stale 0.3.13 box) | The flip by miner signal landed at DAA 1,200 (19:00:5xZ) with the identical line on every node ("epoch 2 ... seed block f0606e20..., threshold 9500 bps, 599 of 599 blue blocks"), one program id per epoch across boxes (epoch 2 0x24304f0788ea9408, epoch 3 0xcc266b4f5dbc3447, epoch 4 0x634018bab5e5f283; the Mac's igneum-pow show agrees and the v3 id differs), 0 PoW rejected all run, the stale box refused by digest at every connect (12 by 19:46Z) and refusing the sixteen-field file at parse ("unknown field program_class_v4_activation_daa"), the floor at 2,400 crossed with v4 in force and nothing to print. Two findings outside the protocol: the plan's CPU engine (5 kH/s a box) cannot move a chain whose genesis difficulty is 2^27 (the CUDA worker fixed it: 14 to 98 MH/s a box), and the 0.3.14 package's igneum-worker-cuda refuses a generator-4 pack ("not a generator version this worker runs (2 or 3)"), so the chain stood at DAA 1,202 from 19:00:5xZ until the release tree's generator-4 Linux worker (97e036e2) went on at 19:15Z; a worker restarted on its old --pack after --exit-on-seed-change also answers every job "epoch seed mismatch" until the pack is re-exported | Home miner, any card: at a class flip the old worker stops dead; the flip is safe only when the generator-4 worker ships in the hive, Mac and Windows packages before the signal window closes, and the app re-exports the pack at every seed change. Rig: the same, times eight. Pool: the pool's node decides the class; a member on an old worker mismatches every share | the cut's rule (publish 2 gated on every worker, not every node) is with the shipper; P1 PASSED is on master |
|
||||
| 12, 18:39:40Z onward: the live devnet's finality paused at 93 voters | On every node read (hub, p1-3080, wave-05) the last lock is checkpoint 6842 at 18:39:36Z; checkpoints keep being determined (6967 at 19:43Z) and no certificate is received or built anywhere after 18:39:40Z. The fleet's boxes never left the live devnet (every live node synced throughout). The hand nodes (node 1, the observer) were down from about 18:2xZ to 19:42Z with the desktop app; every rented node dials only them and the hub (no rented box has an inbound port), so the votes had no path to the aggregators: the star topology that run A (10 blocks/s, 77 to 84% red blocks) shows on Devnet 2 | Home miner: with the hands down, nothing a home node does restores locks; the fix is more public peers, which is the standing fleet's ports rule (a mapped p2p port on every standing box the provider allows) and the hands' move to a box. Rig and pool: the same | the re-read at 19:52Z decides whether the owner is routed (no certificate ten minutes after the hands' return) |
|
||||
|
|
|
|||
Loading…
Reference in a new issue