diff --git a/docs/plans/counter-asic-3-status.md b/docs/plans/counter-asic-3-status.md index b96a26a45..8b4e9ddd0 100644 --- a/docs/plans/counter-asic-3-status.md +++ b/docs/plans/counter-asic-3-status.md @@ -694,6 +694,13 @@ THE MINI'S SECOND STALL, A DIFFERENT CLASS (the shipper, 04:0x UK, once): since KIT 2's FLEET CANARY FAIL, THE ROLL HELD (04:04 UK, the fleet lane on lp-4090-11; the record build-1:/srv/canary/1176efb9/lp-4090-11/, block sha 841fbbdc, FAIL at 02-sync, no shard): the kit installed and ran clean (x86-64-v3, avx512f 0, both --version lines); the node, dialling the three reference hubs, captured four of the reference chain's cuts from the fork point at block 4,893 (the epoch-3 cut) and refused the fifth (dac3dd04) at fix 2's walk bound, verbatim "the walk from dac3dd04... down its selected-parent chain passed 7200 blocks (2 x epoch_blocks) without meeting a block this executor holds a record of", every 60 s since 03:52:49 UK; walled at blocks 16,492 / headers 32,398 against the reference's DAA 41,804. The convergence roll is HELD; no 2.0.2 roll tonight; kit 1 no fallback (the executor wedge). No proved shard, so under the coordinator's condition nothing reached any dl folder: the Windows payload, the installer kit and the hive package stay on the Mac's staging and build-1; PC 2 keeps 2.0.1 (the relay lane's default); the Windows and HiveOS entries are the morning's; the Mac entry stays live as cut (min_supported 2.0.1; a withdrawal changes no user's outcome). The two lanes' class readings, to be settled from the fleet lane's 01-datadir line by 04:30: the fleet lane read the canary against kept datadirs (every island forked at that cut eight epochs back, a roll on kept datadirs heals nothing); the shipper read it as a fresh node from genesis, the reset's exact shape (the executor lagging the header sync by more than two epochs on a fresh IBD). The node stays up on its FAIL for the node lane's read. The fleet block (FAIL, both tarball shas) lands beside the mac block. MAIN'S RULING ON THE FAIL (04:1x UK, relayed in those words to the fleet lane, the shipper and the node lane): (1) the gate's baseline is the 04:0x island count and the 06:00/07:30 refinement runs on it; (2) the reset is a fresh-datadir sync from the three reference hubs on kit 2 everywhere if the hive canary on lp-4090-43 (fresh at 03:58, the three hubs only, bound 90 minutes) PASSes by 05:30, HELD for the morning otherwise (2.0.1 fresh walls at the first cut, the 3,600/4,288 tip-lag class); the mini in scope; (3) the hold's declared-load run and the F07 rows on 2.0.2 reported as not done tonight; (4) software first: the node lane asked for a fix-3 candidate (the walk bound lifted to the fork point, or the capture walking by the reference's own epoch references) with a clock and a canary before any cut, no cut without a read-back, the reset untouched by it, a line by 05:30 or the reason it is a 2.0.3 item. One addition: the shipper runs the mini's fresh-datadir reset at once (the datadir set aside, kit 2 as installed, the three reference hubs only, a pid file, the rule-33 lines read back) as the second data point of the reset's class and so the mini mines before 08:00 if it clears the fifth cut; if it walls it is left with its log for the node lane. The 08:00 line opens on the live heal, then this fault with its class, then the three home machines. +THE CLASS SETTLED FROM THE LINE (the fleet lane, 04:2x UK): lp-4090-11's 01-datadir reads "aside=dn4.aside-canary-5d53a591-022915 entries_in_new_datadir=1", so kit 2's fleet canary was a fresh 2.0.2 node from genesis; what differed from the reset's shape was the dial list: its SEED held seven addresses (the three reference hubs plus four fleet boxes on islands) and its first IBD came from an island on the common prefix, so the executor's first records went on that island's chain and the reference chain's cuts arrived as foreign cuts; the walk-bound class is "a fresh IBD whose executor took its records from an island before the reference chain" (the executor at 16,492 while headers reached 32,398), a harness dial-list fault beside it; the fresh canary's rule from now: the three reference hubs only. The hive canary on lp-4090-43 (the reset's exact shape) at the first cut with the lost-proof class looping every 30 s ("the proofs of 2 carried records are not held yet (the relay has not delivered them): shard record block 3687 shard 0") while its blocks climb (7,440 to 8,115 in four minutes); its bound to about 05:30 UK. THE GATE v2 read back: ~/igneum-fleet/reset-gate-v2.sh (sha256 e34386ec3cbd299e, tools/fleet/ on gpu-fleet at 1e0a1bb30 on the box remote), pid file pids/reset.pid, decisions in reset-decision.txt; its lines: "BASELINE (the kit-2 fleet canary's FAIL at 03:04Z, no roll end): islands=29 | distinct_tips=29 among 33 mining boxes (all 99 nodes: 68 tips)" and "GATE ARMED for 05:00:00Z (06:00 UK); baseline islands=29; the reset needs the hive canary's PASS or gate-override.txt"; the reset, if it fires, kit 2 by its named sha (kit2-sha.txt) onto lp-4090-04's chain, SEED the three reference hubs, empty datadirs, an operator reset; otherwise "RESET HELD for the morning line"; gate-override.txt carries main's word only. +THE MINI'S FRESH-DATADIR RESET (main's addition; the shipper, started 04:12 UK by ~/mini-reset.sh under ~/mini-reset.pid): api/quit bounded, the window host ended by its recorded pid, the island datadir moved aside and kept at ~/Library/Application Support/Igneum/devnet-4.island-9b980b2f- (never deleted; the node lane's veto read lives there: whether the mini's executor state diverged from lp-04's, 2.0.2's acceptance-rule payouts against a 2.0.0 hub's pool-read payouts being a consensus split between versions if so), the installed 2.0.2 DMG as it is, relaunched with IGNEUM_APP_PEERS set to the three reference hubs only (the engine's env override, the packaged list untouched), the sync watched to 75 minutes; the outcome line to come (the DAA reached, the fifth cut passed or the wall line verbatim, the minutes without mining, the chain). +FIX 3 AND KIT 3 (the node lane and the HEAL lane, 04:15 UK): foreign-seed-capture-2 09350682 = kit 2's 5d53a591 plus one commit in igneum/exec/src/service.rs: the walk from a foreign cut block used to stop only at a block this executor held a record of, so on a fork many epochs old every capture walked to the original fork point and the fifth passed the 2 x epoch_blocks bound; now the newest successful capture's scratch state stays as the base of the next capture on that chain (the walk stops at its tip), the walk bound 4 x epoch_blocks, the execution bound per capture unchanged; known-failed test first, igneum-exec 95 green on build-7; the two-node heal with the link held across three cuts in its gate (04:35). KIT 3 = kit3-line 09350682: the node lane's six suites, pair, canary set, a fresh read and the kept-datadir read (a copy of build-1's seed datadir, the 6,162 fork eight epochs behind, dialling the three lp hubs: lp-4090-11's class) on build-9 and build-1; the sha line with the suites by 05:00, the reads by 05:30. The fleet kit of 09350682 cutting on build-1 from 04:13:58 UK (the build-server lane, pid 40732, the seed pair from the warm cache, a never-overwrite guard on the tar step; the record 202-evidence-09350682 with a 203-09350682 alias), fleet kit only; it reaches lp-4090-11 only on the six-suite green line, then the fleet lane's two evidence runs (the kept-datadir heal of the walled datadir, the rule-33 fresh canary dialling the three reference hubs only), neither a roll, the reset as armed. The strike rule (a vetoing node never strikes) already in the HEAL lane's b92aafc7. The 2.0.3 line: d1a0aa8c and 07fdac4f (with P22) six suites green, "igneumd 2.0.3", reads against the reference running; the read by 07:45. +TWO INFRASTRUCTURE FAULTS (the fleet lane's 04:07 UK watch read, under the night rule): (1) hub-1's disk at 99 percent (60 GB overlay, 771 MB free): the holders the 19 GB devnet-3 datadir put aside when devnet-3 was switched off (dead since the 2.0 cut), the LIVE node's 15 GB on 16516 (untouchable), dn4 at 7.1 GB growing, logs (node.log 248 MB, dn3-node.log 70 MB, miner-0.log 32 MB); done at once and reversible: the dead logs gzipped, node.log archived and truncated, dn4-node.log archived beside itself; the lane's default (rm -rf of the one aside directory) not given by the coordinator, a kept datadir deleted for good being main's or the founder's word (with main, deadline 04:45 UK, silence keeps); running instead: the aside compressed in place file by file under a pid file (every byte recoverable), the free GB and hours left at dn4's growth in every watch read, a deletion of the compressed files largest first only if free space falls under 300 MB before the word, said once. (2) lp-4090-02 (a hub) died a second time on the open-file class at processor.rs:267:30 ("Too many open files" opening a consensus log for appending); restarted by pid files with the raised limit; the 02:56 restart path had dropped the ulimit raise, a harness fault fixed once; the class on the node lane's 2.0.3 list with the processor.rs:542 panic. +MAIN'S WORD ON hub-1's DISK (04:4x UK): keep. The devnet-3 aside compressed in place file by file under a pid file; the free GB and hours left in every watch read; the emergency rule (under 300 MB free, the compressed files deleted largest first to keep the LIVE node alive, said once) standing; the compressed archive moved to build-1 under /srv/archive/hub-1-devnet-3/ with a sha list when the fleet lane has the bandwidth after the gate, so the delete on hub-1 becomes a move; the datadir's final deletion a morning item for the founder with the disk numbers. +KIT 3's EVIDENCE KIT CUT (04:18 UK, the build-server lane, run pid 40732 from 04:13:58, the warm cache; read back from build-1 by the coordinator at 04:19): /srv/workers/fleet/09350682-node-lane.tgz sha256 23ef6f477f1bf6fd04df18f41cd1f674094b82057c0bd5db5ff48d69438a3223 (26,979,906 B, the sidecar OK); igneumd 56,630,608 B sha256 34ef5b7df6d456596907c73d45863a1e7955e14c72b79c40e2b93138b871b01f ("igneumd 2.0.2", the commit string 09350682 twice, the fingerprint cbc5bd0aa10585c8 twice, glibc 2.34, kit-isa 0); igneum-miner 66fded0f5ce15bde byte-identical to kit 2's; no lab string; the box's ship ISA gate 0 and 0 at x86-64-v3; the record /srv/artefacts/203-09350682 (a link to 202-evidence-09350682, one record two names) with run-evidence.log and the seed pair; fleet kit only, no package, no push; the script refusing to overwrite an existing fleet file (the skip guard by construction), kit 2's de47da86 untouched. It reaches lp-4090-11 only on the node lane's six-suite green line on 09350682 (05:00). The build-server lane's two cuts done and holding. + A SHARED-DEVNET FACT FROM THE FLEET (not this lane's, with the shipper and the infra lane): the Hetzner live seed 188.245.5.161:26611 is still on the old override object (digest eada4bda) 1 h 40 min after the 0.3.20 sweep (the fleet never touches Hetzner nodes, so it was outside the sweep); the 0.3.21 wipe canary c22-1 took five digest-mismatch rejects from it; an app with the packaged peers is refused at the seed and syncs through node1 and the hub only, a fresh joiner with only the seed cannot join, the 14 voters and the hub are unaffected; the owner puts the floor file ov16-floor-900000.json (sha 294f1f80) and the c4459193 pin on it. 0.3.21's STAGING (the node lane): the order dry-merges onto 55768f88 with nothing moving to 0.3.22; the late-join fix is 52e96c94 (70e4601e rebased onto 55768f88, exec suite 33 green with both new tests); f067f7c1, b0444f51 and 437f0438 merge clean in order; 2e32d5f6's one conflict (DST_ADDRESS beside pool-finish's DST_BINDING in consensus/core/src/finality.rs) kept both; the live-file digest eada4bda after each (every switch at never); the staging waits on the shipper's sweep-end word; the re-pin held. PC 2 DOWN AGAIN (main, 16:5x UK): the founder takes PC 2 down for cable work (PC 1 back but his desk); both PCs out of the sweep's waves, each updates on its poller on return; no PC job to PC 1; the Windows G1 completed before the outage, nothing reruns. 0.3.21's SECOND GATE LINE on 55768f88 (sha256 279b1b690e854fc9): the ten-minute mixed-version gate beside the 5899f603 pair, 13:37:40Z to 13:47:52Z, SUMMARY PASS (one digest b0afb2ee on five nodes; 223 new and 381 old blocks accepted by the old hub, 0 rejected; counts equal at 319, 486 and 604 through both clean joins and the restart step at 13:45:22Z; no panic); the node lane's two lines on 0.3.21's first candidate complete, in plan 6.9 on ca3-v4-node; the fleet's set on it (the bare-child 12 GB line, the wipe, the kept read, the cases) is the fleet's. 0.3.21's FIRST GATE LINE on 55768f88 (sha256 279b1b690e854fc9, the string read back; pairing igneum-pow 8c728ca3 at byte 5): the digest gate 13:35:41Z to 13:37:19Z SUMMARY PASS (a89be8a7 on both binaries with the peers; db9a85f9 refused, no peer; the live file's eada4bda unmoved); the ten-minute mixed-version gate from 13:37:40Z, line about 13:50Z. The 0.3.21 order as the shipper sent it: 55768f88; f067f7c1 and 70e4601e; b0444f51; 6eb21fc9; db28d331; then the re-pin from 8bdcbdd8 on the coordinator's word; suites between, the digest read after every one; the mirror's release-0.3.20-node back at the pin c4459193, release-0.3.21-node open at 55768f88. THE LATE-JOIN COMMIT (N9's second half, the node lane): 70e4601e on the box mirror as branch proof-hold-fix, from c4459193, two files (igneum/exec/src/proving.rs, protocol/flows/src/v10/proving.rs); the gap was the fetch side on the joiner (the served record ran the native check against the joiner's trailing exec state before anything was stored, the check refused it, the proof was never held, the body rule read "not held" for 20 s and failed the IBD); the fix holds the proof by hash before the checks (the pool entry still needs them) and the serve side says when it holds fewer than asked; the exec suite 32 passed at 13:26Z with the known-failed shape first, the flows check green 13:28Z, igneumd on build-1 at the 0321 worktree path built 13:32Z, sha256 17649eeb2f7d1290, string read back; with the testnet lane (the resume form, B alone); it joins the 0.3.21 staging as its own commit. THE WIPE CANARY ON c19-1, c4459193 (sha 45be9b02d1b002f5, string read back): FORM END rc 0 at 13:50:53Z. Wipe synced 13:35:50Z (57 minutes, inside the 98-minute class); mining 13:36:00Z to 13:47:07Z, 66 mined, 66 accepted, 0 rejected, isSynced true at the tip throughout; the hub holds 41 of its blocks in its last 700 with 0 rejects (13:47:09Z); the restart on its kept datadir at 13:47:15Z: the old process stopped at once (the new process's first lock line seven seconds after the marker; the watchdog held nothing, the b7cc37e7 fault closed), synced again at 13:48:39Z after 84 s, 109 templates read with max 3,432 ms and 0 timeouts; the kept read on pool-1's 0.3.17 copy on the same pod passed at 13:38Z (the rewrite line once, a clean second start). The pin's set on c4459193: the digest gate PASS, the mixed-version gate PASS, the wipe canary PASS, the kept read PASS, the restart PASS, the 12 GB line proves and verifies (paid is a race, not a gate); CASES END from c20-1 (about 14:50Z) is the last pin line. THE INTEROP FACT stands from the void run: the 5899f603 hub accepted 235 object-byte-5 blocks from the 8097d600 node with 0 rejected, one digest on all five nodes on the live sixteen-field file. The gates: the digest test and the kaspa-pow vector test (the amended devnet epoch-0 id 1a4230699a6b9c60 must equal, c120d7963abdcd96 must differ, the v3 control unchanged) on the box; the mixed-version Devnet 2 gate (the amended 0.3.20 node beside a 5899f603 node for ten minutes on the live file without the v4 fields) after the Mac build; the fresh-join canary the 0.3.20 cut's | | Main's rulings (7 October, morning) | no generator change to v4 on the live devnet; the record's null is the window model with numbers, sent by the hash lane to the attack-pass lane so AP-F8-1 re-gates against it; a fault beyond the model (a low-entropy source at site 15) stops at the coordinator with the two options priced (a 0.3.19 class amendment before the flip, or the flip held at the floor), nothing shipping without the founder's word; the tighter tail, an acceptance bound on the hot-set share, is a CLASS V5 item (sent to the v5 lane a6410f3b8abefb762 with the 64-seed census as its gate; the bound's number follows from the model) |