Counter ASIC 2.0 status 21:57: the era PC job, the verifier regression gate item, the PC 2 order

This commit is contained in:
igneum-labs 2026-10-05 21:57:48 +00:00
parent 5fc5e100ad
commit 3a6beaa0b8
2 changed files with 9 additions and 1 deletions

View file

@ -101,7 +101,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's
## 8a. Proving v1 rides with it
the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL a223ca9 on 5b0d54f (6dc686a plus the resume fix: the resume path re-arms every slot without a live worker, re-exports the pack, and logs "<card> is not mining 90 s after resume"; the known-failed case of PC 2's 21:25:11Z resume is the unit test the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old, `cargo test -p igneum-app resume` 3 passed; cause: Cmd::Resume re-armed only faulted slots after stop_miners had cleared every restart_at; the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which.
the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL a223ca9 on 5b0d54f, docs-only 8b47073 after it (6dc686a plus the resume fix: the resume path re-arms every slot without a live worker, re-exports the pack, and logs "<card> is not mining 90 s after resume"; the known-failed case of PC 2's 21:25:11Z resume is the unit test the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old, `cargo test -p igneum-app resume` 3 passed; cause: Cmd::Resume re-armed only faulted slots after stop_miners had cleared every restart_at; the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which.
Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths.

View file

@ -434,3 +434,11 @@ G4 run 2 (21:46:36 to 21:51:31 UTC, b105a55 = era + mixer, hot None): every chec
build-20261005-215219 (21:53:01 to 21:55:48Z): Linux build ok (igneumd 49,164,264 bytes sha256 11979b49..., igneum-miner d25a8270..., igneum-app 68007173...), igneum-app tests 78 + 26 + 8 passed; the node stage exit 101: kaspa-consensus 96 passed, 1 failed, processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(..., 487112384, 487129578) in mine_on_all. This is the flake the 0.3.10 cut met on the same base 21d4c73c under the six-package parallel run (release-0.3.10.md: it passed alone twice, job build-20261005-182804, 97 passed), a timing-dependent difficulty in the test helper, not a v3 change (v3 touches no finality code). Per the gate rule the publish stops here until the suites are green: job 2 of 3 (kaspa-consensus alone) is published now; job 3 of 3 (the other five crates) follows; the flake itself goes on the next-cut list (make mine_on_all deterministic under parallel load).
The proving v1 app branch is final at a223ca9 (6dc686a plus the resume fix with the PC 2 case as a unit test); rollout plan 8a updated. PC 2's prover fault is the root-socket class (agg-cost-pc2-1 ran the host as root in WSL2 and left /tmp/sp1-cuda-0.sock owned by root; the app's prover has failed every shard since 21:25:24Z); the proving agent's pc2-socket-fix.ps1 (60 s, miners untouched) runs between my two suite jobs.
## 21:57 the era PC job is done; a 2.2x CPU-verifier regression on ca2-v3 HEAD (gate item); the PC 2 order
Era job run-ca2-era-pc1-20261005 (exit 0, 301 s), both cards restored, the 9070 XT ON the bus at run time (gfx1201, 32 CUs): first rows 18.96 to 19.21 MH/s on the era packs on the 9070 XT, self-test PASS, the 2^24 fingerprints equal to the Mac's (era-4 3ace11ad84c053ae, era-5 a8897d82adceb4a1); the full 5090 and 9070 XT tables with the spread per card follow. "go PC 1" to the mixer daily-build job at 21:57.
GATE ITEM (found by the era agent, two binaries on the same v2 input, checksum 19297e99c7b9a55e, same minute, load 4 to 5): the CPU verifier on ca2-v3 HEAD (88dafbc) takes 1.332 ms per warp and on the era branch 1.310, against readwidth's binary at 0.604 and 0.606: the mixer branch's derive_items / Shape path costs 2.2x at m = 1, on the v2 path the live devnet verifies with. The mixer agent's "loaded box" reading of its 1.31 to 1.36 ms v2 figure was the code, not the load. Ordered: find and fix on ca2-mixer, restore v2 to within 5% of 0.604 ms measured the same way, re-measure v2 / x4 / x8 on the fixed binary (the C19 figures too), then the node agent merges it; the publish waits on it (G3's suite does not catch a slowdown; the 10 ms gate and the pool and IBD figures depend on it). Next-cut rule: a verifier benchmark with a pinned bound in the crate's CI.
PC 2 order: suite job 2 of 3 (kaspa-consensus alone, running), the proving agent's socket fix (60 s; the aggregation-cost restore is skipped), suite job 3 of 3 (the five other crates), the aggregation-cost re-run (20 min), the prover-floor agent's windows (through the proving agent). The proving app branch: a223ca9 is the last code change (8b47073 docs only after it).