diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 81ef0c028..f2def19fa 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -129,7 +129,7 @@ The fee switch H = 210,000 arrives about 18:45Z on 6 October (DAA 137,041 at 22: Josh delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL e0de2ab on 5b0d54f (a223ca9 the resume fix, e0de2ab the Metal --prepare-packs fix; docs-only commits between) (6dc686a plus the resume fix: the resume path re-arms every slot without a live worker, re-exports the pack, and logs " is not mining 90 s after resume"; the known-failed case of PC 2's 21:25:11Z resume is the unit test the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old, `cargo test -p igneum-app resume` 3 passed; cause: Cmd::Resume re-armed only faulted slots after stop_miners had cleared every restart_at; the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. -Merge rule for the ship (C31): the app branch proving-v1 (e0de2ab) rewrote docs/evidence.md row 16 (WITHDRAWN, the 24 GB measurement) and the litepaper's proving sentences; ca2-coord carries the older row 16 and its own litepaper edits, so at the merge take proving-v1's row 16 and its proving sentences, and ca2-coord's everything else; the stale row must not win by accident. fud-close (the ledger closer's main-repo branch: 45 public-text fixes on the site and litepaper, spec 8.3 and 8.8, two CI checks, the relay fixes; a merge-tree onto ca2-coord shows 0 conflicts) is NOT in 0.3.11: the tree closed at 23bc2b2 (its workers, DMG and PC build job carry it) before the branch reached the ship order, and the shipper takes no late branch (the 0.3.10 rule); fud-close heads the next cut's list, with a coupling the next cut must respect: fud-close's worker change for ledger M28 (packfile.h's kernel_sha256 check, host.c refusing a pack that fails it) pairs with the fork-side miner change on ledger-fixes 3d4ec451 that stamps the hashes into program.json; the workers without that miner commit refuse every pack, the miner without the workers is harmless, so both go in one cut or the worker half of M28 is held back. fud-close tip 647b08c (its checks green); ledger-fixes is not yet rebased onto 89dfcb95 (two conflicting files: igneum/miner/src/main.rs, protocol/flows/src/ibd/proof.rs). The fork-side ledger-fixes branch (from release-0.3.6, not 21d4c73c) is NOT in 0.3.11: it rebases onto the 0.3.11 fork for the next cut. Branches merged into the 0.3.11 main tree beside ca2-v3 and ca2-coord (tooling, no consensus): consequences (the reviewer's rows), bash-body-check (7adb1ca, 6805125: tools/ci/bash-body-check.sh with fixtures and a self-test in ci.yml, the convention paragraph in packaging/README-ship.md, `publish-jobs.sh add --kind run` running the bash-body and prover-socket checks before signing; add/add with proving-v1 c2544be on tools/ci/prover-socket-check.sh: take bash-body-check's superset and proving-v1's ci.yml step; tools/amd-prove/pc1-cpu-prove.ps1 goes on the allow list, it starts no GPU server), amd-prove (f1d7a7d), card-lifetime (1fecfe2). Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until Josh runs tools/keys/keygen-k2.sh), ember-tune 8ab9068 (the second-engine rule: no pipe into a second engine, its process tree killed at the end and on the budget, the installed app's miners restarted after; applied to ember-tune-pc1.ps1 and sweep-5090.ps1; tools/ci/second-engine-check.sh in ci.yml, shown to fire on a known-bad playbook and pass the fixed pair; every Cmd::Quit stamped with its source; the AMD gmax offset fix); rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. HiveOS local mode (consequences C41, 23:33 UTC, from the shipper's release-0.3.11.md section 5): `packaging/hive/h-run.sh` line 31 starts the bundled node without `--override-params-file`, so no HiveOS package ever carried the override and a rig in local mode has run on genesis parameters, refused by every devnet peer since the first height switch; the README's "republish with the new override" rule moved nothing. The fix is on branch hive-words (OVERRIDE= in the Flight Sheet's extra config written to override-params.json and passed to igneumd, the node's digest line copied to the rig log, the README row, the selftest), next cut, as ONE HiveOS item with the rigs-mine-only README, IDENTITIES=auto and the M28 worker-miner coupling. The rig installer path (igneum-node.sh writes the file from the manifest) is unaffected. The HiveOS tier is marked "untested on a GPU host, local mode never worked" in the morning summary and in 6b. Miner resubscription (consequences C43, 00:14 UTC, from the shipper's finding at node 1's 23:57:12Z restart): the Mac's miner lost its template subscription when its node restarted and never reconnected (5,300 "Not connected to server" submit errors, the template frozen at 2,350, 26 MH/s burned for about 11 minutes until restart-miners-0311-d937c69d at 00:08:01Z); the app's watchdog cannot see it (its rule fires on no print for 90 s or zero for 60 s while synced, and the miner printed 26 MH/s throughout). Every node restart (each override publish, each update, each crash) exposes the path for every home miner, and a rig in grpc:// mode against a remote node has the same path with nobody watching. Two next-cut items: (1) the miner resubscribes when the template stream ends, or exits with a named code after N consecutive submit failures so the app's restart loop takes it; (2) the app watchdog gains "accepted blocks or a template change in the last 120 s while the node is synced" beside the hash-rate test. The update-now queue (00:14 UTC): an update-now job queues behind a running job (PC 2's waited 11 minutes behind a hung sweep's cap); next cut, the update-now pre-empts or the STATUS line reports the queue wait. Explorer (00:16 UTC, the consequences reviewer): master 30cd292 (the live site from about 00:13:45Z) took the explorer branch's first four commits only; 3e01212 (the supply tile "cap 4,000,000,000, about 3.96 billion ever minted" and the 503-with-reason on a schema gap, C13) is safe to merge on its own, no CI change; d7e797c (the strict public-api-check) alone waits for the check's retry rule (C36). Until 3e01212 lands the explorer tile reads "of 4,000,000,000 by the rule" beside the homepage's "4B hard cap". +Merge rule for the ship (C31): the app branch proving-v1 (e0de2ab) rewrote docs/evidence.md row 16 (WITHDRAWN, the 24 GB measurement) and the litepaper's proving sentences; ca2-coord carries the older row 16 and its own litepaper edits, so at the merge take proving-v1's row 16 and its proving sentences, and ca2-coord's everything else; the stale row must not win by accident. fud-close (the ledger closer's main-repo branch: 45 public-text fixes on the site and litepaper, spec 8.3 and 8.8, two CI checks, the relay fixes; a merge-tree onto ca2-coord shows 0 conflicts) is NOT in 0.3.11: the tree closed at 23bc2b2 (its workers, DMG and PC build job carry it) before the branch reached the ship order, and the shipper takes no late branch (the 0.3.10 rule); fud-close heads the next cut's list, with a coupling the next cut must respect: fud-close's worker change for ledger M28 (packfile.h's kernel_sha256 check, host.c refusing a pack that fails it) pairs with the fork-side miner change on ledger-fixes 3d4ec451 that stamps the hashes into program.json; the workers without that miner commit refuse every pack, the miner without the workers is harmless, so both go in one cut or the worker half of M28 is held back. fud-close tip 647b08c (its checks green); ledger-fixes is not yet rebased onto 89dfcb95 (two conflicting files: igneum/miner/src/main.rs, protocol/flows/src/ibd/proof.rs). The fork-side ledger-fixes branch (from release-0.3.6, not 21d4c73c) is NOT in 0.3.11: it rebases onto the 0.3.11 fork for the next cut. Branches merged into the 0.3.11 main tree beside ca2-v3 and ca2-coord (tooling, no consensus): consequences (the reviewer's rows), bash-body-check (7adb1ca, 6805125: tools/ci/bash-body-check.sh with fixtures and a self-test in ci.yml, the convention paragraph in packaging/README-ship.md, `publish-jobs.sh add --kind run` running the bash-body and prover-socket checks before signing; add/add with proving-v1 c2544be on tools/ci/prover-socket-check.sh: take bash-body-check's superset and proving-v1's ci.yml step; tools/amd-prove/pc1-cpu-prove.ps1 goes on the allow list, it starts no GPU server), amd-prove (f1d7a7d), card-lifetime (1fecfe2). Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until Josh runs tools/keys/keygen-k2.sh), ember-tune 8ab9068 (the second-engine rule: no pipe into a second engine, its process tree killed at the end and on the budget, the installed app's miners restarted after; applied to ember-tune-pc1.ps1 and sweep-5090.ps1; tools/ci/second-engine-check.sh in ci.yml, shown to fire on a known-bad playbook and pass the fixed pair; every Cmd::Quit stamped with its source; the AMD gmax offset fix); rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. HiveOS local mode (consequences C41, 23:33 UTC, from the shipper's release-0.3.11.md section 5): `packaging/hive/h-run.sh` line 31 starts the bundled node without `--override-params-file`, so no HiveOS package ever carried the override and a rig in local mode has run on genesis parameters, refused by every devnet peer since the first height switch; the README's "republish with the new override" rule moved nothing. The fix is on branch hive-words (OVERRIDE= in the Flight Sheet's extra config written to override-params.json and passed to igneumd, the node's digest line copied to the rig log, the README row, the selftest), next cut, as ONE HiveOS item with the rigs-mine-only README, IDENTITIES=auto and the M28 worker-miner coupling. The rig installer path (igneum-node.sh writes the file from the manifest) is unaffected. The HiveOS tier is marked "untested on a GPU host, local mode never worked" in the morning summary and in 6b. Miner resubscription (consequences C43, 00:14 UTC, from the shipper's finding at node 1's 23:57:12Z restart): the Mac's miner lost its template subscription when its node restarted and never reconnected (5,300 "Not connected to server" submit errors, the template frozen at 2,350, 26 MH/s burned for about 11 minutes until restart-miners-0311-d937c69d at 00:08:01Z); the app's watchdog cannot see it (its rule fires on no print for 90 s or zero for 60 s while synced, and the miner printed 26 MH/s throughout). Every node restart (each override publish, each update, each crash) exposes the path for every home miner, and a rig in grpc:// mode against a remote node has the same path with nobody watching. Two next-cut items: (1) the miner resubscribes when the template stream ends, or exits with a named code after N consecutive submit failures so the app's restart loop takes it; (2) the app watchdog gains "accepted blocks or a template change in the last 120 s while the node is synced" beside the hash-rate test. The update-now queue (00:14 UTC): an update-now job queues behind a running job (PC 2's waited 11 minutes behind a hung sweep's cap); next cut, the update-now pre-empts or the STATUS line reports the queue wait. Explorer (00:16 UTC, the consequences reviewer): master 30cd292 (the live site from about 00:13:45Z) took the explorer branch's first four commits only; 3e01212 (the supply tile "cap 4,000,000,000, about 3.96 billion ever minted" and the 503-with-reason on a schema gap, C13) is safe to merge on its own, no CI change; d7e797c (the strict public-api-check) alone waits for the check's retry rule (C36). Until 3e01212 lands the explorer tile reads "of 4,000,000,000 by the rule" beside the homepage's "4B hard cap". The app's /api/state answers "{}" at times (the aggregation-cost agent, 00:27 UTC, from job 4's raw print; full state at 21:01Z, empty at 22:22Z, 22:41Z and 00:18Z on PC 2): engine.rs 180 (`serde_json::to_value(st).unwrap_or(json!({}))`) swallows the serialisation error, and the state carries f64 fields (difficulty, last_reading_age_s, hash_now, hash_avg, template_age_s) that serde_json refuses when non-finite, so one NaN or infinity empties the whole answer for every reader (the window, any script) and the app logs nothing. First item of 0.3.12 (the proving-v1 agent has it as the app owner tonight): finite floats at the source (0 or null where a reading has no value), the error logged with the field name, never an empty object; a unit test with a NaN in each f64 field. ## 9. After the publish: Counter ASIC 3.0