diff --git a/docs/plans/consequences-2026-10-05.md b/docs/plans/consequences-2026-10-05.md index 5b118a147..b366062f7 100644 --- a/docs/plans/consequences-2026-10-05.md +++ b/docs/plans/consequences-2026-10-05.md @@ -70,3 +70,5 @@ Merge note for the integrator (22:3x UTC): branch `consequences` is docs/plans/c | C31 | The copy the ship takes (ca2-coord at 22:3x): litepaper line 562 "12 GB or more proves full shards" beside line 452 "Proving needs an NVIDIA card with 24 GB or more"; evidence row 16 still "A 12 GB card proves one shard in about 20 s, designed" while proving-v1 440fd59 withdrew it; line 562 states the dataset "doubles on a step schedule fixed at genesis (years 4, 12 and 28)" | ca2-coord site/litepaper.html 452 and 562, docs/evidence.md 43; proving-v1 440fd59; D4 | every reader of the litepaper; the integrator; the project lead's D4 | One page says 12 GB and 24 GB for the same thing; the evidence table on the ship branch contradicts the measurement, and the two branches will conflict on evidence.md and litepaper.html at the merge (ship order: ca2-coord before proving-v1), so the stale row can win by accident; and the step schedule is written as a genesis fact while the coordinator's own 22:50 entry calls mapping (b) a recommendation for the project lead (gate 1, D4): a public page should not decide a genesis parameter before he does | On ca2-coord: strike "12 GB or more proves full shards" from line 562 (line 452 is the sentence); take proving-v1's evidence row 16 (WITHDRAWN, 24 GB measured) at the merge and say so in the merge plan; write the growth sentence as the recommendation it is ("the plan is a step schedule ... decided at gate 1") until D4 is taken | coordinator ada8afb62d752b1e2 | closed on ca2-coord a7be43f and 0d9b23d (the 12 GB clause struck; the merge rule "take proving-v1's row 16 and its proving sentences" in the rollout plan; one wording in all five places, "the proposed schedule, fixed at the testnet genesis: 2 GB, doubling at years 4, 12 and 28", the lifetime sentences kept as consequences of the proposal and marked approximate) | | C32 | The 0.3.10 install at 21:49Z cleared PC 1's app jobs folder and with it the AMD kit fetched at 21:23:59Z; the amd-card-test playbook now says its fetch must be republished after any app update | bench-log 39f02ff; status 21:05 ("the jobs folder is cleared by fetch jobs" was the earlier, wrong reading) | every PC job tonight and tomorrow; the 0.3.11 rollout | A class, not one playbook: every fetch-then-run pair (era, hot table, mixer, repro, Ember, the AMD sweep, the prover-floor build) loses its kit when an update lands between the fetch and the run, and the run fails in seconds or, worse, runs against a stale copy. The 0.3.11 update-now reaches PC 2 while the prover-floor agent's 60 to 90 minute server build runs there (go at 22:17, to about 23:50): if that build's working directory is under the app's jobs folder, the update wipes it mid-build and the 12 GB rows slip past the morning | (1) The 0.3.11 update-now is sequenced after the prover-floor build closes, or the build's directory is confirmed outside the jobs folder before the ship; (2) every run playbook begins with a presence check of its kit and fails with "kit missing: republish the fetch after the app update" (the class check: the bash-body sub-agent's CI check gains a rule that a run job naming a kit path tests it first, or the coordinator's queue re-fetches after every update as a rule) | coordinator ada8afb62d752b1e2 (the queue and the ship order) | taken (coordinator: the prover-floor build lives under /opt/igneum-floor in WSL2, outside the jobs folder, but it is the app's job process and an app restart ends it, so the 0.3.11 update-now goes to PC 2 only after floor-build-3 closes, the ship's earlier steps not waiting; the re-fetch rule and the presence-check rule are in the rollout plan beside the one-job rule, the playbook owners carry it at their next publish; the CI side is with the bash-body sub-agent as a kit-path check). CI side closed on bash-body-check e3bd761: tools/ci/kit-path-check.sh in ci.yml and in publish-jobs.sh add --kind run (34 tests pass); every existing kit-using playbook (13 across master, ca2-v3, ca2-analysis, rdna4-telemetry) already checks before use, so the gate guards the shape without a backlog | | C33 | `/api/live` at 22:33Z: 13 of 482 blocks fully proven in 10 minutes (2.7%), 1 prover, median proof lag 46 s, the live node's verifier "Off" with 42 pool entries pending and 0 verified; `/api/stats` (the documented public API) carries no proving field at all | live and stats handlers (`site/api/stats.mjs` FIELDS; the explorer branch 3e01212); the homepage "~60 s to a proof"; evidence row 15 | everyone who reads the public API or the homepage tile; the testnet's first external reader | The public stats API hides the one number that qualifies the tile and row 15: coverage is 2.7% with one prover, and the proof lag is 46 s. A reader can find it only on /api/live. The live node's verifier "Off" (42 pending, 0 verified) is the Mac app node in trust mode or without a host, so the page says "verifier Off" while the chain pays provers: a public-page oddity the morning reader will ask about | `/api/stats` gains a `proving` object from the same live_state (`blocks_10m`, `blocks_fully_proven_10m`, `shards_paid_10m`, `provers_10m`, `median_proof_lag_s`, `active`), documented in docs/api/public-stats.md and in its contract test; the live page's verifier line names which node it reads and why it is off; the homepage tile's "~60 s" caption cites the measured 46 s median and the 2.7% coverage ("the target; today one prover covers 2.7% of blocks at a 46 s median") | explorer a76f60b415859b7b5 (the API); the tile caption: D2 wording for the project lead | closed on explorer d7e797c (/api/stats carries the proving object with coverage_10m, 0.0273 at 22:33Z, in the contract test, the live check and docs/api/public-stats.md with the sentence that coverage is what a third party reads before "every block is proven"; the live page's proving legend and the API note say the observer's node runs its own verifier off and reads paid shards from the chain). The homepage tile caption stays with the project lead (D2) | + +Sweep 11 (22:38 UTC) notes, no new row: gate G4b GREEN (a real Metal miner across a v3 boundary; a second Metal-only fault found and fixed first, 00c55aa: serveDataset keyed the day dataset by day alone and would have hashed v3 over an x1 dataset; the coordinator's next-cut rule: every worker path mines across a boundary in the gate network before a class change ships). The reviewer checked the PCs' side of that class: the CUDA worker and the OpenCL host build each resident pair (program, cache, dataset) from the pack's own memhard.h and key it by (epoch, day, class, era) through pairIsClass (proto-cuda/nvrtc/worker.cpp 451, proto-opencl/host.c 1106 on ca2-v3 fa3c932), so the fault does not reach the PCs at N4. The patched sp1-gpu-server built green on PC 2 at 22:32:29Z with sm_86, sm_89, sm_120 (C26's arch list), the floor sweep (9 points) running; the integration merges (readwidth 30ff674, origin/master 1f0d62c) on ca2-v3 49c7e78 with every check green. All gates but G5 (the ship's build) are green; the ship waits on the merged tip.