release-0.3.21 plan: what rides it, the under-12 GB prove-instead switch and the per-architecture prover server as designs, the gate clock

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-07 13:16:59 +00:00
parent 59fe7d08b4
commit 77350dd832

View file

@ -0,0 +1,30 @@
# Release 0.3.21 (the cut after 0.3.20; staged 7 October 2026, 14:2x BST)
Tree: release-0.3.21 on origin, from release-0.3.20's closed tree (the 0.3.20 cut: pin c4459193, igneum-pow 8c728ca3 at byte 5, the floor-moved file). Rules: docs/plans/release-rules.md (4a every gate on every candidate from its build; 4b warm pods per gate class; 5 the sweep in waves).
## 1. What rides it (the project lead's list, 7 October 2026)
| Item | Owner | State |
|---|---|---|
| driver-check 1982dbc7 (the driver table measured: NVIDIA 617.42, AMD 26.9.2 with the amd.com referer, Intel 32.0.101.9034; the install thread with one elevated prompt; `--drivers` on the manifest) | the Intel lane | in: 39187097, afe4e2eb, 416d1f5d |
| miner-reliability, app half (miner-reliability-21 aa2bb07e, 18 commits: the readiness gate and the ladder, the load clock, the hold owner, the orphan sweep, the FAULT lines to the intake, the signed `cards` job kind, execrpc with `probe`; box 226 + 32 + 8, UI 49) | the reliability lane | merged into release-0.3.21 |
| miner-reliability, miner half (miner-reliability-20 f067f7c1: STATUS while waiting for a template, identities capped from the node's template time, the fetch back-off) | the node lane | staged on release-0.3.21-node at the sweep-end word |
| sub-version 2 (the hash lane's 07a809a7, byte 7, id a788661687db4bb3), pending the F8 census (about 15:00 BST; pass = all 64 seeds under 1.2x of the window model plus the suite and G1 lines); if it fails or slips past 20:00 BST the node ships byte 5 again with the re-pin dropped | the Counter lane, the node lane | held |
| the fork gate from horizon-node (6eb21fc9, six commits on finality.rs) | the node lane | dry-merged clean into c4459193 |
| the node's own rust-toolchain.toml | the node lane | in place on its worktree |
| the under-12 GB prove-instead switch (section 2) | the shipper | to write after the 0.3.20 publish |
| the per-architecture prover server (section 3) | the shipper, the fleet | to write after the 0.3.20 publish |
| 0.3.20's version strings moved to 0.3.21 | the shipper | 0d7b9bd0 |
| master 819d536b merged (the 51-check gate; build-remote routes by load; the hands script's pgrep in the bracket form) | the shipper | c83ca904, 0b75cf52 |
## 2. The under-12 GB prove-instead switch (design, main's routing)
The fact (the fleet's 3080 hour, 7 October 2026): a 10 GB card completes the compressed step alone at the default threshold (8,642 to 8,729 MiB of 9,885) and never beside its miner's 1,547 MiB; no lower threshold fits (the server dies before the compressed step at 524288 and below). 0.3.20 ships the rule "proving needs a 12 GB card; mining continues" (provedefault::prove_refused_under_12gb). 0.3.21 adds the choice: a Settings switch `prove_instead` (off by default) that, on a machine whose only NVIDIA card is under 12 GB, stops that card's miner while the prover runs and restarts it when the prover is off; the tile sentence becomes "proving instead of mining on <card> (10 GB holds one, not both)"; the refusal line stays when the switch is off. Engine: the prover loop's refusal rung reads the switch; when on, it takes the miners hold for that card (the hold owner from miner-reliability) for the prover's lifetime. Test: the refusal lifts with the switch on and the card's miner reads held; a 12 GB card beside it never triggers the hold. UI: the switch under Settings > Proving with the sentence; the view test.
## 3. The per-architecture prover server (design, main's rule)
The fact (the fleet, 7 October 2026): sp1-gpu-server is built for one card's compute capability (sm_86 Ampere, sm_89 Ada, sm_120 Blackwell); a server for the wrong architecture fails every proof in 12 s with "CudaRustError: named symbol not found" and the miner never notices; the 610-series driver proves (the driver floor stays 570 or newer, no upper bound). The app's WSL2 setup (setup-wsl.sh) builds the server on the machine, so it matches by construction; the risk is a copied or stale server. 0.3.21: (1) the prover's setup records the compute capability the server was built for (nvidia-smi --query-gpu=compute_cap) beside the binary; (2) at every prover start the engine reads the card's compute_cap and the record, and refuses with "the prover was built for sm_NN; this card is sm_MM; Set up rebuilds it" when they differ (MF-10's capability check in prover.rs, owed by the reliability lane's register); (3) the fleet's kit ships one server per architecture under bin-<arch> picked by compute_cap at install (the fleet lane). Test: a record and a card that differ refuse with the sentence; equal ones pass; no record passes with a warning line.
## 4. Gate clock (rules 4a and 4b)
App: the gate on build-2 about 19:30 BST (`build-remote.sh --box 2 -- test --release` from app/igneum-app), the UI tests, the Windows cross on build-2, the reliability injector's eight steps on a one-shot pod (the fleet, about 45 minutes, after the sweep). Node: the branch commits at the sweep-end word (about 16:45 BST), suites on build-2 about 16:50 to 17:05, the first candidate binary on build-1 about 17:10 with the digest gate, the mixed-version gate, the kept start, the cases and the wipe all started from the build on the warm pods; the pin about 80 minutes after the last candidate builds. The Mac's own binaries and the DMG under the lock after the pin; the staging in a scratch copy; the publish on green with the clock time.