diff --git a/docs/plans/release-0.3.10.md b/docs/plans/release-0.3.10.md index b825916ea..7e3df4e56 100644 --- a/docs/plans/release-0.3.10.md +++ b/docs/plans/release-0.3.10.md @@ -481,7 +481,7 @@ The next cut (0.3.11), decided by the coordinator on 5 October 2026 night: | C1 (the consequences reviewer, 20:0xZ): the 0.3.10 node's `igneum_exportSegments` (fork `igneum/exec/src/rpc.rs`) writes no `daaScore` and no `feesV1ActivationDaa` per segment, which the 0.3.9 exporter needs to replay both sides of the fee switch (`export/src/main.rs`: "a dump without `daaScore` is accepted only when the switch is never or 0"). Measured here: the handler (90 lines from `rpc.rs` 849 in `vendor/igneum-node-0310`) carries neither key (`daaScore` appears in the fork only in other RPCs, lines 239, 423, 819); the reviewer names `vendor/igneum-node-pv1` at eb32c645 (on 21d4c73c) as the carrier: its `rpc.rs` writes `daaScore` per segment (line 961) and `fees` and `feesV1ActivationDaa` at the top level (line 982); confirmed here at 20:4xZ (`git -C vendor/igneum-node-pv1 grep -n feesV1ActivationDaa -- igneum/exec/src/rpc.rs`: line 982); my first read of that path at 20:1xZ was wrong. So at H = 210,000 (about 19:50Z on 6 October) every 0.3.10 prover's export of a post-H block fails or meters with the wrong table and the fleet's provers go dark, unless the next node cut carries that RPC onto every prover before H, or H is republished later. The mechanism (the proving agent, 20:1xZ): the app's prover calls the exporter with no fee flags, so from H every app prover on 0.3.10 cuts with the prototype table and every statement is vetoed. The two closes: (a) the proving-v1 fork (eb32c645, on 21d4c73c) on every prover before H through 0.3.11; (b) republish H = tip + 86,400 by the fee-switch plan's rule. Decision due 16:00Z on 6 October; the coordinator, the proving agent and the 0.3.11 shipper hold the same line | open, dated: (a) or (b) by 16:00Z on 6 October | | C5 (the same reviewer): the app kills its prover child on quit (`prover.rs` 266 to 272), so the `update-now` of this rollout aborts whichever shard each prover has in flight (up to 37 s each, no payout). Accepted as the cost of the restart; the count is read after the rollout from the provers' "stopped after" lines (section 8) so the coverage numbers measured across the restart are read with it | recorded at the rollout | | `/api/resume` answered ok on the 0.3.9 app and never restarted the miners (PC 2's 5090 worker "off" with the card holding 1.7 GB from a job's resume at 21:25:11Z until its 0.3.10 restart, the iGPU miner too; the Counter ASIC coordinator, 21:4xZ). The class: a resume that reports success without a miner restart. The app should re-check the miner processes after a resume and report a failure | open: next cut | -| PC 2's prover after the 0.3.10 restart: every assigned shard fails at once with `CudaClientError: Connect(PermissionDenied)` (section 8), where 0.3.9's PC 2 proved shards from 17:36Z. The host, the pinned ids and the CUDA backend are detected as before; the connect that fails is the SP1 CUDA prover's client socket. The aggregation-cost agent's job `agg-cost-pc2-2` ran as root in WSL minutes before and its lines show no CUDA or socket change; the 0.3.10 app's prover path changed only to read the program ids (`--mode id`). Unexplained; handed to the coordinator and the Counter ASIC coordinator (PC 2's measurement agents) | open: PC 2 proves nothing until it is found; a restart of the WSL distro or of the CUDA prover service is the first thing to try | +| PC 2's prover after the 0.3.10 restart: every assigned shard failed at once with `CudaClientError: Connect(PermissionDenied)` (section 8) from 21:49:56Z. RESOLVED at 22:01Z by a socket fix on PC 2 (the Counter ASIC coordinator's agents; the SP1 CUDA prover's client socket permission), after which PC 2 proved and was paid every 40 s. The proving agent is adding a log line that names the cause on the app branch for the next cut | closed 22:01Z | | PC 37ba0461 (the US laptop) started the 0.3.10 install at 21:40:53Z, stopped its miners and node at 21:41:16Z and had not come back by 21:56Z: the per-user installer on the owner's machine, nothing to drive remotely | open: the console shows when it returns; its 0.3.10 line and worker start are read then | | `dl/public/igneum-downloads.json` (the unsigned index the site's download page reads) alternates at the edge between the new bytes and the previous ones for over 20 minutes after the deploy (one fetch byte-identical at 21:52Z, the next three not): different edge nodes behind one hostname. The two signed manifests were byte-identical and verified from the first check. The ship's verify step counts it as a failure and refuses to post the console item, so the item was posted with `--from console` | open: the verify should accept the index after the signed manifests pass, or retry it for longer; the site serves the previous version's buttons from a stale edge until it settles | | The console's machine card keeps a machine's LAST non-empty node commit string: PC 1's card read `node 2.1.0-a24ab01a` for 17 minutes after its node had restarted as the PC build (`igneumd/2.1.0`, no commit), while PC 2's card read `2.1.0` at once; the Mac's card kept node 1's a24ab01a after node 1 moved to 21d4c73c. A card's commit string is therefore not a fact about the running node until the app re-reads it; the node log is | open: the card should show the string the app last READ, with its time, or nothing |