diff --git a/docs/plans/release-0.3.11.md b/docs/plans/release-0.3.11.md index 316bef162..900247805 100644 --- a/docs/plans/release-0.3.11.md +++ b/docs/plans/release-0.3.11.md @@ -50,6 +50,33 @@ Every check at cc72f4a: identity 0 hits over 220 files, copied-sources, pinned-g | PC 1 build job (the node and the app, Linux and Windows) | `IGNEUM_WIN_RELEASE=/target-integration/x86_64-pc-windows-gnu/release node tools/build-job.mjs run --node vendor/igneum-node-0311 --target ae432dc7 --targets linux,windows --no-tests` from this worktree, its own zip `build-inputs-20261005224625-29789.zip` | job `build-20261005-224654`, published 22:46:54Z (PC 1 given by the Counter ASIC coordinator at 22:4xZ: Ember's collect job closed, the AMD sweep off tonight). (pending) | | PC 2 suites | waits for the Counter ASIC coordinator's "PC 2 suites go" (agg-cost-pc2-3 on PC 2 until about 22:59Z) | (pending) | +### 3a. PC 1's app down, and the relay task that hit the wrong machine (22:31 to 23:05Z) + +PC 1's installed 0.3.10 app quit at 22:31:06Z (its last upload 22:31:08Z, run `win-ae432dc7-20261005-214041`; Ember's tune job on it +reported "aborted (the app is quitting)"; the quit's sender is C35 for the consequences reviewer) and did not come back, so PC 1 mined +nothing, its build job `build-20261005-224654` could not start and no update-now could reach it. Three relay tasks (#241 at 22:52:31Z, #243 at +22:55:10Z, #245 at 23:03:18Z) went to the relay machine named "PC1" to relaunch the app, the second ending an engine that answered nothing on +`/api/state` and the third ending every igneum process before the launch, as the app's own updater does. + +They hit the wrong machine. The relay's "PC1" is the 1ccfe586 box, the console's PC 2: both PCs carry the hostname DESKTOP-KMCV30N, the only +relay agent runs on the 1ccfe586 box and was named PC1 when the clients were set up, and the relay's PC2 entry reads "never seen". The +intake proves it: PC 2 got two new engine runs, `win-1ccfe586-20261005-225528` and `-230330`, at the exact times of #243 and #245, while PC 1 +has no run after 21:40:41Z; #243's "hung" engine, pid 26696 from 21:49:40Z, was PC 2's healthy 0.3.10 engine (my probe's "no answer" on +`/api/state` was its own fault, no token), and the node, three miners and three workers it found were PC 2's own (a 5090 and the iGPU; PC 1 +would have shown the 9070 XT too). So PC 2, the box the night's measurements run on, was force-restarted at 22:55:28Z and 23:03:30Z, which +killed the aggregation-cost agent's job 3 (re-run owed, 20 min) and ended whatever followed; its app came back each time (after #245: pid 30484, +responding, node, three miners and both workers up, the card climbing at 23:04Z). PC 1 is exactly as it was: engine down since 22:31:06Z, +no relay agent on that box, unreachable tonight; it waits for the project lead in the morning and takes 0.3.11 through the manifest at its relaunch; the +fleet runs short its 141 MH/s until then. Told the Counter ASIC coordinator at 23:05Z; it re-set PC 2's schedule (this cut's combined build +and suite job first, then the prover-floor pair, the aggregation-cost re-run, "PC 2 clear", the M16 job) and ruled that no relay task goes to +"PC1" from anyone without its word. Every relay "PC1" reading tonight was PC 2 (the AMD agent's "9070 XT absent" probes read a box that has +no 9070 XT; the hardware events are being corrected). The relay machine should be renamed PC2 (`node tools/relay.mjs name `) +and the PC 1 box get its own agent (section 11). The Mac cross-build fallback for the Windows exes was announced and not started: the +combined PC 2 job replaced it within the minute. + +| PC 2 combined job | `IGNEUM_WIN_RELEASE=/target-integration/x86_64-pc-windows-gnu/release node tools/build-job.mjs run --node vendor/igneum-node-0311 --target 1ccfe586 --targets linux,windows --node-tests "kaspa-consensus kaspa-consensus-core igneum-exec kaspa-pow igneum-miner kaspa-p2p-flows" --app-tests igneum-app` from this worktree (its own zip `build-inputs-20261005230716-50058.zip`); PC 1's job removed from the jobs file so nothing double-places the exes | (pending: the id, the stages, the suites, the exes) | +| PC 2 suites | waits for the Counter ASIC coordinator's "PC 2 suites go" (agg-cost-pc2-3 on PC 2 until about 22:59Z) | (pending) | + ### 3a. PC 1's app down (22:31 to 22:5xZ) PC 1's installed 0.3.10 app quit at 22:31:06Z (its last upload 22:31:08Z, run `win-ae432dc7-20261005-214041`; Ember's tune job on it