release 0.3.11: the plan's step 2, the digest sweep and the open items; sections in order
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
f7613e5f96
commit
d61ce0ccee
1 changed files with 80 additions and 81 deletions
|
|
@ -48,7 +48,11 @@ Every check at cc72f4a: identity 0 hits over 220 files, copied-sources, pinned-g
|
|||
| `igneum-pow` | `cargo test --release` in `igneum-pow`, under the lock | 22:47:33Z: ok 53 + 4 + 19 + 7, 0 failed (the class v3 vectors, the era draw, the mixer x8, the scratch soundness) |
|
||||
| The prover host and export (the pin unchanged) | `cargo build --release -p igneum-prove-export -p igneum-prove-host` in `proving/igneum-prove` (the worktree needed the vendor links: `vendor/igneum-node-exec` and 45 others symlinked to the main checkout's, beside the real `igneum-node-0311` worktree) | 22:48Z: `--mode id` shard `0x2b1a81cb...`, aggregator `0x474678f3...` (the 0.3.9 pin: no prover drain); the nine real fixtures `--mode native` all ok |
|
||||
| PC 1 build job (the node and the app, Linux and Windows) | `IGNEUM_WIN_RELEASE=<fork>/target-integration/x86_64-pc-windows-gnu/release node tools/build-job.mjs run --node vendor/igneum-node-0311 --target ae432dc7 --targets linux,windows --no-tests` from this worktree, its own zip `build-inputs-20261005224625-29789.zip` | job `build-20261005-224654`, published 22:46:54Z (PC 1 given by the Counter ASIC coordinator at 22:4xZ: Ember's collect job closed, the AMD sweep off tonight). (pending) |
|
||||
| PC 2 suites | waits for the Counter ASIC coordinator's "PC 2 suites go" (agg-cost-pc2-3 on PC 2 until about 22:59Z) | (pending) |
|
||||
| PC 2 combined job | `IGNEUM_WIN_RELEASE=<fork>/target-integration/x86_64-pc-windows-gnu/release node tools/build-job.mjs run --node vendor/igneum-node-0311 --target 1ccfe586 --targets linux,windows --node-tests "kaspa-consensus kaspa-consensus-core igneum-exec kaspa-pow igneum-miner kaspa-p2p-flows" --app-tests igneum-app` from this worktree (its own zip `build-inputs-20261005230716-50058.zip`); PC 1's job removed from the jobs file so nothing double-places the exes | job `build-20261005-230745`, published 23:07:45Z (PC 2 woken), the first job on PC 2's re-set schedule; started 23:09Z, done 23:16:49Z (469 s): the Linux stage, the Windows stage 264 s, the test stage `RESULT test node [kaspa-consensus kaspa-consensus-core igneum-exec kaspa-pow igneum-miner kaspa-p2p-flows] exit 0 54 s` and `RESULT test app/igneum-app [igneum-app] exit 0 6 s`; 9 outputs verified and placed: igneumd.exe be8e83c07aeae5eb6842768071735289f3ef149ed9a7088592d8c54a4c252c08 (51,321,856), igneum-miner.exe 1ba1a249e2a21d52087a81f61f37cd2e1e3273c7ec8809ef9cfaa98f015895ce (10,987,520), igneum-app.exe ab104cc0... (3,065,344, the PC's; the installer's engine is the runner's), Linux igneumd d7a2715e... (49,193,960, glibc 2.39, HiveOS), igneum-miner 09d05ff6..., igneum-app 222f30c0... |
|
||||
| PC 2 suites | the combined job above (the PC 1 job of the row above never started: PC 1's app is down, section 3a) | six node suites exit 0 in 54 s, the app suite exit 0 in 6 s (23:16:49Z) |
|
||||
| The Linux workers for HiveOS | `infra/cross/build-workers-linux.sh` (zig, glibc 2.36 target) under the lock, 23:10Z | igneum-worker-cuda 4aaff27fcb26bc5fe98d2b311b9f95f099c59414a556af32080b82b41e8360db (6,755,568), igneum-worker-opencl 82d90890be36f9b795b218794062e388bfd9c6f7524fe671074cdec84c906259 (306,000), ELF x86-64 dynamic; untested on a GPU host, as the script says |
|
||||
| The HiveOS package | `NODE_OUT=<the zig node> WORKERS_OUT=<those workers> VERSION=0.3.11 packaging/hive/make-hive-package.sh`, 23:11:30Z, then `packaging/ota/publish-public.sh --hive` into `dl/public/` (the ship's deploy carries it) | `igneum-hive-0.3.11.tar.gz` c606a17043c013c411086b4abd0f38a227aa8a7db9d5934157760a3da5a5edc9 (24,496,653); no override inside (section 5) |
|
||||
| The DMG | `NODE=<fork>/target-integration/release/igneumd MINER=... packaging/mac/build-dmg.sh` under the lock, 22:50:0x to 22:50:51Z | `Igneum-Miner-0.3.11.dmg` b7e81d4f6f3af9f9179e29faa72c844b795df757dfb2e4cd1d56cd454e78e1e7 (41,592,041): engine 0.3.11, node 89dfcb95 Mac arm64, the 0.3.9 prover host and export, `igneum-bench` (the Metal worker) rebuilt from this tree (667,145 to 555,808 bytes, signed ad hoc), fingerprints 477bb0ef and ed9c4d2e; read back from the mounted image: `Contents/Resources/igneum-app.json` carries the nine-field `node_override_params` with N4 = N5 = 154,800 (C34) |
|
||||
|
||||
### 3a. PC 1's app down, and the relay task that hit the wrong machine (22:31 to 23:05Z)
|
||||
|
||||
|
|
@ -74,40 +78,6 @@ no 9070 XT; the hardware events are being corrected). The relay machine should b
|
|||
and the PC 1 box get its own agent (section 11). The Mac cross-build fallback for the Windows exes was announced and not started: the
|
||||
combined PC 2 job replaced it within the minute.
|
||||
|
||||
| PC 2 combined job | `IGNEUM_WIN_RELEASE=<fork>/target-integration/x86_64-pc-windows-gnu/release node tools/build-job.mjs run --node vendor/igneum-node-0311 --target 1ccfe586 --targets linux,windows --node-tests "kaspa-consensus kaspa-consensus-core igneum-exec kaspa-pow igneum-miner kaspa-p2p-flows" --app-tests igneum-app` from this worktree (its own zip `build-inputs-20261005230716-50058.zip`); PC 1's job removed from the jobs file so nothing double-places the exes | job `build-20261005-230745`, published 23:07:45Z (PC 2 woken), the first job on PC 2's re-set schedule; started 23:09Z, done 23:16:49Z (469 s): the Linux stage, the Windows stage 264 s, the test stage `RESULT test node [kaspa-consensus kaspa-consensus-core igneum-exec kaspa-pow igneum-miner kaspa-p2p-flows] exit 0 54 s` and `RESULT test app/igneum-app [igneum-app] exit 0 6 s`; 9 outputs verified and placed: igneumd.exe be8e83c07aeae5eb6842768071735289f3ef149ed9a7088592d8c54a4c252c08 (51,321,856), igneum-miner.exe 1ba1a249e2a21d52087a81f61f37cd2e1e3273c7ec8809ef9cfaa98f015895ce (10,987,520), igneum-app.exe ab104cc0... (3,065,344, the PC's; the installer's engine is the runner's), Linux igneumd d7a2715e... (49,193,960, glibc 2.39, HiveOS), igneum-miner 09d05ff6..., igneum-app 222f30c0... |
|
||||
| PC 2 suites | waits for the Counter ASIC coordinator's "PC 2 suites go" (agg-cost-pc2-3 on PC 2 until about 22:59Z) | (pending) |
|
||||
|
||||
### 3a. PC 1's app down (22:31 to 22:5xZ)
|
||||
|
||||
PC 1's installed 0.3.10 app quit at 22:31:06Z (its last upload 22:31:08Z, run `win-ae432dc7-20261005-214041`; Ember's tune job on it
|
||||
reported "aborted (the app is quitting)"; the quit's sender is C35 for the consequences reviewer: the tray Quit, stdin EOF in wrapper mode,
|
||||
or `POST /api/quit`, which Ember's playbook sends when its budget is spent) and did not come back, so PC 1 mined nothing, the build job
|
||||
`build-20261005-224654` could not start (no app to fetch it) and no update-now could reach it. The relay agent on PC 1 was alive (it runs
|
||||
as a logon scheduled task at highest privileges in the user's session), so the relaunch went through it at 22:52:31Z as relay item #241
|
||||
(`node tools/relay.mjs run PC1 ... relaunch-pc1.ps1`): the per-user install's `igneum-app.exe` started through `explorer.exe`, which hands
|
||||
the app the user's own medium-integrity token as a double-click does (a plain `Start-Process` from the elevated agent would have made the
|
||||
app elevated, which its updater never does), with a one-shot scheduled task at limited run level as the fallback; the script prints the
|
||||
process, its session, `app.url` and `/api/state`. Its result (22:5xZ): `igneum-app.exe` was ALIVE, pid 26696 since 21:49:40Z in session 1
|
||||
(`app.url` written 21:49:40Z), and answered nothing on `/api/state` in 60 s: the engine had logged its quit at 22:31:06Z, stopped the miners
|
||||
and the node, and then hung instead of exiting (the C35 fact: a quit that never ends; the sender is still the reviewer's question). The
|
||||
script launched nothing because a process existed. A second task (#243, 22:55:10Z) ends that engine by pid, as the app's own updater ends the
|
||||
old engine, and launches as above. Its result (#244, 22:57:18Z) changed the picture: the old engine was `responding=True` and had a node, two
|
||||
miners and three workers ALIVE under it, so PC 1 had been mining with an engine whose uploads and HTTP answers had stopped (my probe's "no
|
||||
answer" on `/api/state` may be its own fault: no token); the kill left those children running as orphans on their ports; the new engine
|
||||
(pid 30964, 22:55:28Z, session 1) wrote its `app.url` at once and then put nothing into the intake for minutes. The next step, if no run id
|
||||
appears by 23:01:30Z, is the updater's own order as one task: every igneum process ended (miners, workers, node, engine), then the launch.
|
||||
|
||||
The reading (Ember's agent through the Counter ASIC coordinator, 23:00Z): the alive children were Ember's SECOND engine's orphans (its
|
||||
tune job starts a second `igneum-app` with its own two miners on the 5090 and the 9070 XT), left mining when the job was aborted, holding
|
||||
both GPUs; the installed engine's quit at 22:31:06Z then stuck in the jobs runner's abort, waiting for EOF on the script's stdout pipe, whose
|
||||
write end the second engine and its miners had inherited (PowerShell's `Process.Start` inherits every inheritable handle), so EOF never
|
||||
came. The class rule, for every job script that starts a second engine (`relay/playbooks/sweep-5090.ps1` has the same shape): no pipe into
|
||||
it, and its whole tree killed at the end. The installer-build fallback script of 0.3.10 starts no second engine.
|
||||
(pending: the reset's result, the new run id, the first STATUS line)
|
||||
| The Linux workers for HiveOS | `infra/cross/build-workers-linux.sh` (zig, glibc 2.36 target) under the lock, 23:10Z | igneum-worker-cuda 4aaff27fcb26bc5fe98d2b311b9f95f099c59414a556af32080b82b41e8360db (6,755,568), igneum-worker-opencl 82d90890be36f9b795b218794062e388bfd9c6f7524fe671074cdec84c906259 (306,000), ELF x86-64 dynamic; untested on a GPU host, as the script says |
|
||||
| The HiveOS package | `NODE_OUT=<the zig node> WORKERS_OUT=<those workers> VERSION=0.3.11 packaging/hive/make-hive-package.sh`, 23:11:30Z, then `packaging/ota/publish-public.sh --hive` into `dl/public/` (the ship's deploy carries it) | `igneum-hive-0.3.11.tar.gz` c606a17043c013c411086b4abd0f38a227aa8a7db9d5934157760a3da5a5edc9 (24,496,653); no override inside (section 5) |
|
||||
| The DMG | `NODE=<fork>/target-integration/release/igneumd MINER=... packaging/mac/build-dmg.sh` under the lock, 22:50:0x to 22:50:51Z | `Igneum-Miner-0.3.11.dmg` b7e81d4f6f3af9f9179e29faa72c844b795df757dfb2e4cd1d56cd454e78e1e7 (41,592,041): engine 0.3.11, node 89dfcb95 Mac arm64, the 0.3.9 prover host and export, `igneum-bench` (the Metal worker) rebuilt from this tree (667,145 to 555,808 bytes, signed ad hoc), fingerprints 477bb0ef and ed9c4d2e; read back from the mounted image: `Contents/Resources/igneum-app.json` carries the nine-field `node_override_params` with N4 = N5 = 154,800 (C34) |
|
||||
|
||||
## 4. The override object, N4 and N5, the digests
|
||||
|
||||
The object every node runs after the publish (the four live fields plus the five new ones):
|
||||
|
|
@ -134,6 +104,30 @@ Three digests on the 0.3.11 binary, so two sweeps: step 1 moves every node to th
|
|||
nine-field object (0139ab9d...). A node on either side of a sweep is refused by the other side (the handshake), so each sweep is one
|
||||
window, as the fee switch's 8 min 47 s was.
|
||||
|
||||
## 5. The rollout order: two publishes, two sweeps (the reviewer's C39 and C40, the Counter ASIC coordinator's rule, the fee-switch shape)
|
||||
|
||||
Why two: the app writes the manifest's `consensus.override` at the manifest TAKE (`ota.rs` 661, `write_override`) and restarts its node with
|
||||
it at the next safe window, whatever binary is installed; a 0.3.10 `igneumd` refuses a file with `program_class_v3_activation_daa` or the
|
||||
proving v1 fields (`OverrideParams` is `deny_unknown_fields`) and dies at start, and the update then waits for a synced node (`engine.rs`
|
||||
2211) until the slot minute or the 1,800-block rule forces it. One manifest with 0.3.11 AND the nine fields would take every 0.3.10 node
|
||||
down at its next safe window: the 0.3.5 class the fee-switch plan named. And the 0.3.11 binary alone flips the digest (section 4), so the
|
||||
binary move is itself a sweep.
|
||||
|
||||
| Step | What | Digest after |
|
||||
|---|---|---|
|
||||
| 1a | the observer, node 1 and the seed on the 0.3.11 binaries with the four-field file: `IGNEUMD=<fork>/target-integration/release/igneumd IGNEUMD_COMMIT=89dfcb95 infra/devnet/restart-hand-nodes.sh '<the four-field object>'`, then `IGNEUMD_LINUX=<the zig build> IGNEUMD_LINUX_SHA256=63cf490d... infra/devnet/restart-seed.sh '<the same>'`; the apps still on 0.3.10 are refused by them from this moment until each updates | 4d8f8bb6... on the three |
|
||||
| 1b | the manifest: 0.3.11 with `consensus` carried over UNCHANGED (`--activation-height 135200 --deadline-note "finality v3"`, the four-field object, exactly as 0.3.10 shipped), `--public` (the HiveOS package rides along) | |
|
||||
| 1c | update-now: the Mac (its card follows node 1) and the laptop first; PC 2 only on the Counter ASIC coordinator's "PC 2 clear"; PC 1 is down and unreachable (section 3a) and takes 0.3.11 through the manifest at its morning relaunch, refused until then. The watch: every app logs 0.3.11 and its node a DAA score at 4d8f8bb6...; every worker starts clean on the first try with the class-aware pair; the Mac's Metal worker takes the class v3 day from the prepared pack; C32: the agents whose kits sit on PC 2 republish their fetches after its update | 4d8f8bb6... on every reporting node |
|
||||
| 2a | the floor: 154,800 minus the tip's DAA at least 10,800 (holds until DAA 144,000, about 00:45Z); past it N4 = N5 re-pinned to the first multiple of 3,600 at or above tip + 14,400, the packaged line, the DMG and the installer rebuilt; the DAA read sent to the Counter ASIC coordinator before 2b | |
|
||||
| 2b | the hand nodes' and the seed's files switched to the nine-field object and restarted (the same two scripts), the manifest republished with `--override '<the nine-field object>' --activation-height 154800 --deadline-note "program class v3 + proving v1"`, update-now (the same order; PC 2 on "clear" again), the sweep | 0139ab9d... on every node |
|
||||
| 3 | the plan's final sections, the merge to master (the live observer must not read stale: `public-api-check`'s other arm), the push, the report with per-machine times | |
|
||||
|
||||
The HiveOS package carries NO override: `packaging/hive/h-run.sh` line 31 starts the rig's node with `--devnet --appdir --rpclisten --listen` and
|
||||
the peers, no `--override-params-file`, and no HiveOS package has ever carried one, so a rig's bundled node runs on genesis params and is refused
|
||||
by every devnet peer (the HiveOS path is untested on a GPU host since 4 October). "Republish with the new override" therefore needs an
|
||||
`h-run.sh` change (the file written from the Flight Sheet's extra config, as `PEERS=` is), which is the next cut's; tonight's package carries
|
||||
the class-aware binaries only (section 11).
|
||||
|
||||
## 6. The push and CI
|
||||
|
||||
`git push -u origin release-0.3.11` at 3b0262f (23:18:29Z, the credential helper; the pre-push hook's site flip restored), `gh workflow run
|
||||
|
|
@ -172,6 +166,40 @@ Baseline 23:19:40Z: tip DAA 139,642; the observer and node 1 on 21d4c73c at 1f4b
|
|||
| the laptop (PC 37ba0461) | silent since 23:28:28Z | its last upload, "2.25 MH/s, mining \| node 139757 blocks, 0 peers" at 23:28:28Z, is 22 s after `update-now-0311-d937c69d-37ba0461` was published; no job line reached the intake before the silence. Its 0.3.10 install kept it silent 55 minutes (21:41 to 22:36Z), so this is its install in progress until shown otherwise; publish 2 does not wait on it (a 2 MH/s machine whose 0.3.11 reads the nine fields when it returns; on 0.3.10 its node would die on the file until the forced apply, the C39 case for one machine) |
|
||||
| PC 1 | unreachable tonight (section 3a); refused by every peer on 1f4b4425 until its morning relaunch takes 0.3.11 through the manifest | |
|
||||
|
||||
## 8. The rollout, step 2 (the object sweep to 0139ab9d...)
|
||||
|
||||
| Step | Time | Result |
|
||||
|---|---|---|
|
||||
| 2a the floor | 23:55:42Z | tip DAA 140,706; 154,800 - 140,706 = 14,094 >= 10,800, so N4 = N5 = 154,800 stand (the floor holds until DAA 144,000); no re-pin, no rebuild |
|
||||
| 2b the observer | 23:56:59Z (pid 97249) | `/tmp/igneum-devnet/override-v3.json` switched to the nine-field object; `igneumd/2.1.0-89dfcb95`, `Program class v3 from the override file: active from epoch 43 (DAA score 154800 ...)`, `Proving v1 from the override file: ... 154800, 8 blocks a segment, unproven after 600 DAA, aggregator share 1000 bps`, digest 0139ab9dc2992d449ec787d8f021974933631eb55740ab4b6ce9d5c226e72888 |
|
||||
| 2b node 1 | 23:57:12Z (pid 97417) | the same lines and digest; it refused the seed (still 4d8f8bb6...) at 23:57:13Z and registered it at 23:57:42Z (protocol version 15), 29 s after the seed's own restart |
|
||||
| 2b the seed | 23:57:30Z (MainPID 126124) | `/etc/igneum/override-v3.json` the nine-field object; the same binary 63cf490d..., the same lines, digest 0139ab9d... |
|
||||
| 2b the manifest | 23:57:49Z | `publish-manifest.sh --version 0.3.11 --override '<the nine-field object>' --activation-height 154800 --deadline-note "program class v3 + proving v1" --notes '<section 1>' --public --deploy`; the live manifest's `consensus` read back byte-identical at the edge (the nine fields, `activation_height` 154800) |
|
||||
| 2b update-now, the Mac and the laptop | 00:00:30Z (`update-now-0311-switch-d937c69d-37ba0461`, apps woken) | the Mac ran it 00:01:06Z: `0.3.11 is current`, `consensus parameters from the signed manifest: {... nine fields ...}`, `consensus override changed (.../Igneum/app/override.json); the node restarts with it at a safe moment`; the Mac's node card is node 1 (external: pid 0, starts 0, `consensus_digest` empty), already restarted by hand at 23:57:12Z, so there was nothing for the app to restart and its digest is node 1's. The laptop (PC 37ba0461) was not on the air (below) |
|
||||
| 2b update-now, PC 2 | 00:01:10Z (`update-now-0311-switch-1ccfe586`, on the Counter ASIC coordinator's "PC 2 clear": `floor-restore-1` closed 23:57:29Z) | PC 2 fetched the woken file 00:01:41Z, ran the job the same second (`0.3.11 is current`, `consensus override changed (C:\Users\Admin\AppData\Local\igneum\app\override.json); the node restarts with it at a safe moment`), its node restarted at 00:01:42Z (`Consensus params digest: 0139ab9d...`, `igneumd/2.1.0`), the first STATUS 00:02:01Z "0.00 MH/s, waiting, node 141065 blocks, 1 peers, synced", mining at 00:02:31Z, 106.48 MH/s with 830 accepted this run and 0 faults at 00:06:31Z; the run id stays `win-1ccfe586-20261005-235130` (the node restarted, not the engine) |
|
||||
| the window | 23:57:42 to 00:01:42Z | PC 2 (on 4d8f8bb6...) refused the seed (on 0139ab9d...) at 23:57:42Z and 23:58:12Z and mined on the old side alone; node 1 refused the seed once (23:57:13Z). The new side (the observer, node 1, the seed) had no miner on it either: the Mac's miner was off node 1 from 23:57:14Z (the next row), so the new side only relayed PC 2's old-side blocks it had already accepted (`PoW accepted ... daa 140757` at 23:57:32Z was the last) and waited; the two sides rejoined when PC 2's node restarted at 00:01:42Z and PC 2's hash carried the chain. The observer's tip read 140,757 at 00:02:22Z and 141,645 at 00:10:39Z (about 1.7 blocks/s, the catch-up after the rejoin), no stall as in step 1 |
|
||||
| the Mac's miner (a miner finding) | 23:57:14Z to 00:08:34Z | the miner (`igneum-miner mine grpc://127.0.0.1:26610`) lost node 1 at node 1's 23:57:12Z restart and NEVER reconnected: 5,317 `submit error ... Not connected to server` and `template error: Not connected to server` lines, its TEMPLATES line frozen at `templates=2350` with `subscribed=true` while `fetch_errors` climbed (143 at 23:59:34Z, 595 at 00:05:07Z), the card's line "the node is not answering; the miner retries", the STATUS line still "26 MH/s, mining" with the accepted count frozen at 18,117 (773 this run). The retries are template and submit calls on a dead gRPC channel; nothing re-subscribes. Recovery through a job: `publish-jobs.sh add --kind restart --target d937c69d --what miners` (`restart-miners-0311-d937c69d`, 00:08:01Z); the Mac ran it 00:08:31Z (`miners restarted`), the new miner (pid 8567) started 00:08:33Z, its first block was accepted 00:08:34Z, the kernel race settled at 26.9 MH/s 00:09:09Z. Lost: 11 min 20 s of 26 MH/s. In step 1 this was masked: node 1's 23:24:06Z restart was followed by the Mac's own engine restart (the update) at 23:29:05Z, and the Mac was paused anyway |
|
||||
| PC 2 at the epoch boundary | 23:52:22Z | 40 s after PC 2's first 0.3.11 pack the hour turned (`f4d9d3d8...` to `cb5b51cc...`): `hourly program changed: the pair was not prepared (unexpected seeds); the worker compiles inline`, 80 s of `worker fault: seed mismatch: the worker holds another program; preparing the current pair epoch cb5b51cc... day 2`, the worker on the new epoch at 23:53:42Z (113 MH/s), one `did not come back within 180 s` line for the old worker instance at 23:56:01Z. The restart landed inside the minute before the boundary, before the next epoch's prepare had run; the attempt-aware path (`program pack checked ... attempt 0`) recovered it without a job. Also at 23:52:14Z: `prover: block 74094 shard 0: ... Failed to create the CUDA prover impl: ... PermissionDenied (a GPU-server socket ...)`, the 0.3.10 socket class again on the first shard after the update; shards after it proved (`block 93665 shard 0 paid 0.9446373 IGN` at 00:06:09Z) |
|
||||
| the laptop (PC 37ba0461) | absent | no upload since 23:28:28Z (section 7) and no card on the console at 00:08Z; both update-now jobs wait for it in the jobs file; its 0.3.10 app takes 0.3.11 and the nine-field object through the manifest when it is next on the air |
|
||||
| PC 1 (ae432dc7), Sam's Mac (3a9bf309) | pending their relaunch | PC 1 down since 22:31:06Z (section 3a), Sam's Mac quit since 20:47Z on 0.3.9; each takes the manifest at its relaunch: the OLD app writes the nine-field object at the manifest take (`ota.rs` 661) and its old node, if restarted with that file before the 0.3.11 install lands, dies on the unknown fields (`deny_unknown_fields`, the reviewer's C39); the install then starts the 0.3.11 node on the file, so the worst case is one node death inside the update, to be read in the morning's intake |
|
||||
|
||||
## 9. The digest sweep
|
||||
|
||||
| Node | Binary | Digest | Since |
|
||||
|---|---|---|---|
|
||||
| the observer (`/tmp/igneum-devnet/observer-v4`, rpc 26640, json 28640) | `igneumd/2.1.0-89dfcb95` (bd7f043c...) | 0139ab9dc2992d449ec787d8f021974933631eb55740ab4b6ce9d5c226e72888 | 23:56:59Z |
|
||||
| node 1 (`/tmp/igneum-devnet/node1`, rpc 26610) | the same | 0139ab9d... | 23:57:12Z |
|
||||
| the seed (188.245.5.161, `igneumd.service`) | `igneumd/2.1.0-89dfcb95` (63cf490d..., glibc 2.36 target) | 0139ab9d... | 23:57:30Z |
|
||||
| PC 2 (1ccfe586, `devnet-v4`, rpc 26610) | the 0.3.11 installer's `igneumd.exe` be8e83c0... (PC 2's own build job) | 0139ab9d... | 00:01:42Z |
|
||||
| the Mac (d937c69d) | attached to node 1 (no node of its own) | node 1's | 23:57:12Z |
|
||||
| the laptop (37ba0461) | 0.3.10 node, 1f4b4425... at its last upload | pending | silent since 23:28:28Z |
|
||||
| PC 1 (ae432dc7) | 0.3.10 node, 1f4b4425... | pending | down since 22:31:06Z |
|
||||
| Sam's Mac (3a9bf309) | 0.3.9 node a24ab01a | pending | quit since 20:47Z |
|
||||
|
||||
The chain at 00:10:39Z: tip DAA 141,645 on the observer (age 1.4 s), node 1 with 2 peers (the seed and PC 2's side), PC 2 with 1 peer at
|
||||
106 to 109 MH/s, the Mac at 26 MH/s, every live node on 0139ab9d... Epoch 43 (DAA 154,800) is about 3 h 45 min out at 0.965 blocks/s
|
||||
(about 03:55Z); the class v3 and proving v1 lines of every node name it.
|
||||
|
||||
## 10. The next cut
|
||||
|
||||
| Branch | What | Why not 0.3.11 |
|
||||
|
|
@ -183,51 +211,22 @@ Baseline 23:19:40Z: tip DAA 139,642; the observer and node 1 on 21d4c73c at 1f4b
|
|||
| `proving-v1` c36dfea | docs only, after the code tip 22c2363 | its agent's choice: docs follow |
|
||||
| the 0.3.10 list (`release-0.3.10.md` section 11): fork `pack-loop` 05ef0fa3, `job-console`, the rest of `opencl-rdna4`, `opencl-rdna4-telemetry` | | unchanged |
|
||||
|
||||
## 5. The rollout order: two publishes, two sweeps (the reviewer's C39 and C40, the Counter ASIC coordinator's rule, the fee-switch shape)
|
||||
## 11. Open after the cut
|
||||
|
||||
Why two: the app writes the manifest's `consensus.override` at the manifest TAKE (`ota.rs` 661, `write_override`) and restarts its node with
|
||||
it at the next safe window, whatever binary is installed; a 0.3.10 `igneumd` refuses a file with `program_class_v3_activation_daa` or the
|
||||
proving v1 fields (`OverrideParams` is `deny_unknown_fields`) and dies at start, and the update then waits for a synced node (`engine.rs`
|
||||
2211) until the slot minute or the 1,800-block rule forces it. One manifest with 0.3.11 AND the nine fields would take every 0.3.10 node
|
||||
down at its next safe window: the 0.3.5 class the fee-switch plan named. And the 0.3.11 binary alone flips the digest (section 4), so the
|
||||
binary move is itself a sweep.
|
||||
|
||||
| Step | What | Digest after |
|
||||
| Item | What | Owner |
|
||||
|---|---|---|
|
||||
| 1a | the observer, node 1 and the seed on the 0.3.11 binaries with the four-field file: `IGNEUMD=<fork>/target-integration/release/igneumd IGNEUMD_COMMIT=89dfcb95 infra/devnet/restart-hand-nodes.sh '<the four-field object>'`, then `IGNEUMD_LINUX=<the zig build> IGNEUMD_LINUX_SHA256=63cf490d... infra/devnet/restart-seed.sh '<the same>'`; the apps still on 0.3.10 are refused by them from this moment until each updates | 4d8f8bb6... on the three |
|
||||
| 1b | the manifest: 0.3.11 with `consensus` carried over UNCHANGED (`--activation-height 135200 --deadline-note "finality v3"`, the four-field object, exactly as 0.3.10 shipped), `--public` (the HiveOS package rides along) | |
|
||||
| 1c | update-now: the Mac (its card follows node 1) and the laptop first; PC 2 only on the Counter ASIC coordinator's "PC 2 clear"; PC 1 is down and unreachable (section 3a) and takes 0.3.11 through the manifest at its morning relaunch, refused until then. The watch: every app logs 0.3.11 and its node a DAA score at 4d8f8bb6...; every worker starts clean on the first try with the class-aware pair; the Mac's Metal worker takes the class v3 day from the prepared pack; C32: the agents whose kits sit on PC 2 republish their fetches after its update | 4d8f8bb6... on every reporting node |
|
||||
| 2a | the floor: 154,800 minus the tip's DAA at least 10,800 (holds until DAA 144,000, about 00:45Z); past it N4 = N5 re-pinned to the first multiple of 3,600 at or above tip + 14,400, the packaged line, the DMG and the installer rebuilt; the DAA read sent to the Counter ASIC coordinator before 2b | |
|
||||
| 2b | the hand nodes' and the seed's files switched to the nine-field object and restarted (the same two scripts), the manifest republished with `--override '<the nine-field object>' --activation-height 154800 --deadline-note "program class v3 + proving v1"`, update-now (the same order; PC 2 on "clear" again), the sweep | 0139ab9d... on every node |
|
||||
| 3 | the plan's final sections, the merge to master (the live observer must not read stale: `public-api-check`'s other arm), the push, the report with per-machine times | |
|
||||
|
||||
The HiveOS package carries NO override: `packaging/hive/h-run.sh` line 31 starts the rig's node with `--devnet --appdir --rpclisten --listen` and
|
||||
the peers, no `--override-params-file`, and no HiveOS package has ever carried one, so a rig's bundled node runs on genesis params and is refused
|
||||
by every devnet peer (the HiveOS path is untested on a GPU host since 4 October). "Republish with the new override" therefore needs an
|
||||
`h-run.sh` change (the file written from the Flight Sheet's extra config, as `PEERS=` is), which is the next cut's; tonight's package carries
|
||||
the class-aware binaries only (section 11).
|
||||
|
||||
## 10. The next cut
|
||||
|
||||
| Branch | What | Why not 0.3.11 |
|
||||
|---|---|---|
|
||||
| `ember-tune` b671c8b (and the tune behind it) | the C35 fix: `Cmd::Quit(&'static str)` so every "quit:" line names its sender (the window host's stdin, the host gone, `POST /api/quit`, the sweep's end), `elevation_allowed()` = Power control alone (the unattended sweep on PC 1 raised one UAC prompt at 22:30Z under the old rule), no power cap at start under `--sweep`; the quit-source hunk is separable (main.rs 2 lines, server.rs 1 line, engine.rs the Quit arm, `elevation_allowed` and its test) | arrived after the tree closed at 23bc2b2 (the app, the DMG and PC 1's job carry it); not among the branches named for this cut |
|
||||
| `fud-close` (the ledger closer's main branch, a22ba27a60a6f1c64; its ready tip was due about 23:05Z) | 45 public-text fixes on the site and litepaper, spec 8.3 and 8.8, two CI checks, relay fixes; touches `packfile.h` and `host.c`, so taking it means the two Windows workers, the Mac worker and the DMG rebuilt from the merged tip (G5) | offered by the Counter ASIC coordinator at 22:5xZ after the tree closed; not among the branches named for this cut; the fork-side `ledger-fixes` is not in 0.3.11 either |
|
||||
| `explorer` d7e797c (and 3e01212) | `/api/stats` gains `proving`; `tools/ci/public-api-check.mjs` then FAILS when the live API lacks it, and ci.yml runs that check against the live site on every master push, which reads the OLD API until Vercel redeploys after the push (the reviewer's C36) | not in 23bc2b2 (only on the explorer branch); its merge needs the check to retry for a few minutes or to require `proving` only when `observer_updated_at` is newer than the commit |
|
||||
| `proving-v1` c36dfea | docs only, after the code tip 22c2363 | its agent's choice: docs follow |
|
||||
| the 0.3.10 list (`release-0.3.10.md` section 11): fork `pack-loop` 05ef0fa3, `job-console`, the rest of `opencl-rdna4`, `opencl-rdna4-telemetry` | | unchanged |
|
||||
|
||||
## 5. The rollout order, as agreed with the Counter ASIC coordinator (PC 2's scheduler tonight)
|
||||
|
||||
The fee-switch and N3 shape (every node in one sweep, because the digest flips to 0139ab9d...), with PC 1 out of reach:
|
||||
|
||||
| Step | What | Gate |
|
||||
|---|---|---|
|
||||
| the floor | 154,800 minus the tip's DAA at least 10,800 at the publish (holds until DAA 144,000, about 00:45Z); past it the line is re-pinned and the DMG and installer rebuilt | read on the observer before the manifest step |
|
||||
| the manifest | `--public`, `--override '<the nine-field object>' --activation-height 154800 --deadline-note "program class v3 + proving v1"`, the changelog line of section 1 | CI green on the release tip |
|
||||
| update-now | the Mac (attached to node 1) and the laptop (PC 37ba0461, if it is back) first; PC 2 only on the coordinator's "PC 2 clear" (after this cut's build job, the prover-floor pair and the aggregation-cost re-run); PC 1 is unreachable (section 3a) and takes 0.3.11 through the manifest at its morning relaunch | |
|
||||
| the watch | every worker starts clean on the first try with the class-aware pair; the Mac's Metal worker takes the class v3 day from the prepared pack (`--prepare-packs`); the re-fetch rule C32: an update clears the jobs folder, so the agents whose kits sit on PC 2 republish their fetches after it | |
|
||||
| the hand nodes, the seed | the two scripts with the same object (`IGNEUMD` the 0.3.11 Mac node, `IGNEUMD_COMMIT=89dfcb95`; `IGNEUMD_LINUX` the zig build 63cf490d... and its sha) | after the app nodes |
|
||||
| the sweep | every restarted node prints 0139ab9dc2992d449ec787d8f021974933631eb55740ab4b6ce9d5c226e72888; PC 1 joins at its relaunch | |
|
||||
| the HiveOS package | rebuilt from this tree's Linux node (63cf490d...) and the two Linux workers (`infra/cross/build-workers-linux.sh`, zig) with `packaging/hive/make-hive-package.sh VERSION=0.3.11`, and published through the `--public` step. It carries NO override: `packaging/hive/h-run.sh` line 31 starts the rig's node with `--devnet --appdir --rpclisten --listen` and the peers, no `--override-params-file`, and no HiveOS package has ever carried one, so a rig's bundled node runs on genesis params and is refused by every devnet peer (the HiveOS path is untested on a GPU host since 4 October). "Republish with the new override" therefore needs an `h-run.sh` change (the file written from the Flight Sheet's extra config, as `PEERS=` is), which is the next cut's; tonight's package carries the class-aware binaries only (section 11) | the coordinator told |
|
||||
| the plan, the master merge, the push | the live observer must not read stale (`public-api-check`'s other arm); the explorer branch's strict check is not in this tree | |
|
||||
| the miner's dead channel (section 8) | `igneum-miner mine grpc://...` keeps `subscribed=true` after the node restarts and retries template and submit calls on the closed channel for ever (11 min 20 s tonight, 5,317 errors, hash burned at 26 MH/s with 0 accepted); it must re-subscribe on the first `Not connected to server` and the app's card must read "the miner lost the node" instead of "mining, 26 MH/s". Until the fix: every hand restart of node 1 is followed by `publish-jobs.sh add --kind restart --target d937c69d --what miners` (the runbook's step), and the Mac's card is read by its accepted count, not its MH/s | miner (next cut) |
|
||||
| the HiveOS package carries no override | `packaging/hive/h-run.sh` starts the rig's node without `--override-params-file`; a rig on `igneum-hive-0.3.11.tar.gz` runs on genesis params and is refused by every peer; the file must come from the Flight Sheet's extra config as `PEERS=` does (section 5) | packaging (next cut) |
|
||||
| the prepared pack and a restart near the hour (section 8) | an engine restarted in the last minute before an epoch boundary has no prepared pair for the next hour (80 s of seed-mismatch faults on PC 2 at 23:52:22Z, recovered by the attempt-aware path); the prepare should run at engine start for the next epoch too, or the update-now path should wait out the boundary when it is under 2 min away | app |
|
||||
| PC 2's first shard after an update | `Failed to create the CUDA prover impl: ... PermissionDenied (a GPU-server socket ...)` on block 74094 at 23:52:14Z, the 0.3.10 socket class (fixed 22:01Z for the running engine; the root socket outlives the engine restart); the later shards proved. The socket unlink belongs in the engine's prover start, not only in the playbooks (`prover-socket-check.sh` covers the playbooks) | app / proving |
|
||||
| the relay machine named PC1 (section 3a) | it is the 1ccfe586 box: rename it (`node tools/relay.mjs name DESKTOP-KMCV30N PC2`), give the ae432dc7 box its own agent, and no relay task to "PC1" without the Counter ASIC coordinator's word until then | relay (the project lead in the morning for the PC 1 box) |
|
||||
| PC 1 down since 22:31:06Z | relaunch in the morning (the installed app, double-clicked or through the Start menu, never elevated); it takes 0.3.11 and the nine-field object through the manifest, one node death inside the update possible (section 8); the intake's first lines of `win-ae432dc7-...` are the check; `ember-tune` b671c8b carries the C35 fix for the quit's sender | the project lead, then the release engineer |
|
||||
| the laptop (37ba0461) | silent since 23:28:28Z, off the console; its 0.3.10 install kept it silent 55 min once already; when it uploads again its first STATUS line must show 0.3.11, the nine-field object in the manifest log line and 0139ab9d... | watch |
|
||||
| Sam's Mac (3a9bf309) | 0.3.9, quit since 20:47Z; `min_supported` is 0.3.0 so it updates straight to 0.3.11 at its relaunch | watch |
|
||||
| the app's job queue | an update-now queues behind a running job (PC 2's step-1 update waited 11 min behind a hung sweep's 30 min cap, section 7); pre-empt, or report the wait in the STATUS line | app (section 10) |
|
||||
| the public index at the edge | `dl/public/igneum-downloads.json` alternates between the unsigned and the signed copy at the edge for minutes after a deploy, so the ship's verify refuses it once per cut (0.3.10 and 0.3.11 both); `--from console` after it settles is the workaround; the verify should accept either copy while both carry the cut's hashes, or the deploy should purge the edge | ship-app |
|
||||
| the pre-push site flip | the pre-push hook rebuilds `site/` and leaves the tree modified (`git checkout -- site/` after every push); the site build is not idempotent | site |
|
||||
| the explorer branch's strict check (C36) | `public-api-check.mjs` on `explorer` d7e797c fails against the live API until Vercel redeploys; merge it right after a master push, not before | explorer (section 10) |
|
||||
| C1 (the reviewer) | the 0.3.10 `exportSegments` lacks `daaScore` and `feesV1ActivationDaa`; proving-v1's `rpc.rs:982` (eb32c645) carries them; the decision on which the aggregator reads is due 16:00Z 6 October | the Counter ASIC coordinator |
|
||||
| the Metal worker at resume (G4b) | "the pair was not prepared" at the Mac's 23:39:10Z resume (section 7), compiled inline; same class as the PC 2 boundary row above | app |
|
||||
| the 0.3.10 list | `release-0.3.10.md` section 11, unchanged: fork `pack-loop` 05ef0fa3, `job-console`, the rest of `opencl-rdna4`, `opencl-rdna4-telemetry`, the per-job zip pin check | |
|
||||
|
|
|
|||
Loading…
Reference in a new issue