Counter ASIC 2.0 status 22:59: the orphaned second-engine miners, the hang's mechanism, the second-engine playbook rule
This commit is contained in:
parent
44c598c929
commit
1bc86cd7b4
1 changed files with 4 additions and 0 deletions
|
|
@ -621,3 +621,7 @@ PC 1 (Ember's finding, confirmed on the intake at 22:51): the installed app has
|
|||
22:57. Next-cut coupling recorded (the ledger closer): fud-close 647b08c's worker half of M28 (the kernel_sha256 check in packfile.h and host.c) must ship with the fork-side ledger-fixes miner commit 3d4ec451 that stamps the hashes, or every worker refuses every pack; ledger-fixes is not yet rebased onto 89dfcb95 (two conflicting files); round 2's ledger branches stay behind tonight.
|
||||
|
||||
22:59. Next-cut note from the reviewer's merge-tree against 23bc2b2 (0 conflicts for explorer d7e797c, ember-tune b671c8b, rig-install 086008a, pool-v0 425b875, ota-k2 89a76b1, hive-words 2d056e8, fud-close 647b08c and the others already in): explorer d7e797c makes tools/ci/public-api-check.mjs fail when the live /api/stats lacks `proving`, and that check runs on master pushes against the live site, which Vercel redeploys only after the push, so the first master CI after a ship carrying it goes red through no fault; the next cut holds d7e797c or gives the check a retry loop. Waiting now on two watchers: the aggregation-cost close on PC 2 (then the shipper's suites) and PC 1's app relaunch (then the exes, the push and CI).
|
||||
|
||||
## 22:59 PC 1: the hang explained; orphaned miners from the second engine hold both GPUs
|
||||
|
||||
The shipper's relay task #244 (22:55:24Z) counted 1 igneumd, 2 igneum-miner, 2 igneum-worker-cuda and 1 igneum-worker-opencl running although the installed app had stopped its miners at 22:30:20Z and its node at 22:31:06Z: Ember's second engine's children, orphaned when its job was aborted, mining on both GPUs; they would fight the relaunched app's miners and void every number. The relay lane kills the tree (taskkill /F /T on every miner, worker and non-app igneumd) and relaunches the app. The hang (Ember's reading): the installed engine's quit got stuck in the jobs runner's abort, whose reader waits for EOF on the script's stdout pipe; the pipe's write end was inherited by the second engine and its miners (PowerShell's Process.Start with redirection inherits every inheritable handle), so EOF never came and the engine sat "responding" until #243 ended it. Class rule for every playbook that starts a second engine (ember-tune-pc1.ps1, relay/playbooks/sweep-5090.ps1 and any job script of the shipper's): no inherited pipe into the second engine, its whole process tree killed at the end and on abort, the installed app's miners restarted only after; a CI check that fails a playbook starting an engine without those lines (Ember's branch). The quit's sender is still open (the tray excluded by the missing 45-s host timer kill: stdin EOF or POST /api/quit; the event-log collect decides, when the app is back).
|
||||
|
|
|
|||
Loading…
Reference in a new issue