From fa4b64fe8a536380bbff13986f52727a526b2e5d Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Tue, 6 Oct 2026 19:02:01 +0100 Subject: [PATCH] Counter ASIC 3.0 status: PC 1 free, the 9070 XT restored (confirmed off at 17:55Z, mining from 17:57Z), the amd:1 flag back --- docs/plans/counter-asic-3-status.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/plans/counter-asic-3-status.md b/docs/plans/counter-asic-3-status.md index 17eb7c958..7ba7e8bad 100644 --- a/docs/plans/counter-asic-3-status.md +++ b/docs/plans/counter-asic-3-status.md @@ -10,7 +10,7 @@ The test every result is judged against (Josh, 6 October): a chip maker must hav |---|---|---| | Mac M5 Max | mining paused; Metal worker free | measurements under `with-lock.sh measure` only | | PC 2 (1ccfe586, RTX 5090) | HELD for the 0.3.14 shipper's Rust suite from 17:5x UK (main): nothing of 3.0 goes to PC 2 until the shipper reports the suite done (the P2 G6 run waits on it). Earlier: RELEASED at 08:43:24Z after the three jobs (derive 08:26 to 08:29Z, shadow 08:29 to 08:41Z, family 08:41 to 08:43Z, all exit 0; prover ON, app untouched); then the prover-floor agent, then the 0.3.12 release engineer's build. Before that: CLEAR at 08:24:27Z (the clear file carries the proving agent's constraints: prover left ON, no quit or restart, /opt/igneum, /opt/igneum-segal and settings.json untouched); the three 3.0 jobs run one at a time through the mkdir lock (items 8, 2, 6); "PC 2 released" to the proving agent after the last; the prover-floor agent next. Before that: the proving agent's jobs `segments-pc2-pv1` runs b and c (run b claimed nothing on a PowerShell key bug; run c from about 07:53Z); "PC 2 clear" expected about 08:35Z; then three 3.0 jobs one at a time (items 2, 6, 8); the prover-floor agent queues after "PC 2 released" | nothing published until "PC 2 clear"; one job at a time (`/tmp/igneum-devnet/pc2-ca3.lock`); released by message when done | -| PC 1 (ae432dc7, RTX 5090 + RTX 4070 + RX 9070 XT) | Josh's desk; the AMD queue opens on "Ember closed" (reset, family, G2, derive, the 4070, watts, about 25 min), then the Ember agent's 6-minute window, then the 0.3.14 update on Josh's click (after the last job, not between jobs: an update wipes the jobs folder and the kits) | not used today; every AMD row OWED | +| PC 1 (ae432dc7, RTX 5090 + RTX 4070 + RX 9070 XT) | FREE at 18:01:05Z after the queue (reset, family runs a to e, G2, derive, the 4070 ladder, the watts job that failed and left the 9070 XT off, the restore and the flag job); handed to main for Josh's 0.3.14 click and the Ember window. Earlier: Josh's desk; the AMD queue opens on "Ember closed" (reset, family, G2, derive, the 4070, watts, about 25 min), then the Ember agent's 6-minute window, then the 0.3.14 update on Josh's click (after the last job, not between jobs: an update wipes the jobs folder and the kits) | not used today; every AMD row OWED | Standing rule from main (17:5x UTC; CLAUDE.md on master 630b537, `docs/plans/build-server.md`): from the next build, every Linux and Windows cargo build and every Linux test suite from the 3.0 lanes runs on igneum-build-1 through `tools/build-remote.sh` and `tools/cross-remote.sh` from the worktree's crate directory (artefacts in `target-remote/`, `IGNEUM_AGENT=ca3-`, one slot, a 2 h cap; the clean node build 87 s there against 9 to 18 min on the Mac). PC 2 keeps only GPU and Windows-runtime jobs (G6's engine test on the card stays a PC job; the crate suites move). Passed to the node lane, the hash lane and the PC 1 worker. @@ -166,7 +166,7 @@ The live observer runs the shared checkout, which autosync fast-forwards from or |---|---|---| | item 3 (43c3ead) | spec 01 section 1.13.1's era table and `docs/plans/mixer-x4.md` section 2 still said `mixer_mult = 4`; the code (`LoadClass::MX8`, `V3_CLASS`) and spec 1.8.5 say 8 | both lines corrected on ca3-coord, 6 October 2026 | | PC 1 job 2 (family, run-ca3-pc1-amd-family-20261006, 16:56 to 16:59Z) | the probe ran on device 0, the integrated gfx1036, not the 9070 XT (device 1): every row `dev=0 name=` empty, alu 40.59 G steps/s (a one-CU figure); the rows are RDNA 2 iGPU ratios (shifts 1.15, bfe 1.15 native, andn 1.16, popc 1.55, clz 1.75, sel 1.81, shfla and shflx 1.89 through ds_bpermute, perm 2.43 and dot4 2.57 emulated, mm8 none: AMD's OpenCL C compiles no byte-permute, dot4 or WMMA builtin) and the 9070 XT column stays owed | the probe picks the device by name (gfx1201 on the newest AMD platform) and prints it in every row; re-run in the next PC 1 slot | -| the watts job (run-ca3-pc1-amd-watts-20261006, 17:49:46 to 17:52:54Z, FAILED exit 1) | the sampler proved itself (12 of 12 9070 XT watts lines), the 90 s app-state window ran (the card mining at 18.87 MH/s, 203 W, identities 8 before it), every gfx1201 entry was posted off, the card went quiet and the sh256x27 bench started; then the script exited with no APPROW, no LADDER, no `watts error=` and NO RESTORE line in any upload (no `enabled=True` post, no `card_workers_after`): a `finally` that did not run or did not print, so the 9070 XT may have been left OFF | the card restored at once through the identities job (run d: one POST enabled true, identities 8, the state read after); the entry amd:1:gfx1201 (enabled before) gets its flag back with one more POST; the AMD watts row stays OWED; the cause and a fixed script (an unconditional finally with its own first line, no `exit` inside the try, the device dumps out of the report) with the script worker before any re-run | +| the watts job (run-ca3-pc1-amd-watts-20261006, 17:49:46 to 17:52:54Z, FAILED exit 1) | the sampler proved itself (12 of 12 9070 XT watts lines), the 90 s app-state window ran (the card mining at 18.87 MH/s, 203 W, identities 8 before it), every gfx1201 entry was posted off, the card went quiet and the sh256x27 bench started; then the script exited with no APPROW, no LADDER, no `watts error=` and NO RESTORE line in any upload (no `enabled=True` post, no `card_workers_after`): a `finally` that did not run or did not print, so the 9070 XT may have been left OFF | CONFIRMED and restored: api/state at 17:55:10Z read the card enabled false, state off, 23 W idle; the identities job run d posted it back and at 17:57:10Z it mined under pid 3436 at 18.9 MH/s, 195 W, identities 8; the entry amd:1:gfx1201 was set back to enabled at 17:59:05Z (job run-ca3-pc1-amd-flag-amd1-20261006) with the card mining through it (18.86 MH/s, 200 W); PC 1 free at 18:01:05Z; the AMD watts row stays OWED; the cause and a fixed script (an unconditional finally with its own first line, no `exit` inside the try, the device dumps out of the report) with the script worker before any re-run | | the 4070 ladder's restore line | "no entry was enabled before this job" and no wait for the worker, while the card had mined at 28.8 MH/s before it | the intake shows the 4070's worker restarted 15 s after the restore and racing at 30.9 MH/s: the card came back; the script's settings read, not the card, was wrong | | PC 1 family run d (run-ca3-pc1-amd-family-20261006-d, 17:17 to 17:20Z, exit 0; run b had failed because the script's probe parameter was named `$args`, PowerShell's automatic variable, so the splat was empty: renamed, with a rule added to `tools/ci/ps-drive-ref-check.sh` that fires on the old signature and stays quiet on the new) | the probe chose the 9070 XT by name (gfx1201, 32 CUs); every variant built on it: mm8 NATIVE through the WMMA builtin on gfx12 (exact unverified), dot4 native (sudot4), bfe native, the shuffles native (ds_bpermute, ds_swizzle), the byte permute EMULATED (no amd_perm path in AMD's OpenCL C). The step costs are unusable: the card was mining beside the probe, the alu chain read 226 then 195 G steps/s and the ratios swung from 5.9 to 11.8x (run 1) to 0.17 to 1.3x (run 2), the loaded card's scheduler | the 9070 XT column is taken with the card alone (run e, job 1's switch path), after the G2 job | | the identities job, first run (run-ca3-pc1-amd-identities-20261006, 17:02:10Z, failed in 0 s) | my own script: `"... under $appDir: nothing posted"` is a PowerShell 5.1 parse error (`$appDir:` reads as a drive-qualified variable), so the script never started; the script worker checks this shape by hand, CI did not | `${appDir}:`; the class guard `tools/ci/ps-drive-ref-check.sh` added to ci.yml (67 .ps1 files clean; a backtick-escaped `$` in a bash-generating here-string is ignored); the job republished as -b |