igneum/tools/ca3-pc1-amd/README.md

92 lines
31 KiB
Markdown

# Counter ASIC 3.0: the PC 1 AMD jobs (6 October 2026)
Four signed `run` jobs for PC 1 (machine `ae432dc7`: RTX 5090, RTX 4070, RX 9070 XT `amd:gfx1201` on the eGPU, all on the installed app) that fill the OWED AMD rows of `docs/plans/counter-asic-3-status.md` sections 3, 5 and 7, plus the 4070 row of item 8. Prepared on branch `ca3-pc1-amd`; NOTHING here is published by this branch. The coordinator publishes on "go PC 1 AMD", one job at a time, in the order below, and reads each closing report before the next (`node tools/jobs.mjs <id>`, `node tools/jobs.mjs watch <id>`).
## The rule of 6 October 2026, 18:xx UTC: a script never switches the installed app's cards
The watts job (`run-ca3-pc1-amd-watts-20261006`, 17:49:46 to 17:52:54Z, failed exit 1) switched the 9070 XT off itself, and at 17:52:54Z the installed app went from 0.3.13 to 0.3.14 (the old miner run `win-ae432dc7-20261006-152452` ends 17:52:59Z with the gRPC connections closing, the new run `win-ae432dc7-20261006-175303` starts 17:53:03Z with `version=0.3.14`); the relaunch killed the job's process tree 71 s into the bench, powershell.exe left with code 1, the script's `finally` never ran, the card stayed off until a hand POST (the report is 10,776 bytes: not the 200 KB cap; not the runner's abort or time cap, which print their own lines and report `aborted`/`timeout`). Which path relaunched under an active job is open: `tick_update`'s `miner_busy` counts `jobs.active()`, and the runner's `tick` starts a queued `update-now` or `restart` only with no active job, so neither should have; main's to settle from the app's event log.
Main's rule: a POST of `enabled:false` to `/api/cards` from a script is the same class as `/api/pause` from a script. A job that needs a card alone asks the RUNNER: `packaging/ota/publish-jobs.sh add --kind run ... --cards-off <key,key>` (this branch, for the cut after 0.3.14): the runner matches the present live cards by key with or without the device index (`jobrun::cards_off_choices`), switches them off through the app's own card path (`apply_cards`) before the script starts, writes `cards-off: <keys> ... cards restored: <keys> by the runner on any exit` into the job's report, and puts the exact entries back (their own enabled flag, identities and cap) on ANY exit: done, failed, timeout, aborted, and the app quitting (`Jobs::abort` returns the restore and the engine applies it on its own thread before the exit). Tests on the box (`igneum-app` 114 passed): `cards_off_matches_keys_with_and_without_the_device_index_and_restores_exactly` and `a_card_hold_restores_once_on_any_exit` (the known-failed case: the list comes out once whichever path ends the job, never twice). `tools/ci/playbook-quit-check.sh` rule 2 (taken from master `630b537` and extended): any request to `api/cards` from a script fails CI; the pre-rule playbooks of 5 and 6 October are on a dated allow list (`PRE_RULE`) as the record of runs already made, never republished; `relay/playbooks/shard-test.ps1` fails rule 1 on master too (pre-existing, not touched here). Every script in this folder now reads the card's state only (the quiet confirmation by the process list stays); jobs 1, 1b, 2 and 4 are published with `--cards-off`.
## Rules every script keeps
| Rule | How |
|---|---|
| One job at a time on PC 1 | the publish order below; each job's closing report read before the next |
| The card alone only where the fixture needs it | jobs 1, 1b, 2 (run e) and 4 are published with the RUNNER's `--cards-off <key>` (the card under test off through the app's own card path before the script, restored on any exit; the other two cards keep mining); job 1a and job 3 run beside the miners and print `card_state=` on every row; no script posts to `/api/cards` |
| Card keys with or without the device index | the runner's `--cards-off` matches a live card by its key as written or with the index removed (`amd:1:gfx1201` reaches the live `amd:gfx1201`); a key that matches no present card switches nothing and the report says so |
| Quiet confirmed by the process list | the worker process carrying the card's `--device <index>` must be gone (Win32_Process command lines; job 4 also reads `nvidia-smi -i <idx> --query-compute-apps`); an unconfirmed card prints `UNCONFIRMED` and the rows say `card_alone=False` |
| Restored by the runner | the exact entries back on any exit, by the runner, not the script; the scripts' `finally` only ends their sampler by its own pid (never by name: the app runs its own copy of the AMD helper) and prints `RESULT finally_begin` as its first line |
| The installed app untouched | no `/api/quit`, `/api/pause`, `/api/resume`, no manifest, no settings.json write; the only writes are POST `api/cards` on the card under test |
| Copied sources re-stamped | nothing is built on the PC by these jobs (no cargo; the OpenCL kernels are compiled by the driver at run time); `tools/ci/copied-sources-check.sh` passes |
| Prover socket | not applicable (no prover host in these jobs) |
| Every failure a `RESULT ... error=<text>` line, a `SUMMARY {json}` line at the end | all four scripts |
If `--stop-miners` IS added at publish time, jobs 1 and 4 see the card already quiet, say so (`card already quiet before the switch`) and skip the switch: the scripts work either way.
## The kit (one fetch job, before job 1)
`tools/ca3-pc1-amd/make-kit.sh` builds `igneum-ca3-pc1-amd-kit.zip`: `bin/igneum-worker-opencl.exe` (THIS tree's `proto-opencl/host.c`, the installed worker's source plus one `build <ms> clBuildProgram` line, cross-compiled with mingw as `proto-cuda/nvrtc/build-windows.sh` does), `bin/family-probe-cl.exe` (`proto-opencl/family-probe.c`), `src/` (both sources and `cl_dynamic.h` for the sha256 record), `packs/` (mx8-devnet-epoch0, sh256x13, sh256x27, sh256x53, sh256x88, sh64x52, dr736-genesis, dr736-devnet-epoch0, each with its kernel_bound.cl, program.json and vectors), `SHA256SUMS`.
Built 6 October 2026, 15:4x UTC (`IGNEUM_REDIST=/Users/joshm/Projects/igneum-wt-ca2-mixer/proto-cuda/nvrtc/redist`): zip sha256 `a70fce5be672f33c3af08a5699488c09aef2848010f468ec99165fd4ab9ca61f`, 884,381 bytes, 113 files; `bin/igneum-worker-opencl.exe` `a572948e86c23317a80d9d7bb9e4bf70d5f7dfc40470974e2f08da1892016ac6`, `bin/family-probe-cl.exe` `d51a36bcae13ab36c40c2f347e3bdbe28f010499ccf84e423ecc04a9b4f151e3`, `src/host.c` `5e23ac94…`, `src/family-probe.c` `c293be9d…`. The zip is in the session scratchpad (`pc1amd/igneum-ca3-pc1-amd-kit.zip`), not in the repository; re-running `make-kit.sh` rebuilds it and prints a new sha256 (mingw output is not byte-reproducible), and the publish line takes the printed one. Every job prints `RESULT kitfile <path> sha256 <hex>` for the exes and sources it used, so the record closes on the PC side.
```
tools/ca3-pc1-amd/make-kit.sh "$TMPDIR/igneum-ca3-pc1-amd-kit.zip"
packaging/ota/publish-jobs.sh add --kind fetch --target ae432dc7 --id fetch-ca3-pc1-amd-20261006 \
--file "$TMPDIR/igneum-ca3-pc1-amd-kit.zip" --dir jobs --extract --title "CA3 PC 1 AMD kit" --expires-hours 48 --deploy
```
The fetch lands at `<jobs>\fetch-ca3-pc1-amd-20261006\{bin,packs,src,SHA256SUMS}` (jobrun.rs `fetch_base`: the job id under the jobs folder). The wiped-jobs-folder class: an app update clears the jobs folder, so if the 0.3.12 install lands on PC 1 between jobs, republish the fetch (each run job tests the kit before use and fails in seconds with `kit missing` if it is gone).
## The jobs, in order
| # | Script | Publish (after the fetch; `--id` fixed so the read-back is one command) | Length | Card | Rows it fills |
|---|---|---|---|---|---|
| 1 | `pc1-amd-g1-shadow.ps1` | `packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-g1-shadow-20261006-b --cards-off amd:gfx1201 --script tools/ca3-pc1-amd/pc1-amd-g1-shadow.ps1 --timeout-minutes 40 --title "CA3 PC 1: G1 + AMD ladder on the 9070 XT" --deploy` (the app that carries `--cards-off` must be installed on PC 1 first: the cut after 0.3.14) | about 15 min (7 benches of 60 to 90 dispatches of 2^24 at 0.9 to 2 s each, the card switch up to 150 s, the installed-worker cross-check 40 dispatches) | 9070 XT ALONE (its worker stopped; 5090 and 4070 keep mining) | G1 for `mx8+sh256x27` on the AMD vendor (section 5: `RESULT G1 pack=sh256x27 fingerprint=3d2e8245cc084d07 match=yes` turns "AMD vendor NOT RUN" green; the other five packs give the same for the ladder); the 9070 XT column of the item 8 tables (section 3: `RESULT LADDER pack=<p> ops=<N> mhs= watts= uj= gclk_mhz=`), where the card leaves the latency bound by the 5 percent rule; item 1's AMD per-joule row from the control (`RESULT LADDER pack=mx8-devnet-epoch0 ... uj=`: section 3 "Consequences per tier (item 1)", 16 GB AMD row); section 7 rows "item 8: the 9070 XT rows", "item 1: the 9070 XT rows" |
| 0b (first in the queue after "Ember closed") | `pc1-amd-reset.ps1` | `packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-reset-20261006 --script tools/ca3-pc1-amd/pc1-amd-reset.ps1 --timeout-minutes 5 --title "CA3 PC 1: 9070 XT ADLX back to factory, identities read" --deploy` | about 1 min | beside the miners, no card switch | the 9070 XT's ADLX tuning state after Ember run 6 (`gmax 0 plimit 0 factory 0`): `RESULT amd_state before gmax= plimit= factory=`, the factory reset as the Ember playbook did it (ember-tune ca990c2: `igneum-gpu-telemetry --card <ordinal> --reset`, the helper that answers `--tune`: the installed one first, else the newest amd-kit copy, which one printed), `RESULT amd_state after ...`, FAILED unless the after line reads factory 1; then READ ONLY the identities question of the 15:56Z restore: every gfx1201 settings.json entry (`RESULT settings_entry key= enabled= identities=`) and the app's live card list (`RESULT state_card key= identities= state= mhs_now= watts=`), `RESULT identities_answer live_key= identities=`; no POST of identities (the coordinator's call after the read) |
| 1a (the control watts row, read only, any slot) | `pc1-amd-watts-app.ps1` | `packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-watts-app-20261006 --script tools/ca3-pc1-amd/pc1-amd-watts-app.ps1 --timeout-minutes 10 --title "CA3 PC 1: watts of the 9070 XT in the app (class v3 control)" --deploy` | about 2 min | the card ON in the app, nothing switched | item 1's AMD per-joule row from the app's own path: `RESULT APPROW class=v3 mhs_wall= watts_mean= watts_min= watts_max= samples= distinct_watts= uj=` (printed right after the 90 s window, robust to a missing field) and the helper's calibration beside it: `RESULT sampler_vs_app helper_mean= app_mean= ratio=`, the ratio job 1b reads with |
| 1b (the candidate watts row, after the cut that carries `--cards-off`) | `pc1-amd-watts.ps1` | `R=<the ratio from job 1a's sampler_vs_app line>; sed "s/__CALIBRATION_RATIO__/$R/" tools/ca3-pc1-amd/pc1-amd-watts.ps1 > "$TMPDIR/pc1-amd-watts.ps1"; packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-watts-20261006-b --cards-off amd:gfx1201 --script "$TMPDIR/pc1-amd-watts.ps1" --timeout-minutes 15 --title "CA3 PC 1: watts of sh256x27 on the 9070 XT alone" --deploy` (without the sed the row prints `calibration_ratio=owed` and the raw helper watts) | about 3 min | the 9070 XT ALONE through the runner's `--cards-off` | the candidate's AMD cell of the item 8 table: `RESULT LADDER pack=sh256x27 ops=102100 mhs= watts= uj= ... watts_source=adlx calibration_ratio= raw_helper_mean=`; the sampler proven on 12 idle seconds first, no bench without it; a ratio outside 0.95 to 1.05 prints the watts as owed with the reason | the two `watts= uj=` cells job 1 left owed, taken the way main trusts after Ember run 6 (the app's path: engine telemetry 363 of 363 nonzero, /api/state 456 of 456; a direct helper call alone is not trusted): (1) the CONTROL row with the card ON in the app mining class v3, `/api/state` polled at 1 Hz for 90 s (`RESULT APPROW class=v3 mhs_wall= watts_mean= watts_min= watts_max= samples= uj=`; the app's watts come from its own helper at -l 5, so they move every 5 s: `distinct_watts=` says how many values), the helper sampler running beside over the same 90 s (`RESULT sampler_vs_app helper_mean= app_mean= ratio=`); (2) the CANDIDATE row (sh256x27, the card alone through the kit worker, which /api/state cannot see) with the sampler as its only source and the ratio as its calibration (`RESULT LADDER ... watts_source=adlx calibration_ratio=`); a ratio outside 0.95 to 1.05 prints the candidate's watts as owed with the reason. The sampler is proven on the 12 idle seconds first (`RESULT sampler_raw` x3, `RESULT sampler_idle lines= lines_9070_with_watts=`); without a 9070 XT watts line no row is taken (`RESULT watts error=<why>`, status failed). About 5 min: 12 s proof, 90 s app window, the switch, about 90 s of bench, the restore |
| G2 (job 2's slot, after "Ember closed") | `pc1-amd-g2.ps1` | its own small kit first: `tools/ca3-pc1-amd/make-g2-kit.sh "$TMPDIR/igneum-ca3-pc1-amd-g2-kit.zip"` (rebuilt 6 October 16:4x UTC from the merged tree with the generator-4 packs and the merged packfile.h: sha256 `a16165483cbef9c9001897e2e965b95051c4e683940e0dbc9bded002ed0a0e7e`, 198,597 bytes, 35 files; worker exe `7dc3b1ee…`; the 16:2x build `f4c029c7…` carried the generator-3 copies and is superseded), then `packaging/ota/publish-jobs.sh add --kind fetch --target ae432dc7 --id fetch-ca3-pc1-amd-g2-20261006 --file "$TMPDIR/igneum-ca3-pc1-amd-g2-kit.zip" --dir jobs --extract --title "CA3 PC 1 G2 kit" --deploy`, then `packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-g2-20261006 --script tools/ca3-pc1-amd/pc1-amd-g2.ps1 --timeout-minutes 10 --title "CA3 PC 1: gate G2 on the 9070 XT" --deploy` | about 2 min (two serve-mode jobs of 1,024 nonces; the Apple OpenCL dry run took 2.6 s a pack) | beside the miners (a correctness gate) | the "RX 9070 XT" cell of G2 in `docs/plans/counter-asic-3-gate/hash-gates.md` and the AMD line of G2 in status section 5: `RESULT g2 <pack> found <n> of 1024 distinct=<d> sha256 <digest> file <path> exit=<code>` (the hash lane's form) plus `RESULT G2 pack= found= distinct= digest= digest_match=yes|no`; the found lines go to `g2-<pack>.found` in the job folder, never to stdout; the Mac re-check is `tools/ca3-v4/g2-recheck.sh proto-cuda/packs-ca3-v4/<pack> --digest <hex>` (a digest match is 1,024 of 1,024), or `--file` after `publish-jobs.sh add --kind collect --target ae432dc7 --glob "app/jobs/run-ca3-pc1-amd-g2-20261006/g2-*.found"`. Packs: `packs-ca3-v4/mx8-devnet-epoch0` (control, class v3, id 73bcbfe8ccf988f1) and `packs-ca3-v4/v4-devnet-epoch0` (the candidate with the era at the devnet epoch-0 chain seeds; since the program-id fix 7c22d0d generator 4, class "v4", id c120d7963abdcd96, kernels and vectors byte-identical), NOT the first kit's `sh256x27`: that one is the genesis string-seed pack (`IGNEUM_SEED_BYTES_HEX` of 14 bytes) and the worker's job protocol takes only a 64-hex epoch seed, so it cannot be served (the script refuses such a pack with a RESULT line instead of a malformed job). The job line's class token is the pack's own `IGNEUM_PROGRAM_CLASS` (`class=v3` for the control, `class=v4` for the candidate), as `pc2-v4-gates.ps1` does since ca61dec: the worker refuses a token that disagrees with the pack (Mac check 16:41Z: `class=v3` on the v4 pack found 0, the known-bad case). Expected digests (the verifier through g2-recheck.sh on the Mac, equal to the kit worker's serve mode on Apple OpenCL at 16:21Z on the generator-3 copies and again at 16:41Z on the generator-4 packs through the merged host.c and packfile.h with `class=v4`, 1,024 of 1,024 every time; UNCHANGED by the re-export, as the byte-identical kernels require): mx8 `2a1824a2e0834287cf53b33f84cb7322014c85efa9788bece574e4e0730940bd`, v4 `435b976a4de57c5a26c8bada8b3b9e003504687c790d4b856bc56619d3aabd18`; the script carries them and prints `digest_match` itself |
| 2 (run e, the 9070 XT alone) | `pc1-amd-family.ps1` | the family kit of run b is on the PC (probe exe unchanged); publish only `packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-family-20261006-e --cards-off amd:gfx1201 --script tools/ca3-pc1-amd/pc1-amd-family.ps1 --timeout-minutes 20 --title "CA3 PC 1: family step costs on the 9070 XT (run e, the card alone)" --deploy` | about 5 min (the switch up to 150 s, 3 runs alone, the restore, then the gfx1036 and old-platform columns beside the miners) | the 9070 XT ALONE for its three runs (job 1's path: every entry's key forms, process-list confirmation, each entry restored with its own flag and identities in finally; 5090 and 4070 keep mining); the other AMD columns beside the miners | run d (17:17Z) showed the dependent-chain ratio does not survive the card's own miner (alu 226 then 195 G steps/s; ratios 5.9 to 11.8x then 0.17 to 1.3x), so the 9070 XT column is taken alone as the Mac's and the 5090's were; rows carry `card_state=alone` (or `loaded` with `UNCONFIRMED` printed when the worker did not go); mm8 stays exact=unverified (the gfx12 WMMA iu8 16x16x16 fragment layout is in no source at hand; a guessed CPU layout would turn a wrong guess into exact=no) |
| 2 (run b, superseded by run e) | `pc1-amd-family.ps1` | its own kit first: `tools/ca3-pc1-amd/make-family-kit.sh "$TMPDIR/igneum-ca3-pc1-amd-family-kit.zip"` (built 6 October 18:0x UTC: sha256 `c8337e74af5f589d651603eddd3561700219bb5077f531a600a21719bb87ef58`; the probe exe `e7ad93a4d5b4d977…`), then `packaging/ota/publish-jobs.sh add --kind fetch --target ae432dc7 --id fetch-ca3-pc1-amd-family-20261006 --file "$TMPDIR/igneum-ca3-pc1-amd-family-kit.zip" --dir jobs --extract --title "CA3 PC 1 family kit" --deploy`, then `packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-family-20261006-b --script tools/ca3-pc1-amd/pc1-amd-family.ps1 --timeout-minutes 15 --title "CA3 PC 1: family step costs on the 9070 XT (run b, by name)" --deploy` | about 3 min (3 runs on the 9070 XT by name, then 3 each on the older-platform duplicate and the gfx1036 by ordinal; builds dominate) | beside the miners (ratios; every row says `card_state=`) | the "RX 9070 XT" column of the item 6 table (section 3: `RESULT FAMILYBEST name=<f> ratio=<r> exact= path=` per family, `RESULT FAMILY ... variant=` for the form each took: `shfla_bperm` = `ds_bpermute_b32`, the one number that could move R3; `shflx_swz` / `shflx_bperm`; `perm_amd` = `v_perm_b32` through `amd_perm` or `perm_c` emulated; `bfe_amd`; `dot4_amd` = the 5 October sudot4 row re-measured; `mm8_gfx12` / `mm8_gfx11` = the WMMA builtins, exact=unverified); section 7 row "item 6: the 9070 XT step costs" |
| 3 | `pc1-amd-derive.ps1` | `packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-derive-20261006 --script tools/ca3-pc1-amd/pc1-amd-derive.ps1 --timeout-minutes 15 --title "CA3 PC 1: dr736 on the 9070 XT" --deploy` | about 4 min (3 packs x a 60-dispatch pass and a 3-dispatch pass) | beside the miners (build, compile and bit-exact rows; the rate as a ratio to mx8 in the same job) | the item 2 table's 9070 XT cells (section 3: `RESULT DERIVE pack=dr736-genesis build_ms= build_ms_cached= build_ms_over_mx8= dataset_ms= fingerprint= match= mhs= ratio_to_mx8=`): the OpenCL compile cost of the derivation program against mx8's (the NVRTC +1.1 s finding on AMD), the 1 GiB daily build against mx8's 72 to 77 ms (5 October), bit-exactness on the AMD vendor; section 7 row "item 2: the 9070 XT rows" |
| 4 (optional, confirmed by main) | `pc1-4070-shadow.ps1` | `packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-4070-shadow-20261006 --cards-off "nvidia:NVIDIA GeForce RTX 4070" --script tools/ca3-pc1-amd/pc1-4070-shadow.ps1 --timeout-minutes 30 --title "CA3 PC 1: item 8 ladder on the RTX 4070" --deploy` | about 11 min (7 benches of 100 to 200 dispatches at an expected 0.4 to 1 s each, approximate: no 4070 row exists yet; NVRTC per pack; the card switch) | 4070 ALONE (its CUDA worker stopped; 5090 and 9070 XT keep mining) | the "small NVIDIA card binds near 100,000" row of item 8 (section 3, consequences per tier at N = 100,000; section 7 "the 4060-class rows"): `RESULT LADDER pack=<p> ops=<N> mhs= watts= uj= sm_mhz=` with nvidia-smi at 1 Hz on that index; G1 fingerprints on a second NVIDIA card |
Read back: `node tools/jobs.mjs run-ca3-pc1-amd-g1-shadow-20261006` (and `--all` for the 5-minute progress uploads). The run_id on the intake is `job-<id>-ae432dc7`.
## Job 1's first run (run-ca3-pc1-amd-g1-shadow-20261006, 15:46:27 to 15:56:54Z, exit 0) and what it found
G1 on the AMD vendor 7 of 7 (every fingerprint equal to the Mac's and the 5090's, self-test PASS, 96 of 96 lanes), the card to itself (worker pid 7060 gone after 5 s, back as pid 4900 after the restore), the ladder's rates on the kit worker (the installed worker's cross-check on the control 18.922 against 18.923 MH/s): 18.92 MH/s at 930 ops, 19.31 at 49,700, 19.29 at 102,100, 19.11 at 199,600, 19.60 at 330,700, 19.07 for the 64-instruction block at 49,700; OpenCL build 42 to 51 ms for mx8, 418 to 429 ms for the 256-instruction shadow packs, 342 for the 64-instruction one; the daily build 70 to 74 ms. Two faults, both fixed on this branch (commit after e054ed7):
| Fault | Cause | Fix |
|---|---|---|
| every LADDER row `watts=owed samples=0` though the one-shot helper call printed the 9070 XT at 140.0 W | `Start-Process powershell.exe -ArgumentList @(..., $tele, $samples)`: Windows PowerShell 5.1 joins an ArgumentList with spaces and does not quote, and the helper's path holds "Igneum Miner", so the wrapper's `$Exe` was `...\Programs\Igneum`, its `& $Exe -l 1` failed and nothing was written | the wrapper reads its paths from `IGNEUM_SAMPLER_EXE` and `IGNEUM_SAMPLER_OUT` in the environment (inherited by the child, never on a command line), the `-File` path is quoted, the wrapper logs its own start line with `exists=`; the sampler must prove itself in the 12 idle seconds (first three raw lines printed, a 9070 XT line with watts required) or the benches are not run and the job fails loudly with `RESULT watts error=<why>` |
| the restore posted one entry's identities (2) to every key form while settings.json holds three gfx1201 entries (`amd:3:gfx1201` off / 2, `amd:gfx1201` on / 8, `amd:1:gfx1201` on / 2); the worker came back (pid 4900), but if the app's live key is `amd:gfx1201` its identities moved from 8 to 2 | one `$cardIdent` for every form | every entry is restored with ITS OWN enabled flag and identities (and cap, job 4), the derived forms first and the key as written last so an entry's own settings win on the live key; job 4 carries the same shape. OWED: a read of PC 1's settings.json (or the app's card tile) to confirm which gfx1201 key the app runs and whether its identities read 8 or 2 after the 15:56Z restore; if 2 and the app runs `amd:gfx1201`, one `POST api/cards {key: amd:gfx1201, enabled: true, identities: 8}` puts it back (the coordinator's call, not done by any job here) |
## Job 2's first run (run-ca3-pc1-amd-family-20261006, 150 s, exit 0) and what it found
Every probe ran ordinal 0, PC 1's integrated gfx1036 (one CU: alu 40.59 G steps/s against the Mac's 881 and the 5090's 7,941; the 9070 XT, 32 CUs, should read about 1,000 to 2,000): the script's device-list parse ran a second `-match` after the capturing one, which overwrote `$Matches`, so every device read `idx 0` and an empty name, and the probe took a bare ordinal. The rows are a gfx1036 (RDNA 2 iGPU) column, kept: alu 1.00, shifts 1.15, bfe 1.15 native, andn 1.16, popc 1.55, clz 1.75, sel 1.81, shfla and shflx 1.89 through `ds_bpermute`, perm 2.43 emulated, dot4 2.57 emulated, mm8 none. The 9070 XT column stays owed to run b. Fixed (commit after c38dfef): the probe host takes `--device-name gfx1201` (the match on the newest AMD platform by driver version, the kit worker's dedup rule; `RESULT device_choice name= index= platform= driver= cus=`; `RESULT error` and exit 2 when no name matches, a bare ordinal never the default), prints the device's name, CUs and platform on every `RESULT FAMILY` and `FAMILYBEST` line; the script runs the 9070 XT by name first and the older-platform duplicate and the gfx1036 by explicit ordinal after it as their own labelled columns, and fails the job when the 9070 XT gave no row. Mac check (18:0x UTC, `with-lock.sh run`, load 9.9): `--device-name M5` chose the M5 Max with 40 CUs and ran bit-exact; `--device-name gfx1201` on the Mac printed `RESULT error no OpenCL GPU device whose name holds "gfx1201"` and exit 2 (the known-bad case); no device given printed `RESULT error no device chosen`.
What run a's build failures say, by their log lines (the gfx1036 ran, not the 9070 XT, so each refusal is classed by its cause): `amd_perm` fails with "use of undeclared identifier" although the device lists `cl_amd_media_ops2` and `amd_bfe` from the same extension builds, so the byte permute has NO OpenCL C path on AMD's LC compiler and that holds for the 9070 XT too (R1 reads emulated on AMD, 2.43x on the iGPU, the 9070 XT ratio from run b); `__builtin_amdgcn_sudot4` fails with "needs target feature dot8-insts", a target-feature refusal of RDNA 2 (the 5 October log said the same), and the same builtin DID build and run bit-exact on the 9070 XT on 5 October, so the dot4 row on gfx1201 is decided by run b, not by run a; the two WMMA builtins fail with "needs target feature wmma…", again the RDNA 2 target, and run b says whether the LC compiler reaches them on gfx1201 (if it does the mm8 row gets a step cost with exact=unverified; if not, mm8 has no OpenCL C path and stays OWED on AMD); `cl_khr_subgroup_shuffle` and `cl_intel_subgroups` are not offered (warnings, then the undeclared function), while `ds_bpermute` and `ds_swizzle` build, so the shuffle families (R3) are native on AMD through the clang builtin. Consequence for the AMD tier: a byte-permute family costs an AMD miner the emulated sequence on every step until the worker gains an ISA path (inline assembly through the LC compiler, or a HIP build of the probe); the dot4 and mm8 answers for the 9070 XT, and every ratio, are run b's.
## The AMD watts readback on PC 1 (what exists)
`igneum-gpu-telemetry.exe` (proto-opencl/gpu-telemetry.c on master, shipped by `packaging/windows/make-payload.sh` since 0.3.10; the Ember Tune playbook found it in the install folder or under a `jobs\amd-kit-*\kit\` folder) prints one line per AMD card per sample: `amd <n> bus <pci> kind discrete name "AMD Radeon RX 9070 XT" watts <W> temp_c <C> fan_rpm <r> fan_pct <p> mclk_mhz <m> gclk_mhz <g> util_pct <u> source adlx`, then `end <ms> N card(s)`, at `-l 1` once a second with `fflush` per sample. The watts field is ADLX `GPUPower` with `GPUTotalBoardPower` as the fallback (gpu-telemetry.c lines 120 to 121 on master); on 5 October it read the 9070 XT at 198.9 W at 17.73 MH/s (release-0.3.10.md, job `tele-measure-1`), the figure the app's MH per watt line uses. So a board-watts readback EXISTS and job 1 uses it: a second copy of the helper at 1 Hz through a wrapper that stamps each line (UTC, ms) so the samples window onto each bench (after its first 12 s, before its last 2 s, the PC 2 shape), ended by the wrapper's pid in `finally`. The job prints `RESULT tele helper=<path> sha256=<hex> watts_source=adlx` after one sample; if the helper is missing or prints `watts -` for the 9070 XT the rows carry `watts=owed uj=owed` and the SUMMARY says `watts: owed`, never a guess. Whether ADLX `GPUPower` on RDNA 4 is total board power or ASIC power is not verified here: the row says `adlx` and the number is what the app itself reports.
## What the Mac tested (6 October 2026, 15:27 to 15:28 UTC, `with-lock.sh run`, load average 3.9 at the start)
| Check | Result |
|---|---|
| `proto-opencl/family-probe.c` on Apple OpenCL (`cc -framework OpenCL`, M5 Max) | every C-form and emulated row bit-exact on both 32-lane groups in 3 repetitions; ratios to the alu chain beside the Metal probe's (rotr 1.14 against Metal 1.13, shl 0.89 / 0.85, shr 0.89 / 0.86, bfe 0.82 / 0.77, andn 0.84 / 0.75, popc 1.00 / 0.87, clz 1.11 / 1.01, sel 0.82 / 0.76, perm emulated 1.52 / 1.13, shflx and shfla through `__local` 1.96 and 2.04 (Metal native 0.86 and 1.91), dot4 emulated signed 4.66 / 4.73); the Apple `dot(char4,char4)` row exact=no as the 5 October finding said; every AMD builtin skipped (`build=skipped`), the khr, intel and amd_media_ops2 forms `build=failed` with the log line, the run went on; `FAMILYBEST` picked the exact rows |
| `proto-opencl/host.c` with the build line, `--bench-pack` on all eight kit packs on Apple OpenCL | self-test PASS (96 of 96 lanes) and the 2^24 fingerprint at base 0 equal to the recorded one on every pack: mx8-devnet-epoch0 `90f794dd556f7a3b`, sh256x13 `59ac286fe2a5a9ef`, sh256x27 `3d2e8245cc084d07`, sh256x53 `4f824b15cf2b124a`, sh256x88 `0572522e39a94d8a`, sh64x52 `9dd010f79d8ca9f4`, dr736-genesis `50e3eaa779da4f1e`, dr736-devnet-epoch0 `9553f6d5c667205a`; `build` line: mx8 70 ms, shadow packs 70 to 88 ms, dr736 337 to 356 ms (+270 ms on Apple OpenCL: item 2's compile cost shows on this harness too) |
| mingw cross-builds of both exes | clean (`-Wall -Wextra`), 116 KB and 35 KB |
| The kit worker's serve mode on Apple OpenCL (G2 dry runs, 16:21 UTC at load 5.2 to 5.7 on the generator-3 copies; 16:41 UTC at load 27 to 57, a loaded box, on the generator-4 packs with the merged packfile.h and the `class=v4` token: a correctness run, no number taken) | `packs-ca3-v4/mx8-devnet-epoch0` and `v4-devnet-epoch0`: 1,024 of 1,024 found each (2.6 s), nonces 0 to 1,023, the sorted found file's sha256 equal to the verifier's through `tools/ca3-v4/g2-recheck.sh --file` (the coordinator's `igneum-pow` binary): the two digests above |
| `tools/ci/bash-body-check.sh` on the eight scripts | 0 bash bodies, all parse |
| `tools/ci/kit-path-check.sh` on the eight scripts | 6 of 6 kit paths checked before use (the reset and watts-app jobs read no kit) |
| `tools/ci/playbook-quit-check.sh` (rule 2) on the tree | self-test fires on an `api/cards` POST and passes a read of `api/state`; every script here passes; the pre-rule playbooks on the dated allow list |
| `tools/ci/prover-socket-check.sh`, `copied-sources-check.sh`, `no-conflict-markers.sh` | pass |
| The Windows PowerShell 5.1 parse (`windows.yml` parse job) | the real parser runs only on the Windows runner and only over `relay/playbooks` and the other listed folders (not `tools/`); on the Mac the same rule's approximation (`tools/ci/check-workflow-shell.mjs`'s drive-reference regex, applied to these five files by hand) fired once on `"$dev: --stop-miners"` and the line was fixed to `${dev}:`; 0 findings after; brace, paren and bracket depth 0 by the bash-body-check tokenizer |
## Not tested on the Mac (the PC side)
The PowerShell itself never ran (no PowerShell on this Mac): the `api/cards` switch and its process-list confirmation, the wrapper sampler, `Get-CimInstance` command-line matching, the AMD helper's presence and its watts on the 9070 XT, the AMD OpenCL compiler's answer to each builtin (`ds_bpermute`, `ds_swizzle`, `amd_perm`, `amd_bfe`, the WMMA builtins), the driver's kernel cache (pass 2 of job 3), the 4070's nvidia-smi index against the app's `--device`, and the lengths above (estimates from the PC 2 runs and the 5 October 9070 XT rates). A failure of any of those prints a `RESULT ... error=` line and the SUMMARY carries `failed` or `partial`; the card restore in `finally` does not depend on them.