igneum/tools/ca3-pc1-amd/README.md

31 KiB

Counter ASIC 3.0: the PC 1 AMD jobs (6 October 2026)

Four signed run jobs for PC 1 (machine ae432dc7: RTX 5090, RTX 4070, RX 9070 XT amd:gfx1201 on the eGPU, all on the installed app) that fill the OWED AMD rows of docs/plans/counter-asic-3-status.md sections 3, 5 and 7, plus the 4070 row of item 8. Prepared on branch ca3-pc1-amd; NOTHING here is published by this branch. The coordinator publishes on "go PC 1 AMD", one job at a time, in the order below, and reads each closing report before the next (node tools/jobs.mjs <id>, node tools/jobs.mjs watch <id>).

The rule of 6 October 2026, 18:xx UTC: a script never switches the installed app's cards

The watts job (run-ca3-pc1-amd-watts-20261006, 17:49:46 to 17:52:54Z, failed exit 1) switched the 9070 XT off itself, and at 17:52:54Z the installed app went from 0.3.13 to 0.3.14 (the old miner run win-ae432dc7-20261006-152452 ends 17:52:59Z with the gRPC connections closing, the new run win-ae432dc7-20261006-175303 starts 17:53:03Z with version=0.3.14); the relaunch killed the job's process tree 71 s into the bench, powershell.exe left with code 1, the script's finally never ran, the card stayed off until a hand POST (the report is 10,776 bytes: not the 200 KB cap; not the runner's abort or time cap, which print their own lines and report aborted/timeout). Which path relaunched under an active job is open: tick_update's miner_busy counts jobs.active(), and the runner's tick starts a queued update-now or restart only with no active job, so neither should have; main's to settle from the app's event log.

Main's rule: a POST of enabled:false to /api/cards from a script is the same class as /api/pause from a script. A job that needs a card alone asks the RUNNER: packaging/ota/publish-jobs.sh add --kind run ... --cards-off <key,key> (this branch, for the cut after 0.3.14): the runner matches the present live cards by key with or without the device index (jobrun::cards_off_choices), switches them off through the app's own card path (apply_cards) before the script starts, writes cards-off: <keys> ... cards restored: <keys> by the runner on any exit into the job's report, and puts the exact entries back (their own enabled flag, identities and cap) on ANY exit: done, failed, timeout, aborted, and the app quitting (Jobs::abort returns the restore and the engine applies it on its own thread before the exit). Tests on the box (igneum-app 114 passed): cards_off_matches_keys_with_and_without_the_device_index_and_restores_exactly and a_card_hold_restores_once_on_any_exit (the known-failed case: the list comes out once whichever path ends the job, never twice). tools/ci/playbook-quit-check.sh rule 2 (taken from master 630b537 and extended): any request to api/cards from a script fails CI; the pre-rule playbooks of 5 and 6 October are on a dated allow list (PRE_RULE) as the record of runs already made, never republished; relay/playbooks/shard-test.ps1 fails rule 1 on master too (pre-existing, not touched here). Every script in this folder now reads the card's state only (the quiet confirmation by the process list stays); jobs 1, 1b, 2 and 4 are published with --cards-off.

Rules every script keeps

Rule How
One job at a time on PC 1 the publish order below; each job's closing report read before the next
The card alone only where the fixture needs it jobs 1, 1b, 2 (run e) and 4 are published with the RUNNER's --cards-off <key> (the card under test off through the app's own card path before the script, restored on any exit; the other two cards keep mining); job 1a and job 3 run beside the miners and print card_state= on every row; no script posts to /api/cards
Card keys with or without the device index the runner's --cards-off matches a live card by its key as written or with the index removed (amd:1:gfx1201 reaches the live amd:gfx1201); a key that matches no present card switches nothing and the report says so
Quiet confirmed by the process list the worker process carrying the card's --device <index> must be gone (Win32_Process command lines; job 4 also reads nvidia-smi -i <idx> --query-compute-apps); an unconfirmed card prints UNCONFIRMED and the rows say card_alone=False
Restored by the runner the exact entries back on any exit, by the runner, not the script; the scripts' finally only ends their sampler by its own pid (never by name: the app runs its own copy of the AMD helper) and prints RESULT finally_begin as its first line
The installed app untouched no /api/quit, /api/pause, /api/resume, no manifest, no settings.json write; the only writes are POST api/cards on the card under test
Copied sources re-stamped nothing is built on the PC by these jobs (no cargo; the OpenCL kernels are compiled by the driver at run time); tools/ci/copied-sources-check.sh passes
Prover socket not applicable (no prover host in these jobs)
Every failure a RESULT ... error=<text> line, a SUMMARY {json} line at the end all four scripts

If --stop-miners IS added at publish time, jobs 1 and 4 see the card already quiet, say so (card already quiet before the switch) and skip the switch: the scripts work either way.

The kit (one fetch job, before job 1)

tools/ca3-pc1-amd/make-kit.sh builds igneum-ca3-pc1-amd-kit.zip: bin/igneum-worker-opencl.exe (THIS tree's proto-opencl/host.c, the installed worker's source plus one build <ms> clBuildProgram line, cross-compiled with mingw as proto-cuda/nvrtc/build-windows.sh does), bin/family-probe-cl.exe (proto-opencl/family-probe.c), src/ (both sources and cl_dynamic.h for the sha256 record), packs/ (mx8-devnet-epoch0, sh256x13, sh256x27, sh256x53, sh256x88, sh64x52, dr736-genesis, dr736-devnet-epoch0, each with its kernel_bound.cl, program.json and vectors), SHA256SUMS.

Built 6 October 2026, 15:4x UTC (IGNEUM_REDIST=/Users/joshm/Projects/igneum-wt-ca2-mixer/proto-cuda/nvrtc/redist): zip sha256 a70fce5be672f33c3af08a5699488c09aef2848010f468ec99165fd4ab9ca61f, 884,381 bytes, 113 files; bin/igneum-worker-opencl.exe a572948e86c23317a80d9d7bb9e4bf70d5f7dfc40470974e2f08da1892016ac6, bin/family-probe-cl.exe d51a36bcae13ab36c40c2f347e3bdbe28f010499ccf84e423ecc04a9b4f151e3, src/host.c 5e23ac94…, src/family-probe.c c293be9d…. The zip is in the session scratchpad (pc1amd/igneum-ca3-pc1-amd-kit.zip), not in the repository; re-running make-kit.sh rebuilds it and prints a new sha256 (mingw output is not byte-reproducible), and the publish line takes the printed one. Every job prints RESULT kitfile <path> sha256 <hex> for the exes and sources it used, so the record closes on the PC side.

tools/ca3-pc1-amd/make-kit.sh "$TMPDIR/igneum-ca3-pc1-amd-kit.zip"
packaging/ota/publish-jobs.sh add --kind fetch --target ae432dc7 --id fetch-ca3-pc1-amd-20261006 \
  --file "$TMPDIR/igneum-ca3-pc1-amd-kit.zip" --dir jobs --extract --title "CA3 PC 1 AMD kit" --expires-hours 48 --deploy

The fetch lands at <jobs>\fetch-ca3-pc1-amd-20261006\{bin,packs,src,SHA256SUMS} (jobrun.rs fetch_base: the job id under the jobs folder). The wiped-jobs-folder class: an app update clears the jobs folder, so if the 0.3.12 install lands on PC 1 between jobs, republish the fetch (each run job tests the kit before use and fails in seconds with kit missing if it is gone).

The jobs, in order

# Script Publish (after the fetch; --id fixed so the read-back is one command) Length Card Rows it fills
1 pc1-amd-g1-shadow.ps1 packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-g1-shadow-20261006-b --cards-off amd:gfx1201 --script tools/ca3-pc1-amd/pc1-amd-g1-shadow.ps1 --timeout-minutes 40 --title "CA3 PC 1: G1 + AMD ladder on the 9070 XT" --deploy (the app that carries --cards-off must be installed on PC 1 first: the cut after 0.3.14) about 15 min (7 benches of 60 to 90 dispatches of 2^24 at 0.9 to 2 s each, the card switch up to 150 s, the installed-worker cross-check 40 dispatches) 9070 XT ALONE (its worker stopped; 5090 and 4070 keep mining) G1 for mx8+sh256x27 on the AMD vendor (section 5: RESULT G1 pack=sh256x27 fingerprint=3d2e8245cc084d07 match=yes turns "AMD vendor NOT RUN" green; the other five packs give the same for the ladder); the 9070 XT column of the item 8 tables (section 3: RESULT LADDER pack=<p> ops=<N> mhs= watts= uj= gclk_mhz=), where the card leaves the latency bound by the 5 percent rule; item 1's AMD per-joule row from the control (RESULT LADDER pack=mx8-devnet-epoch0 ... uj=: section 3 "Consequences per tier (item 1)", 16 GB AMD row); section 7 rows "item 8: the 9070 XT rows", "item 1: the 9070 XT rows"
0b (first in the queue after "Ember closed") pc1-amd-reset.ps1 packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-reset-20261006 --script tools/ca3-pc1-amd/pc1-amd-reset.ps1 --timeout-minutes 5 --title "CA3 PC 1: 9070 XT ADLX back to factory, identities read" --deploy about 1 min beside the miners, no card switch the 9070 XT's ADLX tuning state after Ember run 6 (gmax 0 plimit 0 factory 0): RESULT amd_state before gmax= plimit= factory=, the factory reset as the Ember playbook did it (ember-tune ca990c2: igneum-gpu-telemetry --card <ordinal> --reset, the helper that answers --tune: the installed one first, else the newest amd-kit copy, which one printed), RESULT amd_state after ..., FAILED unless the after line reads factory 1; then READ ONLY the identities question of the 15:56Z restore: every gfx1201 settings.json entry (RESULT settings_entry key= enabled= identities=) and the app's live card list (RESULT state_card key= identities= state= mhs_now= watts=), RESULT identities_answer live_key= identities=; no POST of identities (the coordinator's call after the read)
1a (the control watts row, read only, any slot) pc1-amd-watts-app.ps1 packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-watts-app-20261006 --script tools/ca3-pc1-amd/pc1-amd-watts-app.ps1 --timeout-minutes 10 --title "CA3 PC 1: watts of the 9070 XT in the app (class v3 control)" --deploy about 2 min the card ON in the app, nothing switched item 1's AMD per-joule row from the app's own path: RESULT APPROW class=v3 mhs_wall= watts_mean= watts_min= watts_max= samples= distinct_watts= uj= (printed right after the 90 s window, robust to a missing field) and the helper's calibration beside it: RESULT sampler_vs_app helper_mean= app_mean= ratio=, the ratio job 1b reads with
1b (the candidate watts row, after the cut that carries --cards-off) pc1-amd-watts.ps1 R=<the ratio from job 1a's sampler_vs_app line>; sed "s/__CALIBRATION_RATIO__/$R/" tools/ca3-pc1-amd/pc1-amd-watts.ps1 > "$TMPDIR/pc1-amd-watts.ps1"; packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-watts-20261006-b --cards-off amd:gfx1201 --script "$TMPDIR/pc1-amd-watts.ps1" --timeout-minutes 15 --title "CA3 PC 1: watts of sh256x27 on the 9070 XT alone" --deploy (without the sed the row prints calibration_ratio=owed and the raw helper watts) about 3 min the 9070 XT ALONE through the runner's --cards-off the candidate's AMD cell of the item 8 table: RESULT LADDER pack=sh256x27 ops=102100 mhs= watts= uj= ... watts_source=adlx calibration_ratio= raw_helper_mean=; the sampler proven on 12 idle seconds first, no bench without it; a ratio outside 0.95 to 1.05 prints the watts as owed with the reason
G2 (job 2's slot, after "Ember closed") pc1-amd-g2.ps1 its own small kit first: tools/ca3-pc1-amd/make-g2-kit.sh "$TMPDIR/igneum-ca3-pc1-amd-g2-kit.zip" (rebuilt 6 October 16:4x UTC from the merged tree with the generator-4 packs and the merged packfile.h: sha256 a16165483cbef9c9001897e2e965b95051c4e683940e0dbc9bded002ed0a0e7e, 198,597 bytes, 35 files; worker exe 7dc3b1ee…; the 16:2x build f4c029c7… carried the generator-3 copies and is superseded), then packaging/ota/publish-jobs.sh add --kind fetch --target ae432dc7 --id fetch-ca3-pc1-amd-g2-20261006 --file "$TMPDIR/igneum-ca3-pc1-amd-g2-kit.zip" --dir jobs --extract --title "CA3 PC 1 G2 kit" --deploy, then packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-g2-20261006 --script tools/ca3-pc1-amd/pc1-amd-g2.ps1 --timeout-minutes 10 --title "CA3 PC 1: gate G2 on the 9070 XT" --deploy about 2 min (two serve-mode jobs of 1,024 nonces; the Apple OpenCL dry run took 2.6 s a pack) beside the miners (a correctness gate) the "RX 9070 XT" cell of G2 in docs/plans/counter-asic-3-gate/hash-gates.md and the AMD line of G2 in status section 5: RESULT g2 <pack> found <n> of 1024 distinct=<d> sha256 <digest> file <path> exit=<code> (the hash lane's form) plus `RESULT G2 pack= found= distinct= digest= digest_match=yes
2 (run e, the 9070 XT alone) pc1-amd-family.ps1 the family kit of run b is on the PC (probe exe unchanged); publish only packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-family-20261006-e --cards-off amd:gfx1201 --script tools/ca3-pc1-amd/pc1-amd-family.ps1 --timeout-minutes 20 --title "CA3 PC 1: family step costs on the 9070 XT (run e, the card alone)" --deploy about 5 min (the switch up to 150 s, 3 runs alone, the restore, then the gfx1036 and old-platform columns beside the miners) the 9070 XT ALONE for its three runs (job 1's path: every entry's key forms, process-list confirmation, each entry restored with its own flag and identities in finally; 5090 and 4070 keep mining); the other AMD columns beside the miners run d (17:17Z) showed the dependent-chain ratio does not survive the card's own miner (alu 226 then 195 G steps/s; ratios 5.9 to 11.8x then 0.17 to 1.3x), so the 9070 XT column is taken alone as the Mac's and the 5090's were; rows carry card_state=alone (or loaded with UNCONFIRMED printed when the worker did not go); mm8 stays exact=unverified (the gfx12 WMMA iu8 16x16x16 fragment layout is in no source at hand; a guessed CPU layout would turn a wrong guess into exact=no)
2 (run b, superseded by run e) pc1-amd-family.ps1 its own kit first: tools/ca3-pc1-amd/make-family-kit.sh "$TMPDIR/igneum-ca3-pc1-amd-family-kit.zip" (built 6 October 18:0x UTC: sha256 c8337e74af5f589d651603eddd3561700219bb5077f531a600a21719bb87ef58; the probe exe e7ad93a4d5b4d977…), then packaging/ota/publish-jobs.sh add --kind fetch --target ae432dc7 --id fetch-ca3-pc1-amd-family-20261006 --file "$TMPDIR/igneum-ca3-pc1-amd-family-kit.zip" --dir jobs --extract --title "CA3 PC 1 family kit" --deploy, then packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-family-20261006-b --script tools/ca3-pc1-amd/pc1-amd-family.ps1 --timeout-minutes 15 --title "CA3 PC 1: family step costs on the 9070 XT (run b, by name)" --deploy about 3 min (3 runs on the 9070 XT by name, then 3 each on the older-platform duplicate and the gfx1036 by ordinal; builds dominate) beside the miners (ratios; every row says card_state=) the "RX 9070 XT" column of the item 6 table (section 3: RESULT FAMILYBEST name=<f> ratio=<r> exact= path= per family, RESULT FAMILY ... variant= for the form each took: shfla_bperm = ds_bpermute_b32, the one number that could move R3; shflx_swz / shflx_bperm; perm_amd = v_perm_b32 through amd_perm or perm_c emulated; bfe_amd; dot4_amd = the 5 October sudot4 row re-measured; mm8_gfx12 / mm8_gfx11 = the WMMA builtins, exact=unverified); section 7 row "item 6: the 9070 XT step costs"
3 pc1-amd-derive.ps1 packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-derive-20261006 --script tools/ca3-pc1-amd/pc1-amd-derive.ps1 --timeout-minutes 15 --title "CA3 PC 1: dr736 on the 9070 XT" --deploy about 4 min (3 packs x a 60-dispatch pass and a 3-dispatch pass) beside the miners (build, compile and bit-exact rows; the rate as a ratio to mx8 in the same job) the item 2 table's 9070 XT cells (section 3: RESULT DERIVE pack=dr736-genesis build_ms= build_ms_cached= build_ms_over_mx8= dataset_ms= fingerprint= match= mhs= ratio_to_mx8=): the OpenCL compile cost of the derivation program against mx8's (the NVRTC +1.1 s finding on AMD), the 1 GiB daily build against mx8's 72 to 77 ms (5 October), bit-exactness on the AMD vendor; section 7 row "item 2: the 9070 XT rows"
4 (optional, confirmed by main) pc1-4070-shadow.ps1 packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-4070-shadow-20261006 --cards-off "nvidia:NVIDIA GeForce RTX 4070" --script tools/ca3-pc1-amd/pc1-4070-shadow.ps1 --timeout-minutes 30 --title "CA3 PC 1: item 8 ladder on the RTX 4070" --deploy about 11 min (7 benches of 100 to 200 dispatches at an expected 0.4 to 1 s each, approximate: no 4070 row exists yet; NVRTC per pack; the card switch) 4070 ALONE (its CUDA worker stopped; 5090 and 9070 XT keep mining) the "small NVIDIA card binds near 100,000" row of item 8 (section 3, consequences per tier at N = 100,000; section 7 "the 4060-class rows"): RESULT LADDER pack=<p> ops=<N> mhs= watts= uj= sm_mhz= with nvidia-smi at 1 Hz on that index; G1 fingerprints on a second NVIDIA card

Read back: node tools/jobs.mjs run-ca3-pc1-amd-g1-shadow-20261006 (and --all for the 5-minute progress uploads). The run_id on the intake is job-<id>-ae432dc7.

Job 1's first run (run-ca3-pc1-amd-g1-shadow-20261006, 15:46:27 to 15:56:54Z, exit 0) and what it found

G1 on the AMD vendor 7 of 7 (every fingerprint equal to the Mac's and the 5090's, self-test PASS, 96 of 96 lanes), the card to itself (worker pid 7060 gone after 5 s, back as pid 4900 after the restore), the ladder's rates on the kit worker (the installed worker's cross-check on the control 18.922 against 18.923 MH/s): 18.92 MH/s at 930 ops, 19.31 at 49,700, 19.29 at 102,100, 19.11 at 199,600, 19.60 at 330,700, 19.07 for the 64-instruction block at 49,700; OpenCL build 42 to 51 ms for mx8, 418 to 429 ms for the 256-instruction shadow packs, 342 for the 64-instruction one; the daily build 70 to 74 ms. Two faults, both fixed on this branch (commit after e054ed7):

Fault Cause Fix
every LADDER row watts=owed samples=0 though the one-shot helper call printed the 9070 XT at 140.0 W Start-Process powershell.exe -ArgumentList @(..., $tele, $samples): Windows PowerShell 5.1 joins an ArgumentList with spaces and does not quote, and the helper's path holds "Igneum Miner", so the wrapper's $Exe was ...\Programs\Igneum, its & $Exe -l 1 failed and nothing was written the wrapper reads its paths from IGNEUM_SAMPLER_EXE and IGNEUM_SAMPLER_OUT in the environment (inherited by the child, never on a command line), the -File path is quoted, the wrapper logs its own start line with exists=; the sampler must prove itself in the 12 idle seconds (first three raw lines printed, a 9070 XT line with watts required) or the benches are not run and the job fails loudly with RESULT watts error=<why>
the restore posted one entry's identities (2) to every key form while settings.json holds three gfx1201 entries (amd:3:gfx1201 off / 2, amd:gfx1201 on / 8, amd:1:gfx1201 on / 2); the worker came back (pid 4900), but if the app's live key is amd:gfx1201 its identities moved from 8 to 2 one $cardIdent for every form every entry is restored with ITS OWN enabled flag and identities (and cap, job 4), the derived forms first and the key as written last so an entry's own settings win on the live key; job 4 carries the same shape. OWED: a read of PC 1's settings.json (or the app's card tile) to confirm which gfx1201 key the app runs and whether its identities read 8 or 2 after the 15:56Z restore; if 2 and the app runs amd:gfx1201, one POST api/cards {key: amd:gfx1201, enabled: true, identities: 8} puts it back (the coordinator's call, not done by any job here)

Job 2's first run (run-ca3-pc1-amd-family-20261006, 150 s, exit 0) and what it found

Every probe ran ordinal 0, PC 1's integrated gfx1036 (one CU: alu 40.59 G steps/s against the Mac's 881 and the 5090's 7,941; the 9070 XT, 32 CUs, should read about 1,000 to 2,000): the script's device-list parse ran a second -match after the capturing one, which overwrote $Matches, so every device read idx 0 and an empty name, and the probe took a bare ordinal. The rows are a gfx1036 (RDNA 2 iGPU) column, kept: alu 1.00, shifts 1.15, bfe 1.15 native, andn 1.16, popc 1.55, clz 1.75, sel 1.81, shfla and shflx 1.89 through ds_bpermute, perm 2.43 emulated, dot4 2.57 emulated, mm8 none. The 9070 XT column stays owed to run b. Fixed (commit after c38dfef): the probe host takes --device-name gfx1201 (the match on the newest AMD platform by driver version, the kit worker's dedup rule; RESULT device_choice name= index= platform= driver= cus=; RESULT error and exit 2 when no name matches, a bare ordinal never the default), prints the device's name, CUs and platform on every RESULT FAMILY and FAMILYBEST line; the script runs the 9070 XT by name first and the older-platform duplicate and the gfx1036 by explicit ordinal after it as their own labelled columns, and fails the job when the 9070 XT gave no row. Mac check (18:0x UTC, with-lock.sh run, load 9.9): --device-name M5 chose the M5 Max with 40 CUs and ran bit-exact; --device-name gfx1201 on the Mac printed RESULT error no OpenCL GPU device whose name holds "gfx1201" and exit 2 (the known-bad case); no device given printed RESULT error no device chosen.

What run a's build failures say, by their log lines (the gfx1036 ran, not the 9070 XT, so each refusal is classed by its cause): amd_perm fails with "use of undeclared identifier" although the device lists cl_amd_media_ops2 and amd_bfe from the same extension builds, so the byte permute has NO OpenCL C path on AMD's LC compiler and that holds for the 9070 XT too (R1 reads emulated on AMD, 2.43x on the iGPU, the 9070 XT ratio from run b); __builtin_amdgcn_sudot4 fails with "needs target feature dot8-insts", a target-feature refusal of RDNA 2 (the 5 October log said the same), and the same builtin DID build and run bit-exact on the 9070 XT on 5 October, so the dot4 row on gfx1201 is decided by run b, not by run a; the two WMMA builtins fail with "needs target feature wmma…", again the RDNA 2 target, and run b says whether the LC compiler reaches them on gfx1201 (if it does the mm8 row gets a step cost with exact=unverified; if not, mm8 has no OpenCL C path and stays OWED on AMD); cl_khr_subgroup_shuffle and cl_intel_subgroups are not offered (warnings, then the undeclared function), while ds_bpermute and ds_swizzle build, so the shuffle families (R3) are native on AMD through the clang builtin. Consequence for the AMD tier: a byte-permute family costs an AMD miner the emulated sequence on every step until the worker gains an ISA path (inline assembly through the LC compiler, or a HIP build of the probe); the dot4 and mm8 answers for the 9070 XT, and every ratio, are run b's.

The AMD watts readback on PC 1 (what exists)

igneum-gpu-telemetry.exe (proto-opencl/gpu-telemetry.c on master, shipped by packaging/windows/make-payload.sh since 0.3.10; the Ember Tune playbook found it in the install folder or under a jobs\amd-kit-*\kit\ folder) prints one line per AMD card per sample: amd <n> bus <pci> kind discrete name "AMD Radeon RX 9070 XT" watts <W> temp_c <C> fan_rpm <r> fan_pct <p> mclk_mhz <m> gclk_mhz <g> util_pct <u> source adlx, then end <ms> N card(s), at -l 1 once a second with fflush per sample. The watts field is ADLX GPUPower with GPUTotalBoardPower as the fallback (gpu-telemetry.c lines 120 to 121 on master); on 5 October it read the 9070 XT at 198.9 W at 17.73 MH/s (release-0.3.10.md, job tele-measure-1), the figure the app's MH per watt line uses. So a board-watts readback EXISTS and job 1 uses it: a second copy of the helper at 1 Hz through a wrapper that stamps each line (UTC, ms) so the samples window onto each bench (after its first 12 s, before its last 2 s, the PC 2 shape), ended by the wrapper's pid in finally. The job prints RESULT tele helper=<path> sha256=<hex> watts_source=adlx after one sample; if the helper is missing or prints watts - for the 9070 XT the rows carry watts=owed uj=owed and the SUMMARY says watts: owed, never a guess. Whether ADLX GPUPower on RDNA 4 is total board power or ASIC power is not verified here: the row says adlx and the number is what the app itself reports.

What the Mac tested (6 October 2026, 15:27 to 15:28 UTC, with-lock.sh run, load average 3.9 at the start)

Check Result
proto-opencl/family-probe.c on Apple OpenCL (cc -framework OpenCL, M5 Max) every C-form and emulated row bit-exact on both 32-lane groups in 3 repetitions; ratios to the alu chain beside the Metal probe's (rotr 1.14 against Metal 1.13, shl 0.89 / 0.85, shr 0.89 / 0.86, bfe 0.82 / 0.77, andn 0.84 / 0.75, popc 1.00 / 0.87, clz 1.11 / 1.01, sel 0.82 / 0.76, perm emulated 1.52 / 1.13, shflx and shfla through __local 1.96 and 2.04 (Metal native 0.86 and 1.91), dot4 emulated signed 4.66 / 4.73); the Apple dot(char4,char4) row exact=no as the 5 October finding said; every AMD builtin skipped (build=skipped), the khr, intel and amd_media_ops2 forms build=failed with the log line, the run went on; FAMILYBEST picked the exact rows
proto-opencl/host.c with the build line, --bench-pack on all eight kit packs on Apple OpenCL self-test PASS (96 of 96 lanes) and the 2^24 fingerprint at base 0 equal to the recorded one on every pack: mx8-devnet-epoch0 90f794dd556f7a3b, sh256x13 59ac286fe2a5a9ef, sh256x27 3d2e8245cc084d07, sh256x53 4f824b15cf2b124a, sh256x88 0572522e39a94d8a, sh64x52 9dd010f79d8ca9f4, dr736-genesis 50e3eaa779da4f1e, dr736-devnet-epoch0 9553f6d5c667205a; build line: mx8 70 ms, shadow packs 70 to 88 ms, dr736 337 to 356 ms (+270 ms on Apple OpenCL: item 2's compile cost shows on this harness too)
mingw cross-builds of both exes clean (-Wall -Wextra), 116 KB and 35 KB
The kit worker's serve mode on Apple OpenCL (G2 dry runs, 16:21 UTC at load 5.2 to 5.7 on the generator-3 copies; 16:41 UTC at load 27 to 57, a loaded box, on the generator-4 packs with the merged packfile.h and the class=v4 token: a correctness run, no number taken) packs-ca3-v4/mx8-devnet-epoch0 and v4-devnet-epoch0: 1,024 of 1,024 found each (2.6 s), nonces 0 to 1,023, the sorted found file's sha256 equal to the verifier's through tools/ca3-v4/g2-recheck.sh --file (the coordinator's igneum-pow binary): the two digests above
tools/ci/bash-body-check.sh on the eight scripts 0 bash bodies, all parse
tools/ci/kit-path-check.sh on the eight scripts 6 of 6 kit paths checked before use (the reset and watts-app jobs read no kit)
tools/ci/playbook-quit-check.sh (rule 2) on the tree self-test fires on an api/cards POST and passes a read of api/state; every script here passes; the pre-rule playbooks on the dated allow list
tools/ci/prover-socket-check.sh, copied-sources-check.sh, no-conflict-markers.sh pass
The Windows PowerShell 5.1 parse (windows.yml parse job) the real parser runs only on the Windows runner and only over relay/playbooks and the other listed folders (not tools/); on the Mac the same rule's approximation (tools/ci/check-workflow-shell.mjs's drive-reference regex, applied to these five files by hand) fired once on "$dev: --stop-miners" and the line was fixed to ${dev}:; 0 findings after; brace, paren and bracket depth 0 by the bash-body-check tokenizer

Not tested on the Mac (the PC side)

The PowerShell itself never ran (no PowerShell on this Mac): the api/cards switch and its process-list confirmation, the wrapper sampler, Get-CimInstance command-line matching, the AMD helper's presence and its watts on the 9070 XT, the AMD OpenCL compiler's answer to each builtin (ds_bpermute, ds_swizzle, amd_perm, amd_bfe, the WMMA builtins), the driver's kernel cache (pass 2 of job 3), the 4070's nvidia-smi index against the app's --device, and the lengths above (estimates from the PC 2 runs and the 5 October 9070 XT rates). A failure of any of those prints a RESULT ... error= line and the SUMMARY carries failed or partial; the card restore in finally does not depend on them.