13 KiB
Counter ASIC 3.0: the PC 1 AMD jobs (6 October 2026)
Four signed run jobs for PC 1 (machine ae432dc7: RTX 5090, RTX 4070, RX 9070 XT amd:gfx1201 on the eGPU, all on the installed app) that fill the OWED AMD rows of docs/plans/counter-asic-3-status.md sections 3, 5 and 7, plus the 4070 row of item 8. Prepared on branch ca3-pc1-amd; NOTHING here is published by this branch. The coordinator publishes on "go PC 1 AMD", one job at a time, in the order below, and reads each closing report before the next (node tools/jobs.mjs <id>, node tools/jobs.mjs watch <id>).
Rules every script keeps
| Rule | How |
|---|---|
| One job at a time on PC 1 | the publish order below; each job's closing report read before the next |
| The card alone only where the fixture needs it | job 1 and job 4 switch ONLY the card under test off through POST <app.url>api/cards (plain run jobs, not --stop-miners: the other two cards keep mining); jobs 2 and 3 run beside the miners and print `card_state=loaded |
Card keys from settings.json, never /api/state, BOTH forms |
every form is posted: the key as written (amd:1:gfx1201, nvidia:1:NVIDIA GeForce RTX 4070 or whatever the file holds), vendor:index, vendor:code; apply_cards (engine.rs) skips a key that matches no live card, so the extra forms are harmless and nothing is written for them |
| Quiet confirmed by the process list | the worker process carrying the card's --device <index> must be gone (Win32_Process command lines; job 4 also reads nvidia-smi -i <idx> --query-compute-apps); an unconfirmed card prints UNCONFIRMED and the rows say card_alone=False |
Restored in finally |
the card posted back enabled (only when settings.json had it enabled), the worker's return polled for 90 s and printed; the sampler ended by its own pid (never by name: the app runs its own copy of the AMD helper) |
| The installed app untouched | no /api/quit, /api/pause, /api/resume, no manifest, no settings.json write; the only writes are POST api/cards on the card under test |
| Copied sources re-stamped | nothing is built on the PC by these jobs (no cargo; the OpenCL kernels are compiled by the driver at run time); tools/ci/copied-sources-check.sh passes |
| Prover socket | not applicable (no prover host in these jobs) |
Every failure a RESULT ... error=<text> line, a SUMMARY {json} line at the end |
all four scripts |
If --stop-miners IS added at publish time, jobs 1 and 4 see the card already quiet, say so (card already quiet before the switch) and skip the switch: the scripts work either way.
The kit (one fetch job, before job 1)
tools/ca3-pc1-amd/make-kit.sh builds igneum-ca3-pc1-amd-kit.zip: bin/igneum-worker-opencl.exe (THIS tree's proto-opencl/host.c, the installed worker's source plus one build <ms> clBuildProgram line, cross-compiled with mingw as proto-cuda/nvrtc/build-windows.sh does), bin/family-probe-cl.exe (proto-opencl/family-probe.c), src/ (both sources and cl_dynamic.h for the sha256 record), packs/ (mx8-devnet-epoch0, sh256x13, sh256x27, sh256x53, sh256x88, sh64x52, dr736-genesis, dr736-devnet-epoch0, each with its kernel_bound.cl, program.json and vectors), SHA256SUMS.
Built 6 October 2026, 15:4x UTC (IGNEUM_REDIST=/Users/joshm/Projects/igneum-wt-ca2-mixer/proto-cuda/nvrtc/redist): zip sha256 a70fce5be672f33c3af08a5699488c09aef2848010f468ec99165fd4ab9ca61f, 884,381 bytes, 113 files; bin/igneum-worker-opencl.exe a572948e86c23317a80d9d7bb9e4bf70d5f7dfc40470974e2f08da1892016ac6, bin/family-probe-cl.exe d51a36bcae13ab36c40c2f347e3bdbe28f010499ccf84e423ecc04a9b4f151e3, src/host.c 5e23ac94…, src/family-probe.c c293be9d…. The zip is in the session scratchpad (pc1amd/igneum-ca3-pc1-amd-kit.zip), not in the repository; re-running make-kit.sh rebuilds it and prints a new sha256 (mingw output is not byte-reproducible), and the publish line takes the printed one. Every job prints RESULT kitfile <path> sha256 <hex> for the exes and sources it used, so the record closes on the PC side.
tools/ca3-pc1-amd/make-kit.sh "$TMPDIR/igneum-ca3-pc1-amd-kit.zip"
packaging/ota/publish-jobs.sh add --kind fetch --target ae432dc7 --id fetch-ca3-pc1-amd-20261006 \
--file "$TMPDIR/igneum-ca3-pc1-amd-kit.zip" --dir jobs --extract --title "CA3 PC 1 AMD kit" --expires-hours 48 --deploy
The fetch lands at <jobs>\fetch-ca3-pc1-amd-20261006\{bin,packs,src,SHA256SUMS} (jobrun.rs fetch_base: the job id under the jobs folder). The wiped-jobs-folder class: an app update clears the jobs folder, so if the 0.3.12 install lands on PC 1 between jobs, republish the fetch (each run job tests the kit before use and fails in seconds with kit missing if it is gone).
The jobs, in order
| # | Script | Publish (after the fetch; --id fixed so the read-back is one command) |
Length | Card | Rows it fills |
|---|---|---|---|---|---|
| 1 | pc1-amd-g1-shadow.ps1 |
packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-g1-shadow-20261006 --script tools/ca3-pc1-amd/pc1-amd-g1-shadow.ps1 --timeout-minutes 40 --title "CA3 PC 1: G1 + AMD ladder on the 9070 XT" --deploy |
about 15 min (7 benches of 60 to 90 dispatches of 2^24 at 0.9 to 2 s each, the card switch up to 150 s, the installed-worker cross-check 40 dispatches) | 9070 XT ALONE (its worker stopped; 5090 and 4070 keep mining) | G1 for mx8+sh256x27 on the AMD vendor (section 5: RESULT G1 pack=sh256x27 fingerprint=3d2e8245cc084d07 match=yes turns "AMD vendor NOT RUN" green; the other five packs give the same for the ladder); the 9070 XT column of the item 8 tables (section 3: RESULT LADDER pack=<p> ops=<N> mhs= watts= uj= gclk_mhz=), where the card leaves the latency bound by the 5 percent rule; item 1's AMD per-joule row from the control (RESULT LADDER pack=mx8-devnet-epoch0 ... uj=: section 3 "Consequences per tier (item 1)", 16 GB AMD row); section 7 rows "item 8: the 9070 XT rows", "item 1: the 9070 XT rows" |
| 2 | pc1-amd-family.ps1 |
packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-family-20261006 --script tools/ca3-pc1-amd/pc1-amd-family.ps1 --timeout-minutes 15 --title "CA3 PC 1: family step costs on the 9070 XT" --deploy |
about 3 min (3 runs x up to 3 AMD devices, 26 variants each, builds dominate) | beside the miners (ratios; every row says card_state=) |
the "RX 9070 XT" column of the item 6 table (section 3: RESULT FAMILYBEST name=<f> ratio=<r> exact= path= per family, RESULT FAMILY ... variant= for the form each took: shfla_bperm = ds_bpermute_b32, the one number that could move R3; shflx_swz / shflx_bperm; perm_amd = v_perm_b32 through amd_perm or perm_c emulated; bfe_amd; dot4_amd = the 5 October sudot4 row re-measured; mm8_gfx12 / mm8_gfx11 = the WMMA builtins, exact=unverified); section 7 row "item 6: the 9070 XT step costs" |
| 3 | pc1-amd-derive.ps1 |
packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-amd-derive-20261006 --script tools/ca3-pc1-amd/pc1-amd-derive.ps1 --timeout-minutes 15 --title "CA3 PC 1: dr736 on the 9070 XT" --deploy |
about 4 min (3 packs x a 60-dispatch pass and a 3-dispatch pass) | beside the miners (build, compile and bit-exact rows; the rate as a ratio to mx8 in the same job) | the item 2 table's 9070 XT cells (section 3: RESULT DERIVE pack=dr736-genesis build_ms= build_ms_cached= build_ms_over_mx8= dataset_ms= fingerprint= match= mhs= ratio_to_mx8=): the OpenCL compile cost of the derivation program against mx8's (the NVRTC +1.1 s finding on AMD), the 1 GiB daily build against mx8's 72 to 77 ms (5 October), bit-exactness on the AMD vendor; section 7 row "item 2: the 9070 XT rows" |
| 4 (optional, confirmed by main) | pc1-4070-shadow.ps1 |
packaging/ota/publish-jobs.sh add --kind run --target ae432dc7 --id run-ca3-pc1-4070-shadow-20261006 --script tools/ca3-pc1-amd/pc1-4070-shadow.ps1 --timeout-minutes 30 --title "CA3 PC 1: item 8 ladder on the RTX 4070" --deploy |
about 11 min (7 benches of 100 to 200 dispatches at an expected 0.4 to 1 s each, approximate: no 4070 row exists yet; NVRTC per pack; the card switch) | 4070 ALONE (its CUDA worker stopped; 5090 and 9070 XT keep mining) | the "small NVIDIA card binds near 100,000" row of item 8 (section 3, consequences per tier at N = 100,000; section 7 "the 4060-class rows"): RESULT LADDER pack=<p> ops=<N> mhs= watts= uj= sm_mhz= with nvidia-smi at 1 Hz on that index; G1 fingerprints on a second NVIDIA card |
Read back: node tools/jobs.mjs run-ca3-pc1-amd-g1-shadow-20261006 (and --all for the 5-minute progress uploads). The run_id on the intake is job-<id>-ae432dc7.
The AMD watts readback on PC 1 (what exists)
igneum-gpu-telemetry.exe (proto-opencl/gpu-telemetry.c on master, shipped by packaging/windows/make-payload.sh since 0.3.10; the Ember Tune playbook found it in the install folder or under a jobs\amd-kit-*\kit\ folder) prints one line per AMD card per sample: amd <n> bus <pci> kind discrete name "AMD Radeon RX 9070 XT" watts <W> temp_c <C> fan_rpm <r> fan_pct <p> mclk_mhz <m> gclk_mhz <g> util_pct <u> source adlx, then end <ms> N card(s), at -l 1 once a second with fflush per sample. The watts field is ADLX GPUPower with GPUTotalBoardPower as the fallback (gpu-telemetry.c lines 120 to 121 on master); on 5 October it read the 9070 XT at 198.9 W at 17.73 MH/s (release-0.3.10.md, job tele-measure-1), the figure the app's MH per watt line uses. So a board-watts readback EXISTS and job 1 uses it: a second copy of the helper at 1 Hz through a wrapper that stamps each line (UTC, ms) so the samples window onto each bench (after its first 12 s, before its last 2 s, the PC 2 shape), ended by the wrapper's pid in finally. The job prints RESULT tele helper=<path> sha256=<hex> watts_source=adlx after one sample; if the helper is missing or prints watts - for the 9070 XT the rows carry watts=owed uj=owed and the SUMMARY says watts: owed, never a guess. Whether ADLX GPUPower on RDNA 4 is total board power or ASIC power is not verified here: the row says adlx and the number is what the app itself reports.
What the Mac tested (6 October 2026, 15:27 to 15:28 UTC, with-lock.sh run, load average 3.9 at the start)
| Check | Result |
|---|---|
proto-opencl/family-probe.c on Apple OpenCL (cc -framework OpenCL, M5 Max) |
every C-form and emulated row bit-exact on both 32-lane groups in 3 repetitions; ratios to the alu chain beside the Metal probe's (rotr 1.14 against Metal 1.13, shl 0.89 / 0.85, shr 0.89 / 0.86, bfe 0.82 / 0.77, andn 0.84 / 0.75, popc 1.00 / 0.87, clz 1.11 / 1.01, sel 0.82 / 0.76, perm emulated 1.52 / 1.13, shflx and shfla through __local 1.96 and 2.04 (Metal native 0.86 and 1.91), dot4 emulated signed 4.66 / 4.73); the Apple dot(char4,char4) row exact=no as the 5 October finding said; every AMD builtin skipped (build=skipped), the khr, intel and amd_media_ops2 forms build=failed with the log line, the run went on; FAMILYBEST picked the exact rows |
proto-opencl/host.c with the build line, --bench-pack on all eight kit packs on Apple OpenCL |
self-test PASS (96 of 96 lanes) and the 2^24 fingerprint at base 0 equal to the recorded one on every pack: mx8-devnet-epoch0 90f794dd556f7a3b, sh256x13 59ac286fe2a5a9ef, sh256x27 3d2e8245cc084d07, sh256x53 4f824b15cf2b124a, sh256x88 0572522e39a94d8a, sh64x52 9dd010f79d8ca9f4, dr736-genesis 50e3eaa779da4f1e, dr736-devnet-epoch0 9553f6d5c667205a; build line: mx8 70 ms, shadow packs 70 to 88 ms, dr736 337 to 356 ms (+270 ms on Apple OpenCL: item 2's compile cost shows on this harness too) |
| mingw cross-builds of both exes | clean (-Wall -Wextra), 116 KB and 35 KB |
tools/ci/bash-body-check.sh on the four scripts |
0 bash bodies, all parse |
tools/ci/kit-path-check.sh on the four scripts |
4 of 4 kit paths checked before use |
tools/ci/prover-socket-check.sh, copied-sources-check.sh, no-conflict-markers.sh |
pass |
The Windows PowerShell 5.1 parse (windows.yml parse job) |
the real parser runs only on the Windows runner and only over relay/playbooks and the other listed folders (not tools/); on the Mac the same rule's approximation (tools/ci/check-workflow-shell.mjs's drive-reference regex, applied to these four files by hand) fired once on "$dev: --stop-miners" and the line was fixed to ${dev}:; 0 findings after; brace, paren and bracket depth 0 by the bash-body-check tokenizer |
Not tested on the Mac (the PC side)
The PowerShell itself never ran (no PowerShell on this Mac): the api/cards switch and its process-list confirmation, the wrapper sampler, Get-CimInstance command-line matching, the AMD helper's presence and its watts on the 9070 XT, the AMD OpenCL compiler's answer to each builtin (ds_bpermute, ds_swizzle, amd_perm, amd_bfe, the WMMA builtins), the driver's kernel cache (pass 2 of job 3), the 4070's nvidia-smi index against the app's --device, and the lengths above (estimates from the PC 2 runs and the 5 October 9070 XT rates). A failure of any of those prints a RESULT ... error= line and the SUMMARY carries failed or partial; the card restore in finally does not depend on them.