Merge branch 'ca3-pc1-amd' into ca3-coord

This commit is contained in:
igneum-labs 2026-10-06 17:05:54 +00:00
commit 5a0894f03f

View file

@ -59,7 +59,7 @@ G1 on the AMD vendor 7 of 7 (every fingerprint equal to the Mac's and the 5090's
Every probe ran ordinal 0, PC 1's integrated gfx1036 (one CU: alu 40.59 G steps/s against the Mac's 881 and the 5090's 7,941; the 9070 XT, 32 CUs, should read about 1,000 to 2,000): the script's device-list parse ran a second `-match` after the capturing one, which overwrote `$Matches`, so every device read `idx 0` and an empty name, and the probe took a bare ordinal. The rows are a gfx1036 (RDNA 2 iGPU) column, kept: alu 1.00, shifts 1.15, bfe 1.15 native, andn 1.16, popc 1.55, clz 1.75, sel 1.81, shfla and shflx 1.89 through `ds_bpermute`, perm 2.43 emulated, dot4 2.57 emulated, mm8 none. The 9070 XT column stays owed to run b. Fixed (commit after c38dfef): the probe host takes `--device-name gfx1201` (the match on the newest AMD platform by driver version, the kit worker's dedup rule; `RESULT device_choice name= index= platform= driver= cus=`; `RESULT error` and exit 2 when no name matches, a bare ordinal never the default), prints the device's name, CUs and platform on every `RESULT FAMILY` and `FAMILYBEST` line; the script runs the 9070 XT by name first and the older-platform duplicate and the gfx1036 by explicit ordinal after it as their own labelled columns, and fails the job when the 9070 XT gave no row. Mac check (18:0x UTC, `with-lock.sh run`, load 9.9): `--device-name M5` chose the M5 Max with 40 CUs and ran bit-exact; `--device-name gfx1201` on the Mac printed `RESULT error no OpenCL GPU device whose name holds "gfx1201"` and exit 2 (the known-bad case); no device given printed `RESULT error no device chosen`. Every probe ran ordinal 0, PC 1's integrated gfx1036 (one CU: alu 40.59 G steps/s against the Mac's 881 and the 5090's 7,941; the 9070 XT, 32 CUs, should read about 1,000 to 2,000): the script's device-list parse ran a second `-match` after the capturing one, which overwrote `$Matches`, so every device read `idx 0` and an empty name, and the probe took a bare ordinal. The rows are a gfx1036 (RDNA 2 iGPU) column, kept: alu 1.00, shifts 1.15, bfe 1.15 native, andn 1.16, popc 1.55, clz 1.75, sel 1.81, shfla and shflx 1.89 through `ds_bpermute`, perm 2.43 emulated, dot4 2.57 emulated, mm8 none. The 9070 XT column stays owed to run b. Fixed (commit after c38dfef): the probe host takes `--device-name gfx1201` (the match on the newest AMD platform by driver version, the kit worker's dedup rule; `RESULT device_choice name= index= platform= driver= cus=`; `RESULT error` and exit 2 when no name matches, a bare ordinal never the default), prints the device's name, CUs and platform on every `RESULT FAMILY` and `FAMILYBEST` line; the script runs the 9070 XT by name first and the older-platform duplicate and the gfx1036 by explicit ordinal after it as their own labelled columns, and fails the job when the 9070 XT gave no row. Mac check (18:0x UTC, `with-lock.sh run`, load 9.9): `--device-name M5` chose the M5 Max with 40 CUs and ran bit-exact; `--device-name gfx1201` on the Mac printed `RESULT error no OpenCL GPU device whose name holds "gfx1201"` and exit 2 (the known-bad case); no device given printed `RESULT error no device chosen`.
The finding on AMD's OpenCL C from run a (the same compiler serves the 9070 XT): `amd_perm` (cl_amd_media_ops2) does not compile, `__builtin_amdgcn_sudot4` does not compile under OpenCL C (the 5 October dot4 row used it through the same probe shape, so the LC compiler's builtin set differs between that build and this; the row is owed a second look), and the WMMA builtins (`__builtin_amdgcn_wmma_i32_16x16x16_iu8_w32_gfx12` and `_w32`) do not compile; `__builtin_amdgcn_ds_bpermute` and `ds_swizzle` do. So on AMD the byte-permute, dot4 and mm8 families are the generic C sequence until a path exists (inline assembly through the LC compiler, or a HIP probe), and the reserve order's R1 (perm) and R8 (mm8) rows read emulated on AMD; the shuffle rows (R3) are native through `ds_bpermute`. Consequence for the AMD tier: a reserve family whose only native path is a builtin AMD's OpenCL C does not expose costs an AMD miner the emulated sequence on every step until the worker gains an ISA path; the 9070 XT ratios of run b say how much. What run a's build failures say, by their log lines (the gfx1036 ran, not the 9070 XT, so each refusal is classed by its cause): `amd_perm` fails with "use of undeclared identifier" although the device lists `cl_amd_media_ops2` and `amd_bfe` from the same extension builds, so the byte permute has NO OpenCL C path on AMD's LC compiler and that holds for the 9070 XT too (R1 reads emulated on AMD, 2.43x on the iGPU, the 9070 XT ratio from run b); `__builtin_amdgcn_sudot4` fails with "needs target feature dot8-insts", a target-feature refusal of RDNA 2 (the 5 October log said the same), and the same builtin DID build and run bit-exact on the 9070 XT on 5 October, so the dot4 row on gfx1201 is decided by run b, not by run a; the two WMMA builtins fail with "needs target feature wmma…", again the RDNA 2 target, and run b says whether the LC compiler reaches them on gfx1201 (if it does the mm8 row gets a step cost with exact=unverified; if not, mm8 has no OpenCL C path and stays OWED on AMD); `cl_khr_subgroup_shuffle` and `cl_intel_subgroups` are not offered (warnings, then the undeclared function), while `ds_bpermute` and `ds_swizzle` build, so the shuffle families (R3) are native on AMD through the clang builtin. Consequence for the AMD tier: a byte-permute family costs an AMD miner the emulated sequence on every step until the worker gains an ISA path (inline assembly through the LC compiler, or a HIP build of the probe); the dot4 and mm8 answers for the 9070 XT, and every ratio, are run b's.
## The AMD watts readback on PC 1 (what exists) ## The AMD watts readback on PC 1 (what exists)