Counter ASIC 2.0 layer 7: RTX 5090 and RX 9070 XT dot4 numbers from PC 1 (job run-dot4-20261005): dp4a 7.45 T/s on the 5090 via inline PTX, v_dot4_i32_iu8 0.66 T/s on the 9070 XT via the clang builtin in Adrenalin OpenCL C, signed emulation 7.1x / 1.46x / 4.7x the ALU step on NVIDIA / AMD / Apple; no PC platform lists cl_khr_integer_dot_product
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
da42719b81
commit
4d0f7ecf9a
2 changed files with 32 additions and 5 deletions
|
|
@ -134,11 +134,26 @@ its own and a variant the platform cannot compile prints a "build failed" row.
|
|||
|---|---|---|---|---|---|---|---|
|
||||
| Apple M5 Max, Metal | 5 October 2026 20:0x UTC, `with-lock.sh measure ./dot4-probe` (swiftc -O), GPU start-to-end time | 879.8 (4.882 ms) | 188.2 (22.82 ms) | 548.2 (7.834 ms) | none exists | signed 4.7x, unsigned 1.6x | yes, all three kernels |
|
||||
| Apple M5 Max, Apple OpenCL 1.2 | same, `with-lock.sh measure ./dot4-probe-cl --device 0`, event time | 871.5 (4.928 ms) | 188.4 (22.80 ms) | not in this probe | `cl_khr_integer_dot_product` not listed; the kernel using `dot(char4, char4)` compiled anyway and ran at 846 G/s but MISMATCHED the CPU reference on every lane checked (Apple's `dot` on char4 is not an integer dot; the extension macro must gate it) | signed 4.6x | alu and dot4e yes; dot4_khr NO |
|
||||
| RTX 5090 (PC 1 or 2), NVIDIA OpenCL inline PTX, and CUDA `__dp4a` | owed: PC job prepared (`relay/playbooks/dot4-probe.ps1`, exe `dot4-probe-cl.exe` cross-compiled, sha256 in the bench-log entry); not published until the coordinator's "go PC" | | | | | | |
|
||||
| RX 9070 XT (PC 1), AMD OpenCL `__builtin_amdgcn_sudot4` | owed, same job (the job runs every listed device: the 5090 on NVIDIA's OpenCL, the 9070 XT, the gfx1036) | | | | | | |
|
||||
| RTX 5090 (PC 1, ae432dc7), NVIDIA OpenCL 3.0 CUDA, driver 617.14 | 5 October 2026 20:29 UTC, job `run-dot4-20261005` (`relay/playbooks/dot4-probe.ps1`, both mining cards switched off in the app first, restored after; `node tools/jobs.mjs run-dot4-20261005`), event time | 8,753.5 (0.491 ms) | 1,239.1 (3.466 ms) | not in the OpenCL probe | 7,453.6 (0.576 ms) via inline PTX `dp4a.s32.s32` | emulation 7.1x; the intrinsic 1.17x (the chain is one dp4a plus 3 ops against 5 ops), so the emulation costs 6.0x the instruction | yes, all three |
|
||||
| RX 9070 XT (PC 1, gfx1201, eGPU), AMD OpenCL 2.0 AMD-APP 3683.0 (PAL,LC) | same job, same time | 701.4 (6.124 ms) | 480.8 (8.932 ms) | not in the OpenCL probe | 664.3 (6.465 ms) via `__builtin_amdgcn_sudot4(true, a, true, b, acc, false)`: the Adrenalin OpenCL C compiler accepts the clang builtin and emits `v_dot4_i32_iu8` | emulation 1.46x; the intrinsic 1.06x, so the emulation costs 1.38x the instruction | yes, all three; the older 3652.0 platform entry for the same card gave 696.2 / 501.7 / 683.6 |
|
||||
| Ryzen 9800X3D gfx1036 (PC 1, integrated RDNA 2, 2 CUs) | same job | 40.6 (105.9 ms) | 15.8 (272.3 ms) | | `sudot4` does not build: "needs target feature dot8-insts" (RDNA 2 has `dot1-insts`' `v_dot4_i32_i8`, the `sdot4` builtin, which the probe did not try) | emulation 2.6x | alu and dot4e yes |
|
||||
| Every PC device | `cl_khr_integer_dot_product` not listed on NVIDIA (OpenCL 3.0) or AMD (2.0); the pragma draws "unknown OpenCL extension" on both and the `dot(char4, char4)` kernel does not build | | | | | | |
|
||||
|
||||
Reading across the three cards. Per dot4 at the hardware rate: the 5090 does 7.45 T dot4/s (one `dp4a` per step, 0.85
|
||||
of its ALU-chain step rate), the 9070 XT 0.66 T (0.95 of its ALU-chain rate), the M5 Max 0.55 T at best (the unsigned
|
||||
emulation; no instruction). On the ALU chain the 5090 is 12.5x the 9070 XT and 10x the M5 Max; on hardware dot4 it is
|
||||
11.2x the 9070 XT, so the family does not widen the AMD gap, and 13.6x the M5 Max, so Apple's emulation widens its gap
|
||||
by 1.4x on this op (approximate: one probe shape, the ratios of best-of-3 numbers). The signed emulation is where the
|
||||
vendors differ most: 7.1x the ALU step on NVIDIA, 4.7x on Apple, 1.46x on AMD (AMD's compiler and byte-permute
|
||||
hardware make the four sign-extended products nearly free; the NVIDIA OpenCL compiler does not pattern-match the
|
||||
emulation into `dp4a`, which the 6x gap between `dot4e` and `dot4_nv` shows). None of this is a hash-rate number: the
|
||||
hash is bound by 128 dependent DRAM reads, and a family at W_new = 4 adds about 21 of these ops per hash per lane
|
||||
against about 1.2 microseconds of memory latency per hash per lane (approximate), so the per-op penalties above turn
|
||||
into hash-rate losses well under 5% on every card, to be measured with the family live.
|
||||
|
||||
Reading of the Mac numbers. The ALU chain's 880 G steps/s on the M5 Max is the integer baseline (5 ops per step
|
||||
counted, so about 4.4 T int ops/s, approximate). A signed dot4 emulated as `int4(as_type<char4>(a))` products costs
|
||||
counted, so about 4.4 T int ops/s, approximate; the 5090's 8,754 G steps/s is about 43.8 T, against the whitepaper's
|
||||
104.8 peak INT32 TOPS which counts a multiply-add as two). A signed dot4 emulated as `int4(as_type<char4>(a))` products costs
|
||||
4.7 of those steps; the unsigned form 1.6 steps. The 3x gap between the two is the sign extension (Metal lowers the
|
||||
unsigned byte extraction to masks that fold into the multiplies, approximate reading of the result, not of the
|
||||
compiled code). Both are far under the 8x bound of section 3, and the hash spends its time on DRAM reads, so a per-lane
|
||||
|
|
@ -149,8 +164,9 @@ Apple OpenCL `dot(char4, char4)` mismatch is the kind of thing the edge vectors
|
|||
|
||||
| Item | State |
|
||||
|---|---|
|
||||
| dp4a throughput on the RTX 5090 (inline PTX through NVIDIA OpenCL; `__dp4a` through CUDA if the PC has nvcc) | owed, PC job prepared, waiting for "go PC" |
|
||||
| `sudot4` on the 9070 XT through Adrenalin's OpenCL C, and whether `cl_khr_integer_dot_product` appears on the 3683.0 platform | owed, same job; the probe prints both |
|
||||
| dp4a throughput on the RTX 5090 through NVIDIA OpenCL inline PTX | measured (section 4); the CUDA `__dp4a` form (`proto-cuda/dot4-probe.cu`) is unrun (no nvcc job tonight) and is a cross-check, not a gap |
|
||||
| `sudot4` on the 9070 XT through Adrenalin's OpenCL C; `cl_khr_integer_dot_product` on the 3683.0 platform | measured: the builtin works and emits the instruction; the extension is not listed and the `dot(char4, char4)` kernel does not build |
|
||||
| `sdot4` (`dot1-insts`) on RDNA 2 (gfx1036) | not tried; the probe only carries `sudot4` |
|
||||
| Metal 4 `matmul2d` uchar x uchar into int on the M5 Max: wrap or saturate at the int32 edge, native or emulated, throughput | owed (a second Metal probe; the API needs a tensor set-up the dot4 probe does not have) |
|
||||
| Metal Feature Set Tables: which Apple GPU families run int8 matmul2d natively | not read tonight |
|
||||
| RDNA 3 and RDNA 4 ISA guides: the instruction text itself (names taken from LLVM and GPUOpen) | AMD's CDN refused the downloads tonight |
|
||||
|
|
|
|||
|
|
@ -1539,3 +1539,14 @@ Apple M5 Max, macOS 26, branch `ca2-analysis` (base `readwidth` 4badcee). The pr
|
|||
| Apple OpenCL 1.2, `dot4_khr` | `acc + dot(as_char4(x), as_char4(y))` under `#pragma OPENCL EXTENSION cl_khr_integer_dot_product : enable` | 5.076 | 846.2 | 1,239 | NO: mismatched the CPU reference on every lane checked in all 3 repetitions |
|
||||
|
||||
Reading: on this GPU a signed-byte dot4 costs 4.7 ALU-chain steps and an unsigned-byte one 1.6; Metal has no dp4a and no integer simdgroup matrix (MSL 4.1 sections 2.4 and 6.9), so these are the honest Apple costs of a per-lane dot4 family, and an unsigned definition is 3x cheaper for Apple at no cost to NVIDIA or AMD (both carry the unsigned form, PTX `dp4a.u32.u32`, AMD `v_dot4_u32_u8`). Apple's OpenCL does not list `cl_khr_integer_dot_product`; its `dot` on `char4` compiled anyway and returned something other than the integer dot (the mismatch), which is why a family's conformance vectors must gate every vendor path on the feature macro, not on "it compiled". Not run here: NVIDIA and AMD. The PC job is prepared and not published (coordinator's rule): `relay/playbooks/dot4-probe.ps1` with `dot4-probe-cl.exe` (proto-opencl/dot4-probe.c cross-compiled with mingw as `x86_64-w64-mingw32-gcc -std=c99 -O2 -static -DIGNEUM_CL_DYNAMIC -DCL_TARGET_OPENCL_VERSION=120 -I proto-cuda/nvrtc/redist/include`, sha256 `5adaeb1aceb03dc41135baabe0b53f1ed5fac891a5b3c3849645b03efe4416f4`, 161,863 bytes); it runs the scalar, KHR, AMD `__builtin_amdgcn_sudot4` and NVIDIA inline-PTX `dp4a` variants on every OpenCL GPU of the machine with the mining cards switched off through `/api/cards` and restored after. The CUDA form (`proto-cuda/dot4-probe.cu`, `__dp4a`) needs nvcc on the PC and is the cross-check.
|
||||
|
||||
**PC 1, 5 October 2026 20:29 UTC, the same probe on the RTX 5090 and the RX 9070 XT** (machine ae432dc7, Windows 11; fetch job `fetch-dot4-20261005` placed `dot4-probe-cl.exe` sha256 `5adaeb1a…6416f4`, run job `run-dot4-20261005` ran `relay/playbooks/dot4-probe.ps1`: the app's `nvidia:0` and `amd:1:gfx1201` cards switched off through `POST api/cards`, the probe run on every OpenCL device, the cards restored with their settings (identities 8 and 2, power cap 80% and none); `node tools/jobs.mjs run-dot4-20261005`; 101 s wall, every kernel under 10 ms; device event time, best of 3, same lanes and steps as the Mac rows):
|
||||
|
||||
| Device, platform | alu, G steps/s (ms) | dot4e signed emulation, G dot4/s (ms) | dot4 instruction, G dot4/s (ms) | `cl_khr_integer_dot_product` | ok |
|
||||
|---|---|---|---|---|---|
|
||||
| RTX 5090, NVIDIA OpenCL 3.0 CUDA, driver 617.14 | 8,753.5 (0.491) | 1,239.1 (3.466), 7.1x the ALU step | 7,453.6 (0.576) via inline PTX `dp4a.s32.s32`, 1.17x the ALU step | not listed; the `dot(char4,char4)` kernel does not build | yes |
|
||||
| RX 9070 XT (gfx1201), AMD-APP 3683.0 (PAL,LC), OpenCL 2.0 | 701.4 (6.124) | 480.8 (8.932), 1.46x | 664.3 (6.465) via `__builtin_amdgcn_sudot4`, 1.06x | not listed; same | yes |
|
||||
| RX 9070 XT, the older 3652.0 platform entry (dup) | 696.2 (6.169) | 501.7 (8.561) | 683.6 (6.283) | not listed | yes |
|
||||
| gfx1036 (integrated RDNA 2, 2 CUs), 3683.0 | 40.6 (105.9) | 15.8 (272.3), 2.6x | `sudot4` does not build: "needs target feature dot8-insts" | not listed | alu and dot4e yes |
|
||||
|
||||
Reading: one `dp4a` on the 5090 costs about one ALU-chain step (7.45 T dot4/s, 0.85 of the chain's 8.75 T steps/s); one `v_dot4_i32_iu8` on the 9070 XT the same (0.66 T, 0.95 of its chain). Emulating the signed dot4 costs 6.0x the instruction on NVIDIA (the OpenCL compiler does not fold the four sign-extended products into `dp4a`) and 1.38x on AMD. Vendor ratios: the 5090 is 12.5x the 9070 XT on the ALU chain and 11.2x on hardware dot4; against the M5 Max's best (unsigned emulation, 0.55 T) it is 10x on the chain and 13.6x on dot4. The hash itself is bound by DRAM reads, so these per-op numbers bound a family's cost and are not hash rates (`docs/analysis/int8-matrix-family.md` section 4). Adrenalin's OpenCL C accepts the clang builtin and emits the instruction on RDNA 4 (the third-party RDNA 3 report of the same route is now confirmed on this card); no PC platform lists the Khronos integer-dot extension. The 5090 SM clock read 2,505 MHz before and after (nvidia-smi; 2,850 MHz while mining in the telemetry entry), so the card was idle for the probe.
|
||||
|
|
|
|||
Loading…
Reference in a new issue