Counter ASIC 3.0 status: item 6's RX 9070 XT column (the card alone; every family 0.75x to 1.30x except perm and mm8; R3 unmoved)
This commit is contained in:
parent
7704ed1181
commit
1fbdf1008e
1 changed files with 3 additions and 1 deletions
|
|
@ -117,6 +117,8 @@ Reading: the AMD card never leaves the latency bound on this ladder (its ALU bud
|
|||
| dot4 (comparison) | 1.60 unsigned / 4.73 signed, emulated | 1.16 | 1.06 (5 October) | |
|
||||
| mm8 (comparison, R8) | OWED (Metal 4 matmul2d) | 2.43 (`mma.m8n8k16.u8`, bit-exact) | OWED | the licensable block |
|
||||
|
||||
The RX 9070 XT column (PC 1 job run-ca3-pc1-amd-family-20261006-e, 17:31:57 to 17:34:47Z, exit 0; the card ALONE: every gfx1201 entry posted off, its worker pid 19788 gone, three runs at 17:32:11 / 15 / 19Z with `card_state=alone`, every entry restored with its own flag and identities and the card mining again 20 s later under pid 13040; the gfx1036 and old-platform columns run after the restore beside the miners; AMD OpenCL 3683.0, gfx1201, 32 CUs; the range below is over the three runs' best-of-3): alu 1,074 to 1,186 G steps/s (ratio 1.00); rotr 0.99 to 1.13; shflx (xor shuffle, ds_bpermute) 0.82 to 1.21; shl 1.00 to 1.02; shr 0.91 to 1.02; bfe 1.00 to 1.09 NATIVE (amd_bfe); andn 0.91 to 1.01; perm 1.73 to 1.93 EMULATED (no byte-permute path in AMD's OpenCL C); popc 0.91 to 1.01; clz 1.19 to 1.30; sel 0.90 to 1.01; shfla (lane + delta, ds_bpermute) 0.75 to 0.84 NATIVE; dot4 1.00 to 1.11 NATIVE (sudot4); mm8 1.68 to 1.83 NATIVE (the WMMA iu8 builtin reaches gfx12; exact UNVERIFIED: the fragment layout is not in any source at hand, so the CPU reference was not attempted rather than guessed). The AMD column of item 6 is CLOSED. Reading: on RDNA 4 every 32-bit datapath family sits within 0.75x to 1.30x of the alu chain, the shuffles cheaper than it (an LDS operation overlapping the dependent chain), so the ds_bpermute number that could have moved R3 leaves the order as proposed; the two dear rows on AMD are the same two as on Apple and NVIDIA, the byte permute and mm8.
|
||||
|
||||
Reading: against the live rotr every candidate is 0.95x to 1.23x on NVIDIA; on Apple only perm (emulated) and shfla cost more than rotr; the 8x emulation bound of 1.13.2 holds everywhere by 4x or more; the 5 percent hash-rate bound at W_new = 4 is argued from the step cost (under 1 percent of ALU time on a read-bound hash), not measured, since no reserve family is live. Proposed order R1 perm, R2 popc and clz, R3 shfla, R4 bfe, R5 shifts, R6 sel, R7 andn, R8 mm8, W_new = 4 each, family n at era n (mm8's era-4 unlock kept as a named exception or moved to era 8: the project lead's call); the full proposed 1.13.2 text with edge vectors per family is `docs/plans/counter-asic-3-reserve.md` section 6. Consequence per tier: an Apple miner pays the emulated perm at 1.13x a step and shfla at 1.91x, under 1 percent of its hash rate at W_new = 4 (argued); an NVIDIA miner pays nothing measurable; an AMD miner's row is owed and its `ds_bpermute_b32` cost is the one number that could move R3; a chip pays a barrel shifter, a byte crossbar, a popcount tree and a 32-lane crossbar per lane, which is the point.
|
||||
|
||||
### Item 6, the Mac rows (interim, ca3-reserve 192a683; the 5090 job waits on PC 2)
|
||||
|
|
@ -252,6 +254,6 @@ Either replaces "under 2x" on the site and in the litepaper once the project lea
|
|||
| 4 + 7 | the live observer restart with both hooks, and the first live detector and vendor-share rows | needs a push to master (the project lead's word) |
|
||||
| 2 | the 2019-class core (O-1.14), which decides dr736 against dr368; the once-a-day NVRTC module for the item function (required before any activation); cryptanalysis of random ARX programs; the loaded-iGPU tier's build with the day program; the 5090 absolutes re-run with the card quiet (ratios stand) | unmeasured; unimplemented |
|
||||
| 8 | the 9070 XT and 4060-class rows (where a small card binds); the 5090 clock rows (an elevated job); the 5090 at a 575 W cap (model only); the Mac package watts (IOReport gives GPU + DRAM, Ember's 38 W approximate); the chip side's k, lane area and 28 nm scaling; the Metal fuzz, edge and stats runs and gates G2, G4 to G6 on the class; the acceptance rule's reading of the block | PC 1 not released; no elevated job; the gates are the next step on the project lead's word |
|
||||
| 6 | mm8 as a chain on Apple (Metal 4 matmul2d; the Mac's Swift toolchain has no tensor API); the 5 percent rule per family with the family live (argued only); the RDNA ISA guides unread (mnemonics from LLVM's tables) | |
|
||||
| 6 | mm8 as a chain on Apple (Metal 4 matmul2d; the Mac's Swift toolchain has no tensor API); mm8's exactness on AMD (the gfx12 WMMA fragment layout; an empirical layout probe would settle it in one job); the 5 percent rule per family with the family live (argued only); the RDNA ISA guides unread (mnemonics from LLVM's tables). The 9070 XT step costs themselves are CLOSED (run e) | |
|
||||
| 1 | the DRAM energy figures are streaming figures applied to random 32-byte reads; the GDDR7 burst and HBM3 tFAW are behind the JEDEC paywall; no chip has been built or torn down | |
|
||||
| all | every AMD RDNA 4 number in this run | PC 1 is the project lead's desk today. 16:0x UTC: the RX 9070 XT is back on PC 1 (amd:gfx1201, 16,304 MiB) beside the 5090 and an RTX 4070; main has asked for the PC 1 jobs to be PREPARED, not published: (1) G1 AMD plus item 8's ladder on the 9070 XT (about 15 min, the card alone), (2) item 6's family step costs on AMD (about 3 min), (3) item 2's dr736 build and compile on AMD (about 3 min), (4) optional: the 4070 ladder (about 10 min); branch ca3-pc1-amd (1f33cc6, e054ed7, merged): the four scripts pass every CI check, the kit zip sha256 a70fce5b... verified, the AMD watts readback is igneum-gpu-telemetry.exe (ADLX board watts); the commands and the row map in tools/ca3-pc1-amd/README.md; published one at a time on "go PC 1 AMD" after the 0.3.13 update and the Ember table run, about 33 min of PC 1 in all |
|
||||
|
|
|
|||
Loading…
Reference in a new issue