diff --git a/docs/analysis/counter-asic-4-research.md b/docs/analysis/counter-asic-4-research.md index 6b3b691be..895e209eb 100644 --- a/docs/analysis/counter-asic-4-research.md +++ b/docs/analysis/counter-asic-4-research.md @@ -292,6 +292,10 @@ What the other rows say, in the identity's terms (the chip's cost per op from th **The one-sentence answer of section 15, corrected on measured rows: nothing on the 5090 reads k above 1; the int8 tile, the one block the public figures put near parity, is 1 to 4 pJ per MAC measured against a 5 nm array's claimed 0.04 to 0.4, so it is the worst lever of all, and the ALU shadow (k 0.3 to 0.8 on 6 to 11 pJ per op) stays the best forcing work the card has.** +### 15.1b The op-mix re-weight's hold, restated on measured rows (06:0x UTC, 8 October 2026) + +The re-opening condition set at 05:2x UTC (the 5090's shuffle and multiply at or under the add's picojoules per op) is not met and is now a measurement: shfl 55.8 pJ against the ARX op's 11.3 (4.9x), mul 13.9 (1.2x), mulhi 39.6 (3.5x). A shuffle-heavy weight table (the algorithm lane's shfl 14 and shfla 8 of 75 against class v4's shfl 8) therefore costs the GPU MORE per instruction at the same count, so the served 3.4x stands on a measured basis and the re-weight stays held. What the chip's `k` on shuffle-heavy work would have to be for 2.9x to be the honest pessimistic column, on the identity at the 1,300 knee (card 2.33 microjoules, `E_mem` 0.466): if the premium stayed at the ALU shadow's 0.652 microjoules (the same watts, fewer instructions), 2.9x needs `E_mem + k F = 0.80`, `k = 0.52`; if the instruction count stayed and the premium rose with the mix (about 29 percent of the counted ops at 4.9x the cost: the premium about 2.1x, 1.4 microjoules, the card 3.1), 2.9x needs `k = 0.42`. So the re-weight's pessimistic column is honest only if a chip pays 0.4 to 0.5 of the GPU's 55.8 pJ per shuffle, 22 to 28 pJ for a 32-lane crossbar move of a 32-bit word, which is above the wire figure of 15.1 (about 20 pJ across 1 mm, approximate) and unmeasured; and the GPU side of that bargain is a premium up to 2x higher per instruction. On measured rows the re-weight is not a candidate; it would re-price only against a measured chip crossbar. + ### 15.2 The tensor shadow, re-read on the identity The 6 October verdict on scheme B ("never as class v5 content") was right for the question it answered: at the 4090's free band (R = 512, 0.23 microjoules) the block forced too few joules and the verifier paid 26x the ALU shadow's cost per joule. The founder's question is a different one: not "is there a cheaper lever" but "is there a lever whose joules the chip cannot undercut". On that question the tensor tile is the best block on the card, because the GPU's marginal 0.056 pJ per MAC is within a factor of about 2 of what a 5 nm MAC array costs anyone (the test chip's 0.04 to 0.1 pJ, claimed), while the ALU shadow's 6.5 to 10.4 pJ per op is 10x to 50x what a fixed SIMD datapath costs. What it would take to make the tensor block carry the SAME premium as the ALU shadow (0.654 microjoules at the lock), on the 5090: