diff --git a/docs/analysis/counter-asic-4-research.md b/docs/analysis/counter-asic-4-research.md index 77b799c0d..bd5a892e4 100644 --- a/docs/analysis/counter-asic-4-research.md +++ b/docs/analysis/counter-asic-4-research.md @@ -373,7 +373,7 @@ The fix of 20.2a held for distinctness and then met the value-level requirement The crypto lane's finding: a product's low bits are biased (P(bit 0) = 1/4, measured exactly), the bias survives the odd stride multiplier, and the stride rotation places the biased bits at address bits R and up, inside the 28-bit item index unless R is 28 or more. The devnet era draws R = 29, which cuts them off, so 31 of 32 devnet-era programs read clean while 6 of 16 drawn-era programs (R from 3 to 22) show a site over 1.04x (2 over 1.2x, the worst 1.51x); under the 2 GiB genesis dataset (D = 29) R = 29 would show it too. The devnet's cleanliness is an accident of its era draw; the chain prevalence is the drawn-era figure; the price to a partial-store chip stays under 0.1 percent of a hash's reads per site, so no chip number moves. The requirement for any class this file proposes (the per-load shadow, the tile block, a re-weighted shadow): (1) a load whose source register's last writer is a product (`mul`, `mulhi`, `mad`) carries biased low bits into the address, and the acceptance must judge it at the VALUE level (the bit bias of the index at the site over the units), not by the index-distinctness ratio alone, which the duplicate-lane test above is; (2) the census of any candidate reads across the drawn eras split by R (3 to 22 against 28 to 31), as 20.2a now does for the per-load class, never the devnet era alone. The per-load class's dynamic rule covers distinctness, not bias; the value-level test is owed and is the same item for class v5's acceptance. Nothing in class v4 or v5 moves on this without main's word. -### 20.3 The card rows (PC 1, the RTX 5090 alone, the installed worker, 04:04 to 04:22 UTC, 8 October 2026; the hash lane's job `run-ca4-pc1-packs-5090-20261008-b`, 250 batches of 2^24 at one warp per block, nvidia-smi at 1 Hz with the three power fields agreeing, the 1,300 MHz lock through the helper) +### 20.3 The card rows (PC 1, the RTX 5090 alone, the installed worker, 04:04:33 to 04:22:54 UTC, 8 October 2026 (05:04 to 05:22 UK); the hash lane's job `run-ca4-pc1-packs-5090-20261008-b`, 250 batches of 2^24 at one warp per block, nvidia-smi at 1 Hz with the three power fields agreeing, the 1,300 MHz lock through the helper) Every pack's self-test PASS at both states (the cache, the dataset, the 96 vector lanes through the bound kernel), so the int8 tile's inline PTX compiles under NVRTC 12.8 on sm_120 and is bit-exact against the Rust verifier on the card, as the Metal reference was on the M5 Max; the 2^24 fingerprints are the same at both states and equal the Mac's where the Mac ran them (sh256x27 3d2e8245cc084d07, shl256x27_v2 ee5d7c71180e5ea7, mm128 270e4ae36b37e9a1, mm512 a1c1ff3148d775d1; mm1430 8e9b7066239d35d1 on CUDA, the Mac row not run). @@ -416,6 +416,6 @@ Energy per hash = the card's measured; the chip's = `E_mem + k x F`, `F` the mea | **tile block mm1430, 1,300 lock** | 2.24 | 0.564 | 3.46x / 2.98x / **2.17x** / 1.73x / 1.42x | 2.53x | the same | USD 100 M | measured | | the per-load shape (the fixed export) | 3.30 / 2.24 | 0.998 / 0.555 | as the class v4 shape within 5 percent | | ALU | USD 200 M if the sound form exists (16.2) | dead as drawn (20.2a-close) | -Reading: at `k = 1` the tile block and the ALU shadow give the same 2.1x to 2.2x, and the tile block's premium at the knee is 14 percent lower for it; the difference is the chip's reachable `k`. On the ALU shadow a fixed SIMD array can plausibly reach `k` 0.3 to 0.5 (the served pessimistic 3.4x to 3.5x); on the int8 tile the GPU's own tensor core is the dense array, measured at 0.048 to 0.091 pJ per MAC, and a chip at the same node cannot be 2x to 3x cheaper per MAC at the same voltage, so the chip's downside column collapses: the tile block's edge at the measured premium reads 2.2x at `k = 1` and under 1.8x at any `k` over 1.3, against the ALU shadow's 3.5x at `k` 0.3. That is the finding of the k hunt made measured on the card: **the tile block does not beat class v4's premium (it matches it at the same rate, 14 percent cheaper at the knee) and it beats class v4's chip edge at the pessimistic end (about 2.2x against 3.5x) and not at `k = 1`.** What stands against it as a class: the verifier (AVX2 0.047 microseconds per tile per unit, `mm1430` 5.8 ms alone and 10.14 ms with the sibling loaded on the box's core, a 0.14 ms miss at load 84 to 95; scalar 13x worse; NEON unwritten), the Apple tier (the Metal emulation at 1,024 tiles costs 35 percent of the M5 Max's rate and 4,096 costs 78 percent, so 11,440 is out of reach without an integer matrix path in Metal), the AMD WMMA layout (unverified; the OpenCL reference path emitted but unmeasured on the 9070 XT), and the compile-ahead (1.2 to 2.1 s per pack on the 5090, inside the budget). +The k column was priced at this point: section 15.2 set the tile count from the 4090's marginal 0.056 pJ per MAC so that the block would carry the ALU shadow's 0.654 microjoules (11,400 tiles per hash, R about 1,430), and the exported pack is 1,430 tiles per iteration, 11,440 per hash; the 5090 reads 0.091 pJ per MAC unlocked and 0.048 at the lock at that point, so the GPU-side cost of the chip-model-v3 5.11 column (2.1x at k = 1, 1.6x at k = 1.5) is now measured at the premium it was priced for (1.067 microjoules unlocked, 0.564 at the lock, against the ALU shadow's 1.10 and 0.652). The open side is Apple: the Metal emulation costs 35 percent of the M5 Max's rate at 1,024 tiles per hash and 78 percent at 4,096, so the 11,440-tile point is out of the Apple tier's reach without an integer matrix path in Metal; nothing served moves (the chip texts rest on the ALU shadow). Reading: at `k = 1` the tile block and the ALU shadow give the same 2.1x to 2.2x, and the tile block's premium at the knee is 14 percent lower for it; the difference is the chip's reachable `k`. On the ALU shadow a fixed SIMD array can plausibly reach `k` 0.3 to 0.5 (the served pessimistic 3.4x to 3.5x); on the int8 tile the GPU's own tensor core is the dense array, measured at 0.048 to 0.091 pJ per MAC, and a chip at the same node cannot be 2x to 3x cheaper per MAC at the same voltage, so the chip's downside column collapses: the tile block's edge at the measured premium reads 2.2x at `k = 1` and under 1.8x at any `k` over 1.3, against the ALU shadow's 3.5x at `k` 0.3. That is the finding of the k hunt made measured on the card: **the tile block does not beat class v4's premium (it matches it at the same rate, 14 percent cheaper at the knee) and it beats class v4's chip edge at the pessimistic end (about 2.2x against 3.5x) and not at `k = 1`.** What stands against it as a class: the verifier (AVX2 0.047 microseconds per tile per unit, `mm1430` 5.8 ms alone and 10.14 ms with the sibling loaded on the box's core, a 0.14 ms miss at load 84 to 95; scalar 13x worse; NEON unwritten), the Apple tier (the Metal emulation at 1,024 tiles costs 35 percent of the M5 Max's rate and 4,096 costs 78 percent, so 11,440 is out of reach without an integer matrix path in Metal), the AMD WMMA layout (unverified; the OpenCL reference path emitted but unmeasured on the 9070 XT), and the compile-ahead (1.2 to 2.1 s per pack on the 5090, inside the budget). **The first sentence, on measured rows (04:22 UTC, 8 October 2026): neither prototype beats class v4's premium; the tile block matches it at the same hash rate (146.7 W against 151.2 W unlocked, 71.5 against 82.8 W at the 1,300 lock) and beats class v4's chip edge only at the pessimistic end (about 2.2x against 3.5x, because a chip's int8 MAC cannot undercut the 5090's measured 0.048 to 0.091 pJ per MAC the way a fixed datapath undercuts its 6.4 to 10.8 pJ per ALU op), not at k = 1 (2.2x either way); the per-load placement is dead as a construction, and its energy rows read 13 to 14 W under the whole block for the same instructions.** The SM-sparse reading is still owed (three failed runs: the race switch, the race list, the rewrite's capture; the fourth exe checked through NVRTC card-free, queued).