Counter ASIC 4.0 research: 20.3c the hot-table rows as a record and the call (the replaced form is a per-watt gift to every miner and a larger one to the stored-dataset chip: a loss for resistance; ldcs a dead lever)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-08 08:13:00 +01:00
parent 591cba7e97
commit fb61ed4b64

View file

@ -466,6 +466,12 @@ The two readings the design asked for. (1) The shape itself: sp170-w32 against b
What this closes in the ranking: rank 2 (the SM-sparse miner kernel) is dead; the premium-free floor of section 0 stays the card's idle plus its memory system plus whatever `E_wait` is at the knee, and the only miner-side lever on `E_card` is the operating point (rank 1). The 5090's four runs of this night (three failed, one read) also gave four repeats of the knee pass: the class v4 premium 133.9 to 151.2 W unlocked and 82.8 to 90.5 W at 1,300 MHz, with every one of the three power fields agreeing within 0.2 W.
### 20.3c The hot-table rows (PC 1, the 5090, the hash lane's job `run-ca4-pc1-hot-ldcs-5090-20261008b`, 36 rows all PASS, bench-log "8 October 2026, the hot-table packs on the RTX 5090", master c09dfee4 at 07:12 UTC): a record, and the call
(1) `--variant ldcs` (streaming dataset loads) equals base on every pack in both states within 0.1 MH/s and 1 W: the cache-policy hint is a dead lever, as 15.1a's L2 row already said. (2) At the 1,300 lock the replaced-form hot packs lose 2 percent of rate where mx8 loses 7.5 (hot32k4 147.4 to 144.4 MH/s, hot64k8 164.2 to 160.9, mx8 137.7 to 127.3) while every pack drops a third of its watts (319 to 215 W): the hot packs are latency-bound on the table. (3) Per watt at the lock: hot64k8 0.734 MH/W, hot32k4 0.671, hot64k4 0.645, hot96k4 0.635, hot64k2 0.630, the mx8 control 0.602 (equal to the efficiency pass's 0.60, so the rows are comparable), the added-form packs 0.52 to 0.54.
The call, on the identity: a hash 22 percent cheaper on the GPU (hot64k8, 1.36 against 1.66 microjoules at the lock) is a LOSS for resistance in the replaced form, not a gain, because the saving sits in the one path where the chip is cheaper still. The GPU saves 0.30 microjoules by moving 64 of its 128 reads from DRAM (8.7 nJ per read measured, 15.1a) to L2 (1.4 nJ): 4.7 nJ per replaced read; the f = 1 chip moves the same 64 reads from GDDR7 (2.0 nJ, modelled) to on-die SRAM (0.2 to 0.5 nJ, approximate), saving about 1.6 nJ per read, 0.10 microjoules of its 0.466, and, because its rate is bound by the memory's activate ceiling (166 MH/s on the 5090's 16 devices), halving its DRAM reads per hash doubles its rate on the same memory and halves its capex per MH/s. The edge per joule at the lock: mx8 1.66 / 0.466 = 3.6x; hot64k8 about 1.36 / 0.36 = 3.8x (modelled chip, approximate); per dollar the chip's 2.8 USD per MH/s becomes about 1.5. This is the 5 October finding (the replaced form "helps the on-die-cache chip, x1.33 at k = 4", counter-asic-2.md decision 5) re-read against the stored-dataset chip with measured GPU energies: the replaced form is a per-watt gift to every miner and a larger one to the chip. The added form ("a" packs, 0.52 to 0.54 MH/W) costs the GPU rate for no chip cost, as chip-model-v3 section 3 said. Nothing moves.
### 20.4 Into the chip model (the rows measured; the chip side modelled: f = 1 GDDR7 0.466 microjoules per hash, one HBM3 stack 0.321)
Energy per hash = the card's measured; the chip's = `E_mem + k x F`, `F` the measured premium per hash, `k` the chip's cost per forced operation over the 5090's at the same state.