docs/analysis/counter-asic-4-research.md (excluded from the public export): edge = (E_card + F) / (E_mem + k F); at zero premium the stored-dataset chip keeps E_card / E_mem, 3.6x on a 5090 locked at 1,400 MHz (measured card, modelled chip), about 2x at that card's idle-plus-memory bound, 1.7x on the M5 Max; class v5 at zero shadow leaves it at 5.1x to 9.1x; the shadow buys 2.1x at k = 1 for the measured 88 W (81.8 W at the knee) and turns against us below 1.8 pJ per op; per-card classes, the refresh as a cost, proof of useful work, memory shaping and the tensor block priced out with their arithmetic; ten designs ranked with the chip edge, the 5090 and 4070 premiums, the verifier cost, agent hours and what breaks each; the recommendation in five sentences; consequences per tier.
proto-cuda/nvrtc/worker.cpp: opt-in SM-sparse variants sp<N> and sp<N>-w<W> (a persistent grid of N blocks of W warps over the dispatch's nonces; the bound kernel rewritten at compile time with exact anchors into a unit function plus a wrapper of the kernel's name; made on demand by name, never in the default race; refused on variant-5 packs and with minBlocks). Gate: mingw cross-compile on igneum-build-1 under lease pool (class measure, label "ca4 research"), exit 0 with -Wall -Wextra, 22:48 UTC; exe sha256 ee8d0e70dd101f125f42c0c7cf07481a794ee18a1317acf68561b37ea18d72be; unrun on a card (the hash lane's PC 1 job run-ca4-pc1-ca4sparse-5090-20261007 is its first run).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Documents-only replay of b18684240 (counter-asic-4) for the box mirror master; left on the branch: proto-cuda/nvrtc/worker.cpp tools/ci/export-exclude.txt