docs: Counter ASIC 2.0, the second set of chip-resistance layers and the plan to measure and decide them

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-05 19:40:33 +00:00
parent 9746391242
commit be4b29507b

View file

@ -0,0 +1,33 @@
# Counter ASIC 2.0
Internal name, set by Josh on 5 October 2026 (evening), for the second set of chip-resistance layers on the lottery hash. The first set is live: a random program every hour from a VDF seed, era parameters drawn every six months, a dataset that grows on a genesis schedule, instruction families that unlock by height, warp-unit CPU verification. The public claim stays as it is: a chip gains under 2x over a GPU, the model is published, the bounty stands. Nothing here is active; every layer is measured first and decided by Josh, and every one that goes in goes in before the public testnet, as a genesis rule or a reserved family.
Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGPU"): the hash is bound by dependent random 4-byte reads, AMD fetches a 64-byte line per read, so AMD sits at a seventh of the 5090 and the three-vendor claim is bit-exact but not fair. Read-only data can also be mirrored into a chip's SRAM (256 MB is about 45 mm2 at a leading node, approximate).
## The layers
| # | Layer | What it takes from a chip | GPU cost | Step |
|---|---|---|---|---|
| 1 | Load width 4, 16 or 64 bytes, every byte folded into the state | a fixed-width memory pipeline tuned to 4-byte reads | none if latency-bound stays | measure (running) |
| 2 | Load width drawn per program from the seed, era-fixed mix | any fixed-width pipeline; vendor balance averages across hours | none | measure (running) |
| 3 | Per-warp scratch in VRAM with read-modify-writes | the SRAM mirror of read-only data buys nothing for written scratch | measured share, expected 10 to 15% at 25% RMW (approximate) | measure (running), then soundness |
| 4 | Table layout drawn per era: item size, stride, interleave | a layout-tuned chip loses at the next draw | none | design, era draw exists |
| 5 | A second table sized to GPU cache, read beside the 1 GiB table | GPU-class SRAM and DRAM latency at once | small | next experiment |
| 6 | Cache growth on the genesis schedule | the SRAM mirror stays unaffordable | none | already in the design; confirm the schedule against SRAM density |
| 7 | Integer matrix ops in the program (INT8 x INT8 into INT32, exact) | matrix hardware at GPU scale | none on NVIDIA and AMD; Apple to check | reserved family, not at launch |
| 8 | Working-set size drawn per program | one memory design cannot fit every hour | none | folds into 4 and 5 |
Not added: divergent data-dependent branches (cost GPUs more than chips), anything floating point (bit-exactness across vendors).
## The plan, in order
1. Experiment on branch `readwidth` (running): variants 1, 2, 3 on the M5 Max, the RTX 5090 and the RX 9070 XT; hash rate, bit-exactness across Metal, CUDA and OpenCL, CPU verifier cost, bytes per hash, latency-bound share, the chip model re-run per variant. Table in the bench log, recommendation in `docs/plans/read-width.md`. Decision rule: the widest read that keeps every card latency-bound with margin on the 5090.
2. Decision 1 (Josh): the width and whether the per-program mix goes in. Consequences: new test vectors, new program id, spec sections on the hash and the litepaper Mining section rewritten, the soundness checks (uniformity, no out-of-bounds, fuzz) re-run on the new class.
3. Experiment 2: layer 5 (the cache-sized second table) on the three cards, same measurements, plus layer 6's schedule checked against SRAM density per node (cite the source).
4. Soundness project for layer 3 with the cryptographer role: what is written is uniform, no short-cut avoids the writes, the verifier's scratch simulation is exact; only then a vector.
5. Decision 2 (Josh): layers 3, 5 and the era draws of 4 and 8, as genesis rules; layer 7 as a named reserved family in the genesis reserve, unlockable by height or by 90% signal.
6. One generator change ships them all at once, before the public testnet, with the chip model and the before-and-after numbers published beside the litepaper claim.
## What stays true at every step
The latency bound is the property that matters, because DRAM latency is the same physics for everyone and a GPU already keeps thousands of loads in flight. Bandwidth is the property to avoid leaning on, because bandwidth per watt is what a custom memory chip buys (Ethash's chips got about 3x that way). Every layer above is checked against that rule by the latency-bound share in the measurements.