From 52d9a110bec60367aabccd88ff2e91cba0e3a774 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 19:40:33 +0000 Subject: [PATCH] docs: Counter ASIC 2.0, the second set of chip-resistance layers and the plan to measure and decide them Co-Authored-By: Claude Fable 5.1 --- docs/plans/counter-asic-2.md | 33 +++++++++++++++++++++++++++++++++ 1 file changed, 33 insertions(+) create mode 100644 docs/plans/counter-asic-2.md diff --git a/docs/plans/counter-asic-2.md b/docs/plans/counter-asic-2.md new file mode 100644 index 00000000..f6598d19 --- /dev/null +++ b/docs/plans/counter-asic-2.md @@ -0,0 +1,33 @@ +# Counter ASIC 2.0 + +Internal name, set by the project lead on 5 October 2026 (evening), for the second set of chip-resistance layers on the lottery hash. The first set is live: a random program every hour from a VDF seed, era parameters drawn every six months, a dataset that grows on a genesis schedule, instruction families that unlock by height, warp-unit CPU verification. The public claim stays as it is: a chip gains under 2x over a GPU, the model is published, the bounty stands. Nothing here is active; every layer is measured first and decided by the project lead, and every one that goes in goes in before the public testnet, as a genesis rule or a reserved family. + +Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGPU"): the hash is bound by dependent random 4-byte reads, AMD fetches a 64-byte line per read, so AMD sits at a seventh of the 5090 and the three-vendor claim is bit-exact but not fair. Read-only data can also be mirrored into a chip's SRAM (256 MB is about 45 mm2 at a leading node, approximate). + +## The layers + +| # | Layer | What it takes from a chip | GPU cost | Step | +|---|---|---|---|---| +| 1 | Load width 4, 16 or 64 bytes, every byte folded into the state | a fixed-width memory pipeline tuned to 4-byte reads | none if latency-bound stays | measure (running) | +| 2 | Load width drawn per program from the seed, era-fixed mix | any fixed-width pipeline; vendor balance averages across hours | none | measure (running) | +| 3 | Per-warp scratch in VRAM with read-modify-writes | the SRAM mirror of read-only data buys nothing for written scratch | measured share, expected 10 to 15% at 25% RMW (approximate) | measure (running), then soundness | +| 4 | Table layout drawn per era: item size, stride, interleave | a layout-tuned chip loses at the next draw | none | design, era draw exists | +| 5 | A second table sized to GPU cache, read beside the 1 GiB table | GPU-class SRAM and DRAM latency at once | small | next experiment | +| 6 | Cache growth on the genesis schedule | the SRAM mirror stays unaffordable | none | already in the design; confirm the schedule against SRAM density | +| 7 | Integer matrix ops in the program (INT8 x INT8 into INT32, exact) | matrix hardware at GPU scale | none on NVIDIA and AMD; Apple to check | reserved family, not at launch | +| 8 | Working-set size drawn per program | one memory design cannot fit every hour | none | folds into 4 and 5 | + +Not added: divergent data-dependent branches (cost GPUs more than chips), anything floating point (bit-exactness across vendors). + +## The plan, in order + +1. Experiment on branch `readwidth` (running): variants 1, 2, 3 on the M5 Max, the RTX 5090 and the RX 9070 XT; hash rate, bit-exactness across Metal, CUDA and OpenCL, CPU verifier cost, bytes per hash, latency-bound share, the chip model re-run per variant. Table in the bench log, recommendation in `docs/plans/read-width.md`. Decision rule: the widest read that keeps every card latency-bound with margin on the 5090. +2. Decision 1 (the project lead): the width and whether the per-program mix goes in. Consequences: new test vectors, new program id, spec sections on the hash and the litepaper Mining section rewritten, the soundness checks (uniformity, no out-of-bounds, fuzz) re-run on the new class. +3. Experiment 2: layer 5 (the cache-sized second table) on the three cards, same measurements, plus layer 6's schedule checked against SRAM density per node (cite the source). +4. Soundness project for layer 3 with the cryptographer role: what is written is uniform, no short-cut avoids the writes, the verifier's scratch simulation is exact; only then a vector. +5. Decision 2 (the project lead): layers 3, 5 and the era draws of 4 and 8, as genesis rules; layer 7 as a named reserved family in the genesis reserve, unlockable by height or by 90% signal. +6. One generator change ships them all at once, before the public testnet, with the chip model and the before-and-after numbers published beside the litepaper claim. + +## What stays true at every step + +The latency bound is the property that matters, because DRAM latency is the same physics for everyone and a GPU already keeps thousands of loads in flight. Bandwidth is the property to avoid leaning on, because bandwidth per watt is what a custom memory chip buys (Ethash's chips got about 3x that way). Every layer above is checked against that rule by the latency-bound share in the measurements.