igneum/docs/plans/counter-asic-2.md
2026-10-05 19:40:33 +00:00

4.4 KiB

Counter ASIC 2.0

Internal name, set by the project lead on 5 October 2026 (evening), for the second set of chip-resistance layers on the lottery hash. The first set is live: a random program every hour from a VDF seed, era parameters drawn every six months, a dataset that grows on a genesis schedule, instruction families that unlock by height, warp-unit CPU verification. The public claim stays as it is: a chip gains under 2x over a GPU, the model is published, the bounty stands. Nothing here is active; every layer is measured first and decided by the project lead, and every one that goes in goes in before the public testnet, as a genesis rule or a reserved family.

Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGPU"): the hash is bound by dependent random 4-byte reads, AMD fetches a 64-byte line per read, so AMD sits at a seventh of the 5090 and the three-vendor claim is bit-exact but not fair. Read-only data can also be mirrored into a chip's SRAM (256 MB is about 45 mm2 at a leading node, approximate).

The layers

# Layer What it takes from a chip GPU cost Step
1 Load width 4, 16 or 64 bytes, every byte folded into the state a fixed-width memory pipeline tuned to 4-byte reads none if latency-bound stays measure (running)
2 Load width drawn per program from the seed, era-fixed mix any fixed-width pipeline; vendor balance averages across hours none measure (running)
3 Per-warp scratch in VRAM with read-modify-writes the SRAM mirror of read-only data buys nothing for written scratch measured share, expected 10 to 15% at 25% RMW (approximate) measure (running), then soundness
4 Table layout drawn per era: item size, stride, interleave a layout-tuned chip loses at the next draw none design, era draw exists
5 A second table sized to GPU cache, read beside the 1 GiB table GPU-class SRAM and DRAM latency at once small next experiment
6 Cache growth on the genesis schedule the SRAM mirror stays unaffordable none already in the design; confirm the schedule against SRAM density
7 Integer matrix ops in the program (INT8 x INT8 into INT32, exact) matrix hardware at GPU scale none on NVIDIA and AMD; Apple to check reserved family, not at launch
8 Working-set size drawn per program one memory design cannot fit every hour none folds into 4 and 5

Not added: divergent data-dependent branches (cost GPUs more than chips), anything floating point (bit-exactness across vendors).

The plan, in order

  1. Experiment on branch readwidth (running): variants 1, 2, 3 on the M5 Max, the RTX 5090 and the RX 9070 XT; hash rate, bit-exactness across Metal, CUDA and OpenCL, CPU verifier cost, bytes per hash, latency-bound share, the chip model re-run per variant. Table in the bench log, recommendation in docs/plans/read-width.md. Decision rule: the widest read that keeps every card latency-bound with margin on the 5090.
  2. Decision 1 (the project lead): the width and whether the per-program mix goes in. Consequences: new test vectors, new program id, spec sections on the hash and the litepaper Mining section rewritten, the soundness checks (uniformity, no out-of-bounds, fuzz) re-run on the new class.
  3. Experiment 2: layer 5 (the cache-sized second table) on the three cards, same measurements, plus layer 6's schedule checked against SRAM density per node (cite the source).
  4. Soundness project for layer 3 with the cryptographer role: what is written is uniform, no short-cut avoids the writes, the verifier's scratch simulation is exact; only then a vector.
  5. Decision 2 (the project lead): layers 3, 5 and the era draws of 4 and 8, as genesis rules; layer 7 as a named reserved family in the genesis reserve, unlockable by height or by 90% signal.
  6. One generator change ships them all at once, before the public testnet, with the chip model and the before-and-after numbers published beside the litepaper claim.

What stays true at every step

The latency bound is the property that matters, because DRAM latency is the same physics for everyone and a GPU already keeps thousands of loads in flight. Bandwidth is the property to avoid leaning on, because bandwidth per watt is what a custom memory chip buys (Ethash's chips got about 3x that way). Every layer above is checked against that rule by the latency-bound share in the measurements.