igneum/docs/plans/counter-asic-2.md

6.6 KiB

Counter ASIC 2.0

Internal name, set by Josh on 5 October 2026 (evening), for the second set of chip-resistance layers on the lottery hash. The first set is live: a random program every hour from a VDF seed, era parameters drawn every six months, a dataset that grows on a genesis schedule, instruction families that unlock by height, warp-unit CPU verification. The public claim stays as it is: a chip gains under 2x over a GPU, the model is published, the bounty stands. Nothing here is active; every layer is measured first and decided by Josh, and every one that goes in goes in before the public testnet, as a genesis rule or a reserved family.

Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGPU"): the hash is bound by dependent random 4-byte reads, AMD fetches a 64-byte line per read, so AMD sits at a seventh of the 5090 and the three-vendor claim is bit-exact but not fair. Read-only data can also be mirrored into a chip's SRAM (256 MB is about 45 mm2 at a leading node, approximate).

The layers

# Layer What it takes from a chip GPU cost Step
1 Load width 4, 16 or 64 bytes, every byte folded into the state a fixed-width memory pipeline tuned to 4-byte reads none if latency-bound stays measure (running)
2 Load width drawn per program from the seed, era-fixed mix any fixed-width pipeline; vendor balance averages across hours none measure (running)
3 Per-warp scratch in VRAM with read-modify-writes the SRAM mirror of read-only data buys nothing for written scratch measured share, expected 10 to 15% at 25% RMW (approximate) measure (running), then soundness
4 Table layout drawn per era: item size, stride, interleave a layout-tuned chip loses at the next draw none design, era draw exists
5 A second table sized to GPU cache, read beside the 1 GiB table GPU-class SRAM and DRAM latency at once small next experiment
6 Cache growth on the genesis schedule the SRAM mirror stays unaffordable none already in the design; confirm the schedule against SRAM density
7 Integer matrix ops in the program (INT8 x INT8 into INT32, exact) matrix hardware at GPU scale none on NVIDIA and AMD; Apple to check reserved family, not at launch
8 Working-set size drawn per program one memory design cannot fit every hour none folds into 4 and 5
9 Epoch length as a signalled era parameter: base 3,600 DAA s, ladder 600 to 7,200, set by 90% signal at a day boundary, lead and T_epoch fixed a per-program bitstream (an FPGA with a hard datapath): at 600 s nothing it compiles ever runs (42 to 160 min per compile, PRflow FPT 2019; hours on large parts, Aldec) compile-ahead 1 s per epoch on the 5090, 0.5 s on the M5 Max (38 s with the race on); one CPU core 600 / epoch_len busy on the VDF reserve-only tonight: docs/plans/epoch-length.md

Decided 5 October 2026 (night), under Josh's delegation for the devnet (docs/plans/counter-asic-2-rollout.md section 6)

# Decision The number that decided it
1 keep v2's 128 x 4 B w16 passes the rule but closes nothing (5090 139.8 against 136.1 MH/s, 9070 XT 17.90 against 18.15); w64 makes the 5090 bandwidth-bound (share 0.58, 37% of stream)
2 out spread across six programs 5.5 to 22.3% per card, over the 5% rule
3 out (scratch share 0); the construct is sound and its tests stay the on-die-cache recompute chip stays at 2.4x at every share under the 6 GB cap
4 and 8 IN: the era draw of stride, interleave and the working-set window (width pinned at 4 B) six-era spread 1.3% on the 5090, 3.2% on the 9070 XT, 0.8% on the M5 Max; bit-exact on all three vendors
5 OUT of v3 (a measured option for 3.0) the added form costs the 5090 13 to 16% and the 9070 XT 16 to 20% (g 0.87 / 0.85 / 0.84 and 0.84 / 0.81 / 0.80 at 32 / 64 / 96 MiB against the 0.97 rule): no card keeps even 32 MiB resident while the dataset streams; the replaced form helps the on-die-cache chip (x1.33 at k = 4)
6 option C: the cache doubles when the dataset doubles the mirror is 128 mm^2 and $46 at N5 by shipped density; the cache's job is to stay above GPU L2
7 reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal native on all three vendors as a tile; dot4 emulation 1.6x on Apple
M16 mixer x8 (x4 measured beside it) into v3 the only measured lever that moves the named chip: x8 0.31x bare, 0.92x with a 3x fixed-function factor (x4: 0.61x, 1.84x); verifier 2.1x v2 per warp (2.79 ms on a loaded core, about 1.3 ms quiet); the daily build unmoved on every card (latency-bound)
9 reserve-only, no change to the devnet's hour the floor 600 s from the slowest compile-ahead (38 s)

Not added: divergent data-dependent branches (cost GPUs more than chips), anything floating point (bit-exactness across vendors).

The plan, in order

  1. Experiment on branch readwidth (running): variants 1, 2, 3 on the M5 Max, the RTX 5090 and the RX 9070 XT; hash rate, bit-exactness across Metal, CUDA and OpenCL, CPU verifier cost, bytes per hash, latency-bound share, the chip model re-run per variant. Table in the bench log, recommendation in docs/plans/read-width.md. Decision rule: the widest read that keeps every card latency-bound with margin on the 5090.
  2. Decision 1 (Josh): the width and whether the per-program mix goes in. Consequences: new test vectors, new program id, spec sections on the hash and the litepaper Mining section rewritten, the soundness checks (uniformity, no out-of-bounds, fuzz) re-run on the new class.
  3. Experiment 2: layer 5 (the cache-sized second table) on the three cards, same measurements, plus layer 6's schedule checked against SRAM density per node (cite the source).
  4. Soundness project for layer 3 with the cryptographer role: what is written is uniform, no short-cut avoids the writes, the verifier's scratch simulation is exact; only then a vector.
  5. Decision 2 (Josh): layers 3, 5 and the era draws of 4 and 8, as genesis rules; layer 7 as a named reserved family in the genesis reserve, unlockable by height or by 90% signal.
  6. One generator change ships them all at once, before the public testnet, with the chip model and the before-and-after numbers published beside the litepaper claim.

What stays true at every step

The latency bound is the property that matters, because DRAM latency is the same physics for everyone and a GPU already keeps thousands of loads in flight. Bandwidth is the property to avoid leaning on, because bandwidth per watt is what a custom memory chip buys (Ethash's chips got about 3x that way). Every layer above is checked against that rule by the latency-bound share in the measurements.