48 lines
6.7 KiB
Markdown
48 lines
6.7 KiB
Markdown
# Counter ASIC 2.0
|
|
|
|
Internal name, set by the project lead on 5 October 2026 (evening), for the second set of chip-resistance layers on the lottery hash. The first set is live: a random program every hour from a VDF seed, era parameters drawn every six months, a dataset that grows on a genesis schedule, instruction families that unlock by height, warp-unit CPU verification. The public claim stays as it is: a chip gains under 2x over a GPU, the model is published, the bounty stands. Nothing here is active; every layer is measured first and decided by the project lead, and every one that goes in goes in before the public testnet, as a genesis rule or a reserved family.
|
|
|
|
Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGPU"): the hash is bound by dependent random 4-byte reads, AMD fetches a 64-byte line per read, so AMD sits at a seventh of the 5090 and the three-vendor claim is bit-exact but not fair. Read-only data can also be mirrored into a chip's SRAM (256 MB is about 45 mm2 at a leading node, approximate).
|
|
|
|
## The layers
|
|
|
|
| # | Layer | What it takes from a chip | GPU cost | Step |
|
|
|---|---|---|---|---|
|
|
| 1 | Load width 4, 16 or 64 bytes, every byte folded into the state | a fixed-width memory pipeline tuned to 4-byte reads | none if latency-bound stays | measure (running) |
|
|
| 2 | Load width drawn per program from the seed, era-fixed mix | any fixed-width pipeline; vendor balance averages across hours | none | measure (running) |
|
|
| 3 | Per-warp scratch in VRAM with read-modify-writes | the SRAM mirror of read-only data buys nothing for written scratch | measured share, expected 10 to 15% at 25% RMW (approximate) | measure (running), then soundness |
|
|
| 4 | Table layout drawn per era: item size, stride, interleave | a layout-tuned chip loses at the next draw | none | design, era draw exists |
|
|
| 5 | A second table sized to GPU cache, read beside the 1 GiB table | GPU-class SRAM and DRAM latency at once | small | next experiment |
|
|
| 6 | Cache growth on the genesis schedule | the SRAM mirror stays unaffordable | none | already in the design; confirm the schedule against SRAM density |
|
|
| 7 | Integer matrix ops in the program (INT8 x INT8 into INT32, exact) | matrix hardware at GPU scale | none on NVIDIA and AMD; Apple to check | reserved family, not at launch |
|
|
| 8 | Working-set size drawn per program | one memory design cannot fit every hour | none | folds into 4 and 5 |
|
|
| 9 | Epoch length as a signalled era parameter: base 3,600 DAA s, ladder 600 to 7,200, set by 90% signal at a day boundary, lead and `T_epoch` fixed | a per-program bitstream (an FPGA with a hard datapath): at 600 s nothing it compiles ever runs (42 to 160 min per compile, PRflow FPT 2019; hours on large parts, Aldec) | compile-ahead 1 s per epoch on the 5090, 0.5 s on the M5 Max (38 s with the race on); one CPU core `600 / epoch_len` busy on the VDF | reserve-only tonight: `docs/plans/epoch-length.md` |
|
|
|
|
## Decided 5 October 2026 (night), under the project lead's delegation for the devnet (`docs/plans/counter-asic-2-rollout.md` section 6)
|
|
|
|
| # | Decision | The number that decided it |
|
|
|---|---|---|
|
|
| 1 | keep v2's 128 x 4 B | w16 passes the rule but closes nothing (5090 139.8 against 136.1 MH/s, 9070 XT 17.90 against 18.15); w64 makes the 5090 bandwidth-bound (share 0.58, 37% of stream) |
|
|
| 2 | out | spread across six programs 5.5 to 22.3% per card, over the 5% rule |
|
|
| 3 | out (scratch share 0); the construct is sound and its tests stay | the on-die-cache recompute chip stays at 2.4x at every share under the 6 GB cap |
|
|
| 4 and 8 | IN: the era draw of stride, interleave and the working-set window (width pinned at 4 B) | six-era spread 1.3% on the 5090, 3.2% on the 9070 XT, 0.8% on the M5 Max; bit-exact on all three vendors |
|
|
| 5 | OUT of v3 (a measured option for 3.0) | the added form costs the 5090 13 to 16% and the 9070 XT 16 to 20% (g 0.87 / 0.85 / 0.84 and 0.84 / 0.81 / 0.80 at 32 / 64 / 96 MiB against the 0.97 rule): no card keeps even 32 MiB resident while the dataset streams; the replaced form helps the on-die-cache chip (x1.33 at k = 4) |
|
|
| 6 | option C: the cache doubles when the dataset doubles | the mirror is 128 mm^2 and $46 at N5 by shipped density; the cache's job is to stay above GPU L2 |
|
|
| 7 | reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal | native on all three vendors as a tile; dot4 emulation 1.6x on Apple |
|
|
| M16 mixer x8 (x4 measured beside it) | into v3 | the only measured lever that moves the named chip: x8 0.31x bare, 0.92x with a 3x fixed-function factor (x4: 0.61x, 1.84x); verifier 2.1x v2 per warp (2.79 ms on a loaded core, about 1.3 ms quiet); the daily build unmoved on every card (latency-bound) |
|
|
| 9 | reserve-only, no change to the devnet's hour | the floor 600 s from the slowest compile-ahead (38 s) |
|
|
|
|
Not added: divergent data-dependent branches (cost GPUs more than chips), anything floating point (bit-exactness across vendors).
|
|
|
|
## The plan, in order
|
|
|
|
1. Experiment on branch `readwidth` (running): variants 1, 2, 3 on the M5 Max, the RTX 5090 and the RX 9070 XT; hash rate, bit-exactness across Metal, CUDA and OpenCL, CPU verifier cost, bytes per hash, latency-bound share, the chip model re-run per variant. Table in the bench log, recommendation in `docs/plans/read-width.md`. Decision rule: the widest read that keeps every card latency-bound with margin on the 5090.
|
|
2. Decision 1 (the project lead): the width and whether the per-program mix goes in. Consequences: new test vectors, new program id, spec sections on the hash and the litepaper Mining section rewritten, the soundness checks (uniformity, no out-of-bounds, fuzz) re-run on the new class.
|
|
3. Experiment 2: layer 5 (the cache-sized second table) on the three cards, same measurements, plus layer 6's schedule checked against SRAM density per node (cite the source).
|
|
4. Soundness project for layer 3 with the cryptographer role: what is written is uniform, no short-cut avoids the writes, the verifier's scratch simulation is exact; only then a vector.
|
|
5. Decision 2 (the project lead): layers 3, 5 and the era draws of 4 and 8, as genesis rules; layer 7 as a named reserved family in the genesis reserve, unlockable by height or by 90% signal.
|
|
6. One generator change ships them all at once, before the public testnet, with the chip model and the before-and-after numbers published beside the litepaper claim.
|
|
|
|
## What stays true at every step
|
|
|
|
The latency bound is the property that matters, because DRAM latency is the same physics for everyone and a GPU already keeps thousands of loads in flight. Bandwidth is the property to avoid leaning on, because bandwidth per watt is what a custom memory chip buys (Ethash's chips got about 3x that way). Every layer above is checked against that rule by the latency-bound share in the measurements.
|