igneum/docs/plans/counter-asic-2.md
igneum-labs 7eed16a29a Pre-public scrub, the text pass (7 October 2026, 19:5x UK): no founder name, personal login, earlier business or personal address in any tracked text file, and a gate check that keeps it so
The sweep (main's item 1): 199 tracked text files, 783 lines. The founder's full name, first name and possessive become "the founder" (sentence starts capitalised); the lowercase operating-system user name in WSL paths and commands becomes <user>; the second owner login becomes "the second owner login"; the three earlier businesses and the two other brands become "the other business", "the earlier entity", "the earlier business" and "another brand"; the Chrome profile rule names the igneum.network profile, not the profile's label. The standing commit login igneum-labs is not a founder term here: the fresh-repository step renames it in the history (docs/plans/history-rewrite.md, tools/repo/fresh-repo.sh).

The patterns never appear in plain text in the tree (a plaintext list would be the hit): tools/ci/founder-strings.b64 (perl regex, tab, a sample per row) is read by tools/ci/founder-strings-check.sh (every tracked text file, perl, known-failed first: the self-test plants each row's sample in a fixture and the hit must name the file), by tools/community/discord-hooks.mjs (the guard's founder and business rows; the test takes its fixtures from the samples) and by tools/repo/fresh-repo.sh (the business names of the rewrite rules). site/forbidden-strings.txt carries the same patterns as b64: lines, decoded case-insensitive by site/scrub.mjs and tools/ci/launch-gates-check.mjs (whose fixture now plants an encoded made-up name). The check runs in the gate's tree checks on every merge.

Not in this commit, by main's word: the 105 commit messages and 40 personal-identity commits that need the history rewrite (listed, not run), and the secrets found by gitleaks over the history (reported with owners).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 18:39:50 +00:00

48 lines
6.7 KiB
Markdown

# Counter ASIC 2.0
Internal name, set by the founder on 5 October 2026 (evening), for the second set of chip-resistance layers on the lottery hash. The first set is live: a random program every hour from a VDF seed, era parameters drawn every six months, a dataset that grows on a genesis schedule, instruction families that unlock by height, warp-unit CPU verification. The public claim stays as it is: a chip gains under 2x over a GPU, the model is published, the bounty stands. Nothing here is active; every layer is measured first and decided by the founder, and every one that goes in goes in before the public testnet, as a genesis rule or a reserved family.
Trigger: the 9070 XT measurement of 5 October (bench-log "the 9070 XT on the eGPU"): the hash is bound by dependent random 4-byte reads, AMD fetches a 64-byte line per read, so AMD sits at a seventh of the 5090 and the three-vendor claim is bit-exact but not fair. Read-only data can also be mirrored into a chip's SRAM (256 MB is about 45 mm2 at a leading node, approximate).
## The layers
| # | Layer | What it takes from a chip | GPU cost | Step |
|---|---|---|---|---|
| 1 | Load width 4, 16 or 64 bytes, every byte folded into the state | a fixed-width memory pipeline tuned to 4-byte reads | none if latency-bound stays | measure (running) |
| 2 | Load width drawn per program from the seed, era-fixed mix | any fixed-width pipeline; vendor balance averages across hours | none | measure (running) |
| 3 | Per-warp scratch in VRAM with read-modify-writes | the SRAM mirror of read-only data buys nothing for written scratch | measured share, expected 10 to 15% at 25% RMW (approximate) | measure (running), then soundness |
| 4 | Table layout drawn per era: item size, stride, interleave | a layout-tuned chip loses at the next draw | none | design, era draw exists |
| 5 | A second table sized to GPU cache, read beside the 1 GiB table | GPU-class SRAM and DRAM latency at once | small | next experiment |
| 6 | Cache growth on the genesis schedule | the SRAM mirror stays unaffordable | none | already in the design; confirm the schedule against SRAM density |
| 7 | Integer matrix ops in the program (INT8 x INT8 into INT32, exact) | matrix hardware at GPU scale | none on NVIDIA and AMD; Apple to check | reserved family, not at launch |
| 8 | Working-set size drawn per program | one memory design cannot fit every hour | none | folds into 4 and 5 |
| 9 | Epoch length as a signalled era parameter: base 3,600 DAA s, ladder 600 to 7,200, set by 90% signal at a day boundary, lead and `T_epoch` fixed | a per-program bitstream (an FPGA with a hard datapath): at 600 s nothing it compiles ever runs (42 to 160 min per compile, PRflow FPT 2019; hours on large parts, Aldec) | compile-ahead 1 s per epoch on the 5090, 0.5 s on the M5 Max (38 s with the race on); one CPU core `600 / epoch_len` busy on the VDF | reserve-only tonight: `docs/plans/epoch-length.md` |
## Decided 5 October 2026 (night), under the founder's delegation for the devnet (`docs/plans/counter-asic-2-rollout.md` section 6)
| # | Decision | The number that decided it |
|---|---|---|
| 1 | keep v2's 128 x 4 B | w16 passes the rule but closes nothing (5090 139.8 against 136.1 MH/s, 9070 XT 17.90 against 18.15); w64 makes the 5090 bandwidth-bound (share 0.58, 37% of stream) |
| 2 | out | spread across six programs 5.5 to 22.3% per card, over the 5% rule |
| 3 | out (scratch share 0); the construct is sound and its tests stay | the on-die-cache recompute chip stays at 2.4x at every share under the 6 GB cap |
| 4 and 8 | IN: the era draw of stride, interleave and the working-set window (width pinned at 4 B) | six-era spread 1.3% on the 5090, 3.2% on the 9070 XT, 0.8% on the M5 Max; bit-exact on all three vendors |
| 5 | OUT of v3 (a measured option for 3.0) | the added form costs the 5090 13 to 16% and the 9070 XT 16 to 20% (g 0.87 / 0.85 / 0.84 and 0.84 / 0.81 / 0.80 at 32 / 64 / 96 MiB against the 0.97 rule): no card keeps even 32 MiB resident while the dataset streams; the replaced form helps the on-die-cache chip (x1.33 at k = 4) |
| 6 | option C: the cache doubles when the dataset doubles | the mirror is 128 mm^2 and $46 at N5 by shipped density; the cache's job is to stay above GPU L2 |
| 7 | reserve R1 = mm8, unsigned, W_new 4, unlock era 4 or 90% signal | native on all three vendors as a tile; dot4 emulation 1.6x on Apple |
| M16 mixer x8 (x4 measured beside it) | into v3 | the only measured lever that moves the named chip: x8 0.31x bare, 0.92x with a 3x fixed-function factor (x4: 0.61x, 1.84x); verifier 2.1x v2 per warp (2.79 ms on a loaded core, about 1.3 ms quiet); the daily build unmoved on every card (latency-bound) |
| 9 | reserve-only, no change to the devnet's hour | the floor 600 s from the slowest compile-ahead (38 s) |
Not added: divergent data-dependent branches (cost GPUs more than chips), anything floating point (bit-exactness across vendors).
## The plan, in order
1. Experiment on branch `readwidth` (running): variants 1, 2, 3 on the M5 Max, the RTX 5090 and the RX 9070 XT; hash rate, bit-exactness across Metal, CUDA and OpenCL, CPU verifier cost, bytes per hash, latency-bound share, the chip model re-run per variant. Table in the bench log, recommendation in `docs/plans/read-width.md`. Decision rule: the widest read that keeps every card latency-bound with margin on the 5090.
2. Decision 1 (the founder): the width and whether the per-program mix goes in. Consequences: new test vectors, new program id, spec sections on the hash and the litepaper Mining section rewritten, the soundness checks (uniformity, no out-of-bounds, fuzz) re-run on the new class.
3. Experiment 2: layer 5 (the cache-sized second table) on the three cards, same measurements, plus layer 6's schedule checked against SRAM density per node (cite the source).
4. Soundness project for layer 3 with the cryptographer role: what is written is uniform, no short-cut avoids the writes, the verifier's scratch simulation is exact; only then a vector.
5. Decision 2 (the founder): layers 3, 5 and the era draws of 4 and 8, as genesis rules; layer 7 as a named reserved family in the genesis reserve, unlockable by height or by 90% signal.
6. One generator change ships them all at once, before the public testnet, with the chip model and the before-and-after numbers published beside the litepaper claim.
## What stays true at every step
The latency bound is the property that matters, because DRAM latency is the same physics for everyone and a GPU already keeps thousands of loads in flight. Bandwidth is the property to avoid leaning on, because bandwidth per watt is what a custom memory chip buys (Ethash's chips got about 3x that way). Every layer above is checked against that rule by the latency-bound share in the measurements.