7.3 KiB
Counter ASIC: the public description in four levels
the project lead, 5 October 2026 (night): "not an information overload". Four levels; the layer names appear only from level 3 down, next to their numbers. Numbers come from the final table of docs/plans/counter-asic-2-status.md; a number still owed is marked [owed: ...], never guessed. the project lead's copy law throughout.
Level 1: one sentence (site hero, litepaper abstract)
Built for graphics cards. A custom chip gains under 2x, and the model is public. (the project lead's decision of 6 October 2026, 17:35 UTC, ledger M1: no device bounty; the claim is backed by the paid independent cryptanalysis (CA 3.0 item 3, four paid reviews) and the public benchmark with M22's metrics; an optional USD 50,000 cryptanalysis prize may follow later, escrowed before it is named.)
Level 2: one site card, one short litepaper section
Three ideas, no layer names, no widths, no SRAM.
The hash rewrites itself. A new program every hour, drawn from the chain. Its memory pattern changes with it. The rules change on a schedule fixed at launch. No release, no vote.
It waits on memory, not maths. Every hash is a chain of random reads into a table too big for a chip to carry. The wait is the same physics for everyone.
Miners hold the switch. Spare defences are written into the rules, switched off. A 90% miner signal turns one on. No fork.
A custom chip gains under 2x. Model published; tested by paid independent cryptanalysis and the public benchmark. [link: the numbers page]
Litepaper only, a fourth paragraph: No hash has stayed free of chips forever. Igneum does not claim to. It claims the gain is small, the response takes a week, and both are measured.
Site card placement: the Mine section of site/index.html beside "no chip can be built for it" (which this card replaces: the claim is a bounded gain, not impossibility). Litepaper placement: site/litepaper.html section mining, replacing the paragraph that begins "Everything above is automatic" and the "What Igneum does not claim" line on chips; the vs RandomX table keeps its rows, with the "Changes over time" row's Igneum cell reading "A new program every hour, its memory pattern and read widths with it; era draws and reserved families on a schedule fixed at genesis".
Level 3: the numbers page (site/bench.html, section "Counter ASIC")
Headline of the chip model (5 October 2026, night): the strongest chip holds the whole 256 MiB cache on-die (about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density, approximate) and computes dataset items on the fly; its gain over the RTX 5090 is 2.4x as the parameters stand, and no write-scratch share within an 8 GB card's budget changes that. The lever that does is the dataset item's mixer cost (x4: 1.8x with a 3x fixed-function factor, verifier 1.6 to 4.8 ms per warp). Decided 5 October 2026 (delegated): the mixer x4 and the cache growth rule enter class v3, so the headline row is the on-die-cache chip against v3 with everything combined. [owed: the combined row from docs/analysis/chip-model-v3.md; if it reads 1.8x, the claim is "under 2x" with the margin stated as thin, and the next levers are named: the mixer x8 and the hot table.]
Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per hash, the latency-bound share (rate over the card's random-read ceiling per load), the CPU verifier per warp, with machine, date and command. The chip model before and after Counter ASIC 2.0 (the m16 model's gain arithmetic at the v2 class and at the v3 class, with the SRAM a mirror needs, cited or approximate as the analysis says). The benchmark terms (spec O-1.17: the leaderboard by card model and the paid independent cryptanalysis's published findings, January 2027; no device bounty by the project lead's decision of 6 October 2026). Here the layers are named next to their numbers: read width, per-program mix, scratch, era layout, working set, hot table, cache schedule, the reserved integer-matrix family.
| Card | v2 MH/s | v3 MH/s (era packs, six eras) | Bytes per hash | Latency-bound share | Verifier ms per warp (v2 / v3, one loaded M5 Max core) | Daily 1 GiB build (v2 / v3) |
|---|---|---|---|---|---|---|
| Apple M5 Max (Metal) | 27.68 | 27.85 to 27.98 (spread 0.5%, the final class) | 512 | 1.06 | 0.61 / 2.08 (3.4x; worst cold 2.15) | 21 / 21 ms |
| RTX 5090 (CUDA) | 137.2 | 135.90 to 137.70 (spread 1.3%, the final class) | 512 | 1.01 | the same verifier | 25 / 23 ms |
| RX 9070 XT (OpenCL) | 18.09 | 18.59 to 19.18 (spread 3.1%, the final class) | 512 | 0.95 | the same verifier | 74 / 75 ms |
Every number measured 5 October 2026 (docs/plans/era-layout.md, docs/plans/mixer-x4.md, docs/bench-log.md); the v3 verifier figure is the x8 mixer on one core at load average 5.5 (the fixed crate; the same session matched readwidth's quiet v2 figure within 1%). Bit-exact: every v3 pack's fingerprint equal on the three vendors.
Chip model, before and after (docs/analysis/chip-model-v3.md): the on-die-cache recompute chip (the whole 256 MiB cache in SRAM, about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density, approximate) against the RTX 5090's measured 136.1 MH/s at 50 T integer op/s: class v2 333 MH/s, 2.4x; class v3 (mixer x8) 41.7 MH/s, 0.31x bare, 0.92x with a 3x fixed-function allowance (approximate), 0.76x at equal silicon. The claim "under 2x" holds with the margin stated: 8% on the allowance (a 3.3x allowance reads 1.0x), 9% on the budget. Next levers, named: the mixer at x16 (the verifier at about 4 ms per warp, inside the 10 ms gate; a 2019-class core unmeasured), a hot table small enough to stay resident beside the streaming dataset (measured and not adopted tonight: 32 to 96 MiB tables cost the GPU 7 to 20% and help the chip).
Levers measured and not adopted (5 October 2026): wider reads (16 and 64 B: no card gains, the 5090 goes bandwidth-bound at 64 B), the per-load width mix (5.5 to 22.3% spread), a per-warp write scratch (the chip keeps it implicitly: 2.4x at every share), the hot table (above). Reserved, switched off: the integer matrix family R1 (mm8, native on all three vendors as a tile; dp4a 1.17x a step on the 5090, 1.06x on the 9070 XT, emulation 1.6x on Apple) and the epoch length (600 s to 2 hours by 90% signal; a per-program FPGA bitstream mines 0% of a 600-s epoch).
AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), 2.2x worse per pound at list prices and 4.9x worse per watt (read-width.md section 4.1, approximate); the card's memory system, not a tuning gap.
Chip model, before and after: [owed: from docs/analysis/sram-mirror.md after the shipped-density correction (256 MiB on-die at about 130 to 165 mm^2 by AMD 3D V-Cache and TSMC N5 macro density, 54 to 83 mm^2 bit-cell-only lower bound), the hot-table and scratch analyses; the on-die-cache recompute chip is a named row per variant].
Level 4: the analysis documents
docs/plans/counter-asic-2.md (the plan and the layer table), docs/plans/read-width.md, docs/plans/era-layout.md, docs/plans/hot-table.md, docs/analysis/scratch-soundness.md, docs/analysis/sram-mirror.md, docs/analysis/int8-matrix-family.md, docs/analysis/m16-recompute-attacker-2026-10-05.md, the bench log entries of 5 October 2026 (night), the specification sections 1.4 to 1.13.